Skip to content

Replication: outbound subscription to a restarted peer wedges connected:false with no reconnect attempt (no SYN), reconcile never re-drives — base-copy + post-restart TLS window #466

Description

@kriszyp

Summary

On a 4-node preprod cluster (harper-pro 5.1.10), when a peer node restarts, some peers fail to re-establish their outbound replication subscription to it and get stuck connected:false with no auto-recovery — no reconnect attempt is ever made again (no socket, no SYN, no scheduled retry timer), and reconcileWorkers/findWedgedNodeUrls never re-drives the entry. The failure is stochastic per node-pair and correlates with reconnecting to a freshly-restarted peer that immediately requests a base-copy resync. Restarting a stuck node just relocates the stuck pair (whack-a-mole).

This is the #420/#424 connection-recovery family and a specific, code-grounded reproduction of #289 (whose root cause was never confirmed). It is distinct from #461 (deploy-orphaning), #454/#453 (the hdb_analytics / system-blob copy-stall watchdog work — though it touches the same base-copy machinery), and #352 (records decode fine here).

Note: the same Client network socket disconnected before secure TLS connection was established error was previously seen in a macOS integration-test harness and written off as a loopback artifact. This live Linux occurrence shows it is a real reconnect-after-restart bug, not a test artifact.

Live evidence (captured read-only while reproducing)

Cluster: 4 nodes, harper-pro 5.1.10. Stuck pair: plr -> iid (plr is stuck on its outbound subscription to iid; iid had restarted and is healthy 9/9). Captured ~13-24 min after the wedge began; state was static across re-polls.

plr cluster_status, iid node block — all three dbs connected:false / lastReceivedStatus:"Waiting", with non-zero (frozen) lastReceivedVersion:

db threadId connected note
harperfast_nextjs 1 false frozen lastReceivedVersion
<cacheDB> 2 false frozen lastReceivedVersion
system 3 false sendingMessage:"Copying", backPressurePercent ~ 99.9998%

plr logs (timestamps intact):

00:16:30  [http/3] Disconnected from wss://<iid>:9933 (db: "system") (code: 1006)
00:16:31  [http/N] Error in connection to wss://<iid>:9933 ... before secure TLS connection was established   (repeats http/1,2,3 -> 00:16:38)
00:16:46  [http/3] Connected to wss://<iid>:9933, db: system
00:16:46  [http/1] Connected to wss://<iid>:9933, db: harperfast_nextjs
00:16:46  [http/2] Connected to wss://<iid>:9933, db: <cacheDB>
00:16:57  [http/14] Peer <iid> requested replication of database system from 2026-06-16..., which predates retained transaction-log history ... forcing a bounded base-copy resync.

LAST iid-related log line: 00:16:57. The retry loop went silent and stayed silent. After 00:16:57 there is nothing iid-related — no further Connected/Disconnected/Failed to connect/Error in connection/Receive watchdog lines for iid.

  • Reconciling replication subscriptions: ZERO in the entire wedge window (40 min observed). findWedgedNodeUrls is not flagging the connected:false plr->iid entry.
  • Receive watchdog NEVER fired for iid — but it fires reliably every 60s for a different peer's system db on the same plr node. So the watchdog mechanism is alive; it simply never armed (or was stopped) for the iid leg.
  • plr maintains healthy ESTABLISHED :9933 sockets to the other two peers — this is iid-subscription-specific, not a global plr replication outage.

OS / libuv (CDP + ss):

  • Zero outbound TCP from plr to iid :9933 in any state (not ESTABLISHED, not SYN-SENT). _getActiveHandles() on workers 1/2/3 shows no socket and no Timeout handle — i.e. no reconnect timer is armed. process.report().libuv shows a single generic timer, no batch of retry timers.
  • The inbound leg from iid (plr serving the base-copy to iid) has a ~2.8 MB stuck send-queue (ss Send-Q 2,888,192), matching the thread-3 system Copying/99.99% backpressure. This is plr's server-accepted send leg — it has no reconnect object and only self-recovers via the receive watchdog / idle-ping terminate.

(The per-connection object fields — reconnectScheduled/intentionallyUnsubscribed/isFinished — could not be read via CDP: harper-pro is an ESM build, the module-scoped connections/connectionReplicationMap maps are not exported or on any global, and require/import() are not wired into the inspector eval context. Confirming those fields directly needs a debug hook exposing the maps, or a Runtime.queryObjects heap harness. See "Next diagnostic".)

Root cause analysis (code-grounded)

All line references are v5.1.10, files replication/replicationConnection.ts and replication/subscriptionManager.ts.

The outbound subscription leg recovered once (00:16:46) and then died again ~00:16:57 during/right after iid's base-copy request, and was never re-driven. Recovery for a connected:false subscription has exactly three drivers, and in this scenario all three are defeated:

1. The close-handler retry can be skipped, and the re-driven connect() can reject without ever rescheduling (PRIMARY, code-confirmed gap)

NodeReplicationConnection.connect() (line 717) schedules its only retry from inside the socket's close handler (line 847) or from forceReconnect() (line 907). Both do:

setTimeout(() => { this.connect(); }, this.retryTime).unref();

connect() is called with no .catch() and is not awaited. connect() first does await createWebSocket(this.url, ...) (line 722). createWebSocket (line 603) is async and can throw/reject before any socket exists — e.g. line 635 throw new Error('Unable to find a valid certificate to use for replication to connect to ...'), or a rejection from SNICallback.initialize() (line 620) while replication TLS secure-contexts are being (re)built. When that happens:

  • the finally clears this.reconnectScheduled = false (line 731) — but
  • the throw propagates out of connect() as an unhandled rejection, the open/error/close listeners are never attached (we threw before line 740), and nothing reschedules another connect(). The connection is left with no socket and no pending timer — exactly the observed state.

The freshly-restarted-peer correlation fits: right after a peer restart there is a window where CA / secure-context state is being rebuilt (monitorNodeCAs -> tls.createSecureContext, and the secureContexts/replicationSecureContext cache in createWebSocket), so createWebSocket is most likely to transiently throw precisely during the post-restart reconnect — and it is per-pair stochastic because it depends on the timing of that rebuild vs. the retry tick.

2. The close handler's reconnectScheduled early-return (line 844)

If reconnectScheduled is true when close fires (set by a forceReconnect, or by a prior schedule whose connect() is mid-flight), the close handler returns early without scheduling its own retry (lines 842-844). Combined with (1) — if that scheduled connect() then rejects in createWebSocket — there is no surviving retry path. reconnectScheduled is only ever cleared in connect()'s finally; nothing re-arms it afterward.

3. The receive watchdog backstop is disabled during backpressure and only re-arms on resume (line 1166)

addPauseReason() calls receiveWatchdog?.stop() when the WS pauses for backpressure (line 1166); it is only restarted in removePauseReason() when pauseReasons returns to 0 (line 1177). During the system base copy plr sat at ~99.99% backpressure, so the watchdog was stopped. Likewise the active sendPing keep-alive is exempt from terminating a paused connection (shouldTerminateIdlePing returns false when pauseReasons > 0, line 227; lastByteActivity is kept fresh at line 1069). So while a leg is paused for backpressure, neither watchdog can fire forceReconnect() — removing the #424 recovery path for exactly this state.

Why findWedgedNodeUrls (the main-thread reconcile backstop) does not catch it

findWedgedNodeUrls (subscriptionManager.ts line 161) only flags an entry when all of:

entry.connected === false && entry.worker && httpWorkers.includes(entry.worker)
  && entry.disconnectedAt != null && now - entry.disconnectedAt >= 30_000
  && isDesired(entry.nodes?.[0], database)

entry.disconnectedAt is set only in disconnectedFromNode (line 606), which runs only from the worker's close handler (line 816) or forceReconnect() (line 887). The entry is first created with no connected and no disconnectedAt field (lines 481-485). So whenever the connected:false transition happens without disconnectedFromNode running — i.e. the connect()-rejected path in (1), where no socket/close ever exists for that attempt — the entry can be left connected:false (from the earlier real disconnect, possibly later re-stamped) without a fresh disconnectedAt, or with state that fails the predicate, so the 30s wedge net never engages. The zero Reconciling log lines confirm the predicate is not matching.

Net: a subscription that drops during a post-restart base-copy can land in connected:false with (a) no scheduled connect() retry, (b) the watchdog stopped by backpressure, and (c) the main-thread reconcile not classifying it wedged — a permanent wedge until process restart.

Confidence

  • Code-confirmed: the setTimeout(() => this.connect()) calls are not .catch()-ed and createWebSocket can reject before listeners attach (mechanism 1); the reconnectScheduled early-return (2); the watchdog stop()-on-pause without re-arm while paused (3); the findWedgedNodeUrls dependency on disconnectedAt which the entry-creation path does not set. These are all directly visible in v5.1.10 source.
  • Hypothesis (not yet directly confirmed on the live host): which of (1)/(2) actually fired for this specific pair, i.e. whether reconnectScheduled is stuck true vs. a connect() rejection swallowed the only retry. The live OS/libuv state (no socket + no timer + connected:false + zero reconcile + watchdog never armed) is consistent with all three gaps, but the per-connection flags could not be read via CDP (ESM eval-surface limitation).

Proposed fix direction (for discussion — no PR yet)

  1. Make connect() self-healing on rejection. Wrap the scheduled setTimeout(() => this.connect()) (lines 847, 907) so a rejection reschedules with backoff instead of vanishing, e.g. this.connect().catch(() => this.scheduleReconnect()), or move the await createWebSocket inside a try/catch in connect() that funnels failures into the same retry path the close handler uses. This closes the primary give-up hole and makes the secure TLS post-restart failure self-recover.
  2. Re-arm the receive watchdog under sustained backpressure (or run a separate liveness check that is not suppressed by pauseReasons), so a leg that dies while paused still triggers forceReconnect(). Today stop()-on-pause removes the #424 net for exactly the base-copy case.
  3. Tighten the main-thread reconcile so it does not depend solely on disconnectedAt. findWedgedNodeUrls should also catch a long-lived connected:false (or never-connected) entry whose worker is live regardless of whether disconnectedAt was stamped — e.g. fall back to a "no positive connected:true confirmation within N x interval" signal — so the backstop covers transitions that bypass disconnectedFromNode.

These are independent and layered (drive-side fix + watchdog backstop + reconcile backstop), matching the existing belt-and-suspenders design.

Next diagnostic (to confirm the hypothesis branch)

Expose the module-scoped connections / connectionReplicationMap maps on a debug global (or add a Runtime.queryObjects-based harness) so a future repro can read, per wedged (peer, db): reconnectScheduled, intentionallyUnsubscribed, isFinished, retryTime, whether this.socket exists, and the main entry's disconnectedAt / worker. That single read distinguishes mechanism (1) from (2) definitively.


Filed from a read-only live investigation (customer redacted: a 4-node preprod cluster). Analysis generated by Claude (Opus 4.8). Holding on a PR pending root-cause confirmation.

Activity

  1. added
    bugSomething isn't working
    area:replicationReplication, cluster sync, peer connections
    on Jun 24, 2026
  2. self-assigned this
    on Jun 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area:replicationReplication, cluster sync, peer connectionsbugSomething isn't working

Type

No type

Fields

Priority

None yet

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions