Summary
On a 4-node preprod cluster (harper-pro 5.1.10), when a peer node restarts, some peers fail to re-establish their outbound replication subscription to it and get stuck connected:false with no auto-recovery — no reconnect attempt is ever made again (no socket, no SYN, no scheduled retry timer), and reconcileWorkers/findWedgedNodeUrls never re-drives the entry. The failure is stochastic per node-pair and correlates with reconnecting to a freshly-restarted peer that immediately requests a base-copy resync. Restarting a stuck node just relocates the stuck pair (whack-a-mole).
This is the #420/#424 connection-recovery family and a specific, code-grounded reproduction of #289 (whose root cause was never confirmed). It is distinct from #461 (deploy-orphaning), #454/#453 (the hdb_analytics / system-blob copy-stall watchdog work — though it touches the same base-copy machinery), and #352 (records decode fine here).
Note: the same Client network socket disconnected before secure TLS connection was established error was previously seen in a macOS integration-test harness and written off as a loopback artifact. This live Linux occurrence shows it is a real reconnect-after-restart bug, not a test artifact.
Live evidence (captured read-only while reproducing)
Cluster: 4 nodes, harper-pro 5.1.10. Stuck pair: plr -> iid (plr is stuck on its outbound subscription to iid; iid had restarted and is healthy 9/9). Captured ~13-24 min after the wedge began; state was static across re-polls.
plr cluster_status, iid node block — all three dbs connected:false / lastReceivedStatus:"Waiting", with non-zero (frozen) lastReceivedVersion:
| db |
threadId |
connected |
note |
harperfast_nextjs |
1 |
false |
frozen lastReceivedVersion |
<cacheDB> |
2 |
false |
frozen lastReceivedVersion |
system |
3 |
false |
sendingMessage:"Copying", backPressurePercent ~ 99.9998% |
plr logs (timestamps intact):
00:16:30 [http/3] Disconnected from wss://<iid>:9933 (db: "system") (code: 1006)
00:16:31 [http/N] Error in connection to wss://<iid>:9933 ... before secure TLS connection was established (repeats http/1,2,3 -> 00:16:38)
00:16:46 [http/3] Connected to wss://<iid>:9933, db: system
00:16:46 [http/1] Connected to wss://<iid>:9933, db: harperfast_nextjs
00:16:46 [http/2] Connected to wss://<iid>:9933, db: <cacheDB>
00:16:57 [http/14] Peer <iid> requested replication of database system from 2026-06-16..., which predates retained transaction-log history ... forcing a bounded base-copy resync.
LAST iid-related log line: 00:16:57. The retry loop went silent and stayed silent. After 00:16:57 there is nothing iid-related — no further Connected/Disconnected/Failed to connect/Error in connection/Receive watchdog lines for iid.
Reconciling replication subscriptions: ZERO in the entire wedge window (40 min observed). findWedgedNodeUrls is not flagging the connected:false plr->iid entry.
Receive watchdog NEVER fired for iid — but it fires reliably every 60s for a different peer's system db on the same plr node. So the watchdog mechanism is alive; it simply never armed (or was stopped) for the iid leg.
- plr maintains healthy ESTABLISHED
:9933 sockets to the other two peers — this is iid-subscription-specific, not a global plr replication outage.
OS / libuv (CDP + ss):
- Zero outbound TCP from plr to iid
:9933 in any state (not ESTABLISHED, not SYN-SENT). _getActiveHandles() on workers 1/2/3 shows no socket and no Timeout handle — i.e. no reconnect timer is armed. process.report().libuv shows a single generic timer, no batch of retry timers.
- The inbound leg from iid (plr serving the base-copy to iid) has a ~2.8 MB stuck send-queue (
ss Send-Q 2,888,192), matching the thread-3 system Copying/99.99% backpressure. This is plr's server-accepted send leg — it has no reconnect object and only self-recovers via the receive watchdog / idle-ping terminate.
(The per-connection object fields — reconnectScheduled/intentionallyUnsubscribed/isFinished — could not be read via CDP: harper-pro is an ESM build, the module-scoped connections/connectionReplicationMap maps are not exported or on any global, and require/import() are not wired into the inspector eval context. Confirming those fields directly needs a debug hook exposing the maps, or a Runtime.queryObjects heap harness. See "Next diagnostic".)
Root cause analysis (code-grounded)
All line references are v5.1.10, files replication/replicationConnection.ts and replication/subscriptionManager.ts.
The outbound subscription leg recovered once (00:16:46) and then died again ~00:16:57 during/right after iid's base-copy request, and was never re-driven. Recovery for a connected:false subscription has exactly three drivers, and in this scenario all three are defeated:
1. The close-handler retry can be skipped, and the re-driven connect() can reject without ever rescheduling (PRIMARY, code-confirmed gap)
NodeReplicationConnection.connect() (line 717) schedules its only retry from inside the socket's close handler (line 847) or from forceReconnect() (line 907). Both do:
setTimeout(() => { this.connect(); }, this.retryTime).unref();
connect() is called with no .catch() and is not awaited. connect() first does await createWebSocket(this.url, ...) (line 722). createWebSocket (line 603) is async and can throw/reject before any socket exists — e.g. line 635 throw new Error('Unable to find a valid certificate to use for replication to connect to ...'), or a rejection from SNICallback.initialize() (line 620) while replication TLS secure-contexts are being (re)built. When that happens:
- the
finally clears this.reconnectScheduled = false (line 731) — but
- the throw propagates out of
connect() as an unhandled rejection, the open/error/close listeners are never attached (we threw before line 740), and nothing reschedules another connect(). The connection is left with no socket and no pending timer — exactly the observed state.
The freshly-restarted-peer correlation fits: right after a peer restart there is a window where CA / secure-context state is being rebuilt (monitorNodeCAs -> tls.createSecureContext, and the secureContexts/replicationSecureContext cache in createWebSocket), so createWebSocket is most likely to transiently throw precisely during the post-restart reconnect — and it is per-pair stochastic because it depends on the timing of that rebuild vs. the retry tick.
2. The close handler's reconnectScheduled early-return (line 844)
If reconnectScheduled is true when close fires (set by a forceReconnect, or by a prior schedule whose connect() is mid-flight), the close handler returns early without scheduling its own retry (lines 842-844). Combined with (1) — if that scheduled connect() then rejects in createWebSocket — there is no surviving retry path. reconnectScheduled is only ever cleared in connect()'s finally; nothing re-arms it afterward.
3. The receive watchdog backstop is disabled during backpressure and only re-arms on resume (line 1166)
addPauseReason() calls receiveWatchdog?.stop() when the WS pauses for backpressure (line 1166); it is only restarted in removePauseReason() when pauseReasons returns to 0 (line 1177). During the system base copy plr sat at ~99.99% backpressure, so the watchdog was stopped. Likewise the active sendPing keep-alive is exempt from terminating a paused connection (shouldTerminateIdlePing returns false when pauseReasons > 0, line 227; lastByteActivity is kept fresh at line 1069). So while a leg is paused for backpressure, neither watchdog can fire forceReconnect() — removing the #424 recovery path for exactly this state.
Why findWedgedNodeUrls (the main-thread reconcile backstop) does not catch it
findWedgedNodeUrls (subscriptionManager.ts line 161) only flags an entry when all of:
entry.connected === false && entry.worker && httpWorkers.includes(entry.worker)
&& entry.disconnectedAt != null && now - entry.disconnectedAt >= 30_000
&& isDesired(entry.nodes?.[0], database)
entry.disconnectedAt is set only in disconnectedFromNode (line 606), which runs only from the worker's close handler (line 816) or forceReconnect() (line 887). The entry is first created with no connected and no disconnectedAt field (lines 481-485). So whenever the connected:false transition happens without disconnectedFromNode running — i.e. the connect()-rejected path in (1), where no socket/close ever exists for that attempt — the entry can be left connected:false (from the earlier real disconnect, possibly later re-stamped) without a fresh disconnectedAt, or with state that fails the predicate, so the 30s wedge net never engages. The zero Reconciling log lines confirm the predicate is not matching.
Net: a subscription that drops during a post-restart base-copy can land in connected:false with (a) no scheduled connect() retry, (b) the watchdog stopped by backpressure, and (c) the main-thread reconcile not classifying it wedged — a permanent wedge until process restart.
Confidence
- Code-confirmed: the
setTimeout(() => this.connect()) calls are not .catch()-ed and createWebSocket can reject before listeners attach (mechanism 1); the reconnectScheduled early-return (2); the watchdog stop()-on-pause without re-arm while paused (3); the findWedgedNodeUrls dependency on disconnectedAt which the entry-creation path does not set. These are all directly visible in v5.1.10 source.
- Hypothesis (not yet directly confirmed on the live host): which of (1)/(2) actually fired for this specific pair, i.e. whether
reconnectScheduled is stuck true vs. a connect() rejection swallowed the only retry. The live OS/libuv state (no socket + no timer + connected:false + zero reconcile + watchdog never armed) is consistent with all three gaps, but the per-connection flags could not be read via CDP (ESM eval-surface limitation).
Proposed fix direction (for discussion — no PR yet)
- Make
connect() self-healing on rejection. Wrap the scheduled setTimeout(() => this.connect()) (lines 847, 907) so a rejection reschedules with backoff instead of vanishing, e.g. this.connect().catch(() => this.scheduleReconnect()), or move the await createWebSocket inside a try/catch in connect() that funnels failures into the same retry path the close handler uses. This closes the primary give-up hole and makes the secure TLS post-restart failure self-recover.
- Re-arm the receive watchdog under sustained backpressure (or run a separate liveness check that is not suppressed by
pauseReasons), so a leg that dies while paused still triggers forceReconnect(). Today stop()-on-pause removes the #424 net for exactly the base-copy case.
- Tighten the main-thread reconcile so it does not depend solely on
disconnectedAt. findWedgedNodeUrls should also catch a long-lived connected:false (or never-connected) entry whose worker is live regardless of whether disconnectedAt was stamped — e.g. fall back to a "no positive connected:true confirmation within N x interval" signal — so the backstop covers transitions that bypass disconnectedFromNode.
These are independent and layered (drive-side fix + watchdog backstop + reconcile backstop), matching the existing belt-and-suspenders design.
Next diagnostic (to confirm the hypothesis branch)
Expose the module-scoped connections / connectionReplicationMap maps on a debug global (or add a Runtime.queryObjects-based harness) so a future repro can read, per wedged (peer, db): reconnectScheduled, intentionallyUnsubscribed, isFinished, retryTime, whether this.socket exists, and the main entry's disconnectedAt / worker. That single read distinguishes mechanism (1) from (2) definitively.
Filed from a read-only live investigation (customer redacted: a 4-node preprod cluster). Analysis generated by Claude (Opus 4.8). Holding on a PR pending root-cause confirmation.
Summary
On a 4-node preprod cluster (harper-pro 5.1.10), when a peer node restarts, some peers fail to re-establish their outbound replication subscription to it and get stuck
connected:falsewith no auto-recovery — no reconnect attempt is ever made again (no socket, no SYN, no scheduled retry timer), andreconcileWorkers/findWedgedNodeUrlsnever re-drives the entry. The failure is stochastic per node-pair and correlates with reconnecting to a freshly-restarted peer that immediately requests a base-copy resync. Restarting a stuck node just relocates the stuck pair (whack-a-mole).This is the
#420/#424connection-recovery family and a specific, code-grounded reproduction of#289(whose root cause was never confirmed). It is distinct from#461(deploy-orphaning),#454/#453(the hdb_analytics / system-blob copy-stall watchdog work — though it touches the same base-copy machinery), and#352(records decode fine here).Live evidence (captured read-only while reproducing)
Cluster: 4 nodes, harper-pro 5.1.10. Stuck pair: plr -> iid (plr is stuck on its outbound subscription to iid; iid had restarted and is healthy 9/9). Captured ~13-24 min after the wedge began; state was static across re-polls.
plr
cluster_status, iid node block — all three dbsconnected:false/lastReceivedStatus:"Waiting", with non-zero (frozen)lastReceivedVersion:harperfast_nextjs<cacheDB>systemsendingMessage:"Copying",backPressurePercent ~ 99.9998%plr logs (timestamps intact):
LAST iid-related log line: 00:16:57. The retry loop went silent and stayed silent. After 00:16:57 there is nothing iid-related — no further
Connected/Disconnected/Failed to connect/Error in connection/Receive watchdoglines for iid.Reconciling replication subscriptions: ZERO in the entire wedge window (40 min observed).findWedgedNodeUrlsis not flagging theconnected:falseplr->iid entry.Receive watchdogNEVER fired for iid — but it fires reliably every 60s for a different peer'ssystemdb on the same plr node. So the watchdog mechanism is alive; it simply never armed (or was stopped) for the iid leg.:9933sockets to the other two peers — this is iid-subscription-specific, not a global plr replication outage.OS / libuv (CDP +
ss)::9933in any state (not ESTABLISHED, not SYN-SENT)._getActiveHandles()on workers 1/2/3 shows no socket and noTimeouthandle — i.e. no reconnect timer is armed.process.report().libuvshows a single generic timer, no batch of retry timers.ssSend-Q 2,888,192), matching the thread-3systemCopying/99.99% backpressure. This is plr's server-accepted send leg — it has no reconnect object and only self-recovers via the receive watchdog / idle-ping terminate.(The per-connection object fields —
reconnectScheduled/intentionallyUnsubscribed/isFinished— could not be read via CDP: harper-pro is an ESM build, the module-scopedconnections/connectionReplicationMapmaps are not exported or on any global, andrequire/import()are not wired into the inspector eval context. Confirming those fields directly needs a debug hook exposing the maps, or aRuntime.queryObjectsheap harness. See "Next diagnostic".)Root cause analysis (code-grounded)
All line references are
v5.1.10, filesreplication/replicationConnection.tsandreplication/subscriptionManager.ts.The outbound subscription leg recovered once (00:16:46) and then died again ~00:16:57 during/right after iid's base-copy request, and was never re-driven. Recovery for a
connected:falsesubscription has exactly three drivers, and in this scenario all three are defeated:1. The close-handler retry can be skipped, and the re-driven
connect()can reject without ever rescheduling (PRIMARY, code-confirmed gap)NodeReplicationConnection.connect()(line 717) schedules its only retry from inside the socket'sclosehandler (line 847) or fromforceReconnect()(line 907). Both do:connect()is called with no.catch()and is not awaited.connect()first doesawait createWebSocket(this.url, ...)(line 722).createWebSocket(line 603) isasyncand can throw/reject before any socket exists — e.g. line 635throw new Error('Unable to find a valid certificate to use for replication to connect to ...'), or a rejection fromSNICallback.initialize()(line 620) while replication TLS secure-contexts are being (re)built. When that happens:finallyclearsthis.reconnectScheduled = false(line 731) — butconnect()as an unhandled rejection, theopen/error/closelisteners are never attached (we threw before line 740), and nothing reschedules anotherconnect(). The connection is left with no socket and no pending timer — exactly the observed state.The freshly-restarted-peer correlation fits: right after a peer restart there is a window where CA / secure-context state is being rebuilt (
monitorNodeCAs->tls.createSecureContext, and thesecureContexts/replicationSecureContextcache increateWebSocket), socreateWebSocketis most likely to transiently throw precisely during the post-restart reconnect — and it is per-pair stochastic because it depends on the timing of that rebuild vs. the retry tick.2. The close handler's
reconnectScheduledearly-return (line 844)If
reconnectScheduledistruewhenclosefires (set by aforceReconnect, or by a prior schedule whoseconnect()is mid-flight), the close handler returns early without scheduling its own retry (lines 842-844). Combined with (1) — if that scheduledconnect()then rejects increateWebSocket— there is no surviving retry path.reconnectScheduledis only ever cleared inconnect()'sfinally; nothing re-arms it afterward.3. The receive watchdog backstop is disabled during backpressure and only re-arms on resume (line 1166)
addPauseReason()callsreceiveWatchdog?.stop()when the WS pauses for backpressure (line 1166); it is only restarted inremovePauseReason()whenpauseReasonsreturns to 0 (line 1177). During thesystembase copy plr sat at ~99.99% backpressure, so the watchdog was stopped. Likewise the activesendPingkeep-alive is exempt from terminating a paused connection (shouldTerminateIdlePingreturns false whenpauseReasons > 0, line 227;lastByteActivityis kept fresh at line 1069). So while a leg is paused for backpressure, neither watchdog can fireforceReconnect()— removing the#424recovery path for exactly this state.Why
findWedgedNodeUrls(the main-thread reconcile backstop) does not catch itfindWedgedNodeUrls(subscriptionManager.ts line 161) only flags an entry when all of:entry.disconnectedAtis set only indisconnectedFromNode(line 606), which runs only from the worker'sclosehandler (line 816) orforceReconnect()(line 887). The entry is first created with noconnectedand nodisconnectedAtfield (lines 481-485). So whenever theconnected:falsetransition happens withoutdisconnectedFromNoderunning — i.e. theconnect()-rejected path in (1), where no socket/close ever exists for that attempt — the entry can be leftconnected:false(from the earlier real disconnect, possibly later re-stamped) without a freshdisconnectedAt, or with state that fails the predicate, so the 30s wedge net never engages. The zeroReconcilinglog lines confirm the predicate is not matching.Net: a subscription that drops during a post-restart base-copy can land in
connected:falsewith (a) no scheduledconnect()retry, (b) the watchdog stopped by backpressure, and (c) the main-thread reconcile not classifying it wedged — a permanent wedge until process restart.Confidence
setTimeout(() => this.connect())calls are not.catch()-ed andcreateWebSocketcan reject before listeners attach (mechanism 1); thereconnectScheduledearly-return (2); the watchdogstop()-on-pause without re-arm while paused (3); thefindWedgedNodeUrlsdependency ondisconnectedAtwhich the entry-creation path does not set. These are all directly visible in v5.1.10 source.reconnectScheduledis stucktruevs. aconnect()rejection swallowed the only retry. The live OS/libuv state (no socket + no timer + connected:false + zero reconcile + watchdog never armed) is consistent with all three gaps, but the per-connection flags could not be read via CDP (ESM eval-surface limitation).Proposed fix direction (for discussion — no PR yet)
connect()self-healing on rejection. Wrap the scheduledsetTimeout(() => this.connect())(lines 847, 907) so a rejection reschedules with backoff instead of vanishing, e.g.this.connect().catch(() => this.scheduleReconnect()), or move theawait createWebSocketinside a try/catch inconnect()that funnels failures into the same retry path theclosehandler uses. This closes the primary give-up hole and makes thesecure TLSpost-restart failure self-recover.pauseReasons), so a leg that dies while paused still triggersforceReconnect(). Todaystop()-on-pause removes the#424net for exactly the base-copy case.disconnectedAt.findWedgedNodeUrlsshould also catch a long-livedconnected:false(or never-connected) entry whose worker is live regardless of whetherdisconnectedAtwas stamped — e.g. fall back to a "no positiveconnected:trueconfirmation within N x interval" signal — so the backstop covers transitions that bypassdisconnectedFromNode.These are independent and layered (drive-side fix + watchdog backstop + reconcile backstop), matching the existing belt-and-suspenders design.
Next diagnostic (to confirm the hypothesis branch)
Expose the module-scoped
connections/connectionReplicationMapmaps on a debug global (or add aRuntime.queryObjects-based harness) so a future repro can read, per wedged(peer, db):reconnectScheduled,intentionallyUnsubscribed,isFinished,retryTime, whetherthis.socketexists, and the main entry'sdisconnectedAt/worker. That single read distinguishes mechanism (1) from (2) definitively.Filed from a read-only live investigation (customer redacted: a 4-node preprod cluster). Analysis generated by Claude (Opus 4.8). Holding on a PR pending root-cause confirmation.