You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Run replication on a dedicated worker pool (replication.threads) #975
Replication shares event loops with application request handling. The sending side, meaning audit streaming, base copy and egress back-pressure, runs on the inbound replication socket. That socket lands on an arbitrary HTTP worker because every worker binds the replication port with SO_REUSEPORT. So moving only outbound subscriptions would leave most of the load where it is.
Inbound listener on the pool only:replicator.start() (server.ws/server.http, replicator.ts:159,202) registers only on pool workers, using core's exclusive port ownership.
Dedicated port required: with neither replication.port nor securePort set, replicator.ts:133-136 falls back to the operations API ports, which the main thread binds exclusively. With replication.threads > 0 in that configuration, refuse at startup with a clear error.
Key custody: the workerData provider in security/keyCustody.ts:111 (http-only today) also covers the pool.
Forwarded operations:OPERATION_REQUEST (replicationConnection.ts:4944) works on pool workers. serverUtilities and the registered operations must be loaded there.
Worker readiness: replace whenWorkerComponentsLoaded (subscriptionManager.ts:2163-2175) with the pool's ready signal.
Worker exit and reconcile (subscriptionManager.ts:1745-1760,1796) work over the pool.
Cache-miss retrieval stays on request threads (Replicator.load, replicator.ts:458-515, nodeNameToRetrievalConnections). Do not bridge it through the pool.
With replication.threads: 2, every inbound and outbound replication socket is on a pool worker (assert via cluster_status connection ownership). HTTP workers hold only cache-miss retrieval connections.
Existing cluster integration suites pass with replication.threads at 0 and at 2. Run the replication cluster suites in CI with the pool on.
Record-lock cluster tests pass with the pool on.
Stress: integrationTests/stress/largeCatchup.test.mjs at 10 GB with concurrent HTTP load. Report request p50/p99 latency and catch-up throughput, pool off vs on.
Part of #435 (W6, dedicated replication threads; design in this comment). Depends on the core worker type: HarperFast/harper#3028.
Problem
Replication shares event loops with application request handling. The sending side, meaning audit streaming, base copy and egress back-pressure, runs on the inbound replication socket. That socket lands on an arbitrary HTTP worker because every worker binds the replication port with SO_REUSEPORT. So moving only outbound subscriptions would leave most of the load where it is.
Scope
When
replication.threads > 0:worker.name === 'http'filters:subscriptionManager.ts:1427,1699,1798andrecordLockTransport.ts:962(httpWorkers()). It selects the replication pool when present and otherwise the non-isolated HTTP workers. The second half is the fix for Replication places subscriptions and record-lock ownership on isolated application workers #974.replicator.start()(server.ws/server.http,replicator.ts:159,202) registers only on pool workers, using core's exclusive port ownership.replication.portnorsecurePortset,replicator.ts:133-136falls back to the operations API ports, which the main thread binds exclusively. Withreplication.threads > 0in that configuration, refuse at startup with a clear error.workerDataprovider insecurity/keyCustody.ts:111(http-only today) also covers the pool.OPERATION_REQUEST(replicationConnection.ts:4944) works on pool workers.serverUtilitiesand the registered operations must be loaded there.whenWorkerComponentsLoaded(subscriptionManager.ts:2163-2175) with the pool's ready signal.subscriptionManager.ts:1745-1760,1796) work over the pool.replication.recordLocks, lock owners are pool workers. Alock()on an HTTP worker relays over the thread mesh (recordLockRpc.ts). Accept that cost; measure it alongside Record locks: measure the Phase 1 cost baseline before the protocol change #824.Replicator.load,replicator.ts:458-515,nodeNameToRetrievalConnections). Do not bridge it through the pool.Acceptance
replication.threads: 2, every inbound and outbound replication socket is on a pool worker (assert viacluster_statusconnection ownership). HTTP workers hold only cache-miss retrieval connections.replication.threadsat 0 and at 2. Run the replication cluster suites in CI with the pool on.integrationTests/stress/largeCatchup.test.mjsat 10 GB with concurrent HTTP load. Report request p50/p99 latency and catch-up throughput, pool off vs on.DESIGN.mddocuments thread placement.🤖 Claude Opus 5.5 on behalf of Kris.