Skip to content

Run replication on a dedicated worker pool (replication.threads) #975

Description

@kriszyp

Part of #435 (W6, dedicated replication threads; design in this comment). Depends on the core worker type: HarperFast/harper#3028.

Problem

Replication shares event loops with application request handling. The sending side, meaning audit streaming, base copy and egress back-pressure, runs on the inbound replication socket. That socket lands on an arbitrary HTTP worker because every worker binds the replication port with SO_REUSEPORT. So moving only outbound subscriptions would leave most of the load where it is.

Scope

When replication.threads > 0:

  • One worker selector replaces the four worker.name === 'http' filters: subscriptionManager.ts:1427,1699,1798 and recordLockTransport.ts:962 (httpWorkers()). It selects the replication pool when present and otherwise the non-isolated HTTP workers. The second half is the fix for Replication places subscriptions and record-lock ownership on isolated application workers #974.
  • Inbound listener on the pool only: replicator.start() (server.ws/server.http, replicator.ts:159,202) registers only on pool workers, using core's exclusive port ownership.
  • Dedicated port required: with neither replication.port nor securePort set, replicator.ts:133-136 falls back to the operations API ports, which the main thread binds exclusively. With replication.threads > 0 in that configuration, refuse at startup with a clear error.
  • Key custody: the workerData provider in security/keyCustody.ts:111 (http-only today) also covers the pool.
  • Forwarded operations: OPERATION_REQUEST (replicationConnection.ts:4944) works on pool workers. serverUtilities and the registered operations must be loaded there.
  • Worker readiness: replace whenWorkerComponentsLoaded (subscriptionManager.ts:2163-2175) with the pool's ready signal.
  • Worker exit and reconcile (subscriptionManager.ts:1745-1760,1796) work over the pool.
  • Record locks: with replication.recordLocks, lock owners are pool workers. A lock() on an HTTP worker relays over the thread mesh (recordLockRpc.ts). Accept that cost; measure it alongside Record locks: measure the Phase 1 cost baseline before the protocol change #824.
  • Cache-miss retrieval stays on request threads (Replicator.load, replicator.ts:458-515, nodeNameToRetrievalConnections). Do not bridge it through the pool.
  • Pace replication sender yields with a worker time budget #959 (sender time budget): apply the 2 ms budget only on pool workers. HTTP workers keep the per-record yield so catch-up cannot starve application requests.

Acceptance

  • With replication.threads: 2, every inbound and outbound replication socket is on a pool worker (assert via cluster_status connection ownership). HTTP workers hold only cache-miss retrieval connections.
  • Existing cluster integration suites pass with replication.threads at 0 and at 2. Run the replication cluster suites in CI with the pool on.
  • Record-lock cluster tests pass with the pool on.
  • Stress: integrationTests/stress/largeCatchup.test.mjs at 10 GB with concurrent HTTP load. Report request p50/p99 latency and catch-up throughput, pool off vs on.
  • DESIGN.md documents thread placement.

🤖 Claude Opus 5.5 on behalf of Kris.

Metadata

Metadata

Assignees

Labels

No labels
No labels

Fields

Priority

P1

Projects

No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions