Repository navigation
Replication W6: Adaptive / dedicated replication threads #435
Description
Activity
- addedenhancementNew feature or requestNew feature or requestarea:replicationReplication, cluster sync, peer connectionsReplication, cluster sync, peer connections
on Jun 20, 2026 - added a commit that references this issue
on Jul 29, 2026 Design for 5.4 (2026-10-05): a static replication worker pool, both directions together
Verified against harper-pro
origin/main(core submoduleda24feb66). Nothing from this issue exists in code yet: noreplication.threadssetting, no thread type, no branch. W1, the dependency listed above, has landed (#814, #843).Why outbound subscriptions alone are not enough
The obvious first step is to move only the outbound subscriptions, because they are already main-thread
postMessagedriven and keyed by thread ID. It would not achieve this issue's goal. An outbound connection only receives and applies. The sending load this issue is about runs on the inbound socket: streaming audit records to a subscribed peer, base copy, and egress back-pressure. So the pool has to own both the inbound replication listener and the outbound subscriptions to change anything.Plan
Core (harper): a general-purpose worker type
- Add
THREAD_TYPES.REPLICATION(todayhdbTerms.ts:984has onlyhttpandjob). Generalize the isolated-application slot machinery (socketRouter.ts:100-109,216-248, from harper#2524) rather than adding a second one-off pool. It already has indices past the HTTP pool, heap-share accounting and scoped restart; Multi-tenant within process: SNI-based instance routing with isolated workers #247 would reuse the same capability. - Restart and deploy wiring:
restartWorkers('http', …)callers (bin/restart.ts,components/operations.js) must restart the pool when needed, since a replication-only change has to reach it. - Readiness: pool workers load databases and tables but no application code, and signal readiness without
threadServer.js. - Port-binding rule: a worker type can own a port exclusively. This generalizes the isolated-worker skip at
threadServer.js:376. - Broadcasts already reach every non-job port (
manageThreads.js:359), so pool workers get schema and other cross-thread broadcasts with no change.
harper-pro: route replication to the pool
- One worker selector replacing the four
worker.name === 'http'filters (subscriptionManager.ts:1427,1699,1798;recordLockTransport.ts:962). It selects the pool when present and otherwise the non-isolated HTTP workers. The second half fixes the isolated-worker leak filed separately. replicator.start()(theserver.ws/server.httpregistration atreplicator.ts:159,202) runs only on pool workers, so only they bind the replication port.- Key custody: material goes only to
name === 'http'workers (security/keyCustody.ts:111). The pool needs it too. - Forwarded operations (
OPERATION_REQUEST,replicationConnection.ts:4944): pool workers loadserverUtilitiesand the registered operations. - Cache-miss retrieval stays on request threads (
Replicator.load,replicator.ts:458-515, which opens its own per-thread connection). Bridging it through the pool would serialize records and blobs across threads for no clear gain. This is the "split streaming from retrieval" item above, done by leaving retrieval where it is. - Record locks: when
replication.recordLocksis on, lock owners live in the pool (the owner must be the thread applying the database's inbound entries,DESIGN.md:106). Everylock()then relays over the thread mesh, measured at about 1 ms under owner load. We accept that cost rather than keeping two placement rules. Record locks: measure the Phase 1 cost baseline before the protocol change #824's cost baseline should measure the pool case.
What needs no change:
- Replicated commits already notify subscribers on every thread (core
transactionBroadcast.ts:66-90). replicateToconfirmations wait on a shared buffer (knownNodes.ts:954-978), not on a connection object.server.nodesis filled per thread by thehdb_nodeswatcher.- Shared-status buffers are process-wide, and ownership is by thread ID.
Configuration and rollout
replication.threads, default 0 in 5.4, which keeps today's behavior (replication on the HTTP workers). Flip the default after soak.- The pool requires a dedicated replication port. With neither
replication.portnorsecurePortset,replicator.ts:133-136falls back to the operations API ports, which main binds exclusively and which cannot be routed to a pool. Withreplication.threads > 0in that configuration, refuse at startup with a clear error rather than silently not using the pool. - macOS sets
noReusePorton HTTP servers, so one pool worker wins the port there. Linux spreads connections across the pool with SO_REUSEPORT.
Relationship to #959 (sender time budget)
#959 replaces the per-record yield with a 2 ms worker-local budget. On a shared HTTP worker that lets a catch-up sender hold the event loop for longer slices, which risks starving application request handling. #959 should land with or after this pool, and on non-pool workers it should keep per-record yields (or apply the budget only on pool workers). With a dedicated pool, the longer slices only delay other replication work.
Acceptance
integrationTests/stress/largeCatchup.test.mjsat 10 GB with concurrent HTTP load: request latency (p50/p99) with and without the pool, and catch-up throughput. This also feeds W13's (Replication W13: Base-copy & catch-up path (consolidation & correctness) #510) open question about bimodal replay throughput.- The back-pressure signal (slot 6, now on the main-thread entry via W1 R3) is no longer inflated by HTTP-induced event-loop stalls. That gives W5 (Adaptive replication routing: dynamic fan-out under egress back-pressure #218) a clean input.
Sub-issues
- A replication worker type: dedicated threads that run Harper's built-in services without application code harper#3028: the core replication worker type
- Run replication on a dedicated worker pool (replication.threads) #975: route replication onto the pool (
replication.threads) - Replication places subscriptions and record-lock ownership on isolated application workers #974: the existing isolated-worker selection bug, which is independent of the pool
🤖 Claude Opus 5.5 on behalf of Kris.
- Add
- added sub-issues
on Oct 5, 2026
Metadata
Metadata
Assignees
Labels
Type
Fields
Priority
Workstream W6 of #430 · adaptive/dedicated thread management
Summary
Replication currently shares HTTP worker event loops, so a bulk copy or an egress-saturated sender steals CPU from request handling and conflates event-loop stalls with genuine egress back-pressure (muddying the #218 signal). Recommendation: dedicated-static replication thread pool first; defer elastic/auto-scaling.
Root cause / current state
subscriptionManager.ts); worker selection is hardcoded toname === 'http'in several places.replicateOverWSis reached viaserver.ws(...)— the HTTP server's upgrade handler — so inbound replication rides the same listener/threads as the ops API.Replicator.load(cache-miss retrieval) couple replication state to request-serving threads.Design direction
'replication'worker pool in the threading layer (configurablereplication.threads).NodeReplicationConnectiontoday).Scope
'replication'worker pool (static, configurable size)(db,node)socket drain/migrationRetires / advances
Dependencies
W1 (clean connection registry) and threading-layer named-pool support. Elastic depends additionally on a runtime pool API + proven W5 drain machinery.
Effort / risk
L / medium (static pool); XL / high (elastic — later).
Acceptance criteria
🤖 Filed by Claude on behalf of Kris.