Skip to content

Replication W6: Adaptive / dedicated replication threads #435

Description

@kriszyp

Workstream W6 of #430 · adaptive/dedicated thread management

Summary

Replication currently shares HTTP worker event loops, so a bulk copy or an egress-saturated sender steals CPU from request handling and conflates event-loop stalls with genuine egress back-pressure (muddying the #218 signal). Recommendation: dedicated-static replication thread pool first; defer elastic/auto-scaling.

Root cause / current state

  • Subscriptions are assigned round-robin across the shared HTTP pool (subscriptionManager.ts); worker selection is hardcoded to name === 'http' in several places.
  • replicateOverWS is reached via server.ws(...) — the HTTP server's upgrade handler — so inbound replication rides the same listener/threads as the ops API.
  • Per-worker connection caches and Replicator.load (cache-miss retrieval) couple replication state to request-serving threads.

Design direction

  1. A named 'replication' worker pool in the threading layer (configurable replication.threads).
  2. Route inbound replication WS upgrades to it (or a second listener bound to the pool).
  3. Decouple cache-miss retrieval connections from streaming connections (they share NodeReplicationConnection today).
  4. Point subscription assignment at the pool.
  5. Defer elastic until (a) a runtime named-pool-spawn API exists (none today) and (b) Adaptive replication routing: dynamic fan-out under egress back-pressure #218's drain/re-route machinery is proven — adaptive routing (W5) delivers most of the load relief without adaptive threading.
  6. Align with Multi-tenant within process: SNI-based instance routing with isolated workers #247 (multi-tenant SNI isolation) so the threading layer grows one general "purpose-pool" capability, not two bespoke ones.

Scope

  • Named 'replication' worker pool (static, configurable size)
  • Route inbound replication WS upgrades to the pool
  • Split streaming connections from cache-miss retrieval connections
  • Subscription assignment targets the replication pool
  • (Later) elastic scale-up/down with live (db,node) socket drain/migration

Retires / advances

Dependencies

W1 (clean connection registry) and threading-layer named-pool support. Elastic depends additionally on a runtime pool API + proven W5 drain machinery.

Effort / risk

L / medium (static pool); XL / high (elastic — later).

Acceptance criteria

  • A bulk copy or saturated sender no longer steals CPU from request handling.
  • The back-pressure signal is not conflated with HTTP-induced event-loop stalls.

🤖 Filed by Claude on behalf of Kris.

Activity

  1. added this to the v5.3 milestone on Jun 20, 2026
  2. added theissue type on Aug 7, 2026
  3. kriszyp commented on Oct 5, 2026

    @kriszyp
    MemberAuthor

    Design for 5.4 (2026-10-05): a static replication worker pool, both directions together

    Verified against harper-pro origin/main (core submodule da24feb66). Nothing from this issue exists in code yet: no replication.threads setting, no thread type, no branch. W1, the dependency listed above, has landed (#814, #843).

    Why outbound subscriptions alone are not enough

    The obvious first step is to move only the outbound subscriptions, because they are already main-thread postMessage driven and keyed by thread ID. It would not achieve this issue's goal. An outbound connection only receives and applies. The sending load this issue is about runs on the inbound socket: streaming audit records to a subscribed peer, base copy, and egress back-pressure. So the pool has to own both the inbound replication listener and the outbound subscriptions to change anything.

    Plan

    Core (harper): a general-purpose worker type

    • Add THREAD_TYPES.REPLICATION (today hdbTerms.ts:984 has only http and job). Generalize the isolated-application slot machinery (socketRouter.ts:100-109,216-248, from harper#2524) rather than adding a second one-off pool. It already has indices past the HTTP pool, heap-share accounting and scoped restart; Multi-tenant within process: SNI-based instance routing with isolated workers #247 would reuse the same capability.
    • Restart and deploy wiring: restartWorkers('http', …) callers (bin/restart.ts, components/operations.js) must restart the pool when needed, since a replication-only change has to reach it.
    • Readiness: pool workers load databases and tables but no application code, and signal readiness without threadServer.js.
    • Port-binding rule: a worker type can own a port exclusively. This generalizes the isolated-worker skip at threadServer.js:376.
    • Broadcasts already reach every non-job port (manageThreads.js:359), so pool workers get schema and other cross-thread broadcasts with no change.

    harper-pro: route replication to the pool

    • One worker selector replacing the four worker.name === 'http' filters (subscriptionManager.ts:1427,1699,1798; recordLockTransport.ts:962). It selects the pool when present and otherwise the non-isolated HTTP workers. The second half fixes the isolated-worker leak filed separately.
    • replicator.start() (the server.ws/server.http registration at replicator.ts:159,202) runs only on pool workers, so only they bind the replication port.
    • Key custody: material goes only to name === 'http' workers (security/keyCustody.ts:111). The pool needs it too.
    • Forwarded operations (OPERATION_REQUEST, replicationConnection.ts:4944): pool workers load serverUtilities and the registered operations.
    • Cache-miss retrieval stays on request threads (Replicator.load, replicator.ts:458-515, which opens its own per-thread connection). Bridging it through the pool would serialize records and blobs across threads for no clear gain. This is the "split streaming from retrieval" item above, done by leaving retrieval where it is.
    • Record locks: when replication.recordLocks is on, lock owners live in the pool (the owner must be the thread applying the database's inbound entries, DESIGN.md:106). Every lock() then relays over the thread mesh, measured at about 1 ms under owner load. We accept that cost rather than keeping two placement rules. Record locks: measure the Phase 1 cost baseline before the protocol change #824's cost baseline should measure the pool case.

    What needs no change:

    • Replicated commits already notify subscribers on every thread (core transactionBroadcast.ts:66-90).
    • replicateTo confirmations wait on a shared buffer (knownNodes.ts:954-978), not on a connection object.
    • server.nodes is filled per thread by the hdb_nodes watcher.
    • Shared-status buffers are process-wide, and ownership is by thread ID.

    Configuration and rollout

    • replication.threads, default 0 in 5.4, which keeps today's behavior (replication on the HTTP workers). Flip the default after soak.
    • The pool requires a dedicated replication port. With neither replication.port nor securePort set, replicator.ts:133-136 falls back to the operations API ports, which main binds exclusively and which cannot be routed to a pool. With replication.threads > 0 in that configuration, refuse at startup with a clear error rather than silently not using the pool.
    • macOS sets noReusePort on HTTP servers, so one pool worker wins the port there. Linux spreads connections across the pool with SO_REUSEPORT.

    Relationship to #959 (sender time budget)

    #959 replaces the per-record yield with a 2 ms worker-local budget. On a shared HTTP worker that lets a catch-up sender hold the event loop for longer slices, which risks starving application request handling. #959 should land with or after this pool, and on non-pool workers it should keep per-record yields (or apply the budget only on pool workers). With a dedicated pool, the longer slices only delay other replication work.

    Acceptance

    Sub-issues

    🤖 Claude Opus 5.5 on behalf of Kris.

  4. modified the milestones: v5.3, v5.4 on Oct 10, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:replicationReplication, cluster sync, peer connectionsenhancementNew feature or request

    Fields

    Priority

    P2

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions