Skip to content

Pace replication sender yields by elapsed time to reduce per-record CPU overhead #958

Description

@kriszyp

The per-subscription sender awaits sendAuditRecord() for every audit entry. Its normal no-backpressure exit and skipAuditRecord() both return new Promise(setImmediate), costing an event-loop turn even for very small records or skipped entries.

Measured on macOS with a local loopback cluster, one source and N subscribing peers, threads.count=4, 10-record write transactions and approximately 500-byte records. Profiling covered only catch-up after writes stopped; all incoming replication connections landed on one source worker. The comparison patch yielded only after 2 ms had elapsed on that worker.

Measurement Baseline 2 ms yield budget
Sender CPU, 4 peers, 3 runs each 9.6–10.3 s (user ~6.9 s, system ~3.1 s) ~3.45 s (user ~2.9 s, system ~0.57 s)
CPU per entry sent ~8.1 µs ~3.2 µs
Replication rate per peer, 4 peers 19.1–19.3k records/s 21.4–21.9k records/s
Replication rate, 1 peer 27.5k records/s 30.8k records/s
Concurrent source write rate 56–59k records/s ~55k records/s

The baseline profile spent 18% in writev and approximately 10% in async-context/immediate/microtask machinery, with high kernel time. Those costs mostly disappeared in the patch; audit reads and GC dominated, and receivers became the replication bottleneck. Concurrent write throughput fell approximately 5% because sending kept up during writes, while total sender work decreased. Measurements are from the dispatcher's experiment; the exact measured build SHA was not provided.

Use a small shared, per-worker time budget for normal and skipped audit sends. Preserve real socket-backpressure and blob-saturation waits. Skipped runs must still yield so their sequence-update timer fires. Inspect the separate full-copy flush yield and preserve its liveness contract. Leave receive pacing and recordConcurrency defaults unchanged.

Acceptance: a regression guard fails on the per-record-yield baseline, verifies budgeted fairness and skipped-sequence timer progress, and replication unit/integration coverage passes.

Priority: P2 — measured CPU/throughput degradation with bounded, recoverable catch-up; no demonstrated data loss or outage. Internal optimization, no new API. No release/backport target committed.

Filed by GPT-5 Codex for the dispatched task.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

Fields

Priority

P2

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions