Repository navigation
docs(deploy): JetStream SyncAlways=true durability vs throughput on commodity storage #84
Description
Activity
- addedarea/apiHTTP handlers, routing, middlewareHTTP handlers, routing, middlewarearea/docsDocumentation, site/, READMEDocumentation, site/, README
on Apr 28, 2026 - addedbugSomething isn't workingSomething isn't workingdocumentationImprovements or additions to documentationImprovements or additions to documentationarea/infraCI, build, deploy, Docker, releaseCI, build, deploy, Docker, releasebreaking-changeBreaking change to public API, CLI, or configBreaking change to public API, CLI, or config
on Apr 28, 2026 - changed the title
[-]discussion: JetStream SyncAlways=true durability vs throughput on commodity storage[/-][+]docs(deploy): JetStream SyncAlways=true durability vs throughput on commodity storage[/+]on May 13, 2026 - addedarea/ingestIngest pipeline (Bento, batching, DLQ)Ingest pipeline (Bento, batching, DLQ)
on May 13, 2026 Edited to reflect need to mention this in documentation as the issue's "item" before publishing docs site.
Cross-link: #139 = code-side knob; #84 = "when and why" docs.
What we're measuring
JetStream calls
fdatasync()(Linux) or the OS equivalent before ACKing everyPublish()underSyncAlways: true. The bench below replicates that exact pattern — 4 KiB write + flush, in a tight loop — and reports percentiles. The number that matters is p99 and max: those are your worst-case publish latency. If p99 is 5 s, your ingest floor is 5 s, andCreateStream's 5-second context deadline trips during boot bursts.We report p50 (typical), p95 (a "bad day"), p99 (the tail), and max (the worst single fsync). Single-thread shows what one publisher sees in isolation. Concurrent shows what happens when multiple writers contend for the same commit cadence — that's the failure mode that bit our CI (and the only way to see it; serial benches look fine on broken setups).
Run during typical system load, not while the box is idle. Idle benches understate real-world tails.
Run the bench
Save once, rerun with different env vars.
cat > /tmp/fsync-bench.py <<'PY' #!/usr/bin/env python3 """WaveHouse fsync benchmark — measures JetStream Publish() tail latency. Picks the honest durability flush automatically: Linux: fdatasync() (plain fsync also honored) macOS: fcntl F_FULLFSYNC (plain fsync() on macOS does NOT flush — it lies) Env vars: BENCH_PATH=/var/lib/wavehouse # default: tempdir WORKERS=8 # default: 1 ITERS=500 # default: 500 """ import os, sys, time, tempfile, platform if sys.version_info < (3, 6): sys.exit(f"Need Python 3.6+, found {sys.version.split()[0]}") if platform.system() == "Darwin": import fcntl SYNC = lambda fd: fcntl.fcntl(fd, 51) # F_FULLFSYNC MODE = "F_FULLFSYNC" else: SYNC = getattr(os, "fdatasync", os.fsync) MODE = "fdatasync" if hasattr(os, "fdatasync") else "fsync" PATH = os.environ.get("BENCH_PATH", tempfile.gettempdir()) WORKERS = int(os.environ.get("WORKERS", 1)) ITERS = int(os.environ.get("ITERS", 500)) def bench(worker_id): os.makedirs(PATH, exist_ok=True) target = os.path.join(PATH, f"fsync-bench-{worker_id}.dat") data, ts = b"x" * 4096, [] try: with open(target, "wb", buffering=0) as f: fd = f.fileno(); f.write(data); SYNC(fd) # warm for _ in range(ITERS): f.write(data); t = time.perf_counter(); SYNC(fd) ts.append((time.perf_counter() - t) * 1000) finally: try: os.unlink(target) except FileNotFoundError: pass return sorted(ts) def pct(ts, q): return ts[max(0, int(len(ts) * q) - 1)] def verdict(p99): if p99 < 1: return "IDEAL" if p99 < 5: return "GOOD" if p99 < 50: return "WORKABLE — watch bursty load" if p99 < 1000: return "MARGINAL — flip to SyncInterval (#139) when it lands" return "BROKEN — CreateStream will time out under load" plural = "s" if WORKERS != 1 else "" print(f"WaveHouse fsync bench [{MODE}, {PATH}, {WORKERS} worker{plural} × {ITERS} iters × 4 KiB write+flush]") if WORKERS == 1: ts = bench(0); p99 = pct(ts, .99) print(f" p50={pct(ts, .5):6.2f}ms p95={pct(ts, .95):6.2f}ms p99={p99:7.2f}ms max={ts[-1]:7.2f}ms") print(f" verdict: {verdict(p99)}") else: import multiprocessing as mp try: mp.set_start_method("fork") except (RuntimeError, ValueError): pass t0 = time.perf_counter() with mp.Pool(WORKERS) as pool: results = pool.map(bench, range(WORKERS)) wall = time.perf_counter() - t0 pooled = sorted(t for r in results for t in r) worst_p99 = max(pct(r, .99) for r in results) worst_max = max(r[-1] for r in results) print(f" pooled p50={pct(pooled, .5):6.2f}ms p95={pct(pooled, .95):6.2f}ms p99={pct(pooled, .99):7.2f}ms max={pooled[-1]:7.2f}ms") print(f" worst-worker p99={worst_p99:7.2f}ms max={worst_max:7.2f}ms") print(f" wall: {wall:.1f}s verdict: {verdict(worst_p99)}") PY
# single-thread: BENCH_PATH=/var/lib/wavehouse python3 /tmp/fsync-bench.py # 8 concurrent workers (the one that finds commit-cadence problems): BENCH_PATH=/var/lib/wavehouse WORKERS=8 python3 /tmp/fsync-bench.py
Sample output:
WaveHouse fsync bench [F_FULLFSYNC, /tmp, 8 workers × 500 iters × 4 KiB write+flush] pooled p50= 12.74ms p95= 19.84ms p99= 25.03ms max= 49.86ms worst-worker p99= 27.19ms max= 49.86ms wall: 8.9s verdict: WORKABLE — watch bursty loadPython availability: macOS ships
/usr/bin/python3since 10.15. Modern Linux distros (Ubuntu 18.04+, Debian 10+, RHEL/Fedora 8+, Alpine, Arch) ship it by default. If you don't have it, install via your package manager or usefio(next section).Alternatives
fio— the industry-standard storage bench. Packaged on every distro (apt/dnf/brew/apk install fio). On Linux this is excellent and honest:fio --name=fsync --directory=/var/lib/wavehouse --rw=write --bs=4k \ --size=64M --fsync=1 --runtime=30 --time_based # add --numjobs=8 --group_reporting for concurrentOn macOS fio calls plain
fsync(), which doesn't actually flush the drive cache (see below). It can incidentally land near the honest number on paced workloads (we measured fio p99=5.7ms vs F_FULLFSYNC p99=5.5ms on this MBP), but the agreement is coincidence, not guarantee. Don't trust fio on Mac for tail-latency planning — use the Python script.pg_test_fsync— Postgres bundles this; runs the same pattern across all sync flavors (fsync,fdatasync,open_sync,open_datasync,fsync_writethrough) and reports ops/sec averages. Useful for "is this disk fast" but it doesn't surface tail percentiles, which is what we actually care about for JetStream.apt install postgresql-contribto get it.What we used for the PVE numbers below: a Python loop nearly identical to the one above — 200 iters of
write(4KiB) + fdatasync()on the path under test, sorted, percentile-reported. Full source in support-infra#1.macOS gotcha
Plain
fsync()on macOS returns success once data is in the drive's volatile cache — it does not force the drive to commit to NAND. Onlyfcntl(fd, F_FULLFSYNC)does. NATS / Postgres / SQLite all use F_FULLFSYNC on macOS; so does our script. Anything under ~1 ms on a Mac is almost certainly not actually flushing. We measured 0.03 ms p99 for plainfsync()vs 5.5 ms p99 forF_FULLFSYNCon the same M4 Pro — ~180× gap.Measured numbers
Status Setup p50 p99 max 🔴 broken PVE VM on ZFS rpool under concurrent CI, in-VM ext4 ( /root)276 ms 16.6 s 24.6 s 🔴 broken PVE VM on ZFS rpool under concurrent CI, containerd dir 52 ms 11.1 s 36.8 s 🟡 marginal PVE VM, ZFS tunings only (arc_max, dirty_data_max), no architectural change — 29 ms ~200 ms ✅ healthy PVE host direct on the same pool, same window 2 ms — 157 ms ✅ healthy PVE LXC on separate-NVMe LVM-Thin (current production) 2 ms <5 ms 4.3 ms ✅ healthy M4 Pro MBP, APFS, F_FULLFSYNC, single-thread 4.3 ms 6.0 ms 8.2 ms ✅ healthy M4 Pro MBP, APFS, F_FULLFSYNC, 8 concurrent 12.6 ms 22.9 ms 34.5 ms ⚠️ dishonestM4 Pro MBP, plain fsync()— macOS lies by default0.03 ms 0.04 ms 0.05 ms Sources: PVE under-burst + host-direct rows from the 2026-04-28 ARC CI burst (support-infra#1, in-VM Python fdatasync loop). PVE LXC row post-LVM-Thin migration, current prod state. ZFS-tunings-only row from
pve/best-practices.md. M4 Pro rows measured 2026-05-14, this comment (script above).Not yet measured (help wanted): AWS EC2 gp3, enterprise NVMe with PLP (Optane / PM9A3 / D7), other hyperscaler block (pd-ssd, Premium SSD v2), ZFS-without-SLOG on different hardware. Run the bench above on any of these and drop the output — I'll add a row.
Verdict
p99 For SyncAlways: true< 1 ms Ideal 1–5 ms Good 5–50 ms Workable — watch bursty load 50 ms – 1 s Marginal — flip to SyncIntervalonce #139 lands> 1 s Broken — fix storage substrate, or wait for #139 If single-thread is fine but concurrent is much worse, your storage has a commit-cadence problem (ZFS-without-SLOG, noisy-neighbor VM host). Bench host AND guest if virtualized: our PVE host saw 2 ms p50 while the VM saw 24 s tails at the same instant — disk was fine, the pathology was the stack.
- added a commit that references this issue
on Jun 11, 2026
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsDone
Context
internal/mq/embedded.go::NewEmbeddedsetsSyncAlways: trueon the embedded NATS server:This is a deliberate, strict durability guarantee: every
Publish()call blocks until the message isfdatasync()'d to non-volatile storage before the publisher gets an ACK. In NATS terminology this is the strongest mode — stronger than the defaultSyncIntervalwhich group-commits every ~100 ms.This issue isn't a bug report. It's a request for the team to make
SyncAlways: truean explicit, auditable choice rather than an implicit default — and to consider whether 100 ms of in-flight durability loss on power loss is genuinely worse than the throughput floor we're paying for it. Background and supporting evidence below.Where this bit us
Wave-RF/support-infra#1 was the canary: WaveHouse's
Testjob (make coverage, which runs the unit tests ininternal/api/dlq_test.goamong others) intermittently failed on the ARCwave-rf-runnersself-hosted pool withcreate stream: context deadline exceededafter 5 seconds.Diagnosis: under concurrent CI burst on the PVE host, individual
fsync()calls inside the runner VM took 11 to 25+ seconds. TheCreateStreamcall inmq.NewEmbeddeddoes multiple sequential metadata fsyncs and blew past the 5 s default deadline. The same cascade applies to everyPublish()call in production — not just the test setup we noticed.The two regimes
NATS JetStream's behavior with
SyncAlways: truelooks fine on:ext4on consumer NVMe (no ZFS)It looks bad on:
What this means for product decisions
The strict durability of
SyncAlways: truetranslates well to managed cloud infrastructure (which is presumably the production target). But the ingest floor the application can sustain is whatever its weakest deployment substrate's worst-case fsync tail is. If anyone ever runs WaveHouse on:…ingest p99 will be visibly bad and
Publish()will occasionally take >5 s. The test failure we saw is the same code path that handles every production message.Things to consider
SyncAlways(or its inverse, aSyncInterval) throughNewEmbedded's options or pull it from config. The default could remainSyncAlways: true— but the team should choose that consciously, not inherit it.SyncInterval(~100 ms group commit). This is what most production NATS deployments use. The durability cost is "≤100 ms of in-flight messages lost on uncontrolled power loss" — small compared to the multi-order-of-magnitude throughput improvement. For a mission-critical workload, opt-in via config toSyncAlways: true.docs/(or inREADME.md) what users can expect from aPublish()ack: "data is durable to disk on the local node" vs "data is durable to disk within 100 ms" vs "data is in NATS memory and will be flushed on next group commit." Tells operators what failure modes they're signing up for.Storageflag (Memory vs File) independently of sync mode. JetStream supports memory-only streams that need no fsync at all, useful for non-critical event streams (telemetry, metrics, transient routing) — different durability needs in the same app.wavehouse storage-checkCLI subcommand so operators can measure their substrate's fsync tail against WaveHouse's calibrated verdict bands before booting. Spec + acceptance criteria + fresh-agent reading order in the "Proposed deliverable" section below.Proposed deliverable:
wavehouse storage-checkCLI subcommandA built-in pre-flight subcommand operators can run before starting WaveHouse to determine whether their storage substrate can sustain
SyncAlways: trueand, if not, whatmq.sync_intervalvalue (the knob from #139) to configure instead. Turns the ad-hoc bench script from the comment on this issue into a calibrated, in-tree, exit-code-driven check.Why ship this in WaveHouse rather than rely on external tools
fio,pg_test_fsync, and friends are excellent general tools but don't emit the WaveHouse-specific verdict (which mode, which interval).fioandpg_test_fsyncdon't handle macOS honestly out of the box. Plainfsync()on macOS does not force the drive cache to NAND; onlyfcntl(fd, F_FULLFSYNC)does. We measured a ~180× gap between the two on consumer NVMe (0.03 ms vs 5.5 ms p99 on an M4 Pro). The bench needs to use the platform's honest flush, and the right choice changes by OS — easy for an in-tree check to get right, hard to ask operators to remember.BROKENverdict should exit non-zero so deployment scripts / Helm preflight / systemdExecStartPrecan refuse to bring WaveHouse up against storage that will time outCreateStreamcalls.WORKABLEandMARGINALfalls is a function of theCreateStreamdeadline and thePublish()SLA. External tools can't track that; an in-tree check can.User experience
Acceptance criteria
wavehouse storage-checksubcommand exists; default-benchescfg.MQ.StoreDir--path <dir>flag overrides the path--workers Nflag for concurrent mode (default1)--iters Nflag (default500)--format text|jsonflag — text default, JSON for scripting / CI--recommendflag prints a suggestedmq.sync_intervalderived from measured worst-worker p99 (formula:2× p99rounded up to the next standard step in{50ms, 100ms, 250ms, 500ms, 1s, 2s})golang.org/x/sys/unix.FcntlInt(fd, unix.F_FULLFSYNC, 0)(verify with side-by-side measurement against plainfsync()on the same path — expect ~100× gap on consumer NVMe)golang.org/x/sys/unix.Fdatasync(fd)0for IDEAL/GOOD/WORKABLE,1for MARGINAL,2for BROKEN--recommendformulat.TempDir(), asserts non-zero ops and a reasonable verdict banddocs/src/content/docs/(e.g.operations/storage-check.md) explaining the check, the verdict bands, and the relationship tomq.sync_intervaldocs/src/content/docs/configuration.md(the page mq: expose JetStream sync_interval as a config knob (throughput vs durability tradeoff) #139 will edit) gets a "see alsowavehouse storage-checkto calibrate this value" cross-referenceImplementation hints
The reference algorithm is the Python script in the comment — port that to Go. Key correctness points the Python version got right that a Go port should preserve:
runtime.GOOS == "darwin"→F_FULLFSYNCviaunix.FcntlInt; everything else →unix.Fdatasync. There is no clean abstraction inos.Filefor this — drop togolang.org/x/sys/unix.t := time.Now(); flush(fd); dur := time.Since(t).errgroup.Groupis fine.max(0, int(n*q) - 1). p99 of n=500 is index 494.os.File.Sync()is honest because that's how it behaves on Linux. It is not on macOS — verify the implementation by running the check on a Mac and comparing F_FULLFSYNC vs plainfsync()on the same path; expect a ~100× gap.Suggested package layout:
internal/storagecheck/bench.go— pure bench logic (no CLI, no logging), testable in isolationinternal/storagecheck/flush_darwin.go+flush_linux.go— per-platform flush function (build tags)internal/storagecheck/verdict.go— verdict-band +--recommendformulacmd/wavehouse/storage_check.go— Cobra subcommand wiringOut of scope
--format=jsoninto their existing observability stack.mq.sync_intervalat runtime. Operator-controlled config only — mq: expose JetStream sync_interval as a config knob (throughput vs durability tradeoff) #139 lands the knob, this issue lands the measurement and recommendation.SyncAlways: trueviability. Broader storage benchmarking isfio's job.Context for an implementing agent
Reading order for someone picking this up cold:
mq.sync_intervalconfig field. The CLI's--recommendoutput is a value for that field; they should land in a consistent state.F_FULLFSYNCrationale with a side-by-side measurement on an M4 Pro.internal/mq/embedded.go::NewEmbedded— the line of code (SyncAlways: true) this whole discussion is about.BROKENlooks like in production: 24 s fsync tails,CreateStream: context deadline exceeded, intermittent CI failures dependent on concurrent storage-pool load. Includes the original Python bench loop that produced the PVE rows of the calibration table.Starting move: port the Python bench loop from step 3 to Go in
internal/storagecheck/, get the platform-conditional flush right (verify F_FULLFSYNC on macOS by side-by-side comparison against plainfsync()), then layer the CLI flags and verdict logic on top.Out of scope (handled in support-infra#1)
The CI failure itself is being mitigated by infrastructure changes on the self-hosted ARC runners (ZFS tunings, planned migration of the github-arc VM disk to a separate non-ZFS LVM-Thin pool that bypasses ZFS commit cadence). Those changes will make the test pass reliably without WaveHouse code changes — but the underlying durability/throughput tradeoff is a WaveHouse-side product decision worth documenting independently.
Related
SyncAlways/SyncIntervalas acfg.MQ.SyncIntervalconfig knob. This issue is the "when and why" + measurement tooling; mq: expose JetStream sync_interval as a config knob (throughput vs durability tradeoff) #139 is the "how to flip it." Ideally land together.Filed by Claude Code on behalf of @EricAndrechek per support-infra#1 investigation.