Skip to content

docs(deploy): JetStream SyncAlways=true durability vs throughput on commodity storage #84

Description

@EricAndrechek

Context

internal/mq/embedded.go::NewEmbedded sets SyncAlways: true on the embedded NATS server:

opts := &natsserver.Options{
    DontListen: true,
    JetStream:  true,
    StoreDir:   storeDir,
    SyncAlways: true, // fsync every JetStream write — publish ACKs only after data is on disk
}

This is a deliberate, strict durability guarantee: every Publish() call blocks until the message is fdatasync()'d to non-volatile storage before the publisher gets an ACK. In NATS terminology this is the strongest mode — stronger than the default SyncInterval which group-commits every ~100 ms.

This issue isn't a bug report. It's a request for the team to make SyncAlways: true an explicit, auditable choice rather than an implicit default — and to consider whether 100 ms of in-flight durability loss on power loss is genuinely worse than the throughput floor we're paying for it. Background and supporting evidence below.

Where this bit us

Wave-RF/support-infra#1 was the canary: WaveHouse's Test job (make coverage, which runs the unit tests in internal/api/dlq_test.go among others) intermittently failed on the ARC wave-rf-runners self-hosted pool with create stream: context deadline exceeded after 5 seconds.

Diagnosis: under concurrent CI burst on the PVE host, individual fsync() calls inside the runner VM took 11 to 25+ seconds. The CreateStream call in mq.NewEmbedded does multiple sequential metadata fsyncs and blew past the 5 s default deadline. The same cascade applies to every Publish() call in production — not just the test setup we noticed.

The two regimes

NATS JetStream's behavior with SyncAlways: true looks fine on:

Environment p99 fsync Why
Cloud block storage (gp3/io2 EBS, GCP pd-ssd, Azure Premium SSD) Sub-millisecond at typical load Backed by hyperscaler storage with battery-backed DRAM cache, parallel commit fabric. Sync writes ack from non-volatile cache, not NAND.
Enterprise NVMe with PLP (Intel Optane, Samsung PM-series, Solidigm D7) <100 µs Power-loss-protection capacitor lets controller ack sync writes from DRAM. fdatasync ≈ memcpy.
Local ext4 on consumer NVMe (no ZFS) 1–10 ms typical jbd2 journal commit, single device, no commit barrier across other workloads. Tail spikes possible under heavy concurrent dirty data, but bounded.

It looks bad on:

Environment p99 fsync Why
ZFS without SLOG, consumer NVMe 100 ms baseline, 5–25 s under concurrent load Every sync write hits ZIL, and ZIL flush is gated by transaction-group commit cadence which serializes across all consumers of the pool. With multiple concurrent writers (CI pods, other VMs), worst-case fsync = worst-case txg commit.
Loopback / qcow2 on ext4 in a VM Highly variable Adds a layer of ext4 jbd2 + journaling, often 10× slower than direct ext4.
Spinning rust 5–50 ms baseline, multi-second tail NAND program path requires mechanical disk seeks.

What this means for product decisions

The strict durability of SyncAlways: true translates well to managed cloud infrastructure (which is presumably the production target). But the ingest floor the application can sustain is whatever its weakest deployment substrate's worst-case fsync tail is. If anyone ever runs WaveHouse on:

  • a self-hosted Proxmox/ZFS deployment without a SLOG (our self-hosted CI),
  • a cheap VPS with non-PLP NVMe on top of a shared QEMU host,
  • nested virtualization or container sandboxes that translate fsync into something slower,

…ingest p99 will be visibly bad and Publish() will occasionally take >5 s. The test failure we saw is the same code path that handles every production message.

Things to consider

  1. Make the sync mode explicit and configurable. Pass SyncAlways (or its inverse, a SyncInterval) through NewEmbedded's options or pull it from config. The default could remain SyncAlways: true — but the team should choose that consciously, not inherit it.
  2. Default to SyncInterval (~100 ms group commit). This is what most production NATS deployments use. The durability cost is "≤100 ms of in-flight messages lost on uncontrolled power loss" — small compared to the multi-order-of-magnitude throughput improvement. For a mission-critical workload, opt-in via config to SyncAlways: true.
  3. Document the durability contract. Whatever default lands, document under docs/ (or in README.md) what users can expect from a Publish() ack: "data is durable to disk on the local node" vs "data is durable to disk within 100 ms" vs "data is in NATS memory and will be flushed on next group commit." Tells operators what failure modes they're signing up for.
  4. Add a benchmark for ingest throughput under different sync modes. Useful for Performance Benchmarks #39 (Performance Benchmarks) and would surface regressions in the future. A simple harness writing 10 k messages and measuring p50/p99 latency under each mode would do.
  5. Consider exposing a Storage flag (Memory vs File) independently of sync mode. JetStream supports memory-only streams that need no fsync at all, useful for non-critical event streams (telemetry, metrics, transient routing) — different durability needs in the same app.
  6. Ship a pre-flight wavehouse storage-check CLI subcommand so operators can measure their substrate's fsync tail against WaveHouse's calibrated verdict bands before booting. Spec + acceptance criteria + fresh-agent reading order in the "Proposed deliverable" section below.

Proposed deliverable: wavehouse storage-check CLI subcommand

A built-in pre-flight subcommand operators can run before starting WaveHouse to determine whether their storage substrate can sustain SyncAlways: true and, if not, what mq.sync_interval value (the knob from #139) to configure instead. Turns the ad-hoc bench script from the comment on this issue into a calibrated, in-tree, exit-code-driven check.

Why ship this in WaveHouse rather than rely on external tools

  1. Operators choosing a sync mode need a measurement calibrated to JetStream's actual pattern — 4 KiB write + flush in a tight loop. fio, pg_test_fsync, and friends are excellent general tools but don't emit the WaveHouse-specific verdict (which mode, which interval).
  2. fio and pg_test_fsync don't handle macOS honestly out of the box. Plain fsync() on macOS does not force the drive cache to NAND; only fcntl(fd, F_FULLFSYNC) does. We measured a ~180× gap between the two on consumer NVMe (0.03 ms vs 5.5 ms p99 on an M4 Pro). The bench needs to use the platform's honest flush, and the right choice changes by OS — easy for an in-tree check to get right, hard to ask operators to remember.
  3. Pre-flight should fail closed. A BROKEN verdict should exit non-zero so deployment scripts / Helm preflight / systemd ExecStartPre can refuse to bring WaveHouse up against storage that will time out CreateStream calls.
  4. Verdict thresholds should track WaveHouse's defaults. Where the line between WORKABLE and MARGINAL falls is a function of the CreateStream deadline and the Publish() SLA. External tools can't track that; an in-tree check can.

User experience

# Default: bench against the configured mq.store_dir
$ wavehouse storage-check
WaveHouse storage check  [fdatasync, /var/lib/wavehouse, 1 worker × 500 iters × 4 KiB write+flush]
  p50=  2.10ms   p95=  4.30ms   p99=   4.50ms   max=    8.20ms
  verdict: GOOD — SyncAlways=true is safe on this storage

# Concurrent mode to surface commit-cadence problems (ZFS-without-SLOG, noisy-neighbor VM hosts)
$ wavehouse storage-check --workers 8
WaveHouse storage check  [fdatasync, /var/lib/wavehouse, 8 workers × 500 iters × 4 KiB write+flush]
  pooled         p50=  3.20ms   p95=  6.70ms   p99=  12.30ms   max=   25.10ms
  worst-worker                                 p99=  15.40ms   max=   25.10ms
  wall: 8.4s   verdict: WORKABLE — watch bursty load

# Against an arbitrary path (e.g. evaluating a candidate StoreDir before pointing config at it)
$ wavehouse storage-check --path /mnt/nvme-candidate

# Recommend a sync_interval derived from measured p99
$ wavehouse storage-check --workers 8 --recommend
  Measured worst-worker p99: 230 ms
  Recommended: mq.sync_interval = 500ms
  Rationale: 2× p99 rounded up to the next standard step; trades up to 500 ms of in-flight
             messages on hard crash for a stable ingest floor. See #139 for the config knob.

Acceptance criteria

  • wavehouse storage-check subcommand exists; default-benches cfg.MQ.StoreDir
  • --path <dir> flag overrides the path
  • --workers N flag for concurrent mode (default 1)
  • --iters N flag (default 500)
  • --format text|json flag — text default, JSON for scripting / CI
  • --recommend flag prints a suggested mq.sync_interval derived from measured worst-worker p99 (formula: 2× p99 rounded up to the next standard step in {50ms, 100ms, 250ms, 500ms, 1s, 2s})
  • macOS uses golang.org/x/sys/unix.FcntlInt(fd, unix.F_FULLFSYNC, 0) (verify with side-by-side measurement against plain fsync() on the same path — expect ~100× gap on consumer NVMe)
  • Linux uses golang.org/x/sys/unix.Fdatasync(fd)
  • Exit codes: 0 for IDEAL/GOOD/WORKABLE, 1 for MARGINAL, 2 for BROKEN
  • Unit-tested: percentile math, verdict-band thresholds, --recommend formula
  • Integration test: runs the bench against t.TempDir(), asserts non-zero ops and a reasonable verdict band
  • Docs page added under docs/src/content/docs/ (e.g. operations/storage-check.md) explaining the check, the verdict bands, and the relationship to mq.sync_interval
  • docs/src/content/docs/configuration.md (the page mq: expose JetStream sync_interval as a config knob (throughput vs durability tradeoff) #139 will edit) gets a "see also wavehouse storage-check to calibrate this value" cross-reference

Implementation hints

The reference algorithm is the Python script in the comment — port that to Go. Key correctness points the Python version got right that a Go port should preserve:

  1. Per-platform flush selection via build tags. runtime.GOOS == "darwin" → F_FULLFSYNC via unix.FcntlInt; everything else → unix.Fdatasync. There is no clean abstraction in os.File for this — drop to golang.org/x/sys/unix.
  2. Warm-up iter. The first write+flush primes the page cache / FS journal and is significantly faster than steady-state. Run it before starting the timer, then discard.
  3. Per-iter timing measures just the flush, not the write. t := time.Now(); flush(fd); dur := time.Since(t).
  4. Concurrent mode contends at the dir/pool level, not the file. Each worker writes to its own file in the same dir. errgroup.Group is fine.
  5. Percentile math: sort ascending; for q in [0,1] take index max(0, int(n*q) - 1). p99 of n=500 is index 494.
  6. The macOS footgun. Many engineers will assume os.File.Sync() is honest because that's how it behaves on Linux. It is not on macOS — verify the implementation by running the check on a Mac and comparing F_FULLFSYNC vs plain fsync() on the same path; expect a ~100× gap.

Suggested package layout:

  • internal/storagecheck/bench.go — pure bench logic (no CLI, no logging), testable in isolation
  • internal/storagecheck/flush_darwin.go + flush_linux.go — per-platform flush function (build tags)
  • internal/storagecheck/verdict.go — verdict-band + --recommend formula
  • cmd/wavehouse/storage_check.go — Cobra subcommand wiring

Out of scope

  • Continuous monitoring (a long-running storage-check daemon). Single-shot pre-flight only. Operators wanting runtime monitoring can wire --format=json into their existing observability stack.
  • Auto-tuning mq.sync_interval at runtime. Operator-controlled config only — mq: expose JetStream sync_interval as a config knob (throughput vs durability tradeoff) #139 lands the knob, this issue lands the measurement and recommendation.
  • Other storage tests (random read, seek time, sequential throughput). Fsync tail is the JetStream-specific failure mode and the only thing affecting SyncAlways: true viability. Broader storage benchmarking is fio's job.
  • Multi-host coordination for clustered deployments. Each node runs the check independently against its local StoreDir.

Context for an implementing agent

Reading order for someone picking this up cold:

  1. This issue (docs(deploy): JetStream SyncAlways=true durability vs throughput on commodity storage #84) — the durability/throughput tradeoff, why operators need a calibrated reading on their storage substrate, what the failure mode looks like in practice.
  2. mq: expose JetStream sync_interval as a config knob (throughput vs durability tradeoff) #139 — the sibling code-side issue tracking the mq.sync_interval config field. The CLI's --recommend output is a value for that field; they should land in a consistent state.
  3. The comment on this issue — the reference Python implementation (port this to Go), the measured-numbers table to calibrate the verdict bands against, and the macOS F_FULLFSYNC rationale with a side-by-side measurement on an M4 Pro.
  4. internal/mq/embedded.go::NewEmbedded — the line of code (SyncAlways: true) this whole discussion is about.
  5. Wave-RF/support-infra#1 — the original infrastructure investigation. Useful for understanding what BROKEN looks like in production: 24 s fsync tails, CreateStream: context deadline exceeded, intermittent CI failures dependent on concurrent storage-pool load. Includes the original Python bench loop that produced the PVE rows of the calibration table.

Starting move: port the Python bench loop from step 3 to Go in internal/storagecheck/, get the platform-conditional flush right (verify F_FULLFSYNC on macOS by side-by-side comparison against plain fsync()), then layer the CLI flags and verdict logic on top.

Out of scope (handled in support-infra#1)

The CI failure itself is being mitigated by infrastructure changes on the self-hosted ARC runners (ZFS tunings, planned migration of the github-arc VM disk to a separate non-ZFS LVM-Thin pool that bypasses ZFS commit cadence). Those changes will make the test pass reliably without WaveHouse code changes — but the underlying durability/throughput tradeoff is a WaveHouse-side product decision worth documenting independently.

Related


Filed by Claude Code on behalf of @EricAndrechek per support-infra#1 investigation.

Activity

  1. added
    bugSomething isn't working
    documentationImprovements or additions to documentation
    area/infraCI, build, deploy, Docker, release
    breaking-changeBreaking change to public API, CLI, or config
    on Apr 28, 2026
  2. changed the title [-]discussion: JetStream SyncAlways=true durability vs throughput on commodity storage[/-] [+]docs(deploy): JetStream SyncAlways=true durability vs throughput on commodity storage[/+] on May 13, 2026
  3. EricAndrechek commented on May 13, 2026

    @EricAndrechek
    MemberAuthor

    Edited to reflect need to mention this in documentation as the issue's "item" before publishing docs site.

  4. EricAndrechek commented on May 14, 2026

    @EricAndrechek
    MemberAuthor

    Cross-link: #139 = code-side knob; #84 = "when and why" docs.

    What we're measuring

    JetStream calls fdatasync() (Linux) or the OS equivalent before ACKing every Publish() under SyncAlways: true. The bench below replicates that exact pattern — 4 KiB write + flush, in a tight loop — and reports percentiles. The number that matters is p99 and max: those are your worst-case publish latency. If p99 is 5 s, your ingest floor is 5 s, and CreateStream's 5-second context deadline trips during boot bursts.

    We report p50 (typical), p95 (a "bad day"), p99 (the tail), and max (the worst single fsync). Single-thread shows what one publisher sees in isolation. Concurrent shows what happens when multiple writers contend for the same commit cadence — that's the failure mode that bit our CI (and the only way to see it; serial benches look fine on broken setups).

    Run during typical system load, not while the box is idle. Idle benches understate real-world tails.

    Run the bench

    Save once, rerun with different env vars.

    cat > /tmp/fsync-bench.py <<'PY'
    #!/usr/bin/env python3
    """WaveHouse fsync benchmark — measures JetStream Publish() tail latency.
    
    Picks the honest durability flush automatically:
      Linux:  fdatasync()         (plain fsync also honored)
      macOS:  fcntl F_FULLFSYNC   (plain fsync() on macOS does NOT flush — it lies)
    
    Env vars:
      BENCH_PATH=/var/lib/wavehouse   # default: tempdir
      WORKERS=8                       # default: 1
      ITERS=500                       # default: 500
    """
    import os, sys, time, tempfile, platform
    
    if sys.version_info < (3, 6):
        sys.exit(f"Need Python 3.6+, found {sys.version.split()[0]}")
    
    if platform.system() == "Darwin":
        import fcntl
        SYNC = lambda fd: fcntl.fcntl(fd, 51)  # F_FULLFSYNC
        MODE = "F_FULLFSYNC"
    else:
        SYNC = getattr(os, "fdatasync", os.fsync)
        MODE = "fdatasync" if hasattr(os, "fdatasync") else "fsync"
    
    PATH = os.environ.get("BENCH_PATH", tempfile.gettempdir())
    WORKERS = int(os.environ.get("WORKERS", 1))
    ITERS = int(os.environ.get("ITERS", 500))
    
    
    def bench(worker_id):
        os.makedirs(PATH, exist_ok=True)
        target = os.path.join(PATH, f"fsync-bench-{worker_id}.dat")
        data, ts = b"x" * 4096, []
        try:
            with open(target, "wb", buffering=0) as f:
                fd = f.fileno(); f.write(data); SYNC(fd)  # warm
                for _ in range(ITERS):
                    f.write(data); t = time.perf_counter(); SYNC(fd)
                    ts.append((time.perf_counter() - t) * 1000)
        finally:
            try: os.unlink(target)
            except FileNotFoundError: pass
        return sorted(ts)
    
    
    def pct(ts, q): return ts[max(0, int(len(ts) * q) - 1)]
    
    
    def verdict(p99):
        if p99 < 1:    return "IDEAL"
        if p99 < 5:    return "GOOD"
        if p99 < 50:   return "WORKABLE — watch bursty load"
        if p99 < 1000: return "MARGINAL — flip to SyncInterval (#139) when it lands"
        return "BROKEN — CreateStream will time out under load"
    
    
    plural = "s" if WORKERS != 1 else ""
    print(f"WaveHouse fsync bench  [{MODE}, {PATH}, {WORKERS} worker{plural} × {ITERS} iters × 4 KiB write+flush]")
    
    if WORKERS == 1:
        ts = bench(0); p99 = pct(ts, .99)
        print(f"  p50={pct(ts, .5):6.2f}ms   p95={pct(ts, .95):6.2f}ms   p99={p99:7.2f}ms   max={ts[-1]:7.2f}ms")
        print(f"  verdict: {verdict(p99)}")
    else:
        import multiprocessing as mp
        try: mp.set_start_method("fork")
        except (RuntimeError, ValueError): pass
        t0 = time.perf_counter()
        with mp.Pool(WORKERS) as pool: results = pool.map(bench, range(WORKERS))
        wall = time.perf_counter() - t0
        pooled = sorted(t for r in results for t in r)
        worst_p99 = max(pct(r, .99) for r in results)
        worst_max = max(r[-1] for r in results)
        print(f"  pooled         p50={pct(pooled, .5):6.2f}ms   p95={pct(pooled, .95):6.2f}ms   p99={pct(pooled, .99):7.2f}ms   max={pooled[-1]:7.2f}ms")
        print(f"  worst-worker                                 p99={worst_p99:7.2f}ms   max={worst_max:7.2f}ms")
        print(f"  wall: {wall:.1f}s   verdict: {verdict(worst_p99)}")
    PY
    # single-thread:
    BENCH_PATH=/var/lib/wavehouse python3 /tmp/fsync-bench.py
    
    # 8 concurrent workers (the one that finds commit-cadence problems):
    BENCH_PATH=/var/lib/wavehouse WORKERS=8 python3 /tmp/fsync-bench.py

    Sample output:

    WaveHouse fsync bench  [F_FULLFSYNC, /tmp, 8 workers × 500 iters × 4 KiB write+flush]
      pooled         p50= 12.74ms   p95= 19.84ms   p99=  25.03ms   max=  49.86ms
      worst-worker                                 p99=  27.19ms   max=  49.86ms
      wall: 8.9s   verdict: WORKABLE — watch bursty load
    

    Python availability: macOS ships /usr/bin/python3 since 10.15. Modern Linux distros (Ubuntu 18.04+, Debian 10+, RHEL/Fedora 8+, Alpine, Arch) ship it by default. If you don't have it, install via your package manager or use fio (next section).

    Alternatives

    fio — the industry-standard storage bench. Packaged on every distro (apt/dnf/brew/apk install fio). On Linux this is excellent and honest:

    fio --name=fsync --directory=/var/lib/wavehouse --rw=write --bs=4k \
        --size=64M --fsync=1 --runtime=30 --time_based
    # add --numjobs=8 --group_reporting for concurrent

    On macOS fio calls plain fsync(), which doesn't actually flush the drive cache (see below). It can incidentally land near the honest number on paced workloads (we measured fio p99=5.7ms vs F_FULLFSYNC p99=5.5ms on this MBP), but the agreement is coincidence, not guarantee. Don't trust fio on Mac for tail-latency planning — use the Python script.

    pg_test_fsync — Postgres bundles this; runs the same pattern across all sync flavors (fsync, fdatasync, open_sync, open_datasync, fsync_writethrough) and reports ops/sec averages. Useful for "is this disk fast" but it doesn't surface tail percentiles, which is what we actually care about for JetStream. apt install postgresql-contrib to get it.

    What we used for the PVE numbers below: a Python loop nearly identical to the one above — 200 iters of write(4KiB) + fdatasync() on the path under test, sorted, percentile-reported. Full source in support-infra#1.

    macOS gotcha

    Plain fsync() on macOS returns success once data is in the drive's volatile cache — it does not force the drive to commit to NAND. Only fcntl(fd, F_FULLFSYNC) does. NATS / Postgres / SQLite all use F_FULLFSYNC on macOS; so does our script. Anything under ~1 ms on a Mac is almost certainly not actually flushing. We measured 0.03 ms p99 for plain fsync() vs 5.5 ms p99 for F_FULLFSYNC on the same M4 Pro — ~180× gap.

    Measured numbers

    Status Setup p50 p99 max
    🔴 broken PVE VM on ZFS rpool under concurrent CI, in-VM ext4 (/root) 276 ms 16.6 s 24.6 s
    🔴 broken PVE VM on ZFS rpool under concurrent CI, containerd dir 52 ms 11.1 s 36.8 s
    🟡 marginal PVE VM, ZFS tunings only (arc_max, dirty_data_max), no architectural change — 29 ms ~200 ms
    ✅ healthy PVE host direct on the same pool, same window 2 ms — 157 ms
    ✅ healthy PVE LXC on separate-NVMe LVM-Thin (current production) 2 ms <5 ms 4.3 ms
    ✅ healthy M4 Pro MBP, APFS, F_FULLFSYNC, single-thread 4.3 ms 6.0 ms 8.2 ms
    ✅ healthy M4 Pro MBP, APFS, F_FULLFSYNC, 8 concurrent 12.6 ms 22.9 ms 34.5 ms
    ⚠️ dishonest M4 Pro MBP, plain fsync() — macOS lies by default 0.03 ms 0.04 ms 0.05 ms

    Sources: PVE under-burst + host-direct rows from the 2026-04-28 ARC CI burst (support-infra#1, in-VM Python fdatasync loop). PVE LXC row post-LVM-Thin migration, current prod state. ZFS-tunings-only row from pve/best-practices.md. M4 Pro rows measured 2026-05-14, this comment (script above).

    Not yet measured (help wanted): AWS EC2 gp3, enterprise NVMe with PLP (Optane / PM9A3 / D7), other hyperscaler block (pd-ssd, Premium SSD v2), ZFS-without-SLOG on different hardware. Run the bench above on any of these and drop the output — I'll add a row.

    Verdict

    p99 For SyncAlways: true
    < 1 ms Ideal
    1–5 ms Good
    5–50 ms Workable — watch bursty load
    50 ms – 1 s Marginal — flip to SyncInterval once #139 lands
    > 1 s Broken — fix storage substrate, or wait for #139

    If single-thread is fine but concurrent is much worse, your storage has a commit-cadence problem (ZFS-without-SLOG, noisy-neighbor VM host). Bench host AND guest if virtualized: our PVE host saw 2 ms p50 while the VM saw 24 s tails at the same instant — disk was fine, the pathology was the stack.

  5. moved this from Backlog to Ready in WaveHouse Task Boardon Jun 10, 2026
  6. moved this from Ready to In progress in WaveHouse Task Boardon Jun 11, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Softwarearea/apiHTTP handlers, routing, middlewarearea/docsDocumentation, site/, READMEarea/infraCI, build, deploy, Docker, releasearea/ingestIngest pipeline (Bento, batching, DLQ)breaking-changeBreaking change to public API, CLI, or configbugSomething isn't workingdocumentationImprovements or additions to documentationenhancementNew feature or request

Type

No type

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions