Problem
internal/mq/embedded.go:NewEmbedded hardcodes SyncAlways: true on the NATS JetStream options, which fsyncs every JetStream write before ACKing the publisher. This is the safest option (zero ack-then-lose risk on crash) but also the throughput floor — every publish pays a disk-sync round-trip, and the gap between SyncAlways: true and SyncAlways: false on commodity NVMe is often 5-10x.
For workloads where the gateway is fronting analytics ingest (the WaveHouse primary use case), occasional message loss on hard crash is a tolerable failure mode in exchange for substantial throughput — but operators currently have no knob to choose.
Proposed Solution
Expose JetStream's sync interval as a cfg.MQ.SyncInterval config field:
- YAML:
mq.sync_interval (duration string, default "0" meaning fsync-on-every-write — preserves current behavior)
- Env:
WH_MQ_SYNC_INTERVAL=2s
- Wire to NATS
Options.SyncInterval (or whichever knob matches the chosen semantics — JetStream's option name changes between server versions)
- Validation at config load: parse duration, reject negative values, accept zero (meaning sync-always)
Default stays at the safe option — operators have to opt into the durability tradeoff explicitly.
Acceptance criteria
Context
TODO comment landed in internal/mq/embedded.go:61 as part of PR #125. Out of scope for the boot-non-fatal fix tracked by #95 but worth pulling into a follow-up. Sibling issue tracks max_bytes_gb upper-bound validation.
Related
Problem
internal/mq/embedded.go:NewEmbeddedhardcodesSyncAlways: trueon the NATS JetStream options, which fsyncs every JetStream write before ACKing the publisher. This is the safest option (zero ack-then-lose risk on crash) but also the throughput floor — every publish pays a disk-sync round-trip, and the gap betweenSyncAlways: trueandSyncAlways: falseon commodity NVMe is often 5-10x.For workloads where the gateway is fronting analytics ingest (the WaveHouse primary use case), occasional message loss on hard crash is a tolerable failure mode in exchange for substantial throughput — but operators currently have no knob to choose.
Proposed Solution
Expose JetStream's sync interval as a
cfg.MQ.SyncIntervalconfig field:mq.sync_interval(duration string, default"0"meaning fsync-on-every-write — preserves current behavior)WH_MQ_SYNC_INTERVAL=2sOptions.SyncInterval(or whichever knob matches the chosen semantics — JetStream's option name changes between server versions)Default stays at the safe option — operators have to opt into the durability tradeoff explicitly.
Acceptance criteria
"0"SyncAlways: truebehaviornatsserver.Optionsdocs/src/content/docs/configuration.md— be explicit that "this means messages ACKed within the interval can be lost on a hard crash"mq.max_bytes_gbvalidation issue — oversubscribed-store + lazy-fsync compounds (the crash-loss window grows with both)Context
TODO comment landed in
internal/mq/embedded.go:61as part of PR #125. Out of scope for the boot-non-fatal fix tracked by #95 but worth pulling into a follow-up. Sibling issue tracksmax_bytes_gbupper-bound validation.Related
SyncAlways=truedurability/throughput tradeoff and per-storage-substrate expectations (the "when and why" companion to this issue's "how"). The two are sibling deliverables: this issue (mq: expose JetStream sync_interval as a config knob (throughput vs durability tradeoff) #139) is the code-side knob; docs(deploy): JetStream SyncAlways=true durability vs throughput on commodity storage #84 is the docs-side guidance that tells admins/ops what number to put in it. Ideal world they land together — flippingSyncIntervalwithout telling operators what they're trading is worse than no knob at all.