Skip to content

docs: correct WAL sync-mode claims and document all-null column handling - #19

Merged
xe-nvdk merged 1 commit into
mainfrom
docs/wal-sync-mode-and-null-columns
Jul 30, 2026
Merged

xe-nvdk merged 1 commit into
mainfrom
docs/wal-sync-mode-and-null-columns

Conversation

@xe-nvdk

@xe-nvdk xe-nvdk commented Jul 29, 2026

Copy link
Copy Markdown
Member

Documentation follow-up to Basekick-Labs/arc#565.

WAL sync modes were documented inaccurately

The docs claimed fdatasync was "50% faster than fsync". That was never measurable — until arc#305, both modes called the same full Sync().

The benchmark table showed it plainly and nobody noticed:

old table
WAL + fdatasync 7.5M rec/s (-21%)
WAL + fsync 7.7M rec/s (-19%)

fdatasync listed as slower than fsync is exactly what measuring two identical code paths plus noise looks like.

Corrected:

  • Collapsed the per-mode rows into one "WAL enabled" row. The cost is in enabling the WAL, not in the mode.
  • Stated why: syncs are batched on a 100 ms ticker, not performed per write, so there are at most ~10/sec regardless of ingest rate. The sync mode selects durability semantics, not speed.
  • Removed the troubleshooting step advising operators to switch to fdatasync to fix slow throughput — that advice cannot help, and sent people chasing the wrong thing.
  • Documented platform support: fdatasync(2) is Linux-only; macOS and Windows fall back to a full fsync and now log it once at startup (fdatasync is unavailable on this platform; using full fsync instead), with fdatasync_supported in the startup log.
  • Added a note that macOS Sync() maps to fsync(2), which does not flush the drive's own write cache — F_FULLFSYNC would be required. Relevant for local development, not Linux production.

MessagePack null handling

Documented the arc#337 behavior on POST /api/v1/write/msgpack:

  • A column may be entirely null within a batch and is preserved rather than dropped — it stays queryable and returns NULLs. Previously such a column vanished, and if it was null in every batch, querying it failed to resolve.
  • A later batch with real values determines the type normally; the all-null batch does not pin it.
  • Column types need only be consistent within a single write — reads union by name across files.
  • time is the exception: a null or non-numeric time is rejected with 400 Bad Request, because a non-timestamp time column makes the partition un-compactable.

I verified the 400 against a running server rather than assuming it — the rejection happens synchronously in normalizeTimestamps, so it is a real HTTP error and not an async flush failure:

HTTP status: 400
{"error":"Invalid MessagePack payload: failed to normalize timestamps: invalid timestamp type in columnar format: <nil>"}

Note

The corrected throughput figures are the existing published numbers with the misleading per-mode split removed — I did not re-benchmark. If you want fresh numbers for 26.09.1, that's worth a separate pass on a quiet machine.

Follows Basekick-Labs/arc#565.

WAL sync modes: the docs claimed fdatasync was "50% faster than fsync".
That was never measurable, because both modes called the same full Sync()
until arc#305 -- and the benchmark table reflected it, showing fdatasync
as SLOWER than fsync (7.5M vs 7.7M rec/s), which is what measuring two
identical code paths plus noise looks like.

Corrected to say what is actually true: syncs are batched on a 100ms
ticker rather than performed per write, so all three modes land within
noise of each other on throughput. The sync mode selects durability
semantics, not speed. Collapsed the per-mode benchmark rows into a single
"WAL enabled" row, and removed the troubleshooting step that advised
switching to fdatasync to fix slow throughput -- that advice cannot help.

Also documented platform support: fdatasync(2) is Linux-only; macOS and
Windows fall back to a full fsync and now log that once at startup. Added
a note that macOS Sync() maps to fsync(2), which does not flush the
drive's write cache.

MessagePack API: documented that a column may be entirely null within a
batch and is preserved rather than dropped (arc#337), that a later batch
with real values determines the type, and that time is the exception --
a null or non-numeric time is rejected with 400, verified against a
running server.
@xe-nvdk
xe-nvdk merged commit fc4cf60 into main Jul 30, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant