You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Evidence (live rig, 2026-09-26, enterprise-shared compose stack, main @ #947)
Leader-crash run with WAL enabled: a writer was killed holding ~3,400 acknowledged records (~3,200 already flushed to parquet, ~200 in the Arrow buffer). On restart, WAL recovery replayed the entire active WAL file — WAL recovery complete entries=3400 — and the queryable count went 19,700 → 23,100 (= +3,400 exactly): the 200 genuinely-lost records were recovered, and the ~3,200 already-durable ones were re-ingested as duplicates.
Why this is two different severities
Tagged measurements (common case): the duplicates are exact copies and compaction dedups on (tags..., time), so they reconcile at the partition's next compaction pass. Consequence: inflated query results between crash-recovery and that pass (which can linger for low-volume partitions below hourly_min_files). Annoying, self-healing.
Tagless measurements: compaction dedup runs only when tag columns exist or the CQ-only arc:dedup_time marker is present (internal/compaction/dedup.go — buildCompactionQuery returns a plain COPY otherwise, deliberately: two same-timestamp tagless rows can be two legitimate events). So for raw tagless ingest, crash-recovery duplicates are never removed. A restart after a hard crash silently double-counts up to a full WAL window (wal.max_size_mb=100 / wal.max_age=1h) of already-durable data, permanently.
There is no safe shortcut on the dedup side: deduping tagless data on time alone would collapse legitimate same-timestamp rows. The fix has to be on the replay side.
Preconditions / exposure
wal.enabled=true (default false) + hard crash (graceful shutdown flushes and purges correctly since #806) + restart with replay. Tagged data heals; tagless does not. Not data loss, not corruption — over-counting.
Proposed fix: flush-watermark checkpointing
Record, per WAL file, a durable watermark of entries whose batches have completed their parquet flush (upload confirmed, not merely handed to the flusher). Recovery replays only entries past the watermark. That fixes tagless permanently and shrinks the tagged inflation window to ~zero.
Care points for the implementer:
Crash-ordering is load-bearing: the watermark must advance strictly AFTER the flush is confirmed durable, never before — advancing early converts this over-counting bug into real data loss (the exact inverse failure). The watermark write itself must be crash-safe (torn-write tolerant).
Hot path: the watermark update sits next to the flush path feeding ~20M rec/s ingest; it must be off the per-record path (per-flush, batched) and benchmarked ABAB against main.
Partial-flush mapping: one WAL file spans many buffers/measurements flushing independently; the watermark needs to be per-batch/segment, not a single file offset, or a slow measurement pins the whole file at 0.
Document the current semantics in the WAL docs + release notes: replay is at-least-once; tagged duplicates reconcile at the next compaction; tagless duplicates currently persist. (Separate small PR.)
Evidence (live rig, 2026-09-26, enterprise-shared compose stack, main @ #947)
Leader-crash run with WAL enabled: a writer was killed holding ~3,400 acknowledged records (~3,200 already flushed to parquet, ~200 in the Arrow buffer). On restart, WAL recovery replayed the entire active WAL file —
WAL recovery complete entries=3400— and the queryable count went 19,700 → 23,100 (= +3,400 exactly): the 200 genuinely-lost records were recovered, and the ~3,200 already-durable ones were re-ingested as duplicates.Why this is two different severities
(tags..., time), so they reconcile at the partition's next compaction pass. Consequence: inflated query results between crash-recovery and that pass (which can linger for low-volume partitions belowhourly_min_files). Annoying, self-healing.arc:dedup_timemarker is present (internal/compaction/dedup.go—buildCompactionQueryreturns a plain COPY otherwise, deliberately: two same-timestamp tagless rows can be two legitimate events). So for raw tagless ingest, crash-recovery duplicates are never removed. A restart after a hard crash silently double-counts up to a full WAL window (wal.max_size_mb=100/wal.max_age=1h) of already-durable data, permanently.There is no safe shortcut on the dedup side: deduping tagless data on time alone would collapse legitimate same-timestamp rows. The fix has to be on the replay side.
Preconditions / exposure
wal.enabled=true(default false) + hard crash (graceful shutdown flushes and purges correctly since #806) + restart with replay. Tagged data heals; tagless does not. Not data loss, not corruption — over-counting.Proposed fix: flush-watermark checkpointing
Record, per WAL file, a durable watermark of entries whose batches have completed their parquet flush (upload confirmed, not merely handed to the flusher). Recovery replays only entries past the watermark. That fixes tagless permanently and shrinks the tagged inflation window to ~zero.
Care points for the implementer:
count(*)should equal HTTP-succeeded with no compaction needed, for both tagged and tagless databases.Interim for 26.09.2
Document the current semantics in the WAL docs + release notes: replay is at-least-once; tagged duplicates reconcile at the next compaction; tagless duplicates currently persist. (Separate small PR.)
Related: #946 (smoke assertion), #945 (fresh-bucket reads), #806/#803 (shutdown ordering), #594 (WAL file deletion).