Skip to content

wal: replay after a crash re-ingests already-flushed records — permanent duplicates for tagless measurements #948

Description

@xe-nvdk

Evidence (live rig, 2026-09-26, enterprise-shared compose stack, main @ #947)

Leader-crash run with WAL enabled: a writer was killed holding ~3,400 acknowledged records (~3,200 already flushed to parquet, ~200 in the Arrow buffer). On restart, WAL recovery replayed the entire active WAL file — WAL recovery complete entries=3400 — and the queryable count went 19,700 → 23,100 (= +3,400 exactly): the 200 genuinely-lost records were recovered, and the ~3,200 already-durable ones were re-ingested as duplicates.

Why this is two different severities

  • Tagged measurements (common case): the duplicates are exact copies and compaction dedups on (tags..., time), so they reconcile at the partition's next compaction pass. Consequence: inflated query results between crash-recovery and that pass (which can linger for low-volume partitions below hourly_min_files). Annoying, self-healing.
  • Tagless measurements: compaction dedup runs only when tag columns exist or the CQ-only arc:dedup_time marker is present (internal/compaction/dedup.go — buildCompactionQuery returns a plain COPY otherwise, deliberately: two same-timestamp tagless rows can be two legitimate events). So for raw tagless ingest, crash-recovery duplicates are never removed. A restart after a hard crash silently double-counts up to a full WAL window (wal.max_size_mb=100 / wal.max_age=1h) of already-durable data, permanently.

There is no safe shortcut on the dedup side: deduping tagless data on time alone would collapse legitimate same-timestamp rows. The fix has to be on the replay side.

Preconditions / exposure

wal.enabled=true (default false) + hard crash (graceful shutdown flushes and purges correctly since #806) + restart with replay. Tagged data heals; tagless does not. Not data loss, not corruption — over-counting.

Proposed fix: flush-watermark checkpointing

Record, per WAL file, a durable watermark of entries whose batches have completed their parquet flush (upload confirmed, not merely handed to the flusher). Recovery replays only entries past the watermark. That fixes tagless permanently and shrinks the tagged inflation window to ~zero.

Care points for the implementer:

Interim for 26.09.2

Document the current semantics in the WAL docs + release notes: replay is at-least-once; tagged duplicates reconcile at the next compaction; tagless duplicates currently persist. (Separate small PR.)

Related: #946 (smoke assertion), #945 (fresh-bucket reads), #806/#803 (shutdown ordering), #594 (WAL file deletion).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions