Skip to content

mq(dlq): shrinking mq.max_bytes_gb silently deletes the oldest dead letters — investigate how the reload should treat a non-empty DLQ #532

Description

@taitelee

Background (verified empirically, nats-server 2.14.5)

Investigated what a hot-reload shrink of mq.max_bytes_gb does to resident stream data (temporary tests against the real NewEmbedded/Resize path, plus a version-matrix experiment):

  • Ingest stream (LimitsPolicy + DiscardNew): safe. Shrinking MaxBytes below resident size deletes nothing — all messages verified byte-for-byte intact, including across a restart with the lowered limit. New publishes are rejected (503 err_code=10077 maximum bytes exceeded, surfaced by the ingest API as HTTP 503 + Retry-After) until the Active Sweeper drains the stream under the new limit, then ingestion resumes on its own. The check-pause-drain orchestration we considered building is native DiscardNew behavior.
  • DLQ stream (LimitsPolicy + DiscardOld): loses data. The same adoption calls EnsureDLQStream with the new budget (max_bytes / 10, cmd/wavehouse/main.go), and DiscardOld limit enforcement synchronously deletes the oldest dead letters until the stream fits. In the control test a shrink wiped 85 of 100 resident messages at the moment of the settings adoption.
  • Version floor: the ingest-stream safety exists only on nats-server >= 2.10.6 (Only drop firstSeq under DiscardOld policy. nats-io/nats-server#4802, "Only drop firstSeq under DiscardOld policy"). On 2.9.x/early 2.10.x the identical shrink deleted 85 of 100 messages from a DiscardNew stream and kept accepting publishes. We pin 2.14.5 in go.mod.

Intended direction: DLQ backed by a ClickHouse table

We want dead letters to end up in their own ClickHouse table rather than living only in the NATS stream. Rows land on the DLQ precisely because a ClickHouse write failed, so ClickHouse cannot be the only home — the NATS stream stays as the bounded buffer for the ClickHouse-is-down case. The shape to design toward:

  • A consumer drains dlq.> into a raw-typed DLQ table (payload/error/target table/timestamp as strings, so a schema mismatch can't fail a second time), removing from NATS only after the ClickHouse write is acked.
  • ClickHouse TTL owns long-term retention; the NATS stream's DiscardOld stays correct for its reduced buffer role (oldest lost only if ClickHouse is down long enough to overflow the bounded buffer).
  • With the NATS side normally near-empty, the shrink-time pruning below becomes mostly moot.

To investigate

  • The ClickHouse-table design above: table schema, drain consumer, replay path from the table, and whether S3 export is worth considering as an alternative/complement for durability independent of ClickHouse health.
  • Until that lands: should the reload hook warn, refuse, or defer the DLQ shrink when the DLQ holds more than the new budget (currently it silently prunes)? A cheap interim option: only ever grow the DLQ limit, or log the number of dead letters about to be dropped before applying.
  • Add a regression test pinning the DiscardNew no-deletion-on-shrink behavior (and the nats-server >= 2.10.6 floor it depends on), since TestEmbeddedNATS_Resize only checks the config round-trip.

Relevant code: internal/mq/embedded.go (Resize, ingestStreamConfig), internal/api/dlq.go (EnsureDLQStream), cmd/wavehouse/main.go (AfterAdopt hook), internal/ingest/sweeper.go.

Metadata

Metadata

Assignees

No one assigned

    Labels

    area/configConfig file, config knobs, hot-reloadarea/streamingSSE / live-query delivery path (/v1/stream)enhancementNew feature or request

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions