Area: mq / ingest: embedded broker retention
Issue. The embedded broker gives each tenant one ingest stream with limits retention. The sweeper reclaims space by purging events that are both acknowledged by the ingest worker's durable consumer and older than the tenant's stream.gap_window_minutes. "Acknowledged" means below the consumer's ack floor: the highest sequence below which every event is acked.
A single event that stays unacked therefore pins the ack floor for the whole stream. For example, a table whose inserts are retried during a table-scoped ClickHouse failure (read-only table, too many parts, missing grant), or a row left unacked because the DLQ is off for its table. While it is pinned, nothing above it can be purged, including every other table's already-written events. The stream grows toward mq.max_bytes_gb, and at the cap DiscardNew refuses new publishes, so every table of that tenant gets 503 because of one table.
How the external NATS mode avoids it (measured in #613's S1 test). Its ingest partitions use interest retention, so each event is deleted as soon as it is acked, whatever the acknowledgement state of its neighbours. SSE replay reads a separate history stream with its own max_age. An unacked event holds only itself.
Direction. Move the embedded broker to the same model: an interest-retention ingest stream plus a history stream with a time limit, per tenant. Then drop the embedded sweeper, making embedded and external share one retention design.
Things to check first:
- the extra disk that storing an event twice costs during the history window;
- the source consumer's attach window (S1 measured rows acked before it attaches never reaching history);
- the upgrade path for existing per-tenant streams.
Related: #613, #532, #138.
Area: mq / ingest: embedded broker retention
Issue. The embedded broker gives each tenant one ingest stream with limits retention. The sweeper reclaims space by purging events that are both acknowledged by the ingest worker's durable consumer and older than the tenant's
stream.gap_window_minutes. "Acknowledged" means below the consumer's ack floor: the highest sequence below which every event is acked.A single event that stays unacked therefore pins the ack floor for the whole stream. For example, a table whose inserts are retried during a table-scoped ClickHouse failure (read-only table, too many parts, missing grant), or a row left unacked because the DLQ is off for its table. While it is pinned, nothing above it can be purged, including every other table's already-written events. The stream grows toward
mq.max_bytes_gb, and at the capDiscardNewrefuses new publishes, so every table of that tenant gets503because of one table.How the external NATS mode avoids it (measured in #613's S1 test). Its ingest partitions use interest retention, so each event is deleted as soon as it is acked, whatever the acknowledgement state of its neighbours. SSE replay reads a separate history stream with its own
max_age. An unacked event holds only itself.Direction. Move the embedded broker to the same model: an interest-retention ingest stream plus a history stream with a time limit, per tenant. Then drop the embedded sweeper, making embedded and external share one retention design.
Things to check first:
Related: #613, #532, #138.