Skip to content

ingest: a lasting one-table failure suspends delivery for every table of its tenant #658

Description

@EricAndrechek

Area: ingest / mq — follow-up to #619

Issue. Since #619, a batch ClickHouse cannot take goes back to the queue with a delayed nak instead of to the DLQ. A nak'd message stays on the consumer's pending list, so it counts toward the tenant's maxAckPending (10,000, internal/ingest/worker.go). For a whole-server outage that is the intended backpressure, since every table is down anyway. For a failure of one table (chconn.TableScoped: TABLE_IS_READ_ONLY, TABLE_IS_PERMANENTLY_READ_ONLY, TOO_MANY_PARTS, TOO_MANY_MUTATIONS, ACCESS_DENIED) it is not. That table backs off alone, but its rows keep arriving and are handed back. Once 10,000 of them are pending, the server stops delivering the tenant's other tables too, and they stop inserting until the failing table is fixed. The failures that last long enough to get there are the common ones: a missing grant, a permanently read-only table.

Before #619, with the DLQ on for the table, those rows were parked and acked, so they never counted toward the limit.

How sure. Inferred from the code and from JetStream's pending accounting; #619 verified in nats-server processNak that a delayed nak keeps the message pending. Not reproduced end to end. #619 documents the behaviour in ingest-pipeline.md (§When ClickHouse cannot take an insert) and in the CHANGELOG.

Options.

  1. Bound it. Once a table's own backoff has been open longer than N, and the DLQ is on for that table, park its rows instead of retrying them. That restores the pre-fix(ingest): retry ClickHouse outages instead of dead-lettering #619 outcome, but only for a failure ClickHouse has already localized to the table. Caveat: TABLE_IS_READ_ONLY on a replicated table during a Keeper outage is table-scoped in form but server-wide in practice, so N must outlast a normal Keeper interruption.
  2. Stop pulling the table instead of cycling its rows. One consumer per tenant can't pause one table. A per-table or per-partition consumer, as in epic(distributed): shared backends and standalone workers for multi-node deployments #613's distributed workers, would let a stuck table exhaust only its own budget.
  3. Accept it as backpressure and alert on it. One table's wavehouse_ingest_retries_total rising steadily while the tenant's other tables go flat is the signature.

Related: #653 is the other half of the same failure: the unacked rows pin the purge floor, so the stream fills. Also #613 and #619.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area/ingestIngest pipeline (Bento, batching, DLQ)enhancementNew feature or request

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions