Skip to content

mq: durable/partition topology checks miss backoff, pause_until, max_waiting, subject_transform, sources #689

Description

@EricAndrechek

Area: mq — external NATS topology validation

WaveHouse's boot-time verifier of an operator-created NATS topology (internal/mq/nats_topology.go) checks most fields of a partition stream and a shard durable against what WaveHouse needs, but not all of them.

  • topologyVerifier.partition (nats_topology.go:421) checks Subjects, Retention, RePublish (via v.republish), Discard, MaxBytes, MaxAge, Storage, Duplicates, NoAck, Sealed, Mirror, MaxMsgsPerSubject/DiscardNewPerSubject, DenyPurge/DenyDelete, PersistMode, Replicas and the wavehouse.dev/partition(s)/wavehouse.dev/shards metadata — but not SubjectTransform or Sources.
  • topologyVerifier.durable (nats_topology.go:508) checks AckPolicy, AckWait, MaxDeliver, MaxAckPending, DeliverPolicy, the filter subject, HeadersOnly, ReplayPolicy, InactiveThreshold, MaxRequestExpires, MaxRequestBatch, PriorityPolicy, PriorityGroups and PinnedTTL — but not Backoff, PauseUntil, or MaxWaiting.

wavehouse mq manifests (internal/mq/nats_manifests.go) never sets any of these six fields on the streams or durables it generates, so a topology produced by the generator is unaffected either way. The gap only matters to an operator who hand-writes or hand-edits the nack Stream/Consumer CRs instead of using the generator: a mismatch in any of these six fields currently passes verifyNATSTopology/awaitNATSTopology silently, with wavehouse_mq_topology_ok reading 1.

Concrete failure scenarios (inferred from reading nats_topology.go and the nats-server consumer/stream config semantics; not run):

  • A partition's subject_transform changes the subject rows are actually stored under. No shard durable's FilterSubjects or the history's republish source then matches the transformed subject — rows are still kept (retention is unaffected), but never delivered to any ingest worker, and the verifier reports nothing wrong.
  • A durable's Backoff is set to something like [60s, 1s]. The server sets the initial ack_wait from backoff[0] (so a check against t.AckWait on cfg.AckWait alone passes), but later redeliveries use later entries in the slice — so with backoff: [60s, 1s], a held row is redelivered after just 1s on its second attempt, not the 60s the durable's own ack_wait implies.

Reproduce/confirm: hand-write a nack Stream CR with a subjectTransform that rewrites the ingest subject, apply it, and check that verifyNATSTopology reports no findings while a published row is never delivered to any shard durable; separately, hand-write a Consumer CR with backoff: [60s, 1s] and confirm the verifier accepts it, then measure the actual redelivery interval of a held row.

Found in review of #624.

Related: #624, #613.

Metadata

Metadata

Assignees

No one assigned

    Labels

    area/configConfig file, config knobs, hot-reloadenhancementNew feature or request

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions