Area: mq — external NATS ownership, mixed-version clusters
Two related gaps in shard ownership, both under internal/ingest (assignment) and internal/mq (the version check), that only show up mid-rollout.
1. A rolling change to N (partitions) or V (shards per partition) can pull the same shard from two processes at once, including shards that did not move.
Ownership is computed independently by each process from its own view of live membership and the configured unit set (assignUnits, internal/ingest/assign.go:17, a capped rendezvous hash). While some processes run the old N or V and others run the new one, the two groups are hashing over different unit sets, so they can disagree about who owns a given shard. docs/src/content/docs/deployment.md:438 and :454 already document that this produces unowned shards during a rollout (wavehouse_ingest_shards_unowned above 0 until it ends), and separately that a moved table (one whose shard count changed) can briefly have two writers. What is not documented or handled is that shards that did not move can also end up claimed by two processes at once during the same window — ownership disagreement is not confined to the newly added units.
When two processes both believe they own a shard, nothing crashes: the server's pinned-consumer mechanism (priority_policy: pinned_client on the shard durable) arbitrates and hands delivery to only one puller at a time, per internal/mq/nats_topology.go's durable checks. So no row is lost, but the assignment layer briefly has two owners for one unit, which is a correctness gap in the ownership model itself, not just a delivery race the server happens to paper over.
-
Scenario: partitions or shards are raised (wavehouse mq manifests --shards <newV>, rolled out pod by pod). Mid-rollout, some ingest processes run the old V, some the new V. Both groups compute assignUnits over their own unit set and can each conclude they own a shard that both the old and new topology still contain.
-
Evidence: measured, with a scratch simulation of assignUnits at two rollout sizes:
| rollout |
units |
pulled by nobody |
pulled by two (pin arbitrates) |
| N=4, V 8→16 |
64 |
3–4 |
3 |
| N=4, V 32→48 |
192 |
2–3 |
4–6 |
2. The version floor WaveHouse enforces is read from whichever server the client happens to be connected to, not the cluster as a whole.
internal/mq/nats_topology.go's topologyVerifier.serverVersion (called from verifyNATSTopology as v.serverVersion(js.Conn().ConnectedServerVersion())) checks the version of one connection's server. In a genuinely mixed-version cluster mid-upgrade, or behind a leafnode where the local server's version differs from the hub's, this can read as uniformly at or above the required minor (recommendedNATSMinor = "2.14.", minNATSServer) when the server that actually resolves a RESET or UNPIN for a given consumer — its leader — is on a different version. WaveHouse requires NATS 2.14 specifically because that is the first release with consumer RESET, which the ownership takeover path depends on (internal/mq/external.go's ResetOrphaned).
- Evidence: inferred from reading
nats_topology.go and the NATS server/leafnode connection model; not run against a mixed-version or leafnode cluster.
Constraints: The design decisions for #624 (see #613) already accept that per-table order is not promised across an N or V change, and that some shards are unowned until a rollout completes — that is documented behavior, not this issue. This issue is about (a) unowned/double-pulled units that are outside that accepted "moved table" case, and (b) the version check's blind spot, not about re-litigating the accepted rollout tradeoffs.
Reproduce/confirm: simulate assignUnits (internal/ingest/assign.go) with two overlapping unit sets representing old-V and new-V topologies and count units assigned in both; for the version floor, connect through a leafnode or a cluster with mixed 2.13/2.14 members and check what ConnectedServerVersion() reports versus what actually answers a RESET.
Found in review of #624.
Related: #624, #613.
Area: mq — external NATS ownership, mixed-version clusters
Two related gaps in shard ownership, both under
internal/ingest(assignment) andinternal/mq(the version check), that only show up mid-rollout.1. A rolling change to N (partitions) or V (shards per partition) can pull the same shard from two processes at once, including shards that did not move.
Ownership is computed independently by each process from its own view of live membership and the configured unit set (
assignUnits,internal/ingest/assign.go:17, a capped rendezvous hash). While some processes run the old N or V and others run the new one, the two groups are hashing over different unit sets, so they can disagree about who owns a given shard.docs/src/content/docs/deployment.md:438and:454already document that this produces unowned shards during a rollout (wavehouse_ingest_shards_unownedabove0until it ends), and separately that a moved table (one whose shard count changed) can briefly have two writers. What is not documented or handled is that shards that did not move can also end up claimed by two processes at once during the same window — ownership disagreement is not confined to the newly added units.When two processes both believe they own a shard, nothing crashes: the server's pinned-consumer mechanism (
priority_policy: pinned_clienton the shard durable) arbitrates and hands delivery to only one puller at a time, perinternal/mq/nats_topology.go's durable checks. So no row is lost, but the assignment layer briefly has two owners for one unit, which is a correctness gap in the ownership model itself, not just a delivery race the server happens to paper over.Scenario: partitions or shards are raised (
wavehouse mq manifests --shards <newV>, rolled out pod by pod). Mid-rollout, some ingest processes run the old V, some the new V. Both groups computeassignUnitsover their own unit set and can each conclude they own a shard that both the old and new topology still contain.Evidence: measured, with a scratch simulation of
assignUnitsat two rollout sizes:2. The version floor WaveHouse enforces is read from whichever server the client happens to be connected to, not the cluster as a whole.
internal/mq/nats_topology.go'stopologyVerifier.serverVersion(called fromverifyNATSTopologyasv.serverVersion(js.Conn().ConnectedServerVersion())) checks the version of one connection's server. In a genuinely mixed-version cluster mid-upgrade, or behind a leafnode where the local server's version differs from the hub's, this can read as uniformly at or above the required minor (recommendedNATSMinor = "2.14.",minNATSServer) when the server that actually resolves a RESET or UNPIN for a given consumer — its leader — is on a different version. WaveHouse requires NATS 2.14 specifically because that is the first release with consumer RESET, which the ownership takeover path depends on (internal/mq/external.go'sResetOrphaned).nats_topology.goand the NATS server/leafnode connection model; not run against a mixed-version or leafnode cluster.Constraints: The design decisions for #624 (see #613) already accept that per-table order is not promised across an N or V change, and that some shards are unowned until a rollout completes — that is documented behavior, not this issue. This issue is about (a) unowned/double-pulled units that are outside that accepted "moved table" case, and (b) the version check's blind spot, not about re-litigating the accepted rollout tradeoffs.
Reproduce/confirm: simulate
assignUnits(internal/ingest/assign.go) with two overlapping unit sets representing old-V and new-V topologies and count units assigned in both; for the version floor, connect through a leafnode or a cluster with mixed 2.13/2.14 members and check whatConnectedServerVersion()reports versus what actually answers a RESET.Found in review of #624.
Related: #624, #613.