Bug report
Describe the bug
The per-socket channel limit in RealtimeChannel.limit_channels/2 keeps all of a tenant's channels under one duplicate key in Realtime.Registry. That makes each join cost O(total channels on the node), and each channel exit cost O(channels of the tenant). When many sockets disconnect at once, cleanup is O(N²) and all of it runs in the registry's single Registry.PIDPartition0 process. Under reconnect load that process falls behind and its backlog keeps growing.
Root cause
lib/realtime_web/channels/realtime_channel.ex#L694-L709:
def limit_channels(tenant, %{transport_pid: pid} = socket) do
key = Tenants.channels_per_client_key(tenant) # {:channel, :clients_per, external_id}
count = Registry.count_match(Realtime.Registry, key, pid)
...
Registry.register(Realtime.Registry, key, pid)
channels_per_client_key/1 is one key per tenant, so every channel of the tenant is registered under the same key. Realtime.Registry is keys: :duplicate with a single partition (application.ex#L172).
- Join:
Registry.count_match/3 builds the match spec {:_, {:_, pattern}} and checks the key in a guard (Elixir 1.19 registry.ex, count_match/4). ETS cannot use the key hash for that, so each join runs :ets.select_count over the whole registry table. The cost grows with all channels on the node, not only the joining socket's channels.
- Exit: when a channel process exits,
Registry.PIDPartition0 handles the :EXIT and runs :ets.match_delete(key_ets, {key, {pid, :_}}). With every entry of the tenant under one key, each delete walks the whole bucket for that key. N channels exiting together means O(N²) work, all done serially in one process.
To Reproduce
Standalone, without Realtime: this script registers channels the way limit_channels/2 does, then compares the current key with a per-transport key. It times one count on a full registry and the time for the registry to empty after every channel exits.
registry_bench.exs
# elixir --erl "+P 2000000" registry_bench.exs <total_channels> <channels_per_socket>
[total, per_socket] = System.argv() |> Enum.map(&String.to_integer/1)
sockets = div(total, per_socket)
defmodule Bench do
def run(label, key_fun, count_fun, sockets, per_socket) do
name = :"reg_#{label}"
{:ok, _} = Registry.start_link(keys: :duplicate, name: name)
parent = self()
transports = for _ <- 1..sockets, do: spawn(fn -> Process.sleep(:infinity) end)
channels =
for t <- transports, _ <- 1..per_socket do
spawn(fn ->
Registry.register(name, key_fun.(t), t)
send(parent, :registered)
receive do: (:stop -> :ok)
end)
end
for _ <- channels, do: receive(do: (:registered -> :ok))
probe = hd(transports)
{count_us, n} = :timer.tc(fn -> Enum.reduce(1..100, 0, fn _, _ -> count_fun.(name, key_fun.(probe), probe) end) end)
{exit_us, _} =
:timer.tc(fn ->
for c <- channels, do: send(c, :stop)
wait_empty(name)
end)
IO.puts("#{label}: count per join #{Float.round(count_us / 100, 1)} us (count=#{n}), mass-exit cleanup #{div(exit_us, 1000)} ms")
for t <- transports, do: Process.exit(t, :kill)
end
defp wait_empty(name) do
if Registry.count(name) == 0, do: :ok, else: (Process.sleep(5); wait_empty(name))
end
end
tenant_key = {:channel, :clients_per, "tenant"}
Bench.run(:tenant_key_count_match, fn _t -> tenant_key end, fn name, key, t -> Registry.count_match(name, key, t) end, sockets, per_socket)
Bench.run(:transport_key_lookup, fn t -> {tenant_key, t} end, fn name, key, _t -> length(Registry.lookup(name, key)) end, sockets, per_socket)
Elixir 1.19.5 / OTP 28 on an Apple Silicon laptop:
| channels (per socket) |
key |
count per join |
registry empty after mass exit |
| 10,000 (50) |
tenant key + count_match (current) |
220 µs |
990 ms |
| 10,000 (50) |
{tenant key, transport_pid} + lookup |
2.7 µs |
39 ms |
| 30,000 (100) |
tenant key + count_match (current) |
1,060 µs |
11,486 ms |
| 30,000 (100) |
{tenant key, transport_pid} + lookup |
5.7 µs |
217 ms |
Tripling the channel count raises cleanup time about 11.6× with the current key, which fits O(N²).
Load test on Realtime
Self-hosted v2.140.7 on 4 vCPU, one tenant, about 32k channels (300 sockets × 98 channels plus 1,400 sockets × 2 channels). We then dropped every socket at once and let the clients reconnect.
- Current code: right after the disconnect, the tenant's key still held 41,484 entries and
Realtime.Registry.PIDPartition0 had a backlog of 14,008 :EXIT messages (current_function was :ets.match_delete). Three minutes later the key held 59,934 entries: stale entries were being added faster than they were removed. In 5 minutes the 98-channel sockets rejoined only 2,142 of 29,400 channels.
- With the per-transport key: the registry size tracked live channels exactly, with no partition backlog. The 98-channel sockets rejoined 21,002 of 29,400 channels in 5 minutes. The 2-channel clients recovered in 5.8 s instead of 9.5 s. After the storm, p95 event age for those clients was 1.6 s instead of 33 s, and failed broadcasts dropped from about 18,000 to about 950.
Expected behavior
The cost of a join, and of cleanup after a channel exits, should depend only on that socket's own channels (at most max_channels_per_client), not on the total channels of the tenant or node.
Proposed fix
Register under {Tenants.channels_per_client_key(tenant), transport_pid} and count with length(Registry.lookup(Realtime.Registry, key)). The limit stays per transport, and the ChannelRateLimitReached log at count + 1 == max is unchanged. lookup/2 and the exit-time match_delete both hash on the key, so each touches at most max_channels_per_client entries. channels_per_client_key/1 is unchanged; it is also used as a limiter id in Tenants.limiter_keys/1, which does not read this registry. PR to follow.
System information
- Realtime: v2.140.7 (load test); same code on
main at f86df8c
- Elixir 1.19.5 / OTP 28
Bug report
Describe the bug
The per-socket channel limit in
RealtimeChannel.limit_channels/2keeps all of a tenant's channels under one duplicate key inRealtime.Registry. That makes each join cost O(total channels on the node), and each channel exit cost O(channels of the tenant). When many sockets disconnect at once, cleanup is O(N²) and all of it runs in the registry's singleRegistry.PIDPartition0process. Under reconnect load that process falls behind and its backlog keeps growing.Root cause
lib/realtime_web/channels/realtime_channel.ex#L694-L709:channels_per_client_key/1is one key per tenant, so every channel of the tenant is registered under the same key.Realtime.Registryiskeys: :duplicatewith a single partition (application.ex#L172).Registry.count_match/3builds the match spec{:_, {:_, pattern}}and checks the key in a guard (Elixir 1.19registry.ex,count_match/4). ETS cannot use the key hash for that, so each join runs:ets.select_countover the whole registry table. The cost grows with all channels on the node, not only the joining socket's channels.Registry.PIDPartition0handles the:EXITand runs:ets.match_delete(key_ets, {key, {pid, :_}}). With every entry of the tenant under one key, each delete walks the whole bucket for that key. N channels exiting together means O(N²) work, all done serially in one process.To Reproduce
Standalone, without Realtime: this script registers channels the way
limit_channels/2does, then compares the current key with a per-transport key. It times one count on a full registry and the time for the registry to empty after every channel exits.registry_bench.exs
Elixir 1.19.5 / OTP 28 on an Apple Silicon laptop:
count_match(current){tenant key, transport_pid}+lookupcount_match(current){tenant key, transport_pid}+lookupTripling the channel count raises cleanup time about 11.6× with the current key, which fits O(N²).
Load test on Realtime
Self-hosted v2.140.7 on 4 vCPU, one tenant, about 32k channels (300 sockets × 98 channels plus 1,400 sockets × 2 channels). We then dropped every socket at once and let the clients reconnect.
Realtime.Registry.PIDPartition0had a backlog of 14,008:EXITmessages (current_functionwas:ets.match_delete). Three minutes later the key held 59,934 entries: stale entries were being added faster than they were removed. In 5 minutes the 98-channel sockets rejoined only 2,142 of 29,400 channels.Expected behavior
The cost of a join, and of cleanup after a channel exits, should depend only on that socket's own channels (at most
max_channels_per_client), not on the total channels of the tenant or node.Proposed fix
Register under
{Tenants.channels_per_client_key(tenant), transport_pid}and count withlength(Registry.lookup(Realtime.Registry, key)). The limit stays per transport, and theChannelRateLimitReachedlog atcount + 1 == maxis unchanged.lookup/2and the exit-timematch_deleteboth hash on the key, so each touches at mostmax_channels_per_cliententries.channels_per_client_key/1is unchanged; it is also used as a limiter id inTenants.limiter_keys/1, which does not read this registry. PR to follow.System information
mainat f86df8c