You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Read authorization dominates join cost during mass reconnects: rejoining takes ~2 min and delays delivery on already-joined channels #2341
With the channel registry fix from #2337 / #2338 applied, read authorization becomes the bottleneck when many private channels reconnect at once. Sockets with many channels take about two minutes to rejoin, and messages to channels they have already rejoined are delayed by 15–22 s (p95) while that happens.
That is one transaction on the tenant pool (authorization.ex#L245-L282): probe insert into realtime.messages, set_config for role and claims, an RLS SELECT of the probe rows, and ROLLBACK AND CHAIN.
During a mass reconnect, normal schedulers sit at 100% with run queues up to about 600. The joining channel process keeps its connection checked out while it waits to be scheduled. Connection hold time is then set by CPU contention rather than by Postgres, and the pool runs dry: checkout queue above 1,200, 0 ready connections.
Queue timeouts feed authorization_errors_per_second_rate (tenants.ex#L379-L395). Its limit is the realtime_connect pool size (#L451-L455), so the breaker opens and further joins fail with IncreaseConnectionPool.
Two things make it worse for sockets with many channels:
Joins are serialized per socket.Phoenix.Socket handles phx_join by calling Phoenix.Channel.Server.join/4, which blocks in a receive until the channel process replies (deps/phoenix/lib/phoenix/socket.ex L739 and lib/phoenix/channel/server.ex L17-L58 at the pinned supabase/phoenix commit 2d728ff). A socket with 98 channels rejoins them one at a time. While a join is pending, the socket process cannot push messages for the channels it has already rejoined.
Failed joins sleep before they reply.join_error/1 sleeps CHANNEL_ERROR_BACKOFF_MS (5 s by default) in the channel process before returning the error (realtime_channel.ex#L1193-L1198). Because of the serialization above, the socket's next join waits for that sleep as well.
Use private channels with RLS policies on realtime.messages.
Connect 300 sockets with 98 private channels each and 1,400 sockets with 2 each (about 32k channels), and publish continuously.
Drop every socket at once and let the clients reconnect and rejoin.
Observed with the stock read authorization:
Sockets with 98 channels needed 116–131 s to reach 99.9% rejoined.
Delivery lag (event age) p95 during the storm was 15–22 s.
Normal schedulers stayed saturated for about 110 s.
For comparison, we ran an experiment where read authorization was decided without the database. It evaluated the decision from the JWT claims, which only works because our policy uses nothing but the claims, so it is not proposed as a fix. Under the same storm:
99.9% rejoined in 8.5 s.
Event age p95 was 165 ms, with no lost messages.
Schedulers were saturated for 7 s.
Settings that made it worse or did not help:
CHANNEL_ERROR_BACKOFF_MS=100 made recovery slower: clients retried sooner while the CPU was already saturated.
max_joins_per_second does not pace joins. When it trips, the channel sends disconnect to the transport (realtime_channel.ex#L193-L195). The whole socket drops and its client rejoins every channel again, so sockets with many channels loop instead of slowing down.
A higher db_queue_target made it worse.
Expected behavior
A mass reconnect should mostly cost CPU for the joins themselves. Each join should not hold a tenant connection for a full authorization transaction, and a socket's slow joins should not delay delivery on its channels that are already joined.
Possible directions
I am not opening a PR for this one, since each option is a design choice for you to make:
Cache read decisions per tenant, topic and token (for example a hash of the claims, kept until exp). Policies are already cached in channel assigns for the JWT's lifetime (see the comment above maybe_assign_policies/3 from docs: Document when RLS policies are evaluated #2294). Extending that cache across rejoins with the same token would let a reconnect storm skip the database for topics already authorized.
Batch a socket's pending joins into one authorization transaction: one probe insert and one RLS SELECT for N topics instead of N transactions.
Evaluate policies without the probe insert, for example through an opt-in policy function that receives the topic and extension. That would remove the write and rollback from every join.
Make the authorization breaker threshold configurable separately from the realtime_connect pool size, so a burst of slow joins does not cut off all authorization for the tenant.
Related: #2337 / #2338 (registry cost of the same storm, which has to be fixed first), and #2339 / #2340 (server-side broadcasts dropped while this breaker is open).
Bug report
Describe the bug
With the channel registry fix from #2337 / #2338 applied, read authorization becomes the bottleneck when many private channels reconnect at once. Sockets with many channels take about two minutes to rejoin, and messages to channels they have already rejoined are delayed by 15–22 s (p95) while that happens.
Where the time goes
Authorization.get_read_authorizations/4(realtime_channel.ex#L1041-L1075).authorization.ex#L245-L282): probe insert intorealtime.messages,set_configfor role and claims, an RLSSELECTof the probe rows, andROLLBACK AND CHAIN.authorization_errors_per_second_rate(tenants.ex#L379-L395). Its limit is therealtime_connectpool size (#L451-L455), so the breaker opens and further joins fail withIncreaseConnectionPool.Two things make it worse for sockets with many channels:
Phoenix.Sockethandlesphx_joinby callingPhoenix.Channel.Server.join/4, which blocks in areceiveuntil the channel process replies (deps/phoenix/lib/phoenix/socket.exL739 andlib/phoenix/channel/server.exL17-L58 at the pinnedsupabase/phoenixcommit2d728ff). A socket with 98 channels rejoins them one at a time. While a join is pending, the socket process cannot push messages for the channels it has already rejoined.join_error/1sleepsCHANNEL_ERROR_BACKOFF_MS(5 s by default) in the channel process before returning the error (realtime_channel.ex#L1193-L1198). Because of the serialization above, the socket's next join waits for that sleep as well.To Reproduce
realtime.messages.Observed with the stock read authorization:
For comparison, we ran an experiment where read authorization was decided without the database. It evaluated the decision from the JWT claims, which only works because our policy uses nothing but the claims, so it is not proposed as a fix. Under the same storm:
Settings that made it worse or did not help:
CHANNEL_ERROR_BACKOFF_MS=100made recovery slower: clients retried sooner while the CPU was already saturated.max_joins_per_seconddoes not pace joins. When it trips, the channel sendsdisconnectto the transport (realtime_channel.ex#L193-L195). The whole socket drops and its client rejoins every channel again, so sockets with many channels loop instead of slowing down.db_queue_targetmade it worse.Expected behavior
A mass reconnect should mostly cost CPU for the joins themselves. Each join should not hold a tenant connection for a full authorization transaction, and a socket's slow joins should not delay delivery on its channels that are already joined.
Possible directions
I am not opening a PR for this one, since each option is a design choice for you to make:
exp). Policies are already cached in channel assigns for the JWT's lifetime (see the comment abovemaybe_assign_policies/3from docs: Document when RLS policies are evaluated #2294). Extending that cache across rejoins with the same token would let a reconnect storm skip the database for topics already authorized.SELECTfor N topics instead of N transactions.realtime_connectpool size, so a burst of slow joins does not cut off all authorization for the tenant.System information
Additional context
Related: #2337 / #2338 (registry cost of the same storm, which has to be fixed first), and #2339 / #2340 (server-side broadcasts dropped while this breaker is open).