Skip to content

Read authorization dominates join cost during mass reconnects: rejoining takes ~2 min and delays delivery on already-joined channels #2341

Description

@virgildotcodes

Bug report

  • I confirm this is a bug with Supabase, not with my own application.
  • I confirm I have searched the Docs, GitHub Discussions, and Discord.

Describe the bug

With the channel registry fix from #2337 / #2338 applied, read authorization becomes the bottleneck when many private channels reconnect at once. Sockets with many channels take about two minutes to rejoin, and messages to channels they have already rejoined are delayed by 15–22 s (p95) while that happens.

Where the time goes

  • Every private join runs Authorization.get_read_authorizations/4 (realtime_channel.ex#L1041-L1075).
  • That is one transaction on the tenant pool (authorization.ex#L245-L282): probe insert into realtime.messages, set_config for role and claims, an RLS SELECT of the probe rows, and ROLLBACK AND CHAIN.
  • During a mass reconnect, normal schedulers sit at 100% with run queues up to about 600. The joining channel process keeps its connection checked out while it waits to be scheduled. Connection hold time is then set by CPU contention rather than by Postgres, and the pool runs dry: checkout queue above 1,200, 0 ready connections.
  • Queue timeouts feed authorization_errors_per_second_rate (tenants.ex#L379-L395). Its limit is the realtime_connect pool size (#L451-L455), so the breaker opens and further joins fail with IncreaseConnectionPool.

Two things make it worse for sockets with many channels:

  • Joins are serialized per socket. Phoenix.Socket handles phx_join by calling Phoenix.Channel.Server.join/4, which blocks in a receive until the channel process replies (deps/phoenix/lib/phoenix/socket.ex L739 and lib/phoenix/channel/server.ex L17-L58 at the pinned supabase/phoenix commit 2d728ff). A socket with 98 channels rejoins them one at a time. While a join is pending, the socket process cannot push messages for the channels it has already rejoined.
  • Failed joins sleep before they reply. join_error/1 sleeps CHANNEL_ERROR_BACKOFF_MS (5 s by default) in the channel process before returning the error (realtime_channel.ex#L1193-L1198). Because of the serialization above, the socket's next join waits for that sleep as well.

To Reproduce

  1. Self-hosted Realtime v2.140.7 with fix: count channels per client under a per-transport registry key #2338 applied, on one 4 vCPU node.
  2. Use private channels with RLS policies on realtime.messages.
  3. Connect 300 sockets with 98 private channels each and 1,400 sockets with 2 each (about 32k channels), and publish continuously.
  4. Drop every socket at once and let the clients reconnect and rejoin.

Observed with the stock read authorization:

  • Sockets with 98 channels needed 116–131 s to reach 99.9% rejoined.
  • Delivery lag (event age) p95 during the storm was 15–22 s.
  • Normal schedulers stayed saturated for about 110 s.

For comparison, we ran an experiment where read authorization was decided without the database. It evaluated the decision from the JWT claims, which only works because our policy uses nothing but the claims, so it is not proposed as a fix. Under the same storm:

  • 99.9% rejoined in 8.5 s.
  • Event age p95 was 165 ms, with no lost messages.
  • Schedulers were saturated for 7 s.

Settings that made it worse or did not help:

  • CHANNEL_ERROR_BACKOFF_MS=100 made recovery slower: clients retried sooner while the CPU was already saturated.
  • max_joins_per_second does not pace joins. When it trips, the channel sends disconnect to the transport (realtime_channel.ex#L193-L195). The whole socket drops and its client rejoins every channel again, so sockets with many channels loop instead of slowing down.
  • A higher db_queue_target made it worse.

Expected behavior

A mass reconnect should mostly cost CPU for the joins themselves. Each join should not hold a tenant connection for a full authorization transaction, and a socket's slow joins should not delay delivery on its channels that are already joined.

Possible directions

I am not opening a PR for this one, since each option is a design choice for you to make:

  1. Cache read decisions per tenant, topic and token (for example a hash of the claims, kept until exp). Policies are already cached in channel assigns for the JWT's lifetime (see the comment above maybe_assign_policies/3 from docs: Document when RLS policies are evaluated #2294). Extending that cache across rejoins with the same token would let a reconnect storm skip the database for topics already authorized.
  2. Batch a socket's pending joins into one authorization transaction: one probe insert and one RLS SELECT for N topics instead of N transactions.
  3. Evaluate policies without the probe insert, for example through an opt-in policy function that receives the topic and extension. That would remove the write and rollback from every join.
  4. Make the authorization breaker threshold configurable separately from the realtime_connect pool size, so a burst of slow joins does not cut off all authorization for the tenant.

System information

Additional context

Related: #2337 / #2338 (registry cost of the same storm, which has to be fixed first), and #2339 / #2340 (server-side broadcasts dropped while this breaker is open).

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions