Version: hackney 4.8.5, Erlang/OTP 29.1.1, Elixir 1.20.4, Linux.
Summary
Requests that ask for socket options through connect_options (for example [nodelay: true]) can be served by pooled TCP connections that were opened without them:
- Prewarmed connections ignore the caller's options.
hackney_pool:prewarm_connections/5 opens connections with connect_options => [] and ssl_options => [] (hackney_pool.erl, around line 1375).
- The plain TCP pool key does not include connect options. The key is
connection_key(Host, Port, Transport) (hackney_pool.erl:131). So a prewarmed or earlier connection is checked out by a later request whose connect_options differ.
Why it matters: TCP_NODELAY
hackney_tcp:connect/4 uses [binary, {active, false}, {packet, raw}] as its base options, with no {nodelay, true}. hackney_ws, hackney_http_connect and hackney_socks5 do set {nodelay, true}.
After upgrading from hackney 1.25 to 4.x we saw:
- WebDriver calls (Wallaby to chromedriver on loopback) go from about 0.8 ms to about 83 ms.
- Every small POST with a body waits for Nagle plus delayed ACK: about 41 ms extra per request on loopback. 1.25 measured about 41 ms total on the same server.
We worked around it by passing connect_options: [nodelay: true] per client. That fixes fresh connections, but once a host has been prewarmed, the prewarmed connections are reused without nodelay and the delay comes back.
Reproduction (plain HTTP)
- Start a local keep-alive HTTP/1.1 server.
- Make a POST with
connect_options: [nodelay: true]. It is fast.
- Trigger prewarm for that host, or exceed the pool so the maintain-prewarm path runs.
- Make the same POST again. It is served by a prewarmed connection and waits about 40 ms.
Suggested fixes, any of which would help
- Set
{nodelay, true} by default in hackney_tcp:connect/4, as the websocket and proxy transports already do, and as 1.x behaved in our measurements.
- Carry the pool's or first request's
connect_options into prewarm_connections.
- Include
connect_options in the TCP pool key, the way the SSL path keys on TLS options, so connections opened with different socket options are not shared.
Happy to test a patch.
Version: hackney 4.8.5, Erlang/OTP 29.1.1, Elixir 1.20.4, Linux.
Summary
Requests that ask for socket options through
connect_options(for example[nodelay: true]) can be served by pooled TCP connections that were opened without them:hackney_pool:prewarm_connections/5opens connections withconnect_options => []andssl_options => [](hackney_pool.erl, around line 1375).connection_key(Host, Port, Transport)(hackney_pool.erl:131). So a prewarmed or earlier connection is checked out by a later request whoseconnect_optionsdiffer.Why it matters: TCP_NODELAY
hackney_tcp:connect/4uses[binary, {active, false}, {packet, raw}]as its base options, with no{nodelay, true}.hackney_ws,hackney_http_connectandhackney_socks5do set{nodelay, true}.After upgrading from hackney 1.25 to 4.x we saw:
We worked around it by passing
connect_options: [nodelay: true]per client. That fixes fresh connections, but once a host has been prewarmed, the prewarmed connections are reused withoutnodelayand the delay comes back.Reproduction (plain HTTP)
connect_options: [nodelay: true]. It is fast.Suggested fixes, any of which would help
{nodelay, true}by default inhackney_tcp:connect/4, as the websocket and proxy transports already do, and as 1.x behaved in our measurements.connect_optionsintoprewarm_connections.connect_optionsin the TCP pool key, the way the SSL path keys on TLS options, so connections opened with different socket options are not shared.Happy to test a patch.