Repository navigation
Implement parallel Happy Eyeballs as described in RFC 8305 #48145
Description
Activity
I originally went for the proper parallel approach as states in RFC8305. I chose not to implement it as I had issues with the code.
Consider that implementing proper RFC8305 will imply that all of sudden Node.js starts opening multiple connections instead of one in case of multiple IP returned by DNS, which is dangerous.
The unwanted timeouts are not properly handled in #47860, which will be available in Node.js 20 once I fix #48000 and #47822.
We can of course disable the feature completely by default, but it was my understanding that this was longly awaited to help beginners struggling with poorly configured dual stack networks.
Consider that implementing proper RFC8305 will imply that all of sudden Node.js starts opening multiple connections instead of one in case of multiple IP returned by DNS, which is dangerous.
Is that a fact? Is there absolutely no way to prevent responding to more than a single address with a TCP ACK packet after having received TCP SYN-ACK?
Nope, we don't control the connection at that level, AFAIK.
@nodejs/libuv Am I wrong?
I think it should be considered disabling this autoSelectFamily by default. It's breaking stuff all over the place. Stabilize first behind a --experimental-auto-select-family or something, but keep this broken stuff out of the upcoming LTS please.
After 20.3.0 all issues should be fixed. Can we please wait for this release before rushing into decisions?
The ENETUNREACH issue mentioned here is still unresolved, right? As I understand it, it can only be fixed with parallel connection attempts.
Not really. We can add a new feature to opt-out A or AAAA in the list of addresses to try when the machine has no IPv4 and IPv6 connectivity respectively.
Note that ENETUNREACH will happen even in previous versions if the DNS only returns AAAA records.
Additionally, to mitigate the problem, a simple fix would be to raise the minimum attempt timeout to 500ms to reduce the possibility that IPv4 fails.
I was interested so I tried reading the RFC as well.
Just giving my 2 cents of what I understand since I saw the discussion...
TBH I was also confused at first how you can make both "non-simultaneous" and "may occur in parallel" connections.
The first paragraph seems conflicting. 😅
But the key is in 2nd paragraph:
A simple implementation can have a fixed delay for how long to wait before starting the next connection attempt.
Which is "non-simultaneous" but also "may occur in parallel".
An alternative/a more sophisticated delay is by using exponential backoff with min & max delay (described in paragraph 3).
Is there absolutely no way to prevent responding to more than a single address with a TCP ACK packet after having received TCP SYN-ACK?
Not in a portable fashion. I'm not even sure you could do it reliably if all you cared about was Linux.
The only way I'm aware of is polling getsockopt(TCP_INFO) and checking the tcp_info.tcpi_state field but that's a) timing sensitive, and b) not a viable approach for an asynchronous runtime like Node.js.
@HinataKah0 Yes, I interpreted in that way as well.
But, as @bnoordhuis said, since we cannot really control the socket once is started, I chose to implement in completely non-overlapped way so that we try to avoid issues and race conditions.
github-actions commented on Dec 8, 2023
On a side note, the implementation in Node.js 20.13.1 also appears to break some node-gyp builds that now fail to download headers from our very own domain nodejs.org (tniessen/node-pqclean#3 (comment)):
gyp http GET https://nodejs.org/download/release/v20.13.1/node-v20.13.1-headers.tar.gz
gyp http fetch GET https://nodejs.org/download/release/v20.13.1/node-v20.13.1-headers.tar.gz attempt 1 failed with ETIMEDOUT
gyp WARN install got an error, rolling back install
gyp ERR! configure error
gyp ERR! stack FetchError: request to https://nodejs.org/download/release/v20.13.1/node-v20.13.1-headers.tar.gz failed, reason:
gyp ERR! stack at ClientRequest.<anonymous> (/usr/lib/node_modules/npm/node_modules/minipass-fetch/lib/index.js:130:14)
gyp ERR! stack at ClientRequest.emit (node:events:519:28)
gyp ERR! stack at _destroy (node:_http_client:880:13)
gyp ERR! stack at onSocketNT (node:_http_client:900:5)
gyp ERR! stack at process.processTicksAndRejections (node:internal/process/task_queues:83:21)
Setting --no-enable-network-family-autoselection seems to fix node-gyp builds.
Rather than disabling the Happy Eyeballs altogether, you could use --network-family-autoselection-attempt-timeout and see if the situation improves.
Meant to comment this on #54359, sorry. Leaving below for posterity.
This has been discussed elsewhere before, e.g. #52216 and https://github.com/orgs/nodejs/discussions/48028.
I very much agree that Node should provide sensible defaults; I can't see how the current defaults are sensible. As a possibly-helpful datapoint, curl and Python are both able to fetch from https://nodejs.org/download/release/v20.17.0/node-v20.17.0-headers.tar.gz on a machine with high latency, and both (AFAICT) have Happy Eyeballs enabled. That same request fails under Node with the error described on these two threads.
I don't know enough about curl/Python/other runtimes to know what they're doing differently, but it's clear that better default behavior is possible, and I hope that it can work its way into Node, even if it has to wait for a future major version.
Another variant of this on v22.12.0 (WSL2, no IPv6 route).
In my case it's not about the 250ms timeout being too short — IPv6 fails instantly with ENETUNREACH (kernel has no route), but autoSelectFamily still doesn't recover. --dns-result-order=ipv4first alone doesn't help either; need --no-network-family-autoselection to fix it.
Minimal repro:
// times out
tls.connect({ host: 'registry.npmjs.org', port: 443, servername: 'registry.npmjs.org' });
// works
tls.connect({ host: 'registry.npmjs.org', port: 443, servername: 'registry.npmjs.org', autoSelectFamily: false, family: 4 });
This seems like a case where ENETUNREACH should be trivially handled as an instant failure — no libuv cancellation needed since the kernel already rejected it. Even without parallel attempts, the algorithm could just immediately move to the next candidate instead of stalling.
Another victim here ;-)
Joplin desktop, Australia -> EU (Koofr), 330ms connect, killed by the 250ms default
laurent22/joplin#16215
I built compliant-eyeballs for this exact gap. It follows RFC 8305 Section 5’s overlapping-attempt rule: trying the next IPv4/IPv6 address doesn’t kill a slow but viable connection. It has TCP/TLS APIs, Node HTTP(S) agents, and an Undici connector, with no runtime dependencies. I compared its behavior with pinned Node.js, Go, curl, and Chromium builds. CI passes on Linux, macOS, and Windows with Node 20, 22, 24, and 26.
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsAwaiting Triage

What is the problem this feature will solve?
autoSelectFamilyis enabled by default in Node.js 20. However, it does not properly implement parallel connection attempts, which are an integral part of the Happy Eyeballs algorithm. This can cause timeouts to occur that did not occur in versions prior to Node.js 20. See, for example, this discussion.What is the feature you are proposing to solve the problem?
Implement the algorithm as described in Section 5 of RFC 8305:
What alternatives have you considered?
Disabling
autoSelectFamilyand restoring the pre-20 behavior, see #47644 (comment).