Repository navigation
'init' action doesn't seem to retry on socket exceptions and network hiccups #3367
Description
Activity
Hi @hossein-nexl 👋🏻
Thanks for reporting this. Yes, I don't believe we generally retry network requests if they fail. We have discussed this internally before, but it would make sense to retry requests in (at least) cases where the error is likely intermittent and doesn't suggest a problem that won't go away. We'll have a look at what we can do to improve this.
Reacted by hanexlAdding another reproduction of this class, on v4.37.0 (
99df26d) — same unhandled'error'event on the download stream, but failing at the TLS handshake rather than mid-transfer, so it appears any transient socket/TLS error during the bundle download hard-crashes the Node process with no retry.Environment: GitHub-hosted larger runner (ubuntu-24.04, image 20260720.247.2), Azure westus.
initdownloadingcodeql-bundle-v2.26.2/codeql-bundle-linux64.tar.zst. Elapsed from "Downloading CodeQL tools" to crash: 56 ms, zero bytes transferred.Downloading CodeQL tools from https://github.com/github/codeql-action/releases/download/codeql-bundle-v2.26.2/codeql-bundle-linux64.tar.zst . This may take a while. Streaming the extraction of the CodeQL bundle. node:events:487 throw er; // Unhandled 'error' event ^ Error: self-signed certificate; if the root CA is installed locally, try running Node.js with --use-system-ca at TLSSocket.onConnectSecure (node:internal/tls/wrap:1748:34) at TLSSocket.emit (node:events:509:28) at TLSSocket._finishInit (node:internal/tls/wrap:1185:8) at ssl.onhandshakedone (node:internal/tls/wrap:966:12) Emitted 'error' event on Writable instance at: at eventHandlers.<computed> (/home/runner/work/_actions/github/codeql-action/99df26d4f13ea111d4ec1a7dddef6063f76b97e9/lib/entry-points.js:82496:28) at ClientRequest.emit (node:events:509:28) at emitErrorEvent (node:_http_client:109:11) at TLSSocket.socketErrorListener (node:_http_client:593:5) at TLSSocket.emit (node:events:509:28) at emitErrorNT (node:internal/streams/destroy:170:8) at emitErrorCloseNT (node:internal/streams/destroy:129:3) at process.processTicksAndRejections (node:internal/process/task_queues:90:21) { code: 'DEPTH_ZERO_SELF_SIGNED_CERT' } Node.js v24.18.0In our case the transient certificate error itself turned out to be infrastructure-side (confirmed and corrected by GitHub Support), but the action-side behavior is the concern: over 15 days / 463 runs we hit this exactly once, and a sibling job in the same workflow run downloaded the same URL successfully — a single retry would have made the failure invisible. Because the crash happens in
init, the analyze step is skipped and the commit silently receives no security analysis.Two asks:
- Handle the download stream's
'error'event so this fails cleanly instead of as an uncaught exception. - Retry the bundle download on transient network/TLS errors. Given the observed frequency (1 in 463 runs), a single retry with short backoff would very likely cover it.
- Handle the download stream's
Using
github/codeql-action/init@v4in our CodeQL Advanced workflow like this:I randomly received the following error today (December 16), which by the look of it was using tag v4.31.8 released on December 12:
A manual retry fixed the issue, but it would be great if the task itself could retry on retriable failures and network hiccups like
ECONNRESET👆