Please confirm the following
Bug Summary
Assisted-by: Claude Opus 5.5. Drafted with AI assistance; the loss measurements come from a running Ascender 25.4.0 deployment, and every code line cited was read at main 640cf5f2.
The generated omfwd action sets no TCP keepalives, so an idle connection to the log aggregator is discarded by an intermediate firewall or load balancer without either end being told, and the next write is the one that fails. Those messages are lost: omfwd over plain TCP is unacknowledged, and rsyslog reconnects without retransmitting. Scheduled workloads make this periodic rather than occasional: the event stream idles between runs, and every gap longer than the path's idle timeout costs the opening writes of the next burst.
The same gap exists upstream; the AWX report is ansible/awx#16711.
Ascender version
25.4.0. The code involved is unchanged on main 640cf5f2.
Select the relevant components
Installation method
kubernetes
Modifications
no
Ansible version
No response
Operating system
No response
Web browser
No response
Steps to reproduce
- Configure external logging with
LOG_AGGREGATOR_PROTOCOL=tcp to a collector reached through a load balancer or firewall that drops idle flows. An AWS Network Load Balancer does so after 350 seconds and doesn't notify the client.
- Let the event stream go quiet for longer than that idle timeout. Scheduled jobs do this between runs.
- Launch a job so events resume, and compare what Ascender recorded (
GET /api/v2/jobs/<id>/job_events/) against what the collector received.
- In the rsyslog log:
CheckConnection detected broken connection: Connection reset by peer, then omfwd: we had a generic or IO error with the remote server.
- Look for a way to keep the flow warm. There is none: no
LOG_AGGREGATOR_* setting reaches a keepalive parameter, and the generated action sets none.
Expected results
An idle connection is probed and kept alive, so it is either healthy or re-established before a message is written into it. rsyslog supports this on omfwd (KeepAlive, KeepAlive.Time, KeepAlive.Interval, KeepAlive.Probes).
Actual results
The first write after the flow is discarded fails, and those messages are gone. rsyslog reconnects, because the generated action sets action.resumeRetryCount="-1", but reconnecting isn't retransmitting: messages the old connection had already accepted aren't resent. A job whose whole output falls inside such a window produces no logs at all while reporting success.
Measured on one deployment: schedules cluster on quarter hours, leaving 9–15 minute quiet gaps; every gap exceeded the load balancer's 350-second timeout; both controller pods logged a reset a few seconds past every quarter hour, unbroken for as long as logs were kept, across pod restarts. That's about 96 loss windows a day. One 36-event run that fell entirely inside a window delivered nothing, while the runs either side of it delivered everything.
Raising the load balancer's idle timeout to 6000 seconds stopped the resets, which confirms the mechanism. But that's a fix in someone else's infrastructure, it isn't available on every path, and the OS default of two hours before the first probe is longer than most intermediaries' idle timeouts.
Additional information
ascender/main/utils/external_logging.py, construct_rsyslog_conf_template(), emits for any non-http* protocol:
'type="omfwd"',
f'target="{host}"',
f'port="{port}"',
f'protocol="{protocol}"',
'action.resumeRetryCount="-1"',
f'action.resumeInterval="{timeout}"',
'template="ascender"',
No KeepAlive* parameter is set, and no setting could supply one. rsyslog's omfwd accepts all four (tools/omfwd.c, action parameter table) and applies them to the stream only when the transport is TCP.
The attached PR registers LOG_AGGREGATOR_TCP_KEEPALIVE (default on) plus _TIME / _INTERVAL / _PROBES (120 s / 30 s / 3), and emits the four parameters on the omfwd action when the protocol is TCP. LOG_AGGREGATOR_TCP_KEEPALIVE=false reproduces the current action byte for byte; HTTPS and UDP render identically.
#965 lists what rsyslog guarantees here ("a pause, not a loss" for an aggregator that's down). An idle-flow reset breaks that promise without the aggregator ever being down, and this closes the common, periodic case. It doesn't make TCP delivery reliable: any other reset still loses what's in flight, because omfwd is unacknowledged. Acknowledged transport is requested separately (RELP, #1099).
Please confirm the following
Bug Summary
Assisted-by: Claude Opus 5.5. Drafted with AI assistance; the loss measurements come from a running Ascender 25.4.0 deployment, and every code line cited was read at
main640cf5f2.The generated
omfwdaction sets no TCP keepalives, so an idle connection to the log aggregator is discarded by an intermediate firewall or load balancer without either end being told, and the next write is the one that fails. Those messages are lost:omfwdover plain TCP is unacknowledged, and rsyslog reconnects without retransmitting. Scheduled workloads make this periodic rather than occasional: the event stream idles between runs, and every gap longer than the path's idle timeout costs the opening writes of the next burst.The same gap exists upstream; the AWX report is ansible/awx#16711.
Ascender version
25.4.0. The code involved is unchanged on
main640cf5f2.Select the relevant components
Installation method
kubernetes
Modifications
no
Ansible version
No response
Operating system
No response
Web browser
No response
Steps to reproduce
LOG_AGGREGATOR_PROTOCOL=tcpto a collector reached through a load balancer or firewall that drops idle flows. An AWS Network Load Balancer does so after 350 seconds and doesn't notify the client.GET /api/v2/jobs/<id>/job_events/) against what the collector received.CheckConnection detected broken connection: Connection reset by peer, thenomfwd: we had a generic or IO error with the remote server.LOG_AGGREGATOR_*setting reaches a keepalive parameter, and the generated action sets none.Expected results
An idle connection is probed and kept alive, so it is either healthy or re-established before a message is written into it. rsyslog supports this on
omfwd(KeepAlive,KeepAlive.Time,KeepAlive.Interval,KeepAlive.Probes).Actual results
The first write after the flow is discarded fails, and those messages are gone. rsyslog reconnects, because the generated action sets
action.resumeRetryCount="-1", but reconnecting isn't retransmitting: messages the old connection had already accepted aren't resent. A job whose whole output falls inside such a window produces no logs at all while reporting success.Measured on one deployment: schedules cluster on quarter hours, leaving 9–15 minute quiet gaps; every gap exceeded the load balancer's 350-second timeout; both controller pods logged a reset a few seconds past every quarter hour, unbroken for as long as logs were kept, across pod restarts. That's about 96 loss windows a day. One 36-event run that fell entirely inside a window delivered nothing, while the runs either side of it delivered everything.
Raising the load balancer's idle timeout to 6000 seconds stopped the resets, which confirms the mechanism. But that's a fix in someone else's infrastructure, it isn't available on every path, and the OS default of two hours before the first probe is longer than most intermediaries' idle timeouts.
Additional information
ascender/main/utils/external_logging.py,construct_rsyslog_conf_template(), emits for any non-http*protocol:No
KeepAlive*parameter is set, and no setting could supply one. rsyslog'somfwdaccepts all four (tools/omfwd.c, action parameter table) and applies them to the stream only when the transport is TCP.The attached PR registers
LOG_AGGREGATOR_TCP_KEEPALIVE(default on) plus_TIME/_INTERVAL/_PROBES(120 s / 30 s / 3), and emits the four parameters on theomfwdaction when the protocol is TCP.LOG_AGGREGATOR_TCP_KEEPALIVE=falsereproduces the current action byte for byte; HTTPS and UDP render identically.#965 lists what rsyslog guarantees here ("a pause, not a loss" for an aggregator that's down). An idle-flow reset breaks that promise without the aggregator ever being down, and this closes the common, periodic case. It doesn't make TCP delivery reliable: any other reset still loses what's in flight, because
omfwdis unacknowledged. Acknowledged transport is requested separately (RELP, #1099).