Skip to content

External logging over TCP loses messages after idle periods because omfwd sets no keepalives #1093

Description

@quanah

Please confirm the following

  • I have checked the current issues for duplicates.
  • I understand that Ascender is open source software provided for free and that I might not receive a timely response.

Bug Summary

Assisted-by: Claude Opus 5.5. Drafted with AI assistance; the loss measurements come from a running Ascender 25.4.0 deployment, and every code line cited was read at main 640cf5f2.

The generated omfwd action sets no TCP keepalives, so an idle connection to the log aggregator is discarded by an intermediate firewall or load balancer without either end being told, and the next write is the one that fails. Those messages are lost: omfwd over plain TCP is unacknowledged, and rsyslog reconnects without retransmitting. Scheduled workloads make this periodic rather than occasional: the event stream idles between runs, and every gap longer than the path's idle timeout costs the opening writes of the next burst.

The same gap exists upstream; the AWX report is ansible/awx#16711.

Ascender version

25.4.0. The code involved is unchanged on main 640cf5f2.

Select the relevant components

  • UI
  • UI (tech preview)
  • API
  • Docs
  • Collection
  • CLI
  • Other

Installation method

kubernetes

Modifications

no

Ansible version

No response

Operating system

No response

Web browser

No response

Steps to reproduce

  1. Configure external logging with LOG_AGGREGATOR_PROTOCOL=tcp to a collector reached through a load balancer or firewall that drops idle flows. An AWS Network Load Balancer does so after 350 seconds and doesn't notify the client.
  2. Let the event stream go quiet for longer than that idle timeout. Scheduled jobs do this between runs.
  3. Launch a job so events resume, and compare what Ascender recorded (GET /api/v2/jobs/<id>/job_events/) against what the collector received.
  4. In the rsyslog log: CheckConnection detected broken connection: Connection reset by peer, then omfwd: we had a generic or IO error with the remote server.
  5. Look for a way to keep the flow warm. There is none: no LOG_AGGREGATOR_* setting reaches a keepalive parameter, and the generated action sets none.

Expected results

An idle connection is probed and kept alive, so it is either healthy or re-established before a message is written into it. rsyslog supports this on omfwd (KeepAlive, KeepAlive.Time, KeepAlive.Interval, KeepAlive.Probes).

Actual results

The first write after the flow is discarded fails, and those messages are gone. rsyslog reconnects, because the generated action sets action.resumeRetryCount="-1", but reconnecting isn't retransmitting: messages the old connection had already accepted aren't resent. A job whose whole output falls inside such a window produces no logs at all while reporting success.

Measured on one deployment: schedules cluster on quarter hours, leaving 9–15 minute quiet gaps; every gap exceeded the load balancer's 350-second timeout; both controller pods logged a reset a few seconds past every quarter hour, unbroken for as long as logs were kept, across pod restarts. That's about 96 loss windows a day. One 36-event run that fell entirely inside a window delivered nothing, while the runs either side of it delivered everything.

Raising the load balancer's idle timeout to 6000 seconds stopped the resets, which confirms the mechanism. But that's a fix in someone else's infrastructure, it isn't available on every path, and the OS default of two hours before the first probe is longer than most intermediaries' idle timeouts.

Additional information

ascender/main/utils/external_logging.py, construct_rsyslog_conf_template(), emits for any non-http* protocol:

'type="omfwd"',
f'target="{host}"',
f'port="{port}"',
f'protocol="{protocol}"',
'action.resumeRetryCount="-1"',
f'action.resumeInterval="{timeout}"',
'template="ascender"',

No KeepAlive* parameter is set, and no setting could supply one. rsyslog's omfwd accepts all four (tools/omfwd.c, action parameter table) and applies them to the stream only when the transport is TCP.

The attached PR registers LOG_AGGREGATOR_TCP_KEEPALIVE (default on) plus _TIME / _INTERVAL / _PROBES (120 s / 30 s / 3), and emits the four parameters on the omfwd action when the protocol is TCP. LOG_AGGREGATOR_TCP_KEEPALIVE=false reproduces the current action byte for byte; HTTPS and UDP render identically.

#965 lists what rsyslog guarantees here ("a pause, not a loss" for an aggregator that's down). An idle-flow reset breaks that promise without the aggregator ever being down, and this closes the common, periodic case. It doesn't make TCP delivery reliable: any other reset still loses what's in flight, because omfwd is unacknowledged. Acknowledged transport is requested separately (RELP, #1099).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions