Skip to content

Expose external logging queue depth and discard counts #1097

Description

@quanah

Please confirm the following

  • I have checked the current issues for duplicates.
  • I understand that Ascender is open source software provided for free and that I might not receive a timely response.

Feature type

Enhancement to Existing Feature

Feature Summary

Feature Summary

Assisted-by: Claude Opus 5.5. Drafted with AI assistance; the counter names were read in rsyslog's source and the binary, and the Ascender code cited was read at main 640cf5f2.

Surface the state of the external logging queue (its depth, and how many messages it has discarded) so an operator can tell whether logs are being delivered. rsyslog already keeps these counters; Ascender neither enables nor exposes them, and the generated configuration leaves no way for an operator to add them. So "the collector has no events for that job" is currently indistinguishable from "the job didn't run". The same request is open upstream as ansible/awx#16717, but Ascender's situation is narrower, because the workaround AWX has doesn't exist here (below).

Select the relevant components: API, Docs, Other

Steps to reproduce

  1. Configure external logging to any destination and run a normal workload.
  2. Ask the two questions an operator has after a missing-logs report: how full is the action queue, and how many messages has it discarded since the pod started.
  3. Look for either answer in /api/v2/metrics, the settings API, the activity stream, or the Ascender log. Neither is anywhere.
  4. Try to add an impstats action yourself under /var/lib/ascender/rsyslog/conf.d/, which the image creates. Today it's never read, because the generated rsyslog.conf dropped its include (filed separately, The generated rsyslog.conf no longer reads /var/lib/ascender/rsyslog/conf.d #1091).

Current results

Nothing reports queue depth and nothing reports discards. A deployment can lose a measurable share of its job events indefinitely with no sign in any Ascender surface: the API still shows every event, because the events reached the database; it's the external copy that's missing.

The loss isn't hypothetical: with queue.discardSeverity="5" hardcoded, every job event is eligible for discard once the queue is 90% full (#1095).

Sugested feature result

Queue depth and cumulative discards are readable from Ascender, at whatever granularity fits the existing metrics. A monitor can then alert on discards rather than on missing logs, which is a signal that arrives too late and is ambiguous when it does.

Additional information

rsyslog already counts this. Every queue registers two resettable counters (runtime/queue.c):

Counter Incremented when
discarded.nf the queue is nearly full and the message is eligible under queue.discardSeverity
discarded.full the queue is full, nothing could be discarded, and the enqueue timeout expired

Both, plus current and peak size, are published by the impstats module, which ships in the base rsyslog package the image already installs (impstats.so; a configuration loading it passes rsyslogd -N1 on rockylinux/rockylinux:9 with rsyslog-8.2510.0). So this asks Ascender to route information rsyslog already produces, not to instrument anything new.

The only workaround is a drop-in impstats action under conf.d/, and today even that doesn't work: the generated rsyslog.conf lost its include in #73, which is filed separately (#1091 ) with a one-line fix. Once that's fixed, an operator can add impstats themselves, as AWX operators can. But nothing documents it, so it stays invisible to anyone who hasn't read the generator, and the numbers still go nowhere an alert can see them.

#965 lists what rsyslog guarantees for external logs, in preparation for the roadmap item "Move external log shipping off the rsyslog process and onto a Python logging handler". That list covers what happens to messages, but not whether anyone can tell it happened. Whatever ships the logs, rsyslog or a handler, a discard that leaves no trace is the part that turns a delivery problem into a monitoring problem.

Shapes this could take, cheapest first. A PR for option 2 follows this issue; the others remain open:

  1. With the include restored (The generated rsyslog.conf no longer reads /var/lib/ascender/rsyslog/conf.d #1091), document conf.d as a supported extension point, with an impstats example, and leave the metrics to the operator.
  2. A setting that enables an impstats action routed into the Ascender log, so the numbers appear where other subsystem messages already do.
  3. Queue depth and discard counters in /api/v2/metrics alongside the other Prometheus series, which makes them alertable without relying on a log pipeline that may itself be what's failing.

Related, not duplicates: #965 (the guarantees list), #987 (spool fallback when the configured path isn't writable).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions