Skip to content

benchmarking: add OOM kills to server telemetry - #2391

Open
Nishanth Kotla (Nishanth29) wants to merge 5 commits into
agent-substrate:mainfrom
Nishanth29:oom-events
Open

Nishanth Kotla (Nishanth29) wants to merge 5 commits into
agent-substrate:mainfrom
Nishanth29:oom-events

Conversation

@Nishanth29

@Nishanth29 Nishanth Kotla (Nishanth29) commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Depends on #2221 and #2386. Only the last commit, 24368cd8 ("benchmarking: add OOM kills to server telemetry"), is new; I'll rebase once those merge.

Part of #1590

Adds an oom_events block to server telemetry:

  • node, pod: OOM kills over the run, from cAdvisor's container_oom_events_total. node counts global OOMs. pod counts kills inside a worker pod (actor sandboxes and the ateom container) and is best effort: a kill whose cgroup is gone before the scrape is missed.
  • worker_restarts, worker_oomkilled_pods, evicted, worker_pods_lost: from the Kubernetes API, read before and after the run. evicted comes from Evicted events, as atecontroller deletes an evicted pod at once.

Other changes:

  • monitoring.yaml keeps container_oom_events_total.
  • The runner's benchmark-workloads Role can also list events.
  • The README explains what each field does and doesn't catch.

Live check (gVisor, 3m):

  • Actor limit 256Mi, 2 users: pod = 6, matching 6 kernel oom-kill lines, one per actor sandbox. Other fields 0; the worker stayed up.

  • No actor limit, 8 users: the kernel killed the whole worker twice at the pod cgroup limit (oom_memcg=/kubepods.slice), and the kubelet evicted it. evicted = 7 (1 eviction plus 6 replacement pods the kubelet refused to start), worker_pods_lost = 1, node = pod = 0, as documented.

  • 80 measurements in the jsonl row.

  • Tests pass

  • Appropriate changes to documentation are included in the PR

@Nishanth29

Copy link
Copy Markdown
Contributor Author

@Nishanth29

Copy link
Copy Markdown
Contributor Author

Da Huang (@git286) Tim Bai (@baizhenyu), adding you as reviewers.

What this uses:

  • Node-level OOMs and worker pod restarts/evictions: cAdvisor and the k8s API (outside Substrate).
  • Kills inside the worker pod (pod, mostly actor sandboxes): cAdvisor too, since there's no Substrate OOM metric to read yet.
  • Not caught: OOMs handled inside the sandbox (a microVM guest, or a single container hitting its own limit under gVisor).

If you'd rather add an actor OOM-kill counter in Substrate (e.g. from the oom_kill read in #2258), pod can switch to that or I can drop pod from this PR until then.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant