Skip to content

Worker: read capacity and free memory from the worker's own cgroup directory - #2845

Merged
amankrx merged 3 commits into
TraceMachina:mainfrom
amankrx:fix/cgroup-path-under-privileged
Sep 30, 2026
Merged

amankrx merged 3 commits into
TraceMachina:mainfrom
amankrx:fix/cgroup-path-under-privileged

Conversation

@amankrx

@amankrx amankrx commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

What and why

capacity.from_cgroup and the free-memory figure on every keepalive read cpu.max, memory.max and memory.current at /sys/fs/cgroup. That is the worker's own cgroup only inside a container with a private cgroup namespace. A privileged container, which is what use_namespaces needs on Ubuntu 24.04 nodes, shares the host's cgroup namespace: /sys/fs/cgroup is the host root with no memory.max at all, and the worker's limits sit under the path /proc/self/cgroup names. Every worker on the bed warned that the cgroup root could not be read, advertised the typed properties instead, and reported no free memory, so the live veto had nothing to go on. The worker now resolves its directory from the 0:: line of /proc/self/cgroup when that directory exists under the mount, and uses the root otherwise, which is what an unprivileged container sees. From there it walks up to the mount root and takes the nearest limit for CPU and for memory separately, since the process's own cgroup is often an unlimited child of the limited one (a systemd scope in a pod, a split cgroup, a privileged container under kubepods.slice); the free-memory figure reads memory.current at the same level as the limit it is measured against. When no level carries a limit the read counts as no limit, the configured properties stand, and the warning names the cgroup it looked at.

How was this verified?

capacity_test: the 0:: line is parsed and cgroup v1 lines are not; the own path is used when it exists under the mount, the root when it does not or when nothing is readable; the limit comes from the nearest limited ancestor and CPU and memory may sit at different levels; a tree that says max at every level, one where nothing is readable, and a limit above the mount root all count as no limit. On the bed a privileged worker now logs Advertising capacity from the cgroup with the pod's 14 cores and 52 GiB.

Risk

A host that previously could not read a limit and kept the operator's typed properties (bare-metal systemd, system.slice/nativelink.service reading max) still keeps them: the walk finding no limit is the same outcome as before. A worker that sits in an unlimited child of a limited parent, which previously advertised the host, now advertises the parent's limit. A container whose /proc/self/cgroup names a path that does not exist under the mount keeps the root read. Nothing changes for cgroup v1 or non-Linux, which never read a cgroup.

AI assistance

An agent drafted the change and I reviewed every line.


This change is Reviewable

@vercel

vercel Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
nativelink Ready Ready Preview Sep 30, 2026 5:45pm UTC
nativelink-aidm Ready Ready Preview Sep 30, 2026 5:45pm UTC

Request Review

Comment thread nativelink-worker/src/capacity.rs
Comment thread nativelink-worker/src/capacity.rs
Comment thread nativelink-worker/src/capacity.rs
…limit anywhere keeps the typed properties, and the failure names the cgroup
@amankrx
amankrx merged commit b4a5468 into TraceMachina:main Sep 30, 2026
45 checks passed

This branch was successfully deployed

2 active deployments
Preview – nativelink — 2467f5e8 Deployed Sep 30, 2026 by vercel[bot]
Preview – nativelink-aidm — 2467f5e8 Deployed Sep 30, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants