Repository navigation
Worker: read capacity and free memory from the worker's own cgroup directory - #2845
Merged
amankrx merged 3 commits intoSep 30, 2026
Merged
Conversation
…rectory, not the mount root
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
corcillo
reviewed
Sep 30, 2026
corcillo
reviewed
Sep 30, 2026
corcillo
reviewed
Sep 30, 2026
amankrx
force-pushed
the
fix/cgroup-path-under-privileged
branch
from
September 30, 2026 17:07
fa2b69a to
9cf562c
Compare
…limit anywhere keeps the typed properties, and the failure names the cgroup
amankrx
force-pushed
the
fix/cgroup-path-under-privileged
branch
from
September 30, 2026 17:21
9cf562c to
feacd5c
Compare
corcillo
approved these changes
Sep 30, 2026
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What and why
capacity.from_cgroupand the free-memory figure on every keepalive readcpu.max,memory.maxandmemory.currentat/sys/fs/cgroup. That is the worker's own cgroup only inside a container with a private cgroup namespace. A privileged container, which is whatuse_namespacesneeds on Ubuntu 24.04 nodes, shares the host's cgroup namespace:/sys/fs/cgroupis the host root with nomemory.maxat all, and the worker's limits sit under the path/proc/self/cgroupnames. Every worker on the bed warned that the cgroup root could not be read, advertised the typed properties instead, and reported no free memory, so the live veto had nothing to go on. The worker now resolves its directory from the0::line of/proc/self/cgroupwhen that directory exists under the mount, and uses the root otherwise, which is what an unprivileged container sees. From there it walks up to the mount root and takes the nearest limit for CPU and for memory separately, since the process's own cgroup is often an unlimited child of the limited one (a systemd scope in a pod, a split cgroup, a privileged container underkubepods.slice); the free-memory figure readsmemory.currentat the same level as the limit it is measured against. When no level carries a limit the read counts as no limit, the configured properties stand, and the warning names the cgroup it looked at.How was this verified?
capacity_test: the0::line is parsed and cgroup v1 lines are not; the own path is used when it exists under the mount, the root when it does not or when nothing is readable; the limit comes from the nearest limited ancestor and CPU and memory may sit at different levels; a tree that saysmaxat every level, one where nothing is readable, and a limit above the mount root all count as no limit. On the bed a privileged worker now logsAdvertising capacity from the cgroupwith the pod's 14 cores and 52 GiB.Risk
A host that previously could not read a limit and kept the operator's typed properties (bare-metal systemd,
system.slice/nativelink.servicereadingmax) still keeps them: the walk finding no limit is the same outcome as before. A worker that sits in an unlimited child of a limited parent, which previously advertised the host, now advertises the parent's limit. A container whose/proc/self/cgroupnames a path that does not exist under the mount keeps the root read. Nothing changes for cgroup v1 or non-Linux, which never read a cgroup.AI assistance
An agent drafted the change and I reviewed every line.
This change is