Skip to content

🐞 Bug: System memory usage is wrong when Arcane runs in Docker inside an LXC container #3741

Description

@andrebrait

Bug Description

When Arcane runs in Docker inside an unprivileged LXC container, the System Overview memory figures do not describe the LXC. applyCgroupLimits in backend/api/ws/system_stats.go is a deliberate no-op inside Docker:

// It is intentionally a no-op inside Docker: Docker's --cpus / --memory flags
// set artificial cgroup constraints that are unrelated to the host totals we
// want to display. gopsutil already reads the correct host values there (via
// the bind-mounted /proc).
if cgroup.IsDockerContainer() {
    return cpuCount, memUsed, memTotal
}

That reasoning is right for Docker on bare metal, where a container's /proc/meminfo really is the machine's. It does not hold when the Docker host is itself an LXC container, which is how the community-scripts Docker LXC and most Proxmox homelabs deploy Arcane.

This is distinct from #913 / #3161 / #2416, which covered Arcane running directly in an LXC and were fixed by the non-Docker branch of this same function. The Docker-inside-LXC path still falls through to gopsutil.

There are two failure modes depending on whether lxcfs files are bound into the container:

  1. Default (no lxcfs binds) — gopsutil reads the physical hypervisor's /proc/meminfo. On a 32 GiB LXC on a 62 GiB host, Arcane reports 62 GiB total and host-wide usage.
  2. With lxcfs binds — lxcfs answers /proc/meminfo against the cgroup of the reading process, so Arcane sees only its own footprint and "system memory used" collapses to near zero. This is becoming more common: it is the standard fix for nested Docker resource visibility, and I have a PR adding it to the community-scripts engine (setup_docker: make LXC resource limits visible to nested containers community-scripts/core#9), which would make it the default for every Docker LXC created by those scripts.

Note the Docker info panel is correct in both cases — dockerClient.Info() is answered by the daemon, which is a plain LXC process and therefore reads the LXC's own values. Only the live system-stats websocket is wrong.

CPU is unaffected: /proc/stat is not cgroup-scoped, so CPU percentages remain right.

Steps To Reproduce

  1. Create an unprivileged Proxmox LXC with features: nesting=1 and a memory limit below the host's (e.g. 32768 MB on a 62 GiB host).
  2. Install Docker inside it and deploy Arcane with the recommended compose.
  3. Open Dashboard → System Overview and compare the memory figures with free -m run inside the LXC.

Expected Behavior

The LXC's total and the LXC's whole memory usage — what free reports inside the LXC, and what Arcane would show if the same LXC were a VM. Usage should include everything in the container: the Docker daemon, sibling containers, and non-Docker processes.

Actual Behavior

Either the hypervisor's totals and usage (case 1), or Arcane's own footprint presented as system usage (case 2). Measured on a 32 GiB LXC with ~109 MiB actually in use, case 2 reports:

LXC truth:        MemFree: 33254596 kB   Cached: 164344 kB
Arcane's view:    MemFree: 33553052 kB   Cached:      0 kB

Proposed Solution

Total is already available and correct — keep using info.MemTotal from dockerClient.Info().

Usage can come from the LXC's own cgroup, and no new bind mount is required, because docker/examples/compose.basic.yaml already ships cgroup: host. With cgroupns=host, /sys/fs/cgroup inside the container is the LXC's cgroup root:

$ docker run --rm --cgroupns=host alpine sh -c 'cat /proc/self/cgroup; cat /sys/fs/cgroup/memory.current'
0::/system.slice/docker-febf150a280b154dd408f619a3fe804e41c92a68bde3ead1f8cb6af8f824c7bc.scope
307421184          # LXC-wide  (its own cgroup reads 1536000)

LXC truth: /sys/fs/cgroup/memory.current = 284708864

The derivation is the part worth getting right. memory.current includes page cache and overstates badly against /proc/meminfo semantics. Same 32 GiB LXC, all four numbers taken together:

/proc/meminfo (the target):
  MemTotal 33554432 kB   MemAvailable 33442505 kB   ->  free-style used = 109 MiB

cgroup:
  memory.current = 273 MiB
  memory.stat: anon 104783872  file 169504768  inactive_file 22294528  active_file 147030016
Formula Result vs target (109 MiB)
memory.current 273 MiB 2.5× too high
memory.current - inactive_file (Docker/cAdvisor convention) 252 MiB 2.3× too high
memory.current - file 111 MiB ✅ within 2 MiB

So memory.current - file (from memory.stat) is the one that matches what free shows inside the LXC.

One trap: do not try to read the total from that cgroup. /sys/fs/cgroup/memory.max at the LXC root reads max — the real limit lives host-side at /sys/fs/cgroup/lxc/<vmid>/memory.max and is invisible from inside the guest. The Docker API value is the correct source.

For detecting "Docker inside LXC" versus "Docker on bare metal", one signal you already have: compare info.MemTotal from the Docker API against the machine total gopsutil reports. On bare metal they agree; in an LXC the daemon reports the LXC's smaller limit while gopsutil reports the hypervisor's. I have not implemented or verified that heuristic, so treat it as a suggestion rather than a recommendation.

Alternatives Considered

As a stopgap I publish the LXC's own /proc/meminfo to a plain file from a small systemd unit inside the LXC (the read is taken by an LXC-level process, so lxcfs scopes it correctly) and bind it over Arcane's /proc/meminfo. That works today and gives exactly the target numbers:

LXC truth:     MemTotal: 2097152 kB   Cached: 81656 kB
Arcane reads:  MemTotal: 2097152 kB   Cached: 81656 kB

But it needs a refresher daemon, a compose bind, and carries up to one refresh interval of staleness. Reading the cgroup directly needs none of that, which is why I would rather see it in Arcane.

Arcane Version

v2.9.0

Installation Method

Docker Compose (Recommended)

Environment Type

Local Docker environment

Database Type

SQLite (default)

Operating System

Debian 13 (unprivileged LXC on Proxmox VE 9.2.10, lxcfs 7.0.0-pve1, Docker 29.7.2)

Additional Context

All figures above were measured on live hardware: unprivileged LXC, features: nesting=1, 32768 MB / 6 cores on a 62 GiB host, plus a second 2048 MB LXC used for the stopgap verification.

Activity

  1. andrebrait commented on Aug 25, 2026

    @andrebrait
    Author

    Reproduced on a completely stock deployment rather than my own setup, in case that is useful for triage.

    Built with the community-scripts Docker LXC (ct/docker.sh, default install, 2 GiB / 2 cores, unprivileged, features: nesting=1,keyctl=1) and then Arcane via their tools/addon/arcane.sh, which fetches your docker/examples/compose.basic.yaml and runs docker compose up -d. Nothing modified.

    LXC truth (2 GiB LXC on a 62 GiB host):
      MemTotal: 2097152 kB   MemFree: 1634840 kB   Cached: 235836 kB   nproc=2
    
    Arcane container sees:
      MemTotal: 65648168 kB  MemFree: 46440020 kB  Cached: 6111572 kB
    
    Docker API (daemon):
      MemTotal=2147483648   NCPU=2
    

    The container sees the entire hypervisor — 31x the LXC's real memory — and host-wide usage that includes every other guest on the machine. The Docker info panel in the same UI reports 2 GiB correctly, so a single dashboard shows two contradictory totals.

    Two things this confirms about the proposal in the issue body:

    1. cgroupns=host is already in effect on a stock install — docker inspect -f '{{.HostConfig.CgroupnsMode}}' arcane returns host, courtesy of the cgroup: host line in your own compose.basic.yaml. No new bind mount is required to reach the LXC's cgroup.

    2. The memory.current - file formula holds here too. Read from inside the stock container:

    memory.current = 484130816
    memory.max     = max            <- confirms the total cannot come from here
    file           = 242016256
    
    target (MemTotal - MemAvailable) = 226 MiB
    memory.current                   = 461 MiB   (2x too high)
    memory.current - file            = 231 MiB   (within 5 MiB)
    

    Worth noting for anyone who finds this issue: installing Arcane natively in the LXC instead of in Docker sidesteps the problem entirely, because an ordinary LXC process reads the correct values straight from /proc/meminfo — that is the !cgroup.IsDockerContainer() path added for #913. The gap is specifically Docker-inside-LXC.

  2. kmendell commented on Aug 25, 2026

    @kmendell
    Member

    Ive spent alot of time trying to figure out lxc cgroups, i thougth i fixed it, but isnt it considered bad practice to run docker inside of an LXC? Ill try to to look again, but im not sure i want to spend abother 8 hours trying to figure this out when its generally frowned upon to run it that way.

  3. added theissue type on Aug 26, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions