Skip to content

ateom: own ate.actor.stats.cpu.time so short activations are counted #2362

Description

@baizhenyu

Context

ate.actor.stats.cpu.time is atelet's counter of actor CPU. atelet adds each activation's increase between the samples it reads from GetActiveWorkloadStats, which lists only actors that are still hosted.

Issue

The counter loses the end of each activation: the CPU from the last sample atelet read to the teardown, up to one atelet poll interval plus one ateom sample interval. An activation shorter than that, the common checkpoint-when-idle case, can lose most or all of its CPU in the metric. A resumed micro-VM's first reading is its baseline, so such an activation can show none. The ate.actor.usage_sampled records keep it, through the final record.

This is a well-known limit of scrape-based collection, and Kubernetes has it too: the kubelet reports container_cpu_usage_seconds_total only while a container exists, so the CPU between the last scrape and the exit is lost, and a container shorter than the scrape interval can be missed entirely. Here the gap is wider, because two intervals stack: atelet reads a sample that is already up to one ateom sample interval old.

Proposal

The ateoms own the counter. Each ateom adds an activation's CPU increase at every reading it takes, initial, periodic, and final, and exports the counter through the existing OTLP relay. No gRPC change is needed, and the final reading's increase is counted whenever the next export runs. atelet keeps the memory gauges and drops the CPU baseline map, its start time, and the silent-worker rule, which also removes the baselines a crashed worker's leftover directory keeps.

Cost

The counter's series carry the ateom's per-pod resource attributes, so there is one series per worker pod per label set instead of one per node, still bounded by pods rather than actors, plus a short-lived series per replaced pod. A collector can drop the pod attributes if that matters. The registry's emitted_by moves from atelet to the ateoms.

Raised in review of #1984. Part of #1748.

Activity

  1. lubingtan commented on Oct 10, 2026

    @lubingtan
    Contributor

    I’m working on this. The change moves ate.actor.stats.cpu.time from atelet polling to the gVisor and microVM ateoms, which count CPU increases at initial, periodic, and final readings and export through the existing OTLP relay. Atelet continues to collect the memory and sampled-actor gauges.

    The changed-package race tests and metrics registry check pass. In a Kind microVM test, an activation used 0.003006 seconds of CPU entirely between two old atelet polls. Both polls returned zero active actors, while the ateom counter increased by exactly 0.003006 seconds. I could not complete the local gVisor end-to-end test because runsc failed to start the golden actor (StartRoot EOF).

    I'll do one last review myself and then open the PR. Let me know if you'd prefer a different approach.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions