Context
ate.actor.stats.cpu.time is atelet's counter of actor CPU. atelet adds each activation's increase between the samples it reads from GetActiveWorkloadStats, which lists only actors that are still hosted.
Issue
The counter loses the end of each activation: the CPU from the last sample atelet read to the teardown, up to one atelet poll interval plus one ateom sample interval. An activation shorter than that, the common checkpoint-when-idle case, can lose most or all of its CPU in the metric. A resumed micro-VM's first reading is its baseline, so such an activation can show none. The ate.actor.usage_sampled records keep it, through the final record.
This is a well-known limit of scrape-based collection, and Kubernetes has it too: the kubelet reports container_cpu_usage_seconds_total only while a container exists, so the CPU between the last scrape and the exit is lost, and a container shorter than the scrape interval can be missed entirely. Here the gap is wider, because two intervals stack: atelet reads a sample that is already up to one ateom sample interval old.
Proposal
The ateoms own the counter. Each ateom adds an activation's CPU increase at every reading it takes, initial, periodic, and final, and exports the counter through the existing OTLP relay. No gRPC change is needed, and the final reading's increase is counted whenever the next export runs. atelet keeps the memory gauges and drops the CPU baseline map, its start time, and the silent-worker rule, which also removes the baselines a crashed worker's leftover directory keeps.
Cost
The counter's series carry the ateom's per-pod resource attributes, so there is one series per worker pod per label set instead of one per node, still bounded by pods rather than actors, plus a short-lived series per replaced pod. A collector can drop the pod attributes if that matters. The registry's emitted_by moves from atelet to the ateoms.
Raised in review of #1984. Part of #1748.
Context
ate.actor.stats.cpu.timeis atelet's counter of actor CPU. atelet adds each activation's increase between the samples it reads fromGetActiveWorkloadStats, which lists only actors that are still hosted.Issue
The counter loses the end of each activation: the CPU from the last sample atelet read to the teardown, up to one atelet poll interval plus one ateom sample interval. An activation shorter than that, the common checkpoint-when-idle case, can lose most or all of its CPU in the metric. A resumed micro-VM's first reading is its baseline, so such an activation can show none. The
ate.actor.usage_sampledrecords keep it, through the final record.This is a well-known limit of scrape-based collection, and Kubernetes has it too: the kubelet reports
container_cpu_usage_seconds_totalonly while a container exists, so the CPU between the last scrape and the exit is lost, and a container shorter than the scrape interval can be missed entirely. Here the gap is wider, because two intervals stack: atelet reads a sample that is already up to one ateom sample interval old.Proposal
The ateoms own the counter. Each ateom adds an activation's CPU increase at every reading it takes, initial, periodic, and final, and exports the counter through the existing OTLP relay. No gRPC change is needed, and the final reading's increase is counted whenever the next export runs. atelet keeps the memory gauges and drops the CPU baseline map, its start time, and the silent-worker rule, which also removes the baselines a crashed worker's leftover directory keeps.
Cost
The counter's series carry the ateom's per-pod resource attributes, so there is one series per worker pod per label set instead of one per node, still bounded by pods rather than actors, plus a short-lived series per replaced pod. A collector can drop the pod attributes if that matters. The registry's
emitted_bymoves from atelet to the ateoms.Raised in review of #1984. Part of #1748.