Repository navigation
benchmarking/locust: harvest server-side telemetry from Prometheus - #1725
Max Smythe (maxsmythe) merged 2 commits into
Conversation
bb6d09b to
659e8f2
Compare
659e8f2 to
b8197ae
Compare
c968a51 to
45d6c08
Compare
45d6c08 to
a69a8eb
Compare
a69a8eb to
8f3967a
Compare
Haowei Cai (Roy) (roycaihw)
left a comment
There was a problem hiding this comment.
LGTM
8f3967a to
d0a8d59
Compare
|
Issue: on microVM, snapshot size and checkpoint count came back null. The queries only matched pages.img (gVisor's memory image), and microVM writes memory-ranges. Fix: match pages.img|memory-ranges. No change for gVisor. Tested on a microVM cluster: checkpoints_in_window now shows 729 (was null). Remaining atelet issue: microVM sizes will still read 2 GiB until #931 lands. cloud-hypervisor writes memory-ranges as a sparse file sized to the full guest RAM (2 GiB) but only fills the pages the guest used. atelet records the file's length (info.Size()), not the bytes actually written, so every sample is 2147483648. atelet's own upload log shows ~294 MiB populated for the same files. #931 fixes the recording; once it lands, sizes and checkpoint_mb_s here will be correct. |
Part of #1590
What this PR does
Locust measures the client side only. This adds a post-run harvest of server-side ground truth from Prometheus, written to a new
server_summary.jsonand summarized as one row instats.jsonl.Proposed Changes
Steady-state scoping
server_telemetry.pyderives the steady-state window fromstats_history.csv: it starts at the first sample at 90% of peak user count and ends at the last one. Ramp-up is excluded from every metric, and the snapshot block also stops at the last full-load sample, so teardown suspends are left out. Under a step-ladder load shape the window covers only the top step.Harvested metrics
Cluster packing, from
ate_workerpool_workers: busy workers (partial+at_capacity) over the pool size at each sample, where the pool size is the sum of all worker states, so it follows scale-ups and scale-downs mid-run. Only the live ateapi is read: after a redeploy the collector keeps re-exporting exited ateapi processes' last values for a few minutes, which would otherwise inflate the counts. Reported as min, p50, p90, p95, p99, max and mean (avg), plus the underlying timeseries so transient spikes remain visible.Kernel pressure, from cAdvisor PSI: CPU, memory and IO stall percentages for the node and for the worker pods, each as the same percentile set over the window.
Snapshots: size mean and p50/p90/p95/p99, checkpoint counts both in-window and cumulative, restore and checkpoint latency mean and p50/p90/p95/p99 from the AteomHerder RPC histograms, and checkpoint throughput as
checkpoint_mb_s. The atelet exports metrics on an interval, so its data reaches Prometheus late: the harvest waits--atelet-lag-s(default 70s, enough for the OTel SDK's 60s default export and a 10s scrape) and reads both window edges half that late.Constraints and failure behavior
The module uses only the standard library, since the locust image is distroless, and every request carries a timeout. An unreachable Prometheus records nulls and does not fail the run.
--prometheus-urloverrides the in-cluster default, and--atelet-lag-ssets the wait for the atelet's last export.Unmeasured fields are
nulland a measured zero is0, consistent with the rest of the runner. Prometheus exposes no byte counter on the restore path, so no restore throughput field is emitted rather than deriving one indirectly.Output
server_summary.jsonholds the full nested artifact.stats.jsonlreceives a singleserver_summaryrow with 45 flat keys for graphing: packing, node PSI and pod PSI at p50/p90/p95/p99, plus the snapshot means, percentiles,checkpoints_in_windowandcheckpoint_mb_s.status.jsonis unchanged.How this was tested
14 unit tests in
test_server_telemetry.pycovering steady-state detection, percentile boundaries, the range-query window guard, malformed Prometheus responses, the packing and checkpoint arithmetic, the per-sample pool size, ignoring exited ateapi series, a missing denominator returning null, snapshot fields returning null rather than zero, and telemetry surviving a missing stats CSV.Verified on 2 user / 2 worker and 4 user / 2 worker sympy runs. All 45 keys matched an independent recomputation from raw Prometheus, and the snapshot block was cross-checked against atelet logs. A run with
--prometheus-urlpointed at an unreachable address completes normally with the affected fields null.Re-harvested 5 past runs (Glutton and SWE-perf, 3 × 60 and 30 × 8 workers) from Prometheus with the new code. Runs with a steady pool match the previous output exactly, and runs that followed a redeploy now read the real pool size on every sample.
References
Agent Substrate: Actor Density Benchmark Specs