Skip to content

benchmarking/locust: harvest server-side telemetry from Prometheus - #1725

Merged
Max Smythe (maxsmythe) merged 2 commits into
agent-substrate:mainfrom
Nishanth29:benchmarking/server-telemetry
Oct 1, 2026
Merged

Max Smythe (maxsmythe) merged 2 commits into
agent-substrate:mainfrom
Nishanth29:benchmarking/server-telemetry

Conversation

@Nishanth29

@Nishanth29 Nishanth Kotla (Nishanth29) commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Part of #1590

What this PR does

Locust measures the client side only. This adds a post-run harvest of server-side ground truth from Prometheus, written to a new server_summary.json and summarized as one row in stats.jsonl.

Proposed Changes

Steady-state scoping

server_telemetry.py derives the steady-state window from stats_history.csv: it starts at the first sample at 90% of peak user count and ends at the last one. Ramp-up is excluded from every metric, and the snapshot block also stops at the last full-load sample, so teardown suspends are left out. Under a step-ladder load shape the window covers only the top step.

Harvested metrics

Cluster packing, from ate_workerpool_workers: busy workers (partial + at_capacity) over the pool size at each sample, where the pool size is the sum of all worker states, so it follows scale-ups and scale-downs mid-run. Only the live ateapi is read: after a redeploy the collector keeps re-exporting exited ateapi processes' last values for a few minutes, which would otherwise inflate the counts. Reported as min, p50, p90, p95, p99, max and mean (avg), plus the underlying timeseries so transient spikes remain visible.

Kernel pressure, from cAdvisor PSI: CPU, memory and IO stall percentages for the node and for the worker pods, each as the same percentile set over the window.

Snapshots: size mean and p50/p90/p95/p99, checkpoint counts both in-window and cumulative, restore and checkpoint latency mean and p50/p90/p95/p99 from the AteomHerder RPC histograms, and checkpoint throughput as checkpoint_mb_s. The atelet exports metrics on an interval, so its data reaches Prometheus late: the harvest waits --atelet-lag-s (default 70s, enough for the OTel SDK's 60s default export and a 10s scrape) and reads both window edges half that late.

Constraints and failure behavior

The module uses only the standard library, since the locust image is distroless, and every request carries a timeout. An unreachable Prometheus records nulls and does not fail the run. --prometheus-url overrides the in-cluster default, and --atelet-lag-s sets the wait for the atelet's last export.

Unmeasured fields are null and a measured zero is 0, consistent with the rest of the runner. Prometheus exposes no byte counter on the restore path, so no restore throughput field is emitted rather than deriving one indirectly.

Output

server_summary.json holds the full nested artifact. stats.jsonl receives a single server_summary row with 45 flat keys for graphing: packing, node PSI and pod PSI at p50/p90/p95/p99, plus the snapshot means, percentiles, checkpoints_in_window and checkpoint_mb_s.

status.json is unchanged.

How this was tested

14 unit tests in test_server_telemetry.py covering steady-state detection, percentile boundaries, the range-query window guard, malformed Prometheus responses, the packing and checkpoint arithmetic, the per-sample pool size, ignoring exited ateapi series, a missing denominator returning null, snapshot fields returning null rather than zero, and telemetry surviving a missing stats CSV.

Verified on 2 user / 2 worker and 4 user / 2 worker sympy runs. All 45 keys matched an independent recomputation from raw Prometheus, and the snapshot block was cross-checked against atelet logs. A run with --prometheus-url pointed at an unreachable address completes normally with the affected fields null.

Re-harvested 5 past runs (Glutton and SWE-perf, 3 × 60 and 30 × 8 workers) from Prometheus with the new code. Runs with a steady pool match the previous output exactly, and runs that followed a redeploy now read the real pool size on every sample.

References

Agent Substrate: Actor Density Benchmark Specs

  • Tests pass
  • Appropriate changes to documentation are included in the PR

Comment thread benchmarking/locust/server_telemetry.py
Comment thread benchmarking/locust/server_telemetry.py Outdated
Comment thread benchmarking/locust/server_telemetry.py Outdated
Comment thread benchmarking/locust/server_telemetry.py
Comment thread benchmarking/locust/runner.py Outdated
Comment thread benchmarking/locust/server_telemetry.py
Comment thread benchmarking/locust/server_telemetry.py Outdated
Comment thread benchmarking/locust/server_telemetry.py

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@Nishanth29

Copy link
Copy Markdown
Contributor Author

Issue: on microVM, snapshot size and checkpoint count came back null. The queries only matched pages.img (gVisor's memory image), and microVM writes memory-ranges.

Fix: match pages.img|memory-ranges. No change for gVisor. Tested on a microVM cluster: checkpoints_in_window now shows 729 (was null).

Remaining atelet issue: microVM sizes will still read 2 GiB until #931 lands. cloud-hypervisor writes memory-ranges as a sparse file sized to the full guest RAM (2 GiB) but only fills the pages the guest used. atelet records the file's length (info.Size()), not the bytes actually written, so every sample is 2147483648. atelet's own upload log shows ~294 MiB populated for the same files.

#931 fixes the recording; once it lands, sizes and checkpoint_mb_s here will be correct.

@maxsmythe Max Smythe (maxsmythe) left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@maxsmythe
Max Smythe (maxsmythe) added this pull request to the merge queue Oct 1, 2026
Merged via the queue into agent-substrate:main with commit 3f789b3 Oct 1, 2026
10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants