Background & Problem
According to the Actor Density Benchmark Specs, key evaluation targets are active actors per node, per vCPU, per GB RAM, and the actor-to-pod (A/P) bin-packing ratio (defined as active concurrent actors / worker pod count).
Currently:
benchmarking/locust/runner.py only records client-side request latencies and does not discover cluster hardware capacity (nodes, vCPUs, RAM allocatable).
- During high-density benchmark runs, client requests may appear successful while server-side actors are queued or experiencing memory pressure. We need ground-truth telemetry from Prometheus (concurrent active/suspended actors, PSI pressure stalls, snapshot sizes, and ateapi throughput) captured alongside client metrics.
Proposed Changes
- Cluster Hardware Discovery: Add RBAC in
locust.yaml allowing the runner to query nodes and compute cluster density frontiers (actors/node, actors/vCPU, actors/GB RAM, and the steady-state A/P ratio: active_actors / worker_pods) written into stats.jsonl.
- Prometheus Telemetry Harvester (
server_telemetry.py): Query in-cluster Prometheus at the end of benchmark trials for ground-truth actor packing, PSI kernel pressure, and snapshot performance, outputting to server_summary.json and stats.jsonl.
- Preserve
status.json contract ({"locust_exit_code": 0, "stats_generated": true}) to maintain compatibility with test harnesses and CI orchestrators.
References
cc Max Smythe (@maxsmythe) Haowei Cai (Roy) (@roycaihw) Aditya Shantanu (@aditya-shantanu)
Background & Problem
According to the Actor Density Benchmark Specs, key evaluation targets are active actors per node, per vCPU, per GB RAM, and the actor-to-pod (A/P) bin-packing ratio (defined as active concurrent actors / worker pod count).
Currently:
benchmarking/locust/runner.pyonly records client-side request latencies and does not discover cluster hardware capacity (nodes, vCPUs, RAM allocatable).Proposed Changes
locust.yamlallowing the runner to querynodesand compute cluster density frontiers (actors/node,actors/vCPU,actors/GB RAM, and the steady-state A/P ratio:active_actors / worker_pods) written intostats.jsonl.server_telemetry.py): Query in-cluster Prometheus at the end of benchmark trials for ground-truth actor packing, PSI kernel pressure, and snapshot performance, outputting toserver_summary.jsonandstats.jsonl.status.jsoncontract ({"locust_exit_code": 0, "stats_generated": true}) to maintain compatibility with test harnesses and CI orchestrators.References
cc Max Smythe (@maxsmythe) Haowei Cai (Roy) (@roycaihw) Aditya Shantanu (@aditya-shantanu)