Repository navigation
[gpum] Report gpu.sm_active from GPM GR engine activity - #57945
martavicentenavarro wants to merge 1 commit into
Conversation
Use GPM_METRIC_GRAPHICS_UTIL, already collected as gpu.gr_engine_active, as a source for gpu.sm_active on physical GPUs without MIG. It measures the percentage of time the GR engine was busy, which closely follows the percentage of time any SM was active. The source has low priority by default, so it's only used when no other source is available, and high priority with the new gpu.prefer_gr_engine_sm_active option. It is not used when gpu.legacy_sm_active is enabled. Equal-priority ties in RemoveDuplicateSamples are now resolved deterministically by collector name. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
There was a problem hiding this comment.
AI review by Codex (OpenAI) - workflow run
patch is correct. No actionable regressions found in the source selection, MIG exclusions, configuration wiring, or deterministic deduplication. Tests were not run in the read-only environment.
Files inventory check summaryFile checks results against ancestor 8e99ce18: Results for datadog-agent_7.86.0~devel.git.328.d5ebf55.pipeline.143810155-1_amd64.deb:No change detected Results for datadog-iot-agent_7.86.0~devel.git.328.d5ebf55.pipeline.143810155-1_amd64.deb:No change detected |
Static quality checks✅ Please find below the results from static quality gates Successful checksInfo
18 successful checks with minimal change (< 2 KiB)
|
Regression DetectorRegression Detector ResultsMetrics dashboard Baseline: 8e99ce1 Optimization Goals: ✅ No significant changes detected
|
| perf | experiment | goal | Δ mean % | Δ mean % CI | trials | links |
|---|---|---|---|---|---|---|
| ➖ | dsd_uds_client_drop_detector_cpu | % cpu utilization | +0.21 | [-0.25, +0.67] | 1 | Logs |
| ➖ | quality_gate_idle_all_features | memory utilization | +0.03 | [-0.05, +0.11] | 1 | Logs bounds checks dashboard |
| ➖ | quality_gate_security_no_fs_load | memory utilization | -0.09 | [-0.16, -0.01] | 1 | Logs bounds checks dashboard |
| ➖ | quality_gate_idle | memory utilization | -0.15 | [-0.19, -0.10] | 1 | Logs bounds checks dashboard |
| ➖ | python_openmetrics | % cpu utilization | -0.18 | [-0.86, +0.49] | 1 | Logs bounds checks dashboard |
| ➖ | quality_gate_security_idle | memory utilization | -0.35 | [-0.38, -0.31] | 1 | Logs bounds checks dashboard |
| ➖ | quality_gate_private_action_runner | memory utilization | -0.35 | [-0.47, -0.22] | 1 | Logs bounds checks dashboard |
| ➖ | dsd_uds_10mb_3k_timestamped_contexts_memory | memory utilization | -0.51 | [-0.73, -0.30] | 1 | Logs |
| ➖ | quality_gate_security_mean_fs_load | memory utilization | -0.55 | [-0.59, -0.50] | 1 | Logs bounds checks dashboard |
| ➖ | quality_gate_logs | % cpu utilization | -0.64 | [-1.50, +0.23] | 1 | Logs bounds checks dashboard |
| ➖ | dsd_uds_10mb_3k_timestamped_contexts_cpu | % cpu utilization | -0.91 | [-1.15, -0.67] | 1 | Logs |
| ➖ | quality_gate_metrics_logs | memory utilization | -0.99 | [-1.23, -0.75] | 1 | Logs bounds checks dashboard |
Bounds Checks: ✅ Passed
| perf | experiment | bounds_check_name | replicates_passed | observed_value | links |
|---|---|---|---|---|---|
| ✅ | python_openmetrics | checks_execution_time | 10/10 | 80.53 ≤ 100 | bounds checks dashboard |
| ✅ | python_openmetrics | cpu_usage | 10/10 | 1333.32 ≤ 1500 | bounds checks dashboard |
| ✅ | python_openmetrics | memory_usage | 10/10 | 4.31GiB ≤ 4.75GiB | bounds checks dashboard |
| ✅ | quality_gate_idle | intake_connections | 10/10 | 4 ≤ 5 | bounds checks dashboard |
| ✅ | quality_gate_idle | memory_usage | 10/10 | 178.18MiB ≤ 181MiB | bounds checks dashboard |
| ✅ | quality_gate_idle | total_bytes_received | 10/10 | 758.77KiB ≤ 819.20KiB | bounds checks dashboard |
| ✅ | quality_gate_idle_all_features | intake_connections | 10/10 | 2 ≤ 5 | bounds checks dashboard |
| ✅ | quality_gate_idle_all_features | memory_usage | 10/10 | 475.58MiB ≤ 542MiB | bounds checks dashboard |
| ✅ | quality_gate_idle_all_features | total_bytes_received | 10/10 | 1.14MiB ≤ 1.25MiB | bounds checks dashboard |
| ✅ | quality_gate_logs | intake_connections | 10/10 | 17 ≤ 40 | bounds checks dashboard |
| ✅ | quality_gate_logs | memory_usage | 10/10 | 215.18MiB ≤ 228MiB | bounds checks dashboard |
| ✅ | quality_gate_logs | missed_bytes | 10/10 | 0B = 0B | bounds checks dashboard |
| ✅ | quality_gate_logs | total_bytes_received | 10/10 | 263.89MiB ≤ 292MiB | bounds checks dashboard |
| ✅ | quality_gate_metrics_logs | cpu_usage | 10/10 | 389.75 ≤ 2000 | bounds checks dashboard |
| ✅ | quality_gate_metrics_logs | intake_connections | 10/10 | 19 ≤ 40 | bounds checks dashboard |
| ✅ | quality_gate_metrics_logs | memory_usage | 10/10 | 447.47MiB ≤ 455MiB | bounds checks dashboard |
| ✅ | quality_gate_metrics_logs | missed_bytes | 10/10 | 0B = 0B | bounds checks dashboard |
| ✅ | quality_gate_metrics_logs | total_bytes_received | 10/10 | 0.95GiB ≤ 1.04GiB | bounds checks dashboard |
| ✅ | quality_gate_private_action_runner | memory_usage | 10/10 | 74.40MiB ≤ 77MiB | bounds checks dashboard |
| ✅ | quality_gate_security_idle | cpu_usage | 10/10 | 30.73 ≤ 100 | bounds checks dashboard |
| ✅ | quality_gate_security_idle | memory_usage | 10/10 | 320.27MiB ≤ 357MiB | bounds checks dashboard |
| ✅ | quality_gate_security_mean_fs_load | cpu_usage | 10/10 | 70.09 ≤ 200 | bounds checks dashboard |
| ✅ | quality_gate_security_mean_fs_load | memory_usage | 10/10 | 310.23MiB ≤ 337MiB | bounds checks dashboard |
| ✅ | quality_gate_security_no_fs_load | cpu_usage | 10/10 | 23.69 ≤ 100 | bounds checks dashboard |
| ✅ | quality_gate_security_no_fs_load | memory_usage | 10/10 | 330.47MiB ≤ 348MiB | bounds checks dashboard |
Explanation
Confidence level: 90.00%
Effect size tolerance: |Δ mean %| ≥ 5.00%
Performance changes are noted in the perf column of each table:
- ✅ = significantly better comparison variant performance
- ❌ = significantly worse comparison variant performance
- ➖ = no significant change in performance
A regression test is an A/B test of target performance in a repeatable rig, where "performance" is measured as "comparison variant minus baseline variant" for an optimization goal (e.g., ingress throughput). Due to intrinsic variability in measuring that goal, we can only estimate its mean value for each experiment; we report uncertainty in that value as a 90.00% confidence interval denoted "Δ mean % CI".
For each experiment, we decide whether a change in performance is a "regression" -- a change worth investigating further -- if all of the following criteria are true:
-
Its estimated |Δ mean %| ≥ 5.00%, indicating the change is big enough to merit a closer look.
-
Its 90.00% confidence interval "Δ mean % CI" does not contain zero, indicating that if our statistical model is accurate, there is at least a 90.00% chance there is a difference in performance between baseline and comparison variants.
-
Its configuration does not mark it "erratic".
CI Pass/Fail Decision
✅ Passed. All Quality Gates passed.
- quality_gate_logs, bounds check missed_bytes: 10/10 replicas passed. Gate passed.
- quality_gate_logs, bounds check memory_usage: 10/10 replicas passed. Gate passed.
- quality_gate_logs, bounds check intake_connections: 10/10 replicas passed. Gate passed.
- quality_gate_logs, bounds check total_bytes_received: 10/10 replicas passed. Gate passed.
- quality_gate_security_mean_fs_load, bounds check cpu_usage: 10/10 replicas passed. Gate passed.
- quality_gate_security_mean_fs_load, bounds check memory_usage: 10/10 replicas passed. Gate passed.
- quality_gate_idle_all_features, bounds check intake_connections: 10/10 replicas passed. Gate passed.
- quality_gate_idle_all_features, bounds check total_bytes_received: 10/10 replicas passed. Gate passed.
- quality_gate_idle_all_features, bounds check memory_usage: 10/10 replicas passed. Gate passed.
- quality_gate_metrics_logs, bounds check memory_usage: 10/10 replicas passed. Gate passed.
- quality_gate_metrics_logs, bounds check total_bytes_received: 10/10 replicas passed. Gate passed.
- quality_gate_metrics_logs, bounds check intake_connections: 10/10 replicas passed. Gate passed.
- quality_gate_metrics_logs, bounds check cpu_usage: 10/10 replicas passed. Gate passed.
- quality_gate_metrics_logs, bounds check missed_bytes: 10/10 replicas passed. Gate passed.
- quality_gate_idle, bounds check memory_usage: 10/10 replicas passed. Gate passed.
- quality_gate_idle, bounds check total_bytes_received: 10/10 replicas passed. Gate passed.
- quality_gate_idle, bounds check intake_connections: 10/10 replicas passed. Gate passed.
- quality_gate_private_action_runner, bounds check memory_usage: 10/10 replicas passed. Gate passed.
- quality_gate_security_no_fs_load, bounds check cpu_usage: 10/10 replicas passed. Gate passed.
- quality_gate_security_no_fs_load, bounds check memory_usage: 10/10 replicas passed. Gate passed.
- quality_gate_security_idle, bounds check memory_usage: 10/10 replicas passed. Gate passed.
- quality_gate_security_idle, bounds check cpu_usage: 10/10 replicas passed. Gate passed.
What does this PR do?
Adds a source for
gpu.sm_activefromGPM_METRIC_GRAPHICS_UTIL(already collected asgpu.gr_engine_active) on physical GPUs that don't have MIG mode enabled.gpu.sm_activesource is available (the NVML sampling source is Medium, and eBPF wins the tie at Low).gpu.prefer_gr_engine_sm_activeoption.gpu.legacy_sm_activeis enabled (the legacy value takes precedence), nor on MIG devices or MIG-enabled GPUs, where it hasn't been validated.RemoveDuplicateSamplesnow resolves equal-priority ties by collector name instead of map iteration order, sosm_activecan't alternate between the ebpf and gpm sources at Low. No other metric has an equal-priority tie across collectors, so their output doesn't change.No new NVML calls: GRAPHICS_UTIL is already queried for
gpu.gr_engine_active.Motivation
gpu.sm_activeshould be the percentage of time at least one SM was active. GRAPHICS_UTIL is a time-based measure of the GR engine being busy, unlikeGPM_METRIC_SM_UTIL, which averages SM activity over all SMs (25% when a quarter of the SMs are busy all the time).This is an alternative to #57928, which derives
sm_activefrom the SM cycle counters and turned out to report the same value asSM_UTIL. Both PRs carry the sameRemoveDuplicateSampleschange.Describe how you validated your changes
Unit tests in
pkg/collector/corechecks/gpu/nvidia.On an H100 80GB HBM3 (driver 595.91.07), GRAPHICS_UTIL was compared with a ground truth computed from per-kernel GPU timestamps (
%globaltimer), as the percentage of time with a kernel running:GRAPHICS_UTIL matches the ground truth within ±0.2 points for kernels of 1 ms or longer. It overestimates by about 2.5 µs per kernel launch, as NVML GPU utilization does. It doesn't count memory copies, which NVML GPU and process utilization (the current sampling source) report as ~99–100% busy.
With the Agent built from this branch (
agent check gpu, three runs 15 s apart, GPU 0):gpu.prefer_gr_engine_sm_active: truegpu.legacy_sm_activetoo--run_time 300 auto --target_sm 60)gr_engine_active)sm_utilization)gr_engine_active)gr_engine_active)gr_engine_active)Additional Notes
sm_activesample is produced on every run and dropped by deduplication when another source wins. This raises theduplicate_metricstelemetry by one per device per run.