Repository navigation
[gpum] Derive gpu.sm_active from GPM SM cycle counters - #57928
martavicentenavarro wants to merge 1 commit into
Conversation
Compute gpu.sm_active on physical GPUs from the raw GPM SM_CYCLES_ELAPSED (248) and SM_CYCLES_ACTIVE (249) counters, as the ratio of their deltas between consecutive samples. NVML returns these counters as cumulative values at Sample2, so the collector keeps the previous reading. The result equals GPM_METRIC_SM_UTIL. The source has low priority by default and high priority with the new gpu.prefer_sm_cycles_sm_active option. It is not computed when gpu.legacy_sm_active is enabled, nor on MIG devices, where the counters haven't been validated. Equal-priority ties in RemoveDuplicateSamples are now resolved deterministically by collector name. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
There was a problem hiding this comment.
AI review by Codex (OpenAI) - workflow run
The patch is correct. No actionable issues were found in the counter calculations, source priority handling, configuration wiring, or tests. Validation was limited to static review; tests were not run in the read-only environment.
Files inventory check summaryFile checks results against ancestor 66e6c51f: Results for datadog-agent_7.86.0~devel.git.306.189db58.pipeline.143779268-1_amd64.deb:No change detected Results for datadog-iot-agent_7.86.0~devel.git.306.189db58.pipeline.143779268-1_amd64.deb:No change detected |
Static quality checks✅ Please find below the results from static quality gates Successful checksInfo
18 successful checks with minimal change (< 2 KiB)
|
Regression DetectorRegression Detector ResultsMetrics dashboard Baseline: 66e6c51 Optimization Goals: ✅ No significant changes detected
|
| perf | experiment | goal | Δ mean % | Δ mean % CI | trials | links |
|---|---|---|---|---|---|---|
| ➖ | dsd_uds_10mb_3k_timestamped_contexts_memory | memory utilization | +1.62 | [+1.40, +1.84] | 1 | Logs |
| ➖ | quality_gate_security_no_fs_load | memory utilization | +0.26 | [+0.19, +0.34] | 1 | Logs bounds checks dashboard |
| ➖ | quality_gate_idle | memory utilization | +0.26 | [+0.21, +0.30] | 1 | Logs bounds checks dashboard |
| ➖ | dsd_uds_10mb_3k_timestamped_contexts_cpu | % cpu utilization | -0.05 | [-0.29, +0.19] | 1 | Logs |
| ➖ | quality_gate_idle_all_features | memory utilization | -0.07 | [-0.15, +0.00] | 1 | Logs bounds checks dashboard |
| ➖ | quality_gate_security_idle | memory utilization | -0.18 | [-0.21, -0.14] | 1 | Logs bounds checks dashboard |
| ➖ | python_openmetrics | % cpu utilization | -0.22 | [-0.90, +0.46] | 1 | Logs bounds checks dashboard |
| ➖ | quality_gate_metrics_logs | memory utilization | -0.38 | [-0.60, -0.15] | 1 | Logs bounds checks dashboard |
| ➖ | quality_gate_logs | % cpu utilization | -0.49 | [-1.34, +0.36] | 1 | Logs bounds checks dashboard |
| ➖ | quality_gate_security_mean_fs_load | memory utilization | -0.58 | [-0.62, -0.55] | 1 | Logs bounds checks dashboard |
| ➖ | quality_gate_private_action_runner | memory utilization | -0.63 | [-0.75, -0.51] | 1 | Logs bounds checks dashboard |
| ➖ | dsd_uds_client_drop_detector_cpu | % cpu utilization | -0.99 | [-1.48, -0.49] | 1 | Logs |
Bounds Checks: ✅ Passed
| perf | experiment | bounds_check_name | replicates_passed | observed_value | links |
|---|---|---|---|---|---|
| ✅ | python_openmetrics | checks_execution_time | 10/10 | 80.05 ≤ 100 | bounds checks dashboard |
| ✅ | python_openmetrics | cpu_usage | 10/10 | 1340.66 ≤ 1500 | bounds checks dashboard |
| ✅ | python_openmetrics | memory_usage | 10/10 | 4.31GiB ≤ 4.75GiB | bounds checks dashboard |
| ✅ | quality_gate_idle | intake_connections | 10/10 | 4 ≤ 5 | bounds checks dashboard |
| ✅ | quality_gate_idle | memory_usage | 10/10 | 179.54MiB ≤ 181MiB | bounds checks dashboard |
| ✅ | quality_gate_idle | total_bytes_received | 10/10 | 774.59KiB ≤ 819.20KiB | bounds checks dashboard |
| ✅ | quality_gate_idle_all_features | intake_connections | 10/10 | 2 ≤ 5 | bounds checks dashboard |
| ✅ | quality_gate_idle_all_features | memory_usage | 10/10 | 471.30MiB ≤ 542MiB | bounds checks dashboard |
| ✅ | quality_gate_idle_all_features | total_bytes_received | 10/10 | 1.14MiB ≤ 1.25MiB | bounds checks dashboard |
| ✅ | quality_gate_logs | intake_connections | 10/10 | 18 ≤ 40 | bounds checks dashboard |
| ✅ | quality_gate_logs | memory_usage | 10/10 | 212.05MiB ≤ 228MiB | bounds checks dashboard |
| ✅ | quality_gate_logs | missed_bytes | 10/10 | 0B = 0B | bounds checks dashboard |
| ✅ | quality_gate_logs | total_bytes_received | 10/10 | 263.88MiB ≤ 292MiB | bounds checks dashboard |
| ✅ | quality_gate_metrics_logs | cpu_usage | 10/10 | 371.56 ≤ 2000 | bounds checks dashboard |
| ✅ | quality_gate_metrics_logs | intake_connections | 10/10 | 20 ≤ 40 | bounds checks dashboard |
| ✅ | quality_gate_metrics_logs | memory_usage | 10/10 | 412.82MiB ≤ 455MiB | bounds checks dashboard |
| ✅ | quality_gate_metrics_logs | missed_bytes | 10/10 | 0B = 0B | bounds checks dashboard |
| ✅ | quality_gate_metrics_logs | total_bytes_received | 10/10 | 0.95GiB ≤ 1.04GiB | bounds checks dashboard |
| ✅ | quality_gate_private_action_runner | memory_usage | 10/10 | 73.62MiB ≤ 77MiB | bounds checks dashboard |
| ✅ | quality_gate_security_idle | cpu_usage | 10/10 | 30.57 ≤ 100 | bounds checks dashboard |
| ✅ | quality_gate_security_idle | memory_usage | 10/10 | 326.95MiB ≤ 357MiB | bounds checks dashboard |
| ✅ | quality_gate_security_mean_fs_load | cpu_usage | 10/10 | 63.11 ≤ 200 | bounds checks dashboard |
| ✅ | quality_gate_security_mean_fs_load | memory_usage | 10/10 | 306.43MiB ≤ 337MiB | bounds checks dashboard |
| ✅ | quality_gate_security_no_fs_load | cpu_usage | 10/10 | 22.62 ≤ 100 | bounds checks dashboard |
| ✅ | quality_gate_security_no_fs_load | memory_usage | 10/10 | 338.43MiB ≤ 348MiB | bounds checks dashboard |
Explanation
Confidence level: 90.00%
Effect size tolerance: |Δ mean %| ≥ 5.00%
Performance changes are noted in the perf column of each table:
- ✅ = significantly better comparison variant performance
- ❌ = significantly worse comparison variant performance
- ➖ = no significant change in performance
A regression test is an A/B test of target performance in a repeatable rig, where "performance" is measured as "comparison variant minus baseline variant" for an optimization goal (e.g., ingress throughput). Due to intrinsic variability in measuring that goal, we can only estimate its mean value for each experiment; we report uncertainty in that value as a 90.00% confidence interval denoted "Δ mean % CI".
For each experiment, we decide whether a change in performance is a "regression" -- a change worth investigating further -- if all of the following criteria are true:
-
Its estimated |Δ mean %| ≥ 5.00%, indicating the change is big enough to merit a closer look.
-
Its 90.00% confidence interval "Δ mean % CI" does not contain zero, indicating that if our statistical model is accurate, there is at least a 90.00% chance there is a difference in performance between baseline and comparison variants.
-
Its configuration does not mark it "erratic".
CI Pass/Fail Decision
✅ Passed. All Quality Gates passed.
- quality_gate_private_action_runner, bounds check memory_usage: 10/10 replicas passed. Gate passed.
- quality_gate_security_idle, bounds check cpu_usage: 10/10 replicas passed. Gate passed.
- quality_gate_security_idle, bounds check memory_usage: 10/10 replicas passed. Gate passed.
- quality_gate_idle, bounds check memory_usage: 10/10 replicas passed. Gate passed.
- quality_gate_idle, bounds check intake_connections: 10/10 replicas passed. Gate passed.
- quality_gate_idle, bounds check total_bytes_received: 10/10 replicas passed. Gate passed.
- quality_gate_idle_all_features, bounds check intake_connections: 10/10 replicas passed. Gate passed.
- quality_gate_idle_all_features, bounds check total_bytes_received: 10/10 replicas passed. Gate passed.
- quality_gate_idle_all_features, bounds check memory_usage: 10/10 replicas passed. Gate passed.
- quality_gate_security_mean_fs_load, bounds check cpu_usage: 10/10 replicas passed. Gate passed.
- quality_gate_security_mean_fs_load, bounds check memory_usage: 10/10 replicas passed. Gate passed.
- quality_gate_metrics_logs, bounds check missed_bytes: 10/10 replicas passed. Gate passed.
- quality_gate_metrics_logs, bounds check total_bytes_received: 10/10 replicas passed. Gate passed.
- quality_gate_metrics_logs, bounds check intake_connections: 10/10 replicas passed. Gate passed.
- quality_gate_metrics_logs, bounds check cpu_usage: 10/10 replicas passed. Gate passed.
- quality_gate_metrics_logs, bounds check memory_usage: 10/10 replicas passed. Gate passed.
- quality_gate_logs, bounds check missed_bytes: 10/10 replicas passed. Gate passed.
- quality_gate_logs, bounds check total_bytes_received: 10/10 replicas passed. Gate passed.
- quality_gate_logs, bounds check intake_connections: 10/10 replicas passed. Gate passed.
- quality_gate_logs, bounds check memory_usage: 10/10 replicas passed. Gate passed.
- quality_gate_security_no_fs_load, bounds check cpu_usage: 10/10 replicas passed. Gate passed.
- quality_gate_security_no_fs_load, bounds check memory_usage: 10/10 replicas passed. Gate passed.
What does this PR do?
Adds a source for
gpu.sm_activecomputed from the GPM countersSM_CYCLES_ELAPSED(248) andSM_CYCLES_ACTIVE(249) on physical GPUs.NVML returns these counters as cumulative values at
Sample2(not as the difference between the two samples), so the GPM collector keeps the previous reading and reports100 * Δactive / Δelapsed.gpu.prefer_sm_cycles_sm_activeoption.gpu.legacy_sm_activeis enabled (the legacy value takes precedence) nor on MIG devices, where the counters haven't been validated.RemoveDuplicateSamplesnow resolves equal-priority ties by collector name instead of map iteration order, sosm_activecan't alternate between the ebpf and gpm sources at Low. No other metric is affected: there were no equal-priority ties across collectors before this change.Motivation
AXT-33
Describe how you validated your changes
Unit tests
Manual validation on an H100 80GB HBM3 (driver 595.91.07, NVML 13.595.91.07) with
agent check gpu:Sample2.sm_active= 0.gpu-burner --run_time 300 auto --target_sm 60):sm_active= 61–62 from the sampling source;gpu.prefer_sm_cycles_sm_active: true:sm_active= 59.3–59.6 from the new source.gr_engine_activeis ~100%.Additional Notes
GPM_METRIC_SM_UTIL(SM activity averaged over all SMs), not the percentage of time any SM was active, so with the option enabledgpu.sm_activematchesgpu.sm_utilization.GpmMetricsGetcalls per GPM device per run, and its sample is deduplicated, which increasesduplicate_metricstelemetry by one per device.