Repository navigation
Conversation
COAT metric reminderThis PR changes
Automated reminder — does not block merging. Updates automatically if this file changes again. |
There was a problem hiding this comment.
AI review by Codex (OpenAI) - workflow run
patch is incorrect: the allowlist additions match the metric emitters, but the new documentation claims a health fallback that does not exist. The added test checks configuration only; no E2E assertion verifies delivery of these metrics.
COAT metric validation VMThe agent package for this pipeline is ready. Create a VM with that package installed using: Need Windows instead? Add See Observing telemetry for how to find the metric once the agent is running. |
2550a21 to
8c675d0
Compare
| # counts/totals. | ||
| # Accepted limitation: zero_metric exclusion (below) drops nodes_reporting==0, so the | ||
| # "no runners reporting" state produces no signal. Only positive counts are observable. | ||
| - name: cluster_checks.nodes_reporting |
There was a problem hiding this comment.
If in the future we find that this limitation is causing issues, we can move this metric to a new separate COAT profile that doesn't exclude zero metrics.
Files inventory check summaryFile checks results against ancestor ba0496db: Results for datadog-agent_7.86.0~devel.git.207.8c675d0.pipeline.143021622-1_amd64.deb:No change detected Results for datadog-iot-agent_7.86.0~devel.git.207.8c675d0.pipeline.143021622-1_amd64.deb:No change detected |
Static quality checks✅ Please find below the results from static quality gates Successful checksInfo
19 successful checks with minimal change (< 2 KiB)
|
|
Docs update PR: DataDog/documentation#40511 |
Regression DetectorRegression Detector ResultsMetrics dashboard Baseline: ba0496d Optimization Goals: ✅ No significant changes detected
|
| perf | experiment | goal | Δ mean % | Δ mean % CI | trials | links |
|---|---|---|---|---|---|---|
| ➖ | quality_gate_logs | % cpu utilization | +0.99 | [+0.10, +1.88] | 1 | Logs bounds checks dashboard |
| ➖ | quality_gate_metrics_logs | memory utilization | +0.62 | [+0.39, +0.84] | 1 | Logs bounds checks dashboard |
| ➖ | quality_gate_private_action_runner | memory utilization | +0.21 | [+0.08, +0.34] | 1 | Logs bounds checks dashboard |
| ➖ | dsd_uds_10mb_3k_timestamped_contexts_memory | memory utilization | +0.11 | [-0.12, +0.34] | 1 | Logs |
| ➖ | python_openmetrics | % cpu utilization | +0.06 | [-0.62, +0.74] | 1 | Logs bounds checks dashboard |
| ➖ | dsd_uds_client_drop_detector_cpu | % cpu utilization | +0.03 | [-0.45, +0.52] | 1 | Logs |
| ➖ | quality_gate_security_mean_fs_load | memory utilization | -0.04 | [-0.08, +0.01] | 1 | Logs bounds checks dashboard |
| ➖ | quality_gate_idle_all_features | memory utilization | -0.06 | [-0.14, +0.01] | 1 | Logs bounds checks dashboard |
| ➖ | quality_gate_idle | memory utilization | -0.13 | [-0.17, -0.09] | 1 | Logs bounds checks dashboard |
| ➖ | quality_gate_security_idle | memory utilization | -0.26 | [-0.30, -0.22] | 1 | Logs bounds checks dashboard |
| ➖ | quality_gate_security_no_fs_load | memory utilization | -0.29 | [-0.37, -0.21] | 1 | Logs bounds checks dashboard |
| ➖ | dsd_uds_10mb_3k_timestamped_contexts_cpu | % cpu utilization | -0.43 | [-0.66, -0.20] | 1 | Logs |
Bounds Checks: ✅ Passed
| perf | experiment | bounds_check_name | replicates_passed | observed_value | links |
|---|---|---|---|---|---|
| ✅ | python_openmetrics | checks_execution_time | 10/10 | 81.40 ≤ 100 | bounds checks dashboard |
| ✅ | python_openmetrics | cpu_usage | 10/10 | 1341.16 ≤ 1500 | bounds checks dashboard |
| ✅ | python_openmetrics | memory_usage | 10/10 | 4.31GiB ≤ 4.75GiB | bounds checks dashboard |
| ✅ | quality_gate_idle | intake_connections | 10/10 | 4 ≤ 5 | bounds checks dashboard |
| ✅ | quality_gate_idle | memory_usage | 10/10 | 177.91MiB ≤ 181MiB | bounds checks dashboard |
| ✅ | quality_gate_idle | total_bytes_received | 10/10 | 770.23KiB ≤ 819.20KiB | bounds checks dashboard |
| ✅ | quality_gate_idle_all_features | intake_connections | 10/10 | 2 ≤ 5 | bounds checks dashboard |
| ✅ | quality_gate_idle_all_features | memory_usage | 10/10 | 475.47MiB ≤ 542MiB | bounds checks dashboard |
| ✅ | quality_gate_idle_all_features | total_bytes_received | 10/10 | 1.13MiB ≤ 1.25MiB | bounds checks dashboard |
| ✅ | quality_gate_logs | intake_connections | 10/10 | 17 ≤ 40 | bounds checks dashboard |
| ✅ | quality_gate_logs | memory_usage | 10/10 | 215.37MiB ≤ 228MiB | bounds checks dashboard |
| ✅ | quality_gate_logs | missed_bytes | 10/10 | 0B = 0B | bounds checks dashboard |
| ✅ | quality_gate_logs | total_bytes_received | 10/10 | 263.89MiB ≤ 292MiB | bounds checks dashboard |
| ✅ | quality_gate_metrics_logs | cpu_usage | 10/10 | 367.04 ≤ 2000 | bounds checks dashboard |
| ✅ | quality_gate_metrics_logs | intake_connections | 10/10 | 19 ≤ 40 | bounds checks dashboard |
| ✅ | quality_gate_metrics_logs | memory_usage | 10/10 | 410.61MiB ≤ 455MiB | bounds checks dashboard |
| ✅ | quality_gate_metrics_logs | missed_bytes | 10/10 | 0B = 0B | bounds checks dashboard |
| ✅ | quality_gate_metrics_logs | total_bytes_received | 10/10 | 0.95GiB ≤ 1.04GiB | bounds checks dashboard |
| ✅ | quality_gate_private_action_runner | memory_usage | 10/10 | 73.80MiB ≤ 77MiB | bounds checks dashboard |
| ✅ | quality_gate_security_idle | cpu_usage | 10/10 | 29.96 ≤ 100 | bounds checks dashboard |
| ✅ | quality_gate_security_idle | memory_usage | 10/10 | 328.12MiB ≤ 357MiB | bounds checks dashboard |
| ✅ | quality_gate_security_mean_fs_load | cpu_usage | 10/10 | 64.32 ≤ 200 | bounds checks dashboard |
| ✅ | quality_gate_security_mean_fs_load | memory_usage | 10/10 | 311.61MiB ≤ 337MiB | bounds checks dashboard |
| ✅ | quality_gate_security_no_fs_load | cpu_usage | 10/10 | 22.74 ≤ 100 | bounds checks dashboard |
| ✅ | quality_gate_security_no_fs_load | memory_usage | 10/10 | 315.37MiB ≤ 348MiB | bounds checks dashboard |
Explanation
Confidence level: 90.00%
Effect size tolerance: |Δ mean %| ≥ 5.00%
Performance changes are noted in the perf column of each table:
- ✅ = significantly better comparison variant performance
- ❌ = significantly worse comparison variant performance
- ➖ = no significant change in performance
A regression test is an A/B test of target performance in a repeatable rig, where "performance" is measured as "comparison variant minus baseline variant" for an optimization goal (e.g., ingress throughput). Due to intrinsic variability in measuring that goal, we can only estimate its mean value for each experiment; we report uncertainty in that value as a 90.00% confidence interval denoted "Δ mean % CI".
For each experiment, we decide whether a change in performance is a "regression" -- a change worth investigating further -- if all of the following criteria are true:
-
Its estimated |Δ mean %| ≥ 5.00%, indicating the change is big enough to merit a closer look.
-
Its 90.00% confidence interval "Δ mean % CI" does not contain zero, indicating that if our statistical model is accurate, there is at least a 90.00% chance there is a difference in performance between baseline and comparison variants.
-
Its configuration does not mark it "erratic".
CI Pass/Fail Decision
✅ Passed. All Quality Gates passed.
- quality_gate_security_no_fs_load, bounds check cpu_usage: 10/10 replicas passed. Gate passed.
- quality_gate_security_no_fs_load, bounds check memory_usage: 10/10 replicas passed. Gate passed.
- quality_gate_idle_all_features, bounds check intake_connections: 10/10 replicas passed. Gate passed.
- quality_gate_idle_all_features, bounds check memory_usage: 10/10 replicas passed. Gate passed.
- quality_gate_idle_all_features, bounds check total_bytes_received: 10/10 replicas passed. Gate passed.
- quality_gate_idle, bounds check total_bytes_received: 10/10 replicas passed. Gate passed.
- quality_gate_idle, bounds check memory_usage: 10/10 replicas passed. Gate passed.
- quality_gate_idle, bounds check intake_connections: 10/10 replicas passed. Gate passed.
- quality_gate_private_action_runner, bounds check memory_usage: 10/10 replicas passed. Gate passed.
- quality_gate_security_mean_fs_load, bounds check cpu_usage: 10/10 replicas passed. Gate passed.
- quality_gate_security_mean_fs_load, bounds check memory_usage: 10/10 replicas passed. Gate passed.
- quality_gate_logs, bounds check intake_connections: 10/10 replicas passed. Gate passed.
- quality_gate_logs, bounds check total_bytes_received: 10/10 replicas passed. Gate passed.
- quality_gate_logs, bounds check memory_usage: 10/10 replicas passed. Gate passed.
- quality_gate_logs, bounds check missed_bytes: 10/10 replicas passed. Gate passed.
- quality_gate_metrics_logs, bounds check memory_usage: 10/10 replicas passed. Gate passed.
- quality_gate_metrics_logs, bounds check intake_connections: 10/10 replicas passed. Gate passed.
- quality_gate_metrics_logs, bounds check cpu_usage: 10/10 replicas passed. Gate passed.
- quality_gate_metrics_logs, bounds check missed_bytes: 10/10 replicas passed. Gate passed.
- quality_gate_metrics_logs, bounds check total_bytes_received: 10/10 replicas passed. Gate passed.
- quality_gate_security_idle, bounds check cpu_usage: 10/10 replicas passed. Gate passed.
- quality_gate_security_idle, bounds check memory_usage: 10/10 replicas passed. Gate passed.
|
/merge |
|
View all feedbacks in Devflow UI.
It will be processed automatically as soon as GitHub reports it as mergeable. View in MergeQueue UI.
devflow unqueued this merge request: It did not become mergeable within the expected time |
What does this PR do?
Onboards
cluster_check.telemtry via COAT to surface richer cross-org cluster-level health signals.Onboarded metrics:
cluster_checks.nodes_reportingendpoint_checks.configs_dispatchedcluster_checks.rebalancing_decisionscluster_checks.successful_rebalancing_movescluster_checks.failed_stats_collectionThe following metrics had already been onboarded:
cluster_checks.configs_danglingcluster_checks.unscheduled_checkcluster_checks.configs_dispatchedcluster_checks.configs_infoExplicitly not onboarded (deferred)
cluster_checks.busynessnode(high-card, dropped). Summing weight scores across runners hides per-runner saturation and scales with runner count. Useful fleet signal would be a max, which COAT cannot produce.cluster_checks.predicted_utilizationWorkersUsed/Workers(checks_distribution.go:36-42). Summing ratios is meaningless:1 runner @1.0 + 9 @0.0and10 runners @0.1both ship1.0. It cannot distinguish a saturated runner from a healthy fleet.Explicitly not onboarded (no real value)
cluster_checks.rebalancing_duration_secondscluster_checks.updating_stats_duration_secondsfailed_stats_collection.Describe how you validated your changes
Unit tests.
Manual Testing:
Created kind cluster:
Created datadog namespace, and created fakeintake:
Deployed the agent and created fakeintake service:
Get pods into running state:
Dispatch confirmed on the DCA:
DCA logs:
fakeintake received the telemetry on the COAT endpoint:
Payloads are base64 → zstd → JSON,
request_type∈ {agent-metrics,message-batch}.The first (embedded, period 900s) flush landed during the dispatcher's 30s warmup, when the
cluster-checks gauges were still 0 and therefore dropped by
zero_metric: true. To obtain apost-warmup flush quickly, the same metric set added in this PR was declared with a short
45s period (still running on the same binary); the embedded wiring itself is covered by the
unit test
TestClusterChecksMetricsInClusterAgentProfile.Additional Notes
zero_metricmetrics),cluser_checks.nodes_reportingwill not be propagated via COAT when its value is 0. To change this, we will have to create a separate profile forcluser_checks.nodes_reportingso that we can reliably distinguish between "no data available" and "no runners reporting". For now, we accept not shipping the metric via COAT when its value is 0.