Skip to content

Phase 2 / S7.E: controller workqueue metrics - #76

Merged
Pal Lakatos-Toth (pallakatos) merged 1 commit into
devfrom
phase2-controller-metrics
Apr 29, 2026
Merged

Pal Lakatos-Toth (pallakatos) merged 1 commit into
devfrom
phase2-controller-metrics

Conversation

@pallakatos

Copy link
Copy Markdown
Collaborator

S7.E — controller Prometheus metrics

Expose /metrics from the controller pod so operators can alert on reconcile health without scraping logs.

  • New controller/src/metrics.rs (counter registration + helper).
  • New controller/src/metrics_server.rs (axum /metrics + /healthz on :9091).
  • 8 error_policy fns wired to record_reconcile_error.
  • Helm chart adds metrics container port.
  • Controller bin 345 → 349 tests. Clippy + fmt + helm lint clean.

Audit: docs/security-audits/2026-04-29-phase2-controller-metrics.md.

Histograms / queue-depth / OTel spans deferred to S7.E.2.

Co-authored-by: Copilot 223556219+Copilot@users.noreply.github.com

Expose Prometheus metrics from the controller pod so operators can
SLO and alert on reconcile health without scraping logs.

* New controller/src/metrics.rs registering two IntCounterVecs:
  azureclaw_controller_reconcile_errors_total{crd_kind, error_class}
  azureclaw_controller_reconcile_retries_total{crd_kind}
* New controller/src/metrics_server.rs — axum server exposing
  /metrics + /healthz on $CONTROLLER_METRICS_ADDR (default :9091).
* All eight error_policy fns wired to record_reconcile_error().
* Helm controller-deployment.yaml declares containerPort 9091.
* Controller Cargo.toml adds axum 0.8.
* 4 new unit tests; controller bin 345 → 349. Clippy / fmt / helm
  lint clean.

Audit: docs/security-audits/2026-04-29-phase2-controller-metrics.md.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@pallakatos
Pal Lakatos-Toth (pallakatos) merged commit 9555272 into dev Apr 29, 2026
14 of 15 checks passed
@pallakatos
Pal Lakatos-Toth (pallakatos) deleted the phase2-controller-metrics branch April 29, 2026 06:46
Pal Lakatos-Toth (pallakatos) added a commit that referenced this pull request May 12, 2026
Expose Prometheus metrics from the controller pod so operators can
SLO and alert on reconcile health without scraping logs.

* New controller/src/metrics.rs registering two IntCounterVecs:
  azureclaw_controller_reconcile_errors_total{crd_kind, error_class}
  azureclaw_controller_reconcile_retries_total{crd_kind}
* New controller/src/metrics_server.rs — axum server exposing
  /metrics + /healthz on $CONTROLLER_METRICS_ADDR (default :9091).
* All eight error_policy fns wired to record_reconcile_error().
* Helm controller-deployment.yaml declares containerPort 9091.
* Controller Cargo.toml adds axum 0.8.
* 4 new unit tests; controller bin 345 → 349. Clippy / fmt / helm
  lint clean.

Audit: docs/security-audits/2026-04-29-phase2-controller-metrics.md.

Co-authored-by: Pal Lakatos-Toth <pallakatos@microsoft.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant