Skip to content

Replication W8: Observability & metrics pipeline #437

Description

@kriszyp

Workstream W8 of #430 · observability & metrics pipeline

Summary

Today there are two half-pipes that don't meet: cluster_status (an ephemeral poll snapshot of health gauges) and hdb_analytics (durable time-series, but only byte counters). Neither gives an operator a divergence alarm — which is why #426 (silent ~39% loss) and #386 (344 lost attachments) went undetected with connected:true. Connect the pipes and add the missing signals.

Tier 1 — extend cluster_status (cheap, high diagnostic value, mostly already filed)

Tier 2 — a real metrics pipeline (trends + alerts + divergence)

  • Promote replication-health gauges into hdb_analytics (it already ingests replication bytes — extend to lag, back-pressure %, blob-failure rate, connection state, cursor position).
  • Prometheus / OpenTelemetry export over that series (none exists today).
  • Divergence detection — periodic cross-peer record/blob-count comparison emitting an alertable metric. This is the connective tissue with W2: it converts every "silent divergence" into a loud, actionable signal before data loss.

Tier 3 — log hygiene

Retires / advances

Dependencies

W1 (registry/last-error surfacing), W2 (gap signal), W4 (per-origin lag). Tier 1 can largely proceed in Phase 1 as quick wins; Tier 2 is the strategic piece.

Effort / risk

M / low.

Acceptance criteria

  • An operator can see per-subscription lag, last connection error, and cursor position.
  • A diverging cluster raises an alert before data loss.
  • Storm-condition logs are bounded and parseable.

🤖 Filed by Claude on behalf of Kris.

Activity

  1. added
    enhancementNew feature or request
    area:replicationReplication, cluster sync, peer connections
    area:metricsMetrics, analytics, monitoring
    on Jun 20, 2026
  2. added this to the v5.2 milestone on Jun 20, 2026
  3. modified the milestones: v5.2, v5.3 on Jul 7, 2026
  4. kriszyp commented on Jul 16, 2026

    @kriszyp
    MemberAuthor

    Three fold-ins from Chris Nelson's monitoring feedback (2026-07-15; reconciled in #532):

    • Name the produce-side distinction explicitly in Tier 1: nothing-to-send (idle) vs cannot-send (peer waiting / backpressured) — receive-side timestamps conflate idle with broken, so lag alone isn't alertable; core knows if txns are actually waiting on a peer. Chris's single most valuable item.
    • time-at-full-buffer counter alongside back-pressure % — raw % oscillates 0–100 on busy meshes and has to be smoothed downstream; a monotonic at-full-buffer counter is directly alertable.
    • link-down-since timestamp per link.

    Also relevant to W8's per-peer labels: the two new replication-observability gap issues #589 (inbound-link awareness / half-open self-report) and #590 (repl cert expiry).

    🤖 Filed by KrAIs on behalf of Kris.

  5. added theissue type on Sep 21, 2026
  6. modified the milestones: v5.3, v5.4 on Oct 10, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:metricsMetrics, analytics, monitoringarea:replicationReplication, cluster sync, peer connectionsenhancementNew feature or request

    Type

    Fields

    Priority

    P2

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions