Repository navigation
Replication W8: Observability & metrics pipeline #437
Copy link
Copy link
Labels
area:metricsMetrics, analytics, monitoringMetrics, analytics, monitoringarea:replicationReplication, cluster sync, peer connectionsReplication, cluster sync, peer connectionsenhancementNew feature or requestNew feature or request
Milestone
Description
Activity
- addedenhancementNew feature or requestNew feature or requestarea:replicationReplication, cluster sync, peer connectionsReplication, cluster sync, peer connectionsarea:metricsMetrics, analytics, monitoringMetrics, analytics, monitoring
on Jun 20, 2026 - added sub-issues
on Jul 7, 2026 Three fold-ins from Chris Nelson's monitoring feedback (2026-07-15; reconciled in #532):
- Name the produce-side distinction explicitly in Tier 1: nothing-to-send (idle) vs cannot-send (peer waiting / backpressured) — receive-side timestamps conflate idle with broken, so lag alone isn't alertable; core knows if txns are actually waiting on a peer. Chris's single most valuable item.
- time-at-full-buffer counter alongside back-pressure % — raw % oscillates 0–100 on busy meshes and has to be smoothed downstream; a monotonic at-full-buffer counter is directly alertable.
- link-down-since timestamp per link.
Also relevant to W8's per-peer labels: the two new replication-observability gap issues #589 (inbound-link awareness / half-open self-report) and #590 (repl cert expiry).
🤖 Filed by KrAIs on behalf of Kris.
- added sub-issues
on Jul 20, 2026
Metadata
Metadata
Assignees
Labels
area:metricsMetrics, analytics, monitoringMetrics, analytics, monitoringarea:replicationReplication, cluster sync, peer connectionsReplication, cluster sync, peer connectionsenhancementNew feature or requestNew feature or request
Type
Fields
Priority
P2
Workstream W8 of #430 · observability & metrics pipeline
Summary
Today there are two half-pipes that don't meet:
cluster_status(an ephemeral poll snapshot of health gauges) andhdb_analytics(durable time-series, but only byte counters). Neither gives an operator a divergence alarm — which is why #426 (silent ~39% loss) and #386 (344 lost attachments) went undetected withconnected:true. Connect the pipes and add the missing signals.Tier 1 — extend
cluster_status(cheap, high diagnostic value, mostly already filed)lastDurableSequenceId/persistedseqIdand a computed lag; fixlastReceivedVersionnot populating on the relay/datapath (it's a non-signal exactly when you need it — see rapid-reconnect stress: follower permanently loses replication backlog after restart churn (data loss; fast-skip ruled out) #426).is_enabled:truewith real membership (cluster_status on a removed node should clearly indicate removal and close sockets #217).blobReplicationFailures).getSystemInfo(cluster_status: avoid full getSystemInfo call for performance #260).Tier 2 — a real metrics pipeline (trends + alerts + divergence)
hdb_analytics(it already ingests replication bytes — extend to lag, back-pressure %, blob-failure rate, connection state, cursor position).Tier 3 — log hygiene
Retires / advances
Dependencies
W1 (registry/last-error surfacing), W2 (gap signal), W4 (per-origin lag). Tier 1 can largely proceed in Phase 1 as quick wins; Tier 2 is the strategic piece.
Effort / risk
M / low.
Acceptance criteria
🤖 Filed by Claude on behalf of Kris.