Skip to content

feat(observability): per-table ingested (uncompressed) vs stored (on-disk) bytes gauges from system.parts #417

Description

@taitelee

Problem

WaveHouse exports no storage/capacity metrics. There is no way to see how much data has gone into ClickHouse versus what it actually occupies on disk — the headline an operator wants is "75 GB ingested, 1.4 GB stored (compressed)", per table and in total, for compression-ratio visibility, disk sizing, and retention decisions.

Today the only route is hand-written SQL against system.parts through POST /v1/admin/query. The exported metric surface (wavehouse_nats_*, wavehouse_pebble_*, wavehouse_bento_events_processed, Go runtime, otelhttp) has no storage dimension, and #94's tier list is latency histograms, error counters, and saturation gauges — storage/capacity isn't covered there either. The #360 incident (disk filled by ClickHouse system-log growth → embedded NATS can't boot → bare 502) showed the same blind spot from the infrastructure side: no exported disk-usage signal of any kind.

Proposed Solution

Follow the existing async-scraper pattern in internal/observability/metrics.go (RegisterSystemMetrics: observable gauges plus one meter.RegisterCallback that scrapes NATS varz and Pebble stats): add a ClickHouse storage scraper that runs one cheap metadata-only query over system.parts (active parts, configured database) and observes two gauges with a table attribute:

  • wavehouse_clickhouse_uncompressed_bytes{table} — sum(data_uncompressed_bytes): the uncompressed size of live column data — the closest system.parts analogue to "bytes ingested".
  • wavehouse_clickhouse_bytes_on_disk{table} — sum(bytes_on_disk): what that data actually occupies (compressed data plus marks/index files).

Instance-wide totals are the sum over the table attribute, and the compression ratio is the quotient — both computed at the dashboard/alert layer, so two gauges are the whole export surface. Per-part granularity stays ad hoc via SQL (a per-part label would be unbounded cardinality, and parts churn constantly with merges).

Wiring: the callback needs the ClickHouse HTTP endpoint and credentials (the same config the query handler and ingest worker already receive); pass them into RegisterSystemMetrics alongside the NATS/Pebble handles and observe with a table attribute per row.

Scope edges, deliberately:

Workaround available today (zero code)

SELECT table,
       formatReadableSize(sum(data_uncompressed_bytes)) AS ingested,
       formatReadableSize(sum(bytes_on_disk))           AS stored
FROM system.parts
WHERE active AND database = currentDatabase()
GROUP BY table

via POST /v1/admin/query (drop the GROUP BY and the table column for the instance-wide total).

Additional Context

Companion to #94 — same observability surface, distinct dimension (storage/capacity rather than latency/error-rate/saturation); filed separately so #94's Tier S → A phasing isn't blocked on it. Scraper pattern established by #11 (NATS/Pebble async scrapers). The admin UI epic (#250) lists dashboards that consume #94 metrics; these gauges feed the same dashboards.

Raised from dogfooding 2026-07-22: wanted "75 GB ingested, 1.4 GB stored (compressed)" visibility per table; issue-tracker sweep found no existing coverage (storage, compression, system.parts all absent).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area/observabilityMetrics, logs, traces, health, profilingenhancementNew feature or request

Type

No type

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions