Problem
In a standalone deployment, WaveHouse runs "black-box" embedded services (NATS JetStream and Pebble KV). Currently, we have visibility into the application layer via traces, but we lack "internal" health metrics (e.g., NATS message backpressure, Pebble WAL size, or memory consumption). Without these, it is difficult to know if the system is scaling or if a bottleneck is forming in the storage/queue layers.
Proposed Solution
Create a background scraper within the internal/observability package using the OpenTelemetry Meter API.
- NATS: Utilize natsServer.Varz() to capture active connections and message throughput.
- Pebble: Wrap the pebble.DB.Metrics() call to observe LSM-tree levels and Write-Ahead Log (WAL) sizes.
- Mechanism: Use meter.RegisterCallback to create asynchronous gauges that poll these internal stats every 15 seconds, ensuring zero performance impact on the hot path of message ingestion.
Alternatives Considered
Any alternative approaches you've thought about.
Additional Context
The implementation should follow the signature:
func RegisterSystemMetrics(natsServer *server.Server, dedup dedupe.Deduplicator) error
Problem
In a standalone deployment, WaveHouse runs "black-box" embedded services (NATS JetStream and Pebble KV). Currently, we have visibility into the application layer via traces, but we lack "internal" health metrics (e.g., NATS message backpressure, Pebble WAL size, or memory consumption). Without these, it is difficult to know if the system is scaling or if a bottleneck is forming in the storage/queue layers.
Proposed Solution
Create a background scraper within the internal/observability package using the OpenTelemetry Meter API.
Alternatives Considered
Any alternative approaches you've thought about.
Additional Context
The implementation should follow the signature:
func RegisterSystemMetrics(natsServer *server.Server, dedup dedupe.Deduplicator) error