Problem
Standard system metrics tell us if the "engine" is running, but not if the "car" is moving. We need custom metrics that track the actual business value and performance of WaveHouse's specific components.
Proposed Solution
Identify and implement the following custom metrics across the stack:
- API:
- wavehouse_api_requests_total (Labeled by status code/method)
- wavehouse_ingest_latency_ms (Histogram for end-to-end ingestion)
- Worker:
- wavehouse_worker_batch_size (Histogram of rows flushed to ClickHouse)
- wavehouse_worker_flush_errors_total (Counter for failed CH inserts)
- Bento:
- wavehouse_bento_transform_latency_ns (Time spent in mapping/pipes)
- wavehouse_bento_output_dropped_total (DLQ trigger count)
Alternatives Considered
None
Additional Context
None
Problem
Standard system metrics tell us if the "engine" is running, but not if the "car" is moving. We need custom metrics that track the actual business value and performance of WaveHouse's specific components.
Proposed Solution
Identify and implement the following custom metrics across the stack:
Alternatives Considered
None
Additional Context
None