You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
WaveHouse exports five custom metrics (wavehouse_bento_events_processed, wavehouse_nats_connections, wavehouse_nats_in_msgs_total, wavehouse_pebble_wal_size, wavehouse_pebble_table_count), Go runtime metrics, and otelhttp HTTP server metrics. Custom spans cover the ingest hot path (IngestHandler.Handle, clickhouse_delete, bento_batch_buffer, SchemaRegistry.Refresh, SSE.PushEvent, WS.PushEvent, WS.ReplayEvent) with W3C TraceContext over NATS headers; PR #88 adds bentoDLQDropped plus retroactive bento_queue_wait/clickhouse_insert spans.
Trace coverage on ingest is reasonable, but metrics coverage is shallow — only counters, no latency histograms, no error rates, no saturation signals. For an API gateway in front of ClickHouse, operators can see throughput but not health: no SLO metric for ingest latency, no leading indicator when the worker falls behind, no visibility into cache effectiveness, dedup hit rate, auth failures, schema-validation rejections, or policy denials. The query path (POST /v1/admin/query, POST /v1/query?table={table}, pipes execution) is entirely untraced.
Note this revisits work nominally closed by the observability buildout: #17 ("Component-Specific Business Metrics") specified wavehouse_ingest_latency_ms, wavehouse_worker_batch_size, and wavehouse_worker_flush_errors_total — none of which are among the metrics actually exported today; #21 ("Instrument ClickHouse-Go for Database Level Tracing") and #18 ("OpenTelemetry spans to primary API and Worker hot paths") didn't reach the query path. This issue is the follow-up that finishes and broadens that surface, not a duplicate — the gap list below is what's missing in practice (observed downstream, see Additional Context).
Proposed Solution
Introduce a wavehouse_*_duration_seconds histogram convention and a wavehouse_*_errors_total counter convention, then instrument the chokepoints. The infra (otel.Meter, OTLP exporter, SigNoz endpoint) is already wired — this is purely adding instruments. Suggested phasing:
ClickHouse latency histogram on INSERT and each query path + error counter labeled operation, clickhouse_code (delivers the wavehouse_worker_flush_errors_total half of Define and Implement Component-Specific Business Metrics #17, plus the query-path coverage it never had).
JetStream consumer-lag observable gauge (num_pending per consumer) pulled from consumer.Info() alongside the existing NATS varz scrape.
Trace gaps (priority order) — these extend the hot-path spans from #18 and the ClickHouse driver instrumentation from #21 to the surfaces they never covered:
Spans on the query handlers (POST /v1/admin/query, POST /v1/query?table={table}) + pipes execution, mirroring IngestHandler.Handle with table/pipe/ cache_hit attributes (Add OpenTelemetry spans to primary API and Worker hot paths #18 covered ingest hot paths only).
Auth-middleware span (JWT verify + JWKS fetch) as a child of the HTTP server span.
Schema-validation span under IngestHandler.Handle.
Cache + dedup lookup spans with hit=true/false.
Policy-evaluation span on authenticated requests that resolve a policy.
Keep counters-only and rely on log scraping for latency/errors — gives no percentiles, no SLO surface, and no leading indicators.
Lean on otelhttp HTTP metrics alone — produces spans but no route/status-labeled histograms, so no user-facing SLO.
Ship everything in one mega-PR — rejected in favor of the Tier S → A → B/C + traces phasing above.
Additional Context
Raised from operational experience integrating WaveHouse downstream.
Follow-up to the closed observability buildout — #17 (Component-Specific Business Metrics), #18 (OTel spans on API/Worker hot paths), #20 (Bridge Bento into traces), #21 (Instrument ClickHouse-Go for DB tracing), #22 (Instrument background system operations) — which delivered the pipeline and ingest-path coverage but not the histograms/error-rates/saturation signals or the query-path/auth/policy spans this issue adds. PR #88 already lands a slice (bentoDLQDropped + bento_queue_wait / clickhouse_insert spans). Also related: #39 (Performance Benchmarks), #44 (Build Info & /version), #84 (a benchmark for ingest under JetStream sync modes overlaps Tier S item 1).
Adding a complementary line item to the roadmap here:
Per-component logger source field. WaveHouse's slog handler already carries trace_id/span_id from active spans (per AGENTS.md design decision #15). What's still missing is a structured component / source field per emitter (e.g. api/ingest_handler, worker/batch_processor, ingest/sweeper, policy/store) so log queries can filter on the producing subsystem without grepping message text.
Mechanically: logger.With("component", "api/ingest_handler") (or equivalent) at each handler/worker entrypoint, propagated through subloggers. Feeds the same OTel pipeline as the other observability work in this issue.
Backing detail from nas-observability dogfooding (jfwoods). The WHissues.md log (2026-05-08) carries a full Tier S/A/B/C metrics + trace-gap breakdown behind this issue. Two counters in particular came out of real incidents and are worth prioritizing within this work:
Schema-validation rejections, labeled by table + reason (unknown_field / type_mismatch / null_violation) — a 23-min dead-letter outage there was invisible because there's no aggregate signal (only per-event warn logs + a reason-less DLQ count).
Dedup hit/miss (wavehouse_ingest_dedup_*_total, by table) — still just the duplicate event skipped log line at ingest.go:142, no counter.
Also called out: an auth-failure counter (by reason), a JetStream consumer-lag gauge (leading indicator before DLQ pressure), and spans on the query handlers (still uninstrumented).
Filed #417 as a storage/capacity companion to this tier list: per-table uncompressed-vs-on-disk bytes gauges scraped from system.parts (the "75 GB ingested, 1.4 GB stored" number), following the metrics.goRegisterCallback scraper pattern. Kept as its own issue since it's a storage dimension rather than latency/error-rate/saturation — cross-linking so the two surfaces stay coordinated. A lifetime bytes-ingested counter on the ingest path was deliberately left out of #417's scope; if we ever want it, it belongs as a Tier-A-style line item here.
Problem
WaveHouse exports five custom metrics (
wavehouse_bento_events_processed,wavehouse_nats_connections,wavehouse_nats_in_msgs_total,wavehouse_pebble_wal_size,wavehouse_pebble_table_count), Go runtime metrics, andotelhttpHTTP server metrics. Custom spans cover the ingest hot path (IngestHandler.Handle,clickhouse_delete,bento_batch_buffer,SchemaRegistry.Refresh,SSE.PushEvent,WS.PushEvent,WS.ReplayEvent) with W3C TraceContext over NATS headers; PR #88 addsbentoDLQDroppedplus retroactivebento_queue_wait/clickhouse_insertspans.Trace coverage on ingest is reasonable, but metrics coverage is shallow — only counters, no latency histograms, no error rates, no saturation signals. For an API gateway in front of ClickHouse, operators can see throughput but not health: no SLO metric for ingest latency, no leading indicator when the worker falls behind, no visibility into cache effectiveness, dedup hit rate, auth failures, schema-validation rejections, or policy denials. The query path (
POST /v1/admin/query,POST /v1/query?table={table}, pipes execution) is entirely untraced.Note this revisits work nominally closed by the observability buildout: #17 ("Component-Specific Business Metrics") specified
wavehouse_ingest_latency_ms,wavehouse_worker_batch_size, andwavehouse_worker_flush_errors_total— none of which are among the metrics actually exported today; #21 ("Instrument ClickHouse-Go for Database Level Tracing") and #18 ("OpenTelemetry spans to primary API and Worker hot paths") didn't reach the query path. This issue is the follow-up that finishes and broadens that surface, not a duplicate — the gap list below is what's missing in practice (observed downstream, see Additional Context).Proposed Solution
Introduce a
wavehouse_*_duration_secondshistogram convention and awavehouse_*_errors_totalcounter convention, then instrument the chokepoints. The infra (otel.Meter, OTLP exporter, SigNoz endpoint) is already wired — this is purely adding instruments. Suggested phasing:Tier S — production SLO essentials (one PR):
wavehouse_ingest_duration_secondsend-to-end histogram, labeledtable,outcome(committed / dlq / dropped); HTTP receive → ClickHouse commit.INSERTand each query path + error counter labeledoperation,clickhouse_code(delivers thewavehouse_worker_flush_errors_totalhalf of Define and Implement Component-Specific Business Metrics #17, plus the query-path coverage it never had).num_pendingper consumer) pulled fromconsumer.Info()alongside the existing NATS varz scrape.route,method,status_class.Tier A — feature health:
5. DLQ depth gauge (companion to #88's
bentoDLQDropped).6. Schema-validation rejection counter labeled
table,reason(unknown_field / type_mismatch / null_violation).7. Auth-failure counter labeled
reason(no_token / bad_signature / expired / jwks_fetch_failed / missing_role_claim).8. Cache hit ratio — L1 (Ristretto) / L2 (Redis) hit/miss/singleflight-collapsed counters.
9. Dedupe hit/miss counter by table.
10. NATS publish-503 counter on the API side (JetStream full →
Retry-After).Tier B — real-time + lifecycle:
11. SSE/WS active-connection gauge by transport +
pushed_events_total+ push-latency histogram (event arrival → wire write).12. Gap-fill replay counter (WS replay starts) + replay duration/message-count histogram.
13. Sweeper
messages_purged_totalcounter + observable gauge for age of oldest un-purged ACKed message.14. JWKS fetch counter / fetch-failure counter / fetch-latency histogram.
Tier C — depth:
15. Policy-denial counter labeled
table,role,denial_reason.16. ClickHouse connection-pool stats (in-use / idle / wait-time) via
clickhouse-goStats().17. Pipes execution counter labeled
pipe_name,outcome.18. Pebble dedup
Has/Putlatency histogram.Trace gaps (priority order) — these extend the hot-path spans from #18 and the ClickHouse driver instrumentation from #21 to the surfaces they never covered:
POST /v1/admin/query,POST /v1/query?table={table}) + pipes execution, mirroringIngestHandler.Handlewithtable/pipe/cache_hitattributes (Add OpenTelemetry spans to primary API and Worker hot paths #18 covered ingest hot paths only).IngestHandler.Handle.hit=true/false.Exec/Querywith SQL fingerprint (not raw SQL), extending thebento_queue_wait+clickhouse_insertmodel (and Instrument ClickHouse-Go for Database Level Tracing #21) to query paths.purged_count/scanned_count(companion to Instrument Background System Operations #22's background-ops instrumentation).Alternatives Considered
otelhttpHTTP metrics alone — produces spans but no route/status-labeled histograms, so no user-facing SLO.Additional Context
Raised from operational experience integrating WaveHouse downstream.
Follow-up to the closed observability buildout — #17 (Component-Specific Business Metrics), #18 (OTel spans on API/Worker hot paths), #20 (Bridge Bento into traces), #21 (Instrument ClickHouse-Go for DB tracing), #22 (Instrument background system operations) — which delivered the pipeline and ingest-path coverage but not the histograms/error-rates/saturation signals or the query-path/auth/policy spans this issue adds. PR #88 already lands a slice (
bentoDLQDropped+bento_queue_wait/clickhouse_insertspans). Also related: #39 (Performance Benchmarks), #44 (Build Info & /version), #84 (a benchmark for ingest under JetStream sync modes overlaps Tier S item 1).