Skip to content

feat(observability): latency histograms, error-rate counters, saturation gauges, query-path traces #94

Description

@jfwoods

Problem

WaveHouse exports five custom metrics (wavehouse_bento_events_processed, wavehouse_nats_connections, wavehouse_nats_in_msgs_total, wavehouse_pebble_wal_size, wavehouse_pebble_table_count), Go runtime metrics, and otelhttp HTTP server metrics. Custom spans cover the ingest hot path (IngestHandler.Handle, clickhouse_delete, bento_batch_buffer, SchemaRegistry.Refresh, SSE.PushEvent, WS.PushEvent, WS.ReplayEvent) with W3C TraceContext over NATS headers; PR #88 adds bentoDLQDropped plus retroactive bento_queue_wait/clickhouse_insert spans.

Trace coverage on ingest is reasonable, but metrics coverage is shallow — only counters, no latency histograms, no error rates, no saturation signals. For an API gateway in front of ClickHouse, operators can see throughput but not health: no SLO metric for ingest latency, no leading indicator when the worker falls behind, no visibility into cache effectiveness, dedup hit rate, auth failures, schema-validation rejections, or policy denials. The query path (POST /v1/admin/query, POST /v1/query?table={table}, pipes execution) is entirely untraced.

Note this revisits work nominally closed by the observability buildout: #17 ("Component-Specific Business Metrics") specified wavehouse_ingest_latency_ms, wavehouse_worker_batch_size, and wavehouse_worker_flush_errors_total — none of which are among the metrics actually exported today; #21 ("Instrument ClickHouse-Go for Database Level Tracing") and #18 ("OpenTelemetry spans to primary API and Worker hot paths") didn't reach the query path. This issue is the follow-up that finishes and broadens that surface, not a duplicate — the gap list below is what's missing in practice (observed downstream, see Additional Context).

Proposed Solution

Introduce a wavehouse_*_duration_seconds histogram convention and a wavehouse_*_errors_total counter convention, then instrument the chokepoints. The infra (otel.Meter, OTLP exporter, SigNoz endpoint) is already wired — this is purely adding instruments. Suggested phasing:

Tier S — production SLO essentials (one PR):

  1. wavehouse_ingest_duration_seconds end-to-end histogram, labeled table, outcome (committed / dlq / dropped); HTTP receive → ClickHouse commit.
  2. ClickHouse latency histogram on INSERT and each query path + error counter labeled operation, clickhouse_code (delivers the wavehouse_worker_flush_errors_total half of Define and Implement Component-Specific Business Metrics #17, plus the query-path coverage it never had).
  3. JetStream consumer-lag observable gauge (num_pending per consumer) pulled from consumer.Info() alongside the existing NATS varz scrape.
  4. HTTP request latency + error-rate histogram labeled route, method, status_class.

Tier A — feature health:
5. DLQ depth gauge (companion to #88's bentoDLQDropped).
6. Schema-validation rejection counter labeled table, reason (unknown_field / type_mismatch / null_violation).
7. Auth-failure counter labeled reason (no_token / bad_signature / expired / jwks_fetch_failed / missing_role_claim).
8. Cache hit ratio — L1 (Ristretto) / L2 (Redis) hit/miss/singleflight-collapsed counters.
9. Dedupe hit/miss counter by table.
10. NATS publish-503 counter on the API side (JetStream full → Retry-After).

Tier B — real-time + lifecycle:
11. SSE/WS active-connection gauge by transport + pushed_events_total + push-latency histogram (event arrival → wire write).
12. Gap-fill replay counter (WS replay starts) + replay duration/message-count histogram.
13. Sweeper messages_purged_total counter + observable gauge for age of oldest un-purged ACKed message.
14. JWKS fetch counter / fetch-failure counter / fetch-latency histogram.

Tier C — depth:
15. Policy-denial counter labeled table, role, denial_reason.
16. ClickHouse connection-pool stats (in-use / idle / wait-time) via clickhouse-go Stats().
17. Pipes execution counter labeled pipe_name, outcome.
18. Pebble dedup Has/Put latency histogram.

Trace gaps (priority order) — these extend the hot-path spans from #18 and the ClickHouse driver instrumentation from #21 to the surfaces they never covered:

  1. Spans on the query handlers (POST /v1/admin/query, POST /v1/query?table={table}) + pipes execution, mirroring IngestHandler.Handle with table/pipe/ cache_hit attributes (Add OpenTelemetry spans to primary API and Worker hot paths #18 covered ingest hot paths only).
  2. Auth-middleware span (JWT verify + JWKS fetch) as a child of the HTTP server span.
  3. Schema-validation span under IngestHandler.Handle.
  4. Cache + dedup lookup spans with hit=true/false.
  5. Policy-evaluation span on authenticated requests that resolve a policy.
  6. ClickHouse driver spans around Exec/Query with SQL fingerprint (not raw SQL), extending the bento_queue_wait + clickhouse_insert model (and Instrument ClickHouse-Go for Database Level Tracing #21) to query paths.
  7. Sweeper batch span per cycle with purged_count/scanned_count (companion to Instrument Background System Operations #22's background-ops instrumentation).

Alternatives Considered

  • Keep counters-only and rely on log scraping for latency/errors — gives no percentiles, no SLO surface, and no leading indicators.
  • Lean on otelhttp HTTP metrics alone — produces spans but no route/status-labeled histograms, so no user-facing SLO.
  • Ship everything in one mega-PR — rejected in favor of the Tier S → A → B/C + traces phasing above.

Additional Context

Raised from operational experience integrating WaveHouse downstream.

Follow-up to the closed observability buildout — #17 (Component-Specific Business Metrics), #18 (OTel spans on API/Worker hot paths), #20 (Bridge Bento into traces), #21 (Instrument ClickHouse-Go for DB tracing), #22 (Instrument background system operations) — which delivered the pipeline and ingest-path coverage but not the histograms/error-rates/saturation signals or the query-path/auth/policy spans this issue adds. PR #88 already lands a slice (bentoDLQDropped + bento_queue_wait / clickhouse_insert spans). Also related: #39 (Performance Benchmarks), #44 (Build Info & /version), #84 (a benchmark for ingest under JetStream sync modes overlaps Tier S item 1).

Activity

  1. added theissue type on May 11, 2026
  2. added
    enhancementNew feature or request
    area/apiHTTP handlers, routing, middleware
    area/ingestIngest pipeline (Bento, batching, DLQ)
    area/cacheLocal / shared / tiered caching
    area/dedupeDeduplication (Pebble, ScyllaDB)
    area/policyAccess control policies (Hasura-style)
    on May 11, 2026
  3. moved this from Backlog to Ready in WaveHouse Task Boardon May 11, 2026
  4. EricAndrechek commented on May 15, 2026

    @EricAndrechek
    Member

    Adding a complementary line item to the roadmap here:

    Per-component logger source field. WaveHouse's slog handler already carries trace_id/span_id from active spans (per AGENTS.md design decision #15). What's still missing is a structured component / source field per emitter (e.g. api/ingest_handler, worker/batch_processor, ingest/sweeper, policy/store) so log queries can filter on the producing subsystem without grepping message text.

    Mechanically: logger.With("component", "api/ingest_handler") (or equivalent) at each handler/worker entrypoint, propagated through subloggers. Feeds the same OTel pipeline as the other observability work in this issue.

  5. moved this from Ready to Backlog in WaveHouse Task Boardon May 18, 2026
  6. moved this from Backlog to In progress in WaveHouse Task Boardon May 24, 2026
  7. moved this from In progress to Ready in WaveHouse Task Boardon Jun 2, 2026
  8. EricAndrechek commented on Jun 4, 2026

    @EricAndrechek
    Member

    Backing detail from nas-observability dogfooding (jfwoods). The WHissues.md log (2026-05-08) carries a full Tier S/A/B/C metrics + trace-gap breakdown behind this issue. Two counters in particular came out of real incidents and are worth prioritizing within this work:

    • Schema-validation rejections, labeled by table + reason (unknown_field / type_mismatch / null_violation) — a 23-min dead-letter outage there was invisible because there's no aggregate signal (only per-event warn logs + a reason-less DLQ count).
    • Dedup hit/miss (wavehouse_ingest_dedup_*_total, by table) — still just the duplicate event skipped log line at ingest.go:142, no counter.

    Also called out: an auth-failure counter (by reason), a JetStream consumer-lag gauge (leading indicator before DLQ pressure), and spans on the query handlers (still uninstrumented).

  9. moved this from Ready to Backlog in WaveHouse Task Boardon Jun 9, 2026
  10. taitelee commented on Jul 22, 2026

    @taitelee
    Member

    Filed #417 as a storage/capacity companion to this tier list: per-table uncompressed-vs-on-disk bytes gauges scraped from system.parts (the "75 GB ingested, 1.4 GB stored" number), following the metrics.go RegisterCallback scraper pattern. Kept as its own issue since it's a storage dimension rather than latency/error-rate/saturation — cross-linking so the two surfaces stay coordinated. A lifetime bytes-ingested counter on the ingest path was deliberately left out of #417's scope; if we ever want it, it belongs as a Tier-A-style line item here.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area/apiHTTP handlers, routing, middlewarearea/cacheLocal / shared / tiered cachingarea/dedupeDeduplication (Pebble, ScyllaDB)area/ingestIngest pipeline (Bento, batching, DLQ)area/observabilityMetrics, logs, traces, health, profilingarea/pipesNamed query pipesarea/policyAccess control policies (Hasura-style)enhancementNew feature or request

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions