Skip to content

system base-copy gated on huge hdb_analytics (analytics.replicate:true) blocks hdb_deployment convergence → deploys fail after any copy #421

Description

@kriszyp

Summary

When analytics.replicate: true, the system database base-copy is gated on the huge hdb_analytics table, so convergence of small critical system tables (notably hdb_deployment) is blocked for a very long time — or indefinitely on a busy node. Any restart-triggered full copy (e.g. an upgrade) therefore leaves replicated component deploys failing to the affected peers long after the connections themselves are healthy.

Observed live on [customer-cluster].harperfabric.com (4-node, 5.1.5) during recovery from the reconnect wedge (harper-pro#420). Distinct mechanism from #420 — that one is the connection never reconnecting; this one is the copy stream never converging.

Evidence (live)

Why this is a bug (not just config)

Even with analytics.replicate: true, a base copy should not let a large, high-churn, largely node-local table (hdb_analytics) block convergence of small operational tables that gate cluster control-plane actions (hdb_deployment, hdb_nodes, etc.). Options for the fix:

  1. Copy critical system tables first — order the system base-copy so hdb_deployment / hdb_nodes / control-plane tables converge before bulk tables like hdb_analytics.
  2. Copy hdb_analytics on a separate, non-blocking stream (or as a lower-priority background copy) so it can't gate the system channel's convergence / awaitDeploymentRow.
  3. Reconsider replicating raw/aggregated analytics at all — it's mostly node-local telemetry; full cross-cluster base-copy of 1.7M rows on every restart is very expensive for little value. At minimum, make analytics.replicate exclude it from the base copy and only stream aggregated rows live.

Operational remediation (for affected clusters now)

  • Set analytics: { replicate: false } (the default) and restart, so analytics stops gating the system copy. Confirm the operator doesn't rely on cross-cluster analytics dashboards first.
  • And/or reduce aggregated-analytics retention (default 1 year) to shrink hdb_analytics.

Acceptance criteria

  • With a large hdb_analytics and analytics.replicate: true, a restart-triggered system copy still converges hdb_deployment (and a deploy_component succeeds to all peers) within a bounded time — i.e. analytics volume does not gate control-plane convergence.

Related

  • harper-pro#420 (reconnect wedge — the connection-level half of this incident; this issue is the copy-convergence half).
  • Config schema: analytics.replicate (default false), raw retention (default 1h), aggregated retention (default 1y).

Investigated via /fabric-investigation (logs + CDP + get_configuration) on the [Customer] preprod incident. Filed by Claude (Opus 4.8).

Metadata

Metadata

Assignees

Labels

area:replicationReplication, cluster sync, peer connectionsbugSomething isn't working

Type

No type

Fields

Priority

None yet

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions