You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
system base-copy gated on huge hdb_analytics (analytics.replicate:true) blocks hdb_deployment convergence → deploys fail after any copy #421
When analytics.replicate: true, the system database base-copy is gated on the huge hdb_analytics table, so convergence of small critical system tables (notably hdb_deployment) is blocked for a very long time — or indefinitely on a busy node. Any restart-triggered full copy (e.g. an upgrade) therefore leaves replicated component deploys failing to the affected peers long after the connections themselves are healthy.
Observed live on [customer-cluster].harperfabric.com (4-node, 5.1.5) during recovery from the reconnect wedge (harper-pro#420). Distinct mechanism from #420 — that one is the connection never reconnecting; this one is the copy stream never converging.
Evidence (live)
get_configuration on the cluster: analytics: { aggregatePeriod: 60, replicate: true }, replication.databases: "*".
hdb_analytics ≈ 1.72M records on the leader (largely node-local: leader 1.72M vs an edge 1.76M — counts diverge, i.e. it's mostly per-node data being force-copied cluster-wide).
The receivers log Resuming interrupted copy of database system ... at table hdb_analytics and stay there; hdb_deployment stays at 32/31 vs the leader's 38 for 20+ minutes with zero movement.
Even with analytics.replicate: true, a base copy should not let a large, high-churn, largely node-local table (hdb_analytics) block convergence of small operational tables that gate cluster control-plane actions (hdb_deployment, hdb_nodes, etc.). Options for the fix:
Copy critical system tables first — order the system base-copy so hdb_deployment / hdb_nodes / control-plane tables converge before bulk tables like hdb_analytics.
Copy hdb_analytics on a separate, non-blocking stream (or as a lower-priority background copy) so it can't gate the system channel's convergence / awaitDeploymentRow.
Reconsider replicating raw/aggregated analytics at all — it's mostly node-local telemetry; full cross-cluster base-copy of 1.7M rows on every restart is very expensive for little value. At minimum, make analytics.replicate exclude it from the base copy and only stream aggregated rows live.
Set analytics: { replicate: false } (the default) and restart, so analytics stops gating the system copy. Confirm the operator doesn't rely on cross-cluster analytics dashboards first.
And/or reduce aggregated-analytics retention (default 1 year) to shrink hdb_analytics.
Acceptance criteria
With a large hdb_analytics and analytics.replicate: true, a restart-triggered system copy still converges hdb_deployment (and a deploy_component succeeds to all peers) within a bounded time — i.e. analytics volume does not gate control-plane convergence.
Related
harper-pro#420 (reconnect wedge — the connection-level half of this incident; this issue is the copy-convergence half).
Summary
When
analytics.replicate: true, thesystemdatabase base-copy is gated on the hugehdb_analyticstable, so convergence of small critical system tables (notablyhdb_deployment) is blocked for a very long time — or indefinitely on a busy node. Any restart-triggered full copy (e.g. an upgrade) therefore leaves replicated component deploys failing to the affected peers long after the connections themselves are healthy.Observed live on
[customer-cluster].harperfabric.com(4-node, 5.1.5) during recovery from the reconnect wedge (harper-pro#420). Distinct mechanism from #420 — that one is the connection never reconnecting; this one is the copy stream never converging.Evidence (live)
get_configurationon the cluster:analytics: { aggregatePeriod: 60, replicate: true },replication.databases: "*".hdb_analytics≈ 1.72M records on the leader (largely node-local: leader 1.72M vs an edge 1.76M — counts diverge, i.e. it's mostly per-node data being force-copied cluster-wide).systemsockets sit at:connected: true,backPressurePercent ≈ 0(NOT backpressure, NOT the Replication wedges permanently after simultaneous cluster restart (reconciler skips open-but-idle sockets) → blocks replicated deploys #420 wedge),lastReceivedStatus: Receiving,recv = None.Resuming interrupted copy of database system ... at table hdb_analyticsand stay there;hdb_deploymentstays at 32/31 vs the leader's 38 for 20+ minutes with zero movement.hdb_analyticsto deliver the newerhdb_deploymentrows, sodeploy_componentto those peers keeps hitting the 120shdb_deployment row did not replicatetimeout (the same surface symptom as the original incident, but a different cause than Replication wedges permanently after simultaneous cluster restart (reconciler skips open-but-idle sockets) → blocks replicated deploys #420).Why this is a bug (not just config)
Even with
analytics.replicate: true, a base copy should not let a large, high-churn, largely node-local table (hdb_analytics) block convergence of small operational tables that gate cluster control-plane actions (hdb_deployment,hdb_nodes, etc.). Options for the fix:hdb_deployment/hdb_nodes/ control-plane tables converge before bulk tables likehdb_analytics.hdb_analyticson a separate, non-blocking stream (or as a lower-priority background copy) so it can't gate thesystemchannel's convergence /awaitDeploymentRow.analytics.replicateexclude it from the base copy and only stream aggregated rows live.Operational remediation (for affected clusters now)
analytics: { replicate: false }(the default) and restart, so analytics stops gating the system copy. Confirm the operator doesn't rely on cross-cluster analytics dashboards first.hdb_analytics.Acceptance criteria
hdb_analyticsandanalytics.replicate: true, a restart-triggered system copy still convergeshdb_deployment(and adeploy_componentsucceeds to all peers) within a bounded time — i.e. analytics volume does not gate control-plane convergence.Related
analytics.replicate(default false), raw retention (default 1h), aggregated retention (default 1y).Investigated via /fabric-investigation (logs + CDP + get_configuration) on the [Customer] preprod incident. Filed by Claude (Opus 4.8).