Skip to content

Investigate: heap corruption / OOM from rapid TLS context churn in monitorNodeCAs during reconnect cycles #288

Description

@kriszyp

Summary

Under aggressive replication reconnect cycling, tls.createSecureContext is called repeatedly inside monitorNodeCAs. This has been observed correlating with heap corruption and OOM in production environments.

Current assessment

Believed to be a symptom of the crash-loop / full-copy OOM cycle (see #286) rather than a standalone bug, but the TLS churn may be an independent contributor worth isolating. Nathan and Devin have looked at this but root cause is not confirmed.

Technical detail

  • monitorNodeCAs calls tls.createSecureContext on every reconnect attempt
  • Under a crash-loop this can happen at high frequency, cycling memory for TLS contexts
  • Observed correlating with heap corruption and OOM process crashes

Relation to other issues

Next step

Confirm whether adding debouncing/caching to monitorNodeCAs's createSecureContext calls reduces or eliminates the heap corruption independent of the full-copy fix.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:replicationReplication, cluster sync, peer connectionsbugSomething isn't working

    Type

    Fields

    Priority

    P2

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions