Skip to content

bug(cache): Redis recovery probe must fit a fresh dial in the op timeout #664

Description

@EricAndrechek

Area: cache (Redis backend, #626)

After a failover behind a stable address, the Redis cache recovers by opening a new connection, since a connection lives at most one minute. rueidis dials that connection under the calling operation's context: mux.go, _pipe → wireFn(ctx), in v1.0.78. So the dial and handshake must fit inside the per-operation Timeout, which defaults to 100 ms.

The breaker's recovery probe runs under the same budget. If a fresh connect plus TLS handshake takes longer than Timeout, as it can across zones or with TLS, every probe times out. The cache then stays bypassed, and its owed invalidations stay pending, although the new primary is healthy.

Proposal: give the background probe a max(Timeout, DialTimeout) budget. Add a test in which a dial is slower than Timeout but faster than DialTimeout, and assert the breaker still closes.

#626 documents this as a caveat for now.

Activity

  1. EricAndrechek commented on Sep 26, 2026

    @EricAndrechek
    MemberAuthor

    #626 (now at 0eea1cb) addresses part of this.

    What changed.

    • A probe's first SET <prefix>:probe gets 2×DialTimeout + Timeout. rueidis bounds the dial (TCP plus TLS) by DialTimeout, then gives the HELLO/AUTH handshake a fresh DialTimeout of its own, so a reconnect can take about 2×DialTimeout.
    • If that write takes longer than Timeout, the probe sends the same SET again under Timeout alone, and only the repeat decides. So the breaker closes only when the server answers a write within Timeout. An earlier version without the repeat made the breaker flap against a slow server that still answered.
    • Two tests pin this, each mutation-checked:
      • TestRedis_ProbeFitsASlowReconnect: a TLS reconnect of about 1.4 s, with DialTimeout 1 s, still closes the breaker.
      • TestRedis_ProbeKeepsASlowServerBypassed: with every reply delayed past Timeout, the breaker stays open.

    What remains (measured with a proxy trace):

    • A standalone rueidis client keeps up to 4 connections, 1 << min(log2 GOMAXPROCS, 2), and each keyed command picks one at random.
    • The probe (and its repeat) reconnects only the connection it lands on. Other connections still redial under the operation's 100 ms Timeout. A repeat that lands on a cold connection costs one more BreakerOpenFor cycle.
    • Cluster mode keeps one connection per node, and the probe reaches one node.

    Options:

    • Use one connection per server (PipelineMultiplex), which costs some throughput.
    • Or make the probe reach every connection, or every node.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area/cacheLocal / shared / tiered cachingbugSomething isn't working

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions