Skip to content

Whole-keyspace ownership and whole-command forwarding for a plugin keyspace #886

Description

@cb1kenobi

The multi-node path for a plugin keyspace: one node owns the whole keyspace, and every other node
forwards entire commands to it. Mandatory before any migration-facing product, because without
it a load-balanced multi-node deployment makes the common Redis uses wrong by construction.

Reuse the operator mechanics, not the home map

An earlier version of this issue recommended generalizing the shipped record-lock home map, on the
observation that a map with exactly one member resolves every key to that node and so is a
whole-keyspace owner. That recommendation was wrong and has been withdrawn.

homeMap() is keyed by database, and its member list governs every table in that database.
Forcing it to one member to obtain a Redis owner would therefore also centralize ordinary record
locks for the whole application database onto the Redis owner node: enabling Redis for one
application would silently re-home unrelated table locks. It is also registered only for replicated
databases, its transport interface is a closed method set, and its freshness barrier refuses
non-replicated tables.

What to reuse is the machinery, in a domain of its own: the generation counter, the
homeIncarnation continuity datum, the staged-then-activate quiescence sequence, the digest
agreement check, and the external-fencing requirement — inside a separately named ownership domain
scoped to the plugin keyspace
, leaving the database's own locking topology untouched.

The ownership contract

  • Whole-keyspace, not per-key. Per-key homing breaks cross-key atomicity, which is exactly why
    the lock work's rendezvous hash over key names is the wrong unit here. A later scale-out step must
    adopt CROSSSLOT plus hash tags rather than pretend.
  • Immutable binding, independent of live membership. Not derived from who is currently reachable.
  • Enforced on every path, including the owner-side handler, with no local fallback. A non-owner
    that cannot reach the owner fails the command. It does not serve it locally.
  • Whole commands forward, reads and cursors included. Note the lock work forwards arbitration
    only
    and then writes locally, so it is precedent for the transport but not for the forwarding
    semantics this needs.
  • Owner changes only by explicit operator quiescence.
  • No automatic failover. Fencing does not exist to support one, and a failover without fencing is
    a correctness bug rather than a degraded mode. The record-lock work reached the same conclusion
    independently and recorded it as a decision, not a gap.

Dependencies

The peer-scoped transport in this epic, and the core enlistment and decision primitives from the
parent epic.

Activity

  1. added
    enhancementNew feature or request
    area:replicationReplication, cluster sync, peer connections
    feature:plugin-substratePlugin substrate primitives; Redis endpoint and job queue are the consumers.
    on Sep 21, 2026
  2. added theissue type on Sep 21, 2026
  3. cb1kenobi commented on Sep 21, 2026

    @cb1kenobi
    MemberAuthor

    From the planning review (Cursor Grok leg), adopted: forwarding needs a session model, not just a transport, and it is a precondition for calling this a product.

    Two things have no owner today:

    • The principal. AUTH alice on a non-owner, then a forwarded GET, runs on the owner with an empty user (every read authorization fails) or with the plugin's connection identity (alice's permissions never apply). Tracked with the authorization contract in Specify the authorization contract for protocol plugins: a read-only principal can write through an adapter harper#2726.
    • Session state: selected database, queued MULTI buffer, pub/sub mode. Keep it on the accepting node and a MULTI / SET / EXEC can run across different nodes. Keep it on the owner and every blocked client and subscriber holds owner-side resources.

    One sticky owner-side session per connection is the answer both point to, and it belongs in this issue's acceptance criteria rather than being discovered during implementation.

    The review also independently corroborated the home-map correction already recorded above, by a different route: on a non-replicated keyspace a non-owner either 503s, because there is no replicated log for the barrier to wait on, or it lock-arbitrates and then writes locally, which produces a second accepted copy of the same key. That is the invariant this epic exists to protect, so the correction stands on two reviews rather than one.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:replicationReplication, cluster sync, peer connectionsenhancementNew feature or requestfeature:plugin-substratePlugin substrate primitives; Redis endpoint and job queue are the consumers.

    Fields

    Priority

    P3

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions