Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
145 changes: 123 additions & 22 deletions pages/clustering/high-availability/ha-commands-reference.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -20,12 +20,21 @@ setup, the choice no longer matters.
</Callout>

<Callout type="info">
All queries can be run on any coordinator. If currently the coordinator is not a
leader, the query will be automatically forwarded to the current leader and
executed there. This is because the Raft protocol specifies that only the
leader should accept changes in the cluster. The only exception is [`YIELD
LEADERSHIP`](#yield-leadership), which must be run directly on the current
leader.
All queries can be run on any coordinator. If the coordinator you are connected
to is not the leader, the query is automatically forwarded to the current leader
and executed there. This is because the Raft protocol specifies that only the
leader should accept changes in the cluster, and because only the leader holds an
up-to-date view of the cluster.

This also holds for the read-only cluster queries — [`SHOW
INSTANCES`](#show-instances), [`SHOW COORDINATOR
SETTINGS`](#coordinator-runtime-settings) and [`SHOW REPLICATION
LAG`](#show-replication-lag) — which are always answered by the leader, and for
[`YIELD LEADERSHIP`](#yield-leadership). Followers never answer them from their
own local state, so you never see a stale or partial picture of the cluster. If
the leader cannot be reached, these queries return **no rows** together with a
warning notification instead of a degraded result — see [Error
handling](#error-handling).
</Callout>

### `ADD COORDINATOR`
Expand Down Expand Up @@ -88,8 +97,8 @@ REMOVE COORDINATOR coordinatorId;

- Leader coordinator **cannot** remove itself. To remove the leader, first
trigger a leadership change with [`YIELD
LEADERSHIP`](#yield-leadership) and then run `REMOVE COORDINATOR` against the
new leader.
LEADERSHIP`](#yield-leadership) and then run `REMOVE COORDINATOR` once a new
leader has been elected.

{<h4 className="custom-header"> Example </h4>}

Expand Down Expand Up @@ -313,16 +322,17 @@ informational notification that the request was submitted.

{<h4 className="custom-header"> Behavior </h4>}

- Must be run **on the coordinator that is currently the leader**. Unlike most
other cluster management queries, this query is **not forwarded** to the
leader — running it on a follower fails with:

> Only the current leader can yield the leadership!

- Can be run on **any coordinator**. If the coordinator is a follower, the query
is forwarded to the current leader, which yields its leadership. You no longer
have to find the leader first.
- Running it on a data instance fails with:

> Only coordinator can run YIELD LEADERSHIP query.

- A coordinator that Raft already elected as leader but that has not yet
finished taking over the cluster still yields its leadership. This makes the
query usable as an escape hatch exactly when it is needed most — when a
freshly elected leader is stuck and you want another coordinator to take over.
- The request is handed over to Raft and processed **asynchronously**. The query
returns as soon as the request is submitted, not when the new leader is
elected.
Expand All @@ -334,15 +344,24 @@ informational notification that the request was submitted.
checks toward all data instances. If no MAIN is found at that point, it
performs a failover.

{<h4 className="custom-header"> Failure modes </h4>}

| Error message | Meaning |
| ------------- | ------- |
| `Yielding leadership failed since the instance is not leader anymore!` | The request reached a coordinator that is no longer the leader — leadership changed in the meantime. Retry. |
| `Tried to forward the request to the current leader but the leader couldn't be found!` | There is currently no known leader to forward the request to (for example, an election is in progress). Retry once a leader is elected. |
| `Request forwarded to the leader but leader failed with request processing! Check logs on the leader to find out what happened!` | The leader was reached but failed to process the request. Inspect the leader's logs. |

{<h4 className="custom-header"> Implications </h4>}

- This changes only the **coordinator** leadership. Data instances keep their
MAIN and REPLICA roles — this is not a data failover, and client queries
against MAIN and REPLICAs are unaffected.
- During the short election window, cluster management queries (e.g. `SHOW
INSTANCES`, registration queries) may temporarily fail or report instances as
`down` because there is no leader to serve them. Retry once the new leader is
elected.
INSTANCES`, registration queries) may temporarily fail because there is no
leader to serve them. Read queries such as `SHOW INSTANCES` return no rows and
a warning notification rather than a stale picture of the cluster. Retry once
the new leader is elected.
- At least one other healthy coordinator must be able to take over. In a
single-coordinator cluster, or when the other coordinators are down, the same
coordinator remains (or becomes again) the leader.
Expand Down Expand Up @@ -386,14 +405,42 @@ SHOW INSTANCES;
{<h4 className="custom-header"> Output includes </h4>}

1. Network endpoints (bolt, coordinator, management)
2. Health state
3. Role: MAIN, REPLICA, LEADER, FOLLOWER, or UNKNOWN
2. Health state (`up` or `down`)
3. Role: MAIN, REPLICA, LEADER or FOLLOWER
4. Time since last health ping

{<h4 className="custom-header"> Behavior on followers </h4>}
A data instance that is currently `down` keeps the role recorded in the Raft log
(`main` or `replica`) instead of being reported with an unknown role, so you can
still tell which instance the cluster considers MAIN while it is unreachable.

{<h4 className="custom-header"> Behavior </h4>}

The query is **strongly consistent**: the result always comes from the leader
coordinator, which is the only coordinator with an up-to-date view of the
cluster.

1. If you are connected to the leader, it answers directly.
2. If you are connected to a follower, the follower forwards the request to the
leader and returns the leader's result.
3. If the leader cannot be reached, the query returns an **empty result set**
together with a `LeaderNotReachable` warning notification:

1. Follower attempts to query the leader for accurate state.
2. If leader unavailable, follower reports all servers as `"down"`.
> Couldn't reach the leader coordinator, so the state of the cluster is
> unknown. Please retry the query.

This happens when no leader is currently elected, when the connection to the
leader is broken, or when the coordinator you are connected to was just
elected leader but has not finished taking over the cluster yet.

<Callout type="warning">
**Behavior change in Memgraph 3.13:** previously, a follower that could not
reach the leader fell back to reporting the cluster from its own local Raft
state, with health reported as `unknown`. It now returns no rows and a warning
notification instead, so a partial or stale cluster picture can never be
mistaken for the real one. If you have tooling or health checks that parse
`SHOW INSTANCES`, treat an empty result as "cluster state unknown, retry" rather
than as "no instances registered".
</Callout>


### `SHOW INSTANCE`
Expand Down Expand Up @@ -423,6 +470,23 @@ Shows replication lag (in committed transactions) for all instances.
SHOW REPLICATION LAG;
```

{<h4 className="custom-header"> Behavior </h4>}

The lag data is collected by the leader coordinator from the current MAIN, so the
query is answered by the leader — a follower forwards the request and returns the
leader's result.

Whenever the lag cannot be determined, the query returns **no rows** and a
warning notification explaining why, so you know whether retrying will help:

| Notification code | Message | Meaning |
| ----------------- | ------- | ------- |
| `LeaderNotReachable` | Couldn't reach the leader coordinator, so the replication lag is unknown. Please retry the query. | No leader could be contacted (for example, an election is in progress). |
| `ReplicationLagUnavailable` | The leader coordinator hasn't finished taking over the cluster, so the replication lag is unknown. Please retry the query. | A new leader was elected but has not finished reconciling the cluster. |
| `ReplicationLagUnavailable` | No instance is currently main, so there is no replication lag to report. | The cluster has no MAIN — promote one with `SET INSTANCE ... TO MAIN`, or wait for failover. |
| `ReplicationLagUnavailable` | The current main didn't respond, so the replication lag is unknown. Check whether the main is up. | The MAIN did not answer the leader's request. |
| `ReplicationLagUnavailable` | The instance the leader considers main reports that it is a replica, so the replication lag is unknown. Please retry the query once the cluster state is reconciled. | The leader's view is stale; the [reconciliation loop](/clustering/high-availability/how-high-availability-works#how-the-reconciliation-loop-works) will fix it. |

{<h4 className="custom-header"> Implications </h4>}

- Lag values survive restarts (stored in snapshots + WAL).
Expand Down Expand Up @@ -490,6 +554,14 @@ cluster without downtime. Use `SET COORDINATOR SETTING` to modify a value and
`SHOW COORDINATOR SETTINGS` to inspect all current values. Changes propagate
automatically to every coordinator in the cluster.

Both queries can be run on any coordinator and are served by the leader — a
follower forwards the request and returns the leader's answer. If the leader
cannot be reached, `SHOW COORDINATOR SETTINGS` returns **no rows** together with a
`LeaderNotReachable` warning notification:

> Couldn't reach the leader coordinator, so the coordinator settings are unknown.
> Please retry the query.

### `instance_health_check_frequency_sec`

How often the coordinator pings data instances, in seconds.
Expand Down Expand Up @@ -653,6 +725,35 @@ promote, demote, add coordinator), the error message will indicate:

> Writing to Raft log failed. Please retry the operation.

### When there is no leader to serve the query

Because every cluster query is served by the leader, queries fail (or return
nothing) while the cluster has no usable leader — most commonly during a leader
election, or right after one, while the new leader is still taking over the
cluster.

State-changing queries (`ADD COORDINATOR`, `REMOVE COORDINATOR`, `UPDATE
CONFIG`, `REGISTER INSTANCE`, `UNREGISTER INSTANCE`, `SET INSTANCE ... TO MAIN`,
`DEMOTE INSTANCE`, `FORCE RESET CLUSTER STATE`) fail with an explicit error:

> Couldn't &lt;operation&gt; since coordinator is not a leader! Try contacting
> other coordinators as there might be leader election happening or other
> coordinators are down.

When the leader is known but the request still could not be executed there, the
message instead names the current leader's id and Bolt address so you can
connect to it directly. If the request could not be forwarded at all, the error
is:

> Tried to forward the request to the current leader but the leader couldn't be
> found!

Read queries (`SHOW INSTANCES`, `SHOW COORDINATOR SETTINGS`, `SHOW REPLICATION
LAG`) do not fail — they return an **empty result set** with a warning
notification (`LeaderNotReachable`, or `ReplicationLagUnavailable` for `SHOW
REPLICATION LAG`). In both cases the operation is safe to retry once a leader is
available.


## Troubleshooting commands

Expand Down
54 changes: 39 additions & 15 deletions pages/clustering/high-availability/how-high-availability-works.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -161,6 +161,8 @@ Below is a cleaned-up categorization.
| RPC | Purpose | Description |
| ------------------------ | -------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------- |
| `ShowInstancesRpc` | Follower requests cluster state from leader. | Sent by a follower coordinator to the leader coordinator when a user executes `SHOW INSTANCES` through the follower. |
| `ShowCoordSettingsRpc` | Follower requests coordinator settings. | Sent by a follower coordinator to the leader coordinator when a user executes `SHOW COORDINATOR SETTINGS` through the follower. |
| `YieldLeadershipRpc` | Follower asks the leader to step down. | Sent by a follower coordinator to the leader coordinator when a user executes `YIELD LEADERSHIP` through the follower. |
| `AddCoordinatorRpc` | Follower requests adding coordinator. | Sent by a follower coordinator to the leader coordinator when a user executes `ADD COORDINATOR` through the follower. |
| `RemoveCoordinatorRpc` | Follower requests removing coordinator. | Sent by a follower coordinator to the leader coordinator when a user executes `REMOVE COORDINATOR` through the follower. |
| `RegisterInstanceRpc` | Follower requests registering an instance. | Sent by a follower coordinator to the leader coordinator when a user executes `REGISTER INSTANCE` through the follower. |
Expand Down Expand Up @@ -546,21 +548,43 @@ modes](/clustering/replication/how-replication-works#replication-modes).

## Actions on follower coordinators

Follower coordinators operate in a restricted mode.
They can **only execute** the `SHOW INSTANCES` command.

All state-changing operations are disabled on followers, including:

- Registering data instances
- Unregistering data instances
- Demoting an instance
- Promoting an instance to MAIN
- Forcing a cluster state reset
- Yielding leadership

These operations are permitted **only on the leader coordinator**. Note that
[`YIELD LEADERSHIP`](/clustering/high-availability/ha-commands-reference#yield-leadership)
is not forwarded to the leader — it fails on a follower instead.
Follower coordinators never execute cluster operations themselves. Instead, they
act as a transparent entry point: every cluster query you run on a follower is
**forwarded to the current leader**, executed there, and the leader's answer is
returned to you. This holds both for state-changing operations (registering and
unregistering data instances, promoting and demoting instances, adding and
removing coordinators, updating configuration, forcing a cluster state reset,
[yielding
leadership](/clustering/high-availability/ha-commands-reference#yield-leadership))
and for read-only ones (`SHOW INSTANCES`, `SHOW COORDINATOR SETTINGS`, `SHOW
REPLICATION LAG`, and `bolt+routing` routing table requests).

As a result you can point your tooling at any coordinator without first having to
discover which one is the leader.

### Why followers never answer from local state

A follower's own Raft state is not enough to describe the cluster: health of the
data instances is only known to the leader, which is the coordinator that pings
them. For this reason followers do not fall back to a local, partial answer when
the leader cannot be reached. Instead:

- Read queries return an **empty result set** together with a warning
notification (`LeaderNotReachable`, or `ReplicationLagUnavailable` for `SHOW
REPLICATION LAG`).
- State-changing queries fail with an error telling you that the coordinator is
not the leader, or that the leader could not be found.

Both cases mean "cluster state unknown — retry", and both are expected during the
brief window of a leader election or while a newly elected leader is still taking
over the cluster. See [Error
handling](/clustering/high-availability/ha-commands-reference#when-there-is-no-leader-to-serve-the-query)
for the exact messages.

The one exception is the `bolt+routing` routing table: a coordinator that Raft
elected as leader answers routing requests from its own Raft state even before it
has finished taking over the cluster, so that clients can keep routing queries
during the leadership transition.

## Raft-first operations and the reconciliation loop

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -64,6 +64,8 @@ The routing protocol works as follows:
receives the latest routing table.
- If it contacts a **follower**, the follower forwards the request to the leader
and returns the leader’s result.
- If the leader cannot be reached at all, an **empty routing table** is returned
and the driver retries against another coordinator.

Because leader state is synchronized via Raft, routing information is always
accurate.
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -202,7 +202,8 @@ SET INSTANCE instance_3 TO MAIN;

## Check cluster state

Connect to the leader coordinator and check cluster state with `SHOW INSTANCES`;
Connect to any coordinator and check cluster state with `SHOW INSTANCES`. The
query is always answered by the leader, so followers report the same state:

| name | bolt_server | coordinator_server | management_server | health | role | last_succ_resp_ms |
| ------------- | -------------- | ------------------ | ----------------- | ------ | -------- | ---------------- |
Expand Down
20 changes: 20 additions & 0 deletions pages/release-notes.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -58,6 +58,15 @@ guide.
you need the previous Python behavior, or update consumers for the new
structured / ISO-8601 JSON forms.
[#4443](https://github.com/memgraph/memgraph/pull/4443)
- `SHOW INSTANCES` is now strongly consistent and always answered by the leader
coordinator. A follower that cannot reach the leader no longer falls back to
reporting the cluster from its own local Raft state with `unknown` health —
it returns an empty result set with a `LeaderNotReachable` warning
notification instead. Tooling that parses `SHOW INSTANCES` should treat an
empty result as "cluster state unknown, retry" rather than "no instances
registered". A data instance that is down now reports the role recorded in the
Raft log (`main` / `replica`) instead of an unknown role.
[#4492](https://github.com/memgraph/memgraph/pull/4492)
- `TERMINATE TRANSACTIONS` now requires transaction ids to parse in full. Ids
with trailing characters previously terminated the transaction matching the
numeric prefix, and unparseable ids were reported back as an attempt on
Expand Down Expand Up @@ -137,6 +146,17 @@ guide.
statements). Under index-heavy workloads this removes GC-correlated latency
spikes on index and constraint creation.
[#4468](https://github.com/memgraph/memgraph/pull/4468)
- `YIELD LEADERSHIP` and `SHOW COORDINATOR SETTINGS` can now be run on any
coordinator — followers forward them to the leader instead of failing or
answering from local state. When no leader can serve a cluster query,
Memgraph now says so explicitly: read queries (`SHOW INSTANCES`, `SHOW
COORDINATOR SETTINGS`, `SHOW REPLICATION LAG`) return no rows with a
`LeaderNotReachable` warning, `SHOW REPLICATION LAG` additionally reports why
the lag is unavailable via `ReplicationLagUnavailable` (no current main, main
unresponsive, leader still taking over, stale leader view), and
`ADD COORDINATOR`, `REMOVE COORDINATOR` and `UPDATE CONFIG` fail with a clear
"coordinator is not a leader" error instead of a misleading one.
[#4492](https://github.com/memgraph/memgraph/pull/4492)

{<h4 className="custom-header">✨ New features</h4>}

Expand Down