Skip to content

feat: orchestrate etcd downgrade across minor version boundaries - #92

Draft
louiseschmidtgen wants to merge 2 commits into
mainfrom
etcd-downgrade-safety
Draft

louiseschmidtgen wants to merge 2 commits into
mainfrom
etcd-downgrade-safety

Conversation

@louiseschmidtgen

Copy link
Copy Markdown
Contributor

Summary

Downgrading the k8s snap across an etcd minor version boundary (e.g. 1.37 -> 1.36, etcd 3.7 -> 3.6) left etcd permanently crash-looping: the older etcd binary refuses to start against a data directory stamped with a newer storage version, and nothing ran etcd's required downgrade validate / downgrade enable protocol before the binary was swapped.

This PR adds two complementary mechanisms to make such downgrades safe.

1. k8s x-etcd prepare-downgrade (pre-refresh preparation)

A new hidden CLI command intended to be called from the snap pre-refresh hook, while the current (newer) etcd binary is still running:

  • reads the target revision's etcd version from the mounted revision's bom.json
  • if the target downgrades etcd by one minor version, runs downgrade validate + downgrade enable and waits for the local member's storage version to migrate
  • idempotent (ErrDowngradeInProcess is treated as success), so concurrent invocations on multiple nodes refreshing at the same time are safe

2. Startup recovery (fallback)

In the k8sd onStart hook, before services are started:

  • detect whether the etcd data directory was written by a newer etcd than the bundled binary, by reading the storage version directly from the bbolt backend (meta bucket, storageVersion key) — no running etcd needed
  • if so (e.g. after snap revert, which runs no refresh hooks, or a refresh from a revision without the pre-refresh preparation), start a compatible etcd binary from a previous snap revision, complete the downgrade protocol, then start the bundled etcd

Verification

  • Unit tests for version parsing/comparison, bbolt storage-version detection, and the validate/enable/wait flow (with a fake etcd client).
  • End-to-end with real etcd 3.7.1 and 3.6.13 binaries (in a Linux container):
    • confirmed the bug: 3.6.13 refuses to start on a 3.7 data dir (version "3.7.0" is not supported)
    • confirmed the pre-refresh protocol (validate + enable + wait) lets 3.6.13 start cleanly
    • confirmed the startup recovery (find previous-revision 3.7 binary, run protocol, swap to 3.6) end-to-end

Companion k8s-snap PR wires the pre-refresh hook and adds an integration test.

Fixes #89


Note: go build ./... / the full test suite require the dqlite CGO toolchain (Linux); verified via go vet and targeted builds in a Linux container. The pre-existing pkg/snap test build failure (dqlite headers) is unrelated and present on main.

Downgrading the k8s snap across an etcd minor version boundary (e.g. 1.37 ->
1.36, etcd 3.7 -> 3.6) previously left etcd crash-looping: the older etcd
binary refuses to start against a data directory stamped with a newer storage
version, and nothing ran etcd's required `downgrade validate`/`downgrade
enable` protocol before the binary was swapped.

Add two complementary mechanisms:

* `k8s x-etcd prepare-downgrade --target-revision <rev>`: intended to be called
  from the snap pre-refresh hook while the current (newer) etcd binary is still
  running. It validates and enables the cluster downgrade to the target
  revision's etcd version (read from the mounted revision's bom.json) and waits
  for the local member's storage version to migrate. The operation is
  idempotent so concurrent invocations on multiple nodes are safe.

* Startup recovery in the k8sd onStart hook: before starting services, detect
  (by reading the storage version directly from the etcd bbolt backend) whether
  the data directory was written by a newer etcd than the bundled binary. If so
  - e.g. after `snap revert`, which runs no refresh hooks, or a refresh from a
  revision without the pre-refresh preparation - start a compatible etcd binary
  from a previous snap revision, complete the downgrade protocol, and only then
  start the bundled etcd.

Fixes #89

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: louiseschmidtgen <louise.schmidtgen@canonical.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: louiseschmidtgen <louise.schmidtgen@canonical.com>
louiseschmidtgen added a commit to canonical/k8s-snap that referenced this pull request Sep 17, 2026
Temporarily point the k8sd component at the companion k8sd PR branch so the
integration tests exercise the etcd downgrade fix end-to-end. Revert to 'main'
before merging (after canonical/k8sd#92 lands).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: louiseschmidtgen <louise.schmidtgen@canonical.com>
@bschimke95

Copy link
Copy Markdown
Collaborator

@louiseschmidtgen should we close this?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

bug: downgrading across an etcd minor version crashes etcd - missing etcdctl downgrade enable orchestration

2 participants