Repository navigation
feat: orchestrate etcd downgrade across minor version boundaries - #92
Draft
louiseschmidtgen wants to merge 2 commits into
Draft
louiseschmidtgen wants to merge 2 commits into
louiseschmidtgen wants to merge 2 commits into
Conversation
Downgrading the k8s snap across an etcd minor version boundary (e.g. 1.37 -> 1.36, etcd 3.7 -> 3.6) previously left etcd crash-looping: the older etcd binary refuses to start against a data directory stamped with a newer storage version, and nothing ran etcd's required `downgrade validate`/`downgrade enable` protocol before the binary was swapped. Add two complementary mechanisms: * `k8s x-etcd prepare-downgrade --target-revision <rev>`: intended to be called from the snap pre-refresh hook while the current (newer) etcd binary is still running. It validates and enables the cluster downgrade to the target revision's etcd version (read from the mounted revision's bom.json) and waits for the local member's storage version to migrate. The operation is idempotent so concurrent invocations on multiple nodes are safe. * Startup recovery in the k8sd onStart hook: before starting services, detect (by reading the storage version directly from the etcd bbolt backend) whether the data directory was written by a newer etcd than the bundled binary. If so - e.g. after `snap revert`, which runs no refresh hooks, or a refresh from a revision without the pre-refresh preparation - start a compatible etcd binary from a previous snap revision, complete the downgrade protocol, and only then start the bundled etcd. Fixes #89 Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: louiseschmidtgen <louise.schmidtgen@canonical.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: louiseschmidtgen <louise.schmidtgen@canonical.com>
louiseschmidtgen
added a commit
to canonical/k8s-snap
that referenced
this pull request
Sep 17, 2026
Temporarily point the k8sd component at the companion k8sd PR branch so the integration tests exercise the etcd downgrade fix end-to-end. Revert to 'main' before merging (after canonical/k8sd#92 lands). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: louiseschmidtgen <louise.schmidtgen@canonical.com>
Collaborator
|
@louiseschmidtgen should we close this? |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Downgrading the k8s snap across an etcd minor version boundary (e.g.
1.37 -> 1.36, etcd3.7 -> 3.6) left etcd permanently crash-looping: the older etcd binary refuses to start against a data directory stamped with a newer storage version, and nothing ran etcd's requireddowngrade validate/downgrade enableprotocol before the binary was swapped.This PR adds two complementary mechanisms to make such downgrades safe.
1.
k8s x-etcd prepare-downgrade(pre-refresh preparation)A new hidden CLI command intended to be called from the snap
pre-refreshhook, while the current (newer) etcd binary is still running:bom.jsondowngrade validate+downgrade enableand waits for the local member's storage version to migrateErrDowngradeInProcessis treated as success), so concurrent invocations on multiple nodes refreshing at the same time are safe2. Startup recovery (fallback)
In the k8sd
onStarthook, before services are started:metabucket,storageVersionkey) — no running etcd neededsnap revert, which runs no refresh hooks, or a refresh from a revision without the pre-refresh preparation), start a compatible etcd binary from a previous snap revision, complete the downgrade protocol, then start the bundled etcdVerification
3.7.1and3.6.13binaries (in a Linux container):3.6.13refuses to start on a3.7data dir (version "3.7.0" is not supported)validate+enable+ wait) lets3.6.13start cleanly3.7binary, run protocol, swap to3.6) end-to-endCompanion k8s-snap PR wires the
pre-refreshhook and adds an integration test.Fixes #89
Note:
go build ./.../ the full test suite require the dqlite CGO toolchain (Linux); verified viago vetand targeted builds in a Linux container. The pre-existingpkg/snaptest build failure (dqlite headers) is unrelated and present onmain.