Repository navigation
ROOTFS snapshot fidelity: micro-VM capture and restore, API opened for micro-VM templates - #2448
Conversation
Atelet refused ROOTFS in its checkpoint, restore, and paused-upload validation while the micro-VM runtime gains the ability to capture it. Which fidelities a sandbox runtime can serve is the runtime's knowledge, not atelet's: atelet now accepts every level of the ladder and forwards it, and each ateom rejects what it cannot capture. The paused-upload path still carried a conversion from the days when a template had separate pause and commit scopes: a MEMORY pause could be uploaded as a VOLUMES snapshot by dropping to the durable-volume tars. Pause and suspend now capture the template's one preferred fidelity, a paused actor's template cannot change, and golden actors are never paused, so the fidelity a pause recorded always equals the one the suspend asks for. Replace the conversion with that equality check and refuse any mismatch as drift, rather than guessing. The manifest still records the durable-volume subset ateom reported; nothing in atelet acts on it anymore.
A ROOTFS snapshot holds the durable volumes plus each container's root filesystem changes, without guest memory: the actor cold-boots on restore but keeps what it wrote to its rootfs. The micro-VM runtime already had every piece. A MEMORY checkpoint tars the host-side rootfs upper of each container alongside the Cloud Hypervisor snapshot, and a MEMORY restore untars those uppers before assembling the overlays. ROOTFS is the same checkpoint without the VM snapshot, and the same upper restore followed by the VOLUMES cold-boot path instead of a VMM relaunch. Checkpoint takes the upper tars for ROOTFS as well as MEMORY and needs no durable volume to have something to capture. Restore gains a ROOTFS arm that re-materializes the uppers and cold-boots on top of them; the cold boot learns to keep a pre-populated upper instead of wiping it, which holds across its retry too since a guest that never reached its agent wrote nothing. Only a MEMORY restore resumes the guest's CPU counters, so only it counts usage from the first reading; the lower fidelities count from zero. Which fidelities a runtime serves is now the runtime's declaration: the shared request validator takes the supported set, the micro-VM runtime passes all three, and gVisor passes VOLUMES and MEMORY because its checkpoint image holds memory and rootfs changes together. gVisor rejects a ROOTFS checkpoint or restore before taking the actor lock, naming the levels it does serve.
The micro-VM runtime now captures and restores ROOTFS snapshots, so the public API stops rejecting the level outright. What decides whether a template may ask for it is its sandbox class: gVisor keeps rootfs changes inside its memory checkpoint and cannot serve ROOTFS, while micro-VM can. The check moves from a field-level rule on preferred_fidelity to the template-level hook that already holds the other cross-field rules, and rejects ROOTFS unless sandbox_config.sandbox_class is SANDBOX_CLASS_MICROVM. An actor can only be repointed to a template with an identical sandbox config, so a ROOTFS template never ends up under a gVisor actor. The control plane needs no change: a ROOTFS snapshot restores at ROOTFS, and a repointed actor drops to VOLUMES as it does for MEMORY, since a rootfs upper is only valid over the image it was written on. Tests pin both, along with a pause recording ROOTFS. Docs and the micro-VM counter demo describe the level; the demo keeps MEMORY because its in-RAM counter is the point.
The lifecycle tests could not tell a ROOTFS snapshot from a VOLUMES one: the counter app kept one counter in memory and one in a durable volume, and both fidelities lose the first and keep the second. Give the app a third counter in a file on its root filesystem and print it in the response. MEMORY and ROOTFS snapshots carry it across pause and suspend, VOLUMES resets it, so every row now states what it expects of rootfs writes and the existing MEMORY and VOLUMES rows gain an assertion that those runtimes handle them as documented. Add micro-VM-only ROOTFS rows to the durable-dir lifecycle tests, plain and suspended from PAUSED: the memory counter restarts, the file and rootfs counters continue, and the suspend records a ROOTFS snapshot. The template-repoint test gains a ROOTFS row too, pinning that the repoint drops the restore to VOLUMES and resets the rootfs counter, that a pause under the new template restores at ROOTFS again, and that the revert cold-boots. gVisor cannot serve ROOTFS yet, so the new rows skip off micro-VM clusters.
38bb71f to
f3eee6c
Compare
| // container's rootfs upper: put those back before the overlays are | ||
| // assembled, then cold-boot on top of them instead of a pristine upper. | ||
| tUpper := time.Now() | ||
| if err := untarRootfsUpper(rootfsUpperDir(p.actorDirs), restoreDir, containerNames(p.containers)); err != nil { |
There was a problem hiding this comment.
🤖 should-fix 🟡 – untarRootfsUpper runs os.RemoveAll on the upper dir before cleanupSandboxState has dropped any leftover overlay mounts, which only happens later inside coldBootActor. teardownActor documents that the overlays must be unmounted before the upperdir is removed. The MEMORY and VOLUMES paths both clean up first. Consider calling s.cleanupSandboxState(ctx, p.actorUID) before the untar, as restoreMemoryFidelity does.
There was a problem hiding this comment.
Suggested fix: clean up once in the ROOTFS arm before the untar, matching the MEMORY path.
case ateompb.SnapshotFidelity_SNAPSHOT_FIDELITY_ROOTFS:
s.cleanupSandboxState(ctx, p.actorUID)
if err := untarRootfsUpper(...); err != nil { ... }There was a problem hiding this comment.
Fixed in a84e095, by moving the upper staging into coldBootActor rather than calling cleanup from the ROOTFS arm. The cold boot now runs cleanupSandboxState first (it moved ahead of buildActorContainers, which only touches the atelet-owned bundle dirs, not the sandbox dirs cleanup clears), and only then stages the upper, so the overlay mounts are gone before untarRootfsUpper removes the dir, matching restoreMemoryFidelity and teardownActor. A retry re-cleans and re-stages from the same snapshot dir, which also removes the "keep what is there" reasoning.
| // A ROOTFS restore has already re-materialized the upper and boots on it; | ||
| // that holds across a retry too, since a guest that never reached its agent | ||
| // wrote nothing and cleanupSandboxState only drops the overlay mounts. | ||
| if !p.keepRootfsUpper { |
There was a problem hiding this comment.
🤖 should-fix 🟡 – The new ROOTFS behavior in ateom-microvm has no unit test: the keepRootfsUpper skip here, the ROOTFS restore arm, and shipsRootfs in checkpoint. Only the micro-VM e2e rows cover it. Could the upper-dir decision and the per-fidelity file set be pulled into small helpers with table-driven tests, so a regression shows up without a KVM lane?
There was a problem hiding this comment.
Done in a84e095. The ladder is now two predicates in fidelity.go, includesRootfs and includesMemory, used by the checkpoint capture set and the restore resume decision, with TestFidelityLadder pinning all four values. The upper-dir decision is stageRootfsUpper(actorDirs, snapshotDir, containers) in rootfsupper.go: empty dir means pristine, otherwise re-materialize from the tars. TestStageRootfsUpper covers both branches on real directories, including that a stale previous activation is wiped either way. The keepRootfsUpper flag is gone; actorBootParams carries rootfsUpperSnapshotDir instead.
| // A ROOTFS snapshot holds no guest state either, but it carries each | ||
| // container's rootfs upper: put those back before the overlays are | ||
| // assembled, then cold-boot on top of them instead of a pristine upper. | ||
| tUpper := time.Now() |
There was a problem hiding this comment.
🤖 question 🟢 – This untar runs serially before the cold boot. restoreMemoryFidelity runs the same untar in a goroutine overlapped with bundle preparation. For an actor with a large upper, the ROOTFS resume pays the full untar time. Is overlapping it worth doing here too?
There was a problem hiding this comment.
Yes, and it is now overlapped. In a84e095 the staging runs as a background goroutine inside coldBootActor, started right after cleanupSandboxState and joined just before stageMergedRootfs, so it hides behind buildActorContainers and the resource-envelope check exactly as the MEMORY path hides its untar behind bundle preparation. Same drain-on-error discipline as there, so an early return never leaves a writer racing a retry. The separate rootfs_upper phase observation went away with the serial untar; the restore record now carries total like VOLUMES.
| // - Rootfs upper tars (ROOTFS and MEMORY): host-backed like the durable | ||
| // volumes — the memory snapshot does not carry rootfs writes. Under | ||
| // VOLUMES the workload cold-starts on restore, discarding rootfs state. | ||
| shipsRootfs := scope == ateompb.SnapshotFidelity_SNAPSHOT_FIDELITY_ROOTFS || |
There was a problem hiding this comment.
🤖 nit 🟢 – "This level includes rootfs" is spelled out separately here, in the switch above, and in resumesGuest in restore. A small helper such as includesRootfs(f) would keep the ladder logic in one place.
There was a problem hiding this comment.
Done: includesRootfs and includesMemory in fidelity.go are the single place the ladder is spelled out; checkpoint uses the first for its capture set and restore uses the second for resumesGuest. The precondition switch in checkpoint keeps its explicit cases since it is about what each level needs present, not what it includes.
| if value.GetSnapshotConfig().GetPreferredFidelity() == ateapipb.SnapshotFidelity_SNAPSHOT_FIDELITY_ROOTFS && | ||
| value.GetSandboxConfig().GetSandboxClass() != ateapipb.SandboxClass_SANDBOX_CLASS_MICROVM { |
There was a problem hiding this comment.
Self-note: we won't be able to check this once we move sandbox class to the workerpool.
There was a problem hiding this comment.
Agreed. Once the sandbox class lives on the WorkerPool the template no longer knows it at admission, so this check has to move to where the class is first known: scheduling or the resume workflow, rejecting a ROOTFS actor that lands on a gVisor pool. The runtime-side guard in ateom-gvisor stays as the backstop either way, so nothing can produce a ROOTFS checkpoint there in the meantime.
f3eee6c to
dc8b90e
Compare
The ROOTFS restore untarred the rootfs uppers before handing off to the cold boot, which only later cleared stale sandbox state. A previous activation's overlay mounts could therefore still sit on the upper dir when the untar removed it, the ordering teardown documents as unsafe. The untar also ran serially, so a large upper added its full extraction time to the resume, where the MEMORY path overlaps the same work with bundle preparation. Make the cold boot own the upper dir. It clears sandbox state first, moved ahead of bundle preparation, which touches only the atelet-owned bundle dirs, then stages the upper in the background and joins it right before the overlay mounts need it: pristine for a plain boot, re-materialized from the snapshot's tars when the boot parameters name a snapshot dir. A retry stages again from that dir, so nothing depends on what an earlier attempt left behind, and the ROOTFS restore arm shrinks to naming the dir. Spell the ladder out once: includesRootfs and includesMemory decide the checkpoint's capture set and whether a restore resumes the guest, and stageRootfsUpper is the one place a cold boot decides between a pristine and a restored upper. Both have table tests, so a regression in what each fidelity carries shows up without a KVM lane.
dc8b90e to
a84e095
Compare
Step III of #2102, tracked in #2447. Step I (#2309) removed
on_pause; step II (#2364) introduced theSnapshotFidelityladder with ROOTFS defined but rejected everywhere.ROOTFS becomes real on micro-VM. A ROOTFS snapshot holds the durable volumes plus each container's root-filesystem changes, without guest memory: the actor cold-boots on resume but keeps what it wrote to its rootfs. gVisor keeps rootfs changes inside its memory checkpoint and cannot serve it yet; it now declares that explicitly and a gVisor template that asks for ROOTFS is rejected at admission.
Four commits, meant to be read in order. Each builds and passes tests on its own, and ROOTFS stays unreachable from the public API until the third.
1. atelet: forward ROOTFS and stop narrowing paused uploads
Atelet accepts every level of the ladder and forwards it; which levels a runtime serves is the runtime's knowledge. The paused-upload path still converted a MEMORY pause into a VOLUMES upload, a leftover from the days of separate pause and commit scopes. Pause and suspend now capture one fidelity, a paused actor's template cannot change, and golden actors are never paused, so captured and requested fidelity are always equal; the conversion is replaced by an equality check that refuses any mismatch as drift.
2. ateom-microvm: capture and restore ROOTFS
Every piece already existed: MEMORY checkpoints tar the host-side rootfs upper per container alongside the Cloud Hypervisor snapshot, and MEMORY restores untar those uppers before assembling the overlays. ROOTFS is the same checkpoint without the VM snapshot, and the upper restore followed by the VOLUMES cold-boot path with a new
keepRootfsUpperboot parameter so the cold boot does not wipe the restored upper (safe across its retry). Only a MEMORY restore resumes the guest's CPU counters; ROOTFS counts from zero like VOLUMES.resources.ValidateSnapshotFidelitynow takes the runtime's supported set: micro-VM passes all three, gVisor passes VOLUMES and MEMORY and rejects ROOTFS before taking the actor lock, naming the levels it serves.3. ateapi: allow ROOTFS on micro-VM templates
The field-level rejection moves to the template-level cross-field hook: ROOTFS requires
sandbox_config.sandbox_class: SANDBOX_CLASS_MICROVM. An actor can only be repointed to a template with an identical sandbox config, so a ROOTFS template never ends up under a gVisor actor. The control plane needs no logic change: a ROOTFS snapshot restores at ROOTFS, and a repointed actor drops to VOLUMES as it does for MEMORY, since a rootfs upper is only valid over the image it was written on.4. e2e: prove it on the counter demo
The counter app gains a third counter in a file on its root filesystem, printed in the response, so a ROOTFS snapshot is distinguishable from VOLUMES. All lifecycle rows state their rootfs expectation (MEMORY and ROOTFS keep it, VOLUMES resets it). New micro-VM-only ROOTFS rows cover pause, suspend, suspend from PAUSED, and the template repoint and revert. On the gVisor CI lane the ROOTFS rows skip; the new rootfs assertions on MEMORY and VOLUMES rows run there as a side check.
Behavior changes
SNAPSHOT_FIDELITY_ROOTFSis accepted on micro-VM templates and rejected on gVisor ones with a message naming the requirement. No field or wire changes; thepreferred_fidelityfield-level custom validation tag is removed.FailedPrecondition; previously a MEMORY capture could be downgraded to VOLUMES. Unreachable in practice, as explained in commit 1.rootfswas already a member).Not in this PR
required_fidelity,ResumeActorResponse.resumed_fidelity, the resume fallback rule, gVisor ROOTFS (pending the gVisor team's separable rootfs capture), and restoring a MEMORY snapshot as a ROOTFS cold boot (the artifacts already allow it; atelet would need the rootfs file subset in the manifest to skip the memory image).Verification
Unit tests for atelet, ateom-microvm, ateom-gvisor, resources, ateattr, ateomphaselog, ateapi (validation, controlapi, functional), kubectl-ate, and apitool pass; the ateom tests are linux-tagged and were type-checked for linux locally.
golangci-lint, kube-api-linter, python-protos, boilerplate pass; codegen shows no drift.
The micro-VM counter demo was run manually on a GKE cluster with a nested-virtualization node pool and behaved as expected. The ROOTFS e2e rows need a micro-VM lane; the gVisor lane runs the rest.
Breaking change
Not backward compatible
🤖 Generated with Claude Code