Skip to content

ROOTFS snapshot fidelity: micro-VM capture and restore, API opened for micro-VM templates - #2448

Merged
Zoe Zhao (zoez7) merged 5 commits into
agent-substrate:mainfrom
dberkov:rootfs-fidelity
Oct 11, 2026
Merged

Zoe Zhao (zoez7) merged 5 commits into
agent-substrate:mainfrom
dberkov:rootfs-fidelity

Conversation

@dberkov

Copy link
Copy Markdown
Collaborator

Step III of #2102, tracked in #2447. Step I (#2309) removed on_pause; step II (#2364) introduced the SnapshotFidelity ladder with ROOTFS defined but rejected everywhere.

ROOTFS becomes real on micro-VM. A ROOTFS snapshot holds the durable volumes plus each container's root-filesystem changes, without guest memory: the actor cold-boots on resume but keeps what it wrote to its rootfs. gVisor keeps rootfs changes inside its memory checkpoint and cannot serve it yet; it now declares that explicitly and a gVisor template that asks for ROOTFS is rejected at admission.

Four commits, meant to be read in order. Each builds and passes tests on its own, and ROOTFS stays unreachable from the public API until the third.

1. atelet: forward ROOTFS and stop narrowing paused uploads

Atelet accepts every level of the ladder and forwards it; which levels a runtime serves is the runtime's knowledge. The paused-upload path still converted a MEMORY pause into a VOLUMES upload, a leftover from the days of separate pause and commit scopes. Pause and suspend now capture one fidelity, a paused actor's template cannot change, and golden actors are never paused, so captured and requested fidelity are always equal; the conversion is replaced by an equality check that refuses any mismatch as drift.

2. ateom-microvm: capture and restore ROOTFS

Every piece already existed: MEMORY checkpoints tar the host-side rootfs upper per container alongside the Cloud Hypervisor snapshot, and MEMORY restores untar those uppers before assembling the overlays. ROOTFS is the same checkpoint without the VM snapshot, and the upper restore followed by the VOLUMES cold-boot path with a new keepRootfsUpper boot parameter so the cold boot does not wipe the restored upper (safe across its retry). Only a MEMORY restore resumes the guest's CPU counters; ROOTFS counts from zero like VOLUMES.

resources.ValidateSnapshotFidelity now takes the runtime's supported set: micro-VM passes all three, gVisor passes VOLUMES and MEMORY and rejects ROOTFS before taking the actor lock, naming the levels it serves.

3. ateapi: allow ROOTFS on micro-VM templates

The field-level rejection moves to the template-level cross-field hook: ROOTFS requires sandbox_config.sandbox_class: SANDBOX_CLASS_MICROVM. An actor can only be repointed to a template with an identical sandbox config, so a ROOTFS template never ends up under a gVisor actor. The control plane needs no logic change: a ROOTFS snapshot restores at ROOTFS, and a repointed actor drops to VOLUMES as it does for MEMORY, since a rootfs upper is only valid over the image it was written on.

4. e2e: prove it on the counter demo

The counter app gains a third counter in a file on its root filesystem, printed in the response, so a ROOTFS snapshot is distinguishable from VOLUMES. All lifecycle rows state their rootfs expectation (MEMORY and ROOTFS keep it, VOLUMES resets it). New micro-VM-only ROOTFS rows cover pause, suspend, suspend from PAUSED, and the template repoint and revert. On the gVisor CI lane the ROOTFS rows skip; the new rootfs assertions on MEMORY and VOLUMES rows run there as a side check.

Behavior changes

  • API: SNAPSHOT_FIDELITY_ROOTFS is accepted on micro-VM templates and rejected on gVisor ones with a message naming the requirement. No field or wire changes; the preferred_fidelity field-level custom validation tag is removed.
  • Atelet: a paused-upload whose desired fidelity differs from the captured one is now FailedPrecondition; previously a MEMORY capture could be downgraded to VOLUMES. Unreachable in practice, as explained in commit 1.
  • Unchanged: MEMORY and VOLUMES capture and restore, the golden-actor MEMORY override, metric labels (rootfs was already a member).
  • No backward compatibility is kept, per the pre-v1.0 policy.

Not in this PR

required_fidelity, ResumeActorResponse.resumed_fidelity, the resume fallback rule, gVisor ROOTFS (pending the gVisor team's separable rootfs capture), and restoring a MEMORY snapshot as a ROOTFS cold boot (the artifacts already allow it; atelet would need the rootfs file subset in the manifest to skip the memory image).

Verification

  • Unit tests for atelet, ateom-microvm, ateom-gvisor, resources, ateattr, ateomphaselog, ateapi (validation, controlapi, functional), kubectl-ate, and apitool pass; the ateom tests are linux-tagged and were type-checked for linux locally.

  • golangci-lint, kube-api-linter, python-protos, boilerplate pass; codegen shows no drift.

  • The micro-VM counter demo was run manually on a GKE cluster with a nested-virtualization node pool and behaved as expected. The ROOTFS e2e rows need a micro-VM lane; the gVisor lane runs the rest.

  • Breaking change

  • Not backward compatible

🤖 Generated with Claude Code

Atelet refused ROOTFS in its checkpoint, restore, and paused-upload
validation while the micro-VM runtime gains the ability to capture it.
Which fidelities a sandbox runtime can serve is the runtime's knowledge,
not atelet's: atelet now accepts every level of the ladder and forwards it,
and each ateom rejects what it cannot capture.

The paused-upload path still carried a conversion from the days when a
template had separate pause and commit scopes: a MEMORY pause could be
uploaded as a VOLUMES snapshot by dropping to the durable-volume tars.
Pause and suspend now capture the template's one preferred fidelity, a
paused actor's template cannot change, and golden actors are never paused,
so the fidelity a pause recorded always equals the one the suspend asks for.
Replace the conversion with that equality check and refuse any mismatch as
drift, rather than guessing. The manifest still records the durable-volume
subset ateom reported; nothing in atelet acts on it anymore.
A ROOTFS snapshot holds the durable volumes plus each container's root
filesystem changes, without guest memory: the actor cold-boots on restore
but keeps what it wrote to its rootfs. The micro-VM runtime already had
every piece. A MEMORY checkpoint tars the host-side rootfs upper of each
container alongside the Cloud Hypervisor snapshot, and a MEMORY restore
untars those uppers before assembling the overlays. ROOTFS is the same
checkpoint without the VM snapshot, and the same upper restore followed by
the VOLUMES cold-boot path instead of a VMM relaunch.

Checkpoint takes the upper tars for ROOTFS as well as MEMORY and needs no
durable volume to have something to capture. Restore gains a ROOTFS arm
that re-materializes the uppers and cold-boots on top of them; the cold
boot learns to keep a pre-populated upper instead of wiping it, which holds
across its retry too since a guest that never reached its agent wrote
nothing. Only a MEMORY restore resumes the guest's CPU counters, so only it
counts usage from the first reading; the lower fidelities count from zero.

Which fidelities a runtime serves is now the runtime's declaration: the
shared request validator takes the supported set, the micro-VM runtime
passes all three, and gVisor passes VOLUMES and MEMORY because its
checkpoint image holds memory and rootfs changes together. gVisor rejects a
ROOTFS checkpoint or restore before taking the actor lock, naming the
levels it does serve.
The micro-VM runtime now captures and restores ROOTFS snapshots, so the
public API stops rejecting the level outright. What decides whether a
template may ask for it is its sandbox class: gVisor keeps rootfs changes
inside its memory checkpoint and cannot serve ROOTFS, while micro-VM can.
The check moves from a field-level rule on preferred_fidelity to the
template-level hook that already holds the other cross-field rules, and
rejects ROOTFS unless sandbox_config.sandbox_class is SANDBOX_CLASS_MICROVM.
An actor can only be repointed to a template with an identical sandbox
config, so a ROOTFS template never ends up under a gVisor actor.

The control plane needs no change: a ROOTFS snapshot restores at ROOTFS,
and a repointed actor drops to VOLUMES as it does for MEMORY, since a rootfs
upper is only valid over the image it was written on. Tests pin both, along
with a pause recording ROOTFS. Docs and the micro-VM counter demo describe
the level; the demo keeps MEMORY because its in-RAM counter is the point.
The lifecycle tests could not tell a ROOTFS snapshot from a VOLUMES one:
the counter app kept one counter in memory and one in a durable volume,
and both fidelities lose the first and keep the second. Give the app a
third counter in a file on its root filesystem and print it in the
response. MEMORY and ROOTFS snapshots carry it across pause and suspend,
VOLUMES resets it, so every row now states what it expects of rootfs
writes and the existing MEMORY and VOLUMES rows gain an assertion that
those runtimes handle them as documented.

Add micro-VM-only ROOTFS rows to the durable-dir lifecycle tests, plain
and suspended from PAUSED: the memory counter restarts, the file and
rootfs counters continue, and the suspend records a ROOTFS snapshot. The
template-repoint test gains a ROOTFS row too, pinning that the repoint
drops the restore to VOLUMES and resets the rootfs counter, that a pause
under the new template restores at ROOTFS again, and that the revert
cold-boots. gVisor cannot serve ROOTFS yet, so the new rows skip off
micro-VM clusters.
Comment thread cmd/ateom-microvm/restore.go Outdated
// container's rootfs upper: put those back before the overlays are
// assembled, then cold-boot on top of them instead of a pristine upper.
tUpper := time.Now()
if err := untarRootfsUpper(rootfsUpperDir(p.actorDirs), restoreDir, containerNames(p.containers)); err != nil {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 should-fix 🟡 – untarRootfsUpper runs os.RemoveAll on the upper dir before cleanupSandboxState has dropped any leftover overlay mounts, which only happens later inside coldBootActor. teardownActor documents that the overlays must be unmounted before the upperdir is removed. The MEMORY and VOLUMES paths both clean up first. Consider calling s.cleanupSandboxState(ctx, p.actorUID) before the untar, as restoreMemoryFidelity does.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested fix: clean up once in the ROOTFS arm before the untar, matching the MEMORY path.

case ateompb.SnapshotFidelity_SNAPSHOT_FIDELITY_ROOTFS:
	s.cleanupSandboxState(ctx, p.actorUID)
	if err := untarRootfsUpper(...); err != nil { ... }

@dberkov Dmitry Berkovich (dberkov) Oct 11, 2026 •

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in a84e095, by moving the upper staging into coldBootActor rather than calling cleanup from the ROOTFS arm. The cold boot now runs cleanupSandboxState first (it moved ahead of buildActorContainers, which only touches the atelet-owned bundle dirs, not the sandbox dirs cleanup clears), and only then stages the upper, so the overlay mounts are gone before untarRootfsUpper removes the dir, matching restoreMemoryFidelity and teardownActor. A retry re-cleans and re-stages from the same snapshot dir, which also removes the "keep what is there" reasoning.

Comment thread cmd/ateom-microvm/run.go Outdated
// A ROOTFS restore has already re-materialized the upper and boots on it;
// that holds across a retry too, since a guest that never reached its agent
// wrote nothing and cleanupSandboxState only drops the overlay mounts.
if !p.keepRootfsUpper {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 should-fix 🟡 – The new ROOTFS behavior in ateom-microvm has no unit test: the keepRootfsUpper skip here, the ROOTFS restore arm, and shipsRootfs in checkpoint. Only the micro-VM e2e rows cover it. Could the upper-dir decision and the per-fidelity file set be pulled into small helpers with table-driven tests, so a regression shows up without a KVM lane?

@dberkov Dmitry Berkovich (dberkov) Oct 11, 2026 •

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in a84e095. The ladder is now two predicates in fidelity.go, includesRootfs and includesMemory, used by the checkpoint capture set and the restore resume decision, with TestFidelityLadder pinning all four values. The upper-dir decision is stageRootfsUpper(actorDirs, snapshotDir, containers) in rootfsupper.go: empty dir means pristine, otherwise re-materialize from the tars. TestStageRootfsUpper covers both branches on real directories, including that a stale previous activation is wiped either way. The keepRootfsUpper flag is gone; actorBootParams carries rootfsUpperSnapshotDir instead.

Comment thread cmd/ateom-microvm/restore.go Outdated
// A ROOTFS snapshot holds no guest state either, but it carries each
// container's rootfs upper: put those back before the overlays are
// assembled, then cold-boot on top of them instead of a pristine upper.
tUpper := time.Now()

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 question 🟢 – This untar runs serially before the cold boot. restoreMemoryFidelity runs the same untar in a goroutine overlapped with bundle preparation. For an actor with a large upper, the ROOTFS resume pays the full untar time. Is overlapping it worth doing here too?

@dberkov Dmitry Berkovich (dberkov) Oct 11, 2026 •

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, and it is now overlapped. In a84e095 the staging runs as a background goroutine inside coldBootActor, started right after cleanupSandboxState and joined just before stageMergedRootfs, so it hides behind buildActorContainers and the resource-envelope check exactly as the MEMORY path hides its untar behind bundle preparation. Same drain-on-error discipline as there, so an early return never leaves a writer racing a retry. The separate rootfs_upper phase observation went away with the serial untar; the restore record now carries total like VOLUMES.

Comment thread cmd/ateom-microvm/checkpoint.go Outdated
// - Rootfs upper tars (ROOTFS and MEMORY): host-backed like the durable
// volumes — the memory snapshot does not carry rootfs writes. Under
// VOLUMES the workload cold-starts on restore, discarding rootfs state.
shipsRootfs := scope == ateompb.SnapshotFidelity_SNAPSHOT_FIDELITY_ROOTFS ||

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 nit 🟢 – "This level includes rootfs" is spelled out separately here, in the switch above, and in resumesGuest in restore. A small helper such as includesRootfs(f) would keep the ladder logic in one place.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done: includesRootfs and includesMemory in fidelity.go are the single place the ladder is spelled out; checkpoint uses the first for its capture set and restore uses the second for resumesGuest. The precondition switch in checkpoint keeps its explicit cases since it is about what each level needs present, not what it includes.

Comment on lines +71 to +72
if value.GetSnapshotConfig().GetPreferredFidelity() == ateapipb.SnapshotFidelity_SNAPSHOT_FIDELITY_ROOTFS &&
value.GetSandboxConfig().GetSandboxClass() != ateapipb.SandboxClass_SANDBOX_CLASS_MICROVM {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Self-note: we won't be able to check this once we move sandbox class to the workerpool.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed. Once the sandbox class lives on the WorkerPool the template no longer knows it at admission, so this check has to move to where the class is first known: scheduling or the resume workflow, rejecting a ROOTFS actor that lands on a gVisor pool. The runtime-side guard in ateom-gvisor stays as the backstop either way, so nothing can produce a ROOTFS checkpoint there in the meantime.

The ROOTFS restore untarred the rootfs uppers before handing off to the
cold boot, which only later cleared stale sandbox state. A previous
activation's overlay mounts could therefore still sit on the upper dir
when the untar removed it, the ordering teardown documents as unsafe. The
untar also ran serially, so a large upper added its full extraction time
to the resume, where the MEMORY path overlaps the same work with bundle
preparation.

Make the cold boot own the upper dir. It clears sandbox state first, moved
ahead of bundle preparation, which touches only the atelet-owned bundle
dirs, then stages the upper in the background and joins it right before
the overlay mounts need it: pristine for a plain boot, re-materialized from
the snapshot's tars when the boot parameters name a snapshot dir. A retry
stages again from that dir, so nothing depends on what an earlier attempt
left behind, and the ROOTFS restore arm shrinks to naming the dir.

Spell the ladder out once: includesRootfs and includesMemory decide the
checkpoint's capture set and whether a restore resumes the guest, and
stageRootfsUpper is the one place a cold boot decides between a pristine
and a restored upper. Both have table tests, so a regression in what each
fidelity carries shows up without a KVM lane.
@zoez7
Zoe Zhao (zoez7) added this pull request to the merge queue Oct 11, 2026
Merged via the queue into agent-substrate:main with commit b73ee7f Oct 11, 2026
14 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants