Skip to content

ateapi: remove SnapshotConfig.on_pause, pause captures the on_commit scope - #2309

Merged
Zoe Zhao (zoez7) merged 1 commit into
agent-substrate:mainfrom
dberkov:remove-on-pause
Oct 7, 2026
Merged

Zoe Zhao (zoez7) merged 1 commit into
agent-substrate:mainfrom
dberkov:remove-on-pause

Conversation

@dberkov

Copy link
Copy Markdown
Collaborator

Part of #2102, milestone 1 as tracked in #2286.

What changes

  • SnapshotConfig.on_pause is removed from the public API. Every snapshot of an actor, the node-local checkpoint a Pause takes and the snapshot a Suspend uploads, now captures the template's on_commit scope. The "on_commit must be a subset of on_pause" custom validation goes with it.
  • Resume of a local snapshot restores at the scope recorded on LocalSnapshot at pause time rather than re-reading a template field. There is no fallback for an unset recorded scope: pause finalization always writes it.
  • The paused-origin suspend precheck (Data pause under a Full-committing template) is removed. An actor's template can only change while SUSPENDED, so the recorded pause scope and the template's on_commit can no longer disagree, and atelet still enforces the rule when uploading a local checkpoint.
  • Demos, benchmarking manifests, the glossary, and the API guide drop onPause.

Behavior changes

  • Breaking API change. Templates that set onPause are rejected. Per the pre-v1.0 policy there is no compatibility shim.
  • Templates with onCommit: SNAPSHOT_CONTENT_SCOPE_DATA now drop process memory across Pause as well as Suspend. Previously the default onPause: FULL kept memory across a Pause. The e2e lifecycle table reflects this.
  • Unchanged: golden actors still commit Full, LocalSnapshot.content_scope and ExternalSnapshot.content_scope keep their meaning, and the atelet/ateom wire protocol is untouched.

Follow-ups (separate PRs)

  • Rename on_commit to capture_scope, replace SnapshotContentScope with the ordered SnapshotScope, add minimum_scope and ResumeActorResponse.scope, and the resume fallback rule.
  • Data-plane ROOTFS scope support.

Verification

  • go test ./cmd/ateapi/... ./cmd/kubectl-ate/... ./internal/ateattr/... and tools/apitool tests pass.
  • hack/update/codegen.sh is idempotent; gofmt, golangci-lint, kube-api-linter, proto-fmt, metrics, python-protos, and licenses verifiers pass.
  • e2e suites compile; not run against a cluster in this PR.

🤖 Generated with Claude Code

…scope

A template used to carry two snapshot scopes: on_pause for the node-local
checkpoint a Pause takes and on_commit for the snapshot a Suspend uploads.
The lifecycle work that replaces these with a single capture scope plus a
minimum resume scope has no place for a per-operation setting, and the
pause/suspend distinction itself is going away. Start by collapsing the two:
every snapshot of an actor, local or uploaded, now captures on_commit.

Resume of a local snapshot restores at the scope recorded on LocalSnapshot
at pause time instead of re-reading a template field, and no longer falls
back to the template when that scope is unset; pause finalization always
records it. The paused-origin suspend precheck that rejected a Data pause
under a Full-committing template is gone with it: an actor's template can
only change while SUSPENDED, so the recorded scope and the template's
on_commit can no longer disagree, and atelet still enforces the rule when it
uploads a local checkpoint.

Templates that set on_pause fail validation; demos, benchmarking manifests,
and docs drop the field. Templates that commit Data now lose process memory
across Pause as well, where before the default Full on_pause preserved it.
@zoez7
Zoe Zhao (zoez7) added this pull request to the merge queue Oct 7, 2026
Merged via the queue into agent-substrate:main with commit 44a201b Oct 7, 2026
20 of 22 checks passed
Chenyi Wang (chw120) pushed a commit to chw120/substrate that referenced this pull request Oct 9, 2026
…ferred_fidelity (agent-substrate#2364)

**This is a rename, not a behavior change.** `SnapshotConfig.onCommit`
becomes `preferredFidelity`, and the two-valued `SnapshotContentScope`
enum becomes the ordered `SnapshotFidelity` enum, with the same
vocabulary carried through atelet, ateom, and the metric label. `MEMORY`
is exactly what `FULL` was and `VOLUMES` is exactly what `DATA` was:
what gets captured on pause and suspend, what gets restored, the
golden-actor override, and the paused-upload narrowing are all
unchanged. The point is to put the lifecycle-v2 names in place before
step III adds `required_fidelity` and `resumed_fidelity` on top of them.
The one addition is `ROOTFS`, defined in the enum so it is complete but
rejected everywhere, since no runtime can produce it yet.

Part of agent-substrate#2102, step II as tracked in agent-substrate#2351. Step I (agent-substrate#2309) removed
`on_pause`.

Two commits, meant to be read in order.

## 1. ateapi: `SnapshotFidelity` and `preferred_fidelity`

- New ordered enum `SnapshotFidelity { VOLUMES = 1, ROOTFS = 2, MEMORY =
3 }`, each level including the ones below it. `SnapshotContentScope` is
deleted. MEMORY is the old FULL and VOLUMES the old DATA, so behavior is
unchanged.
- `SnapshotConfig.on_commit` becomes `preferred_fidelity`, documented as
a best-effort ceiling: the system never packages more than this level
and may package less when a layer cannot be captured or is too expensive
at the time. A process resumed in place keeps its memory regardless.
Defaults to MEMORY. Field numbers rearranged to `storage_location = 1`,
`preferred_fidelity = 2`.
- `ExternalSnapshot` and `LocalSnapshot` record `fidelity` instead of
`content_scope`.
- ROOTFS is in the enum so it is complete, but no sandbox runtime
captures rootfs without memory yet, so a new field-level validator
rejects it at template admission. Actors only take their fidelity from a
template, so none can carry it.
- kubectl-ate tag output, demos, benchmarking manifests, glossary, API
guide, and architecture doc use the new vocabulary. The obsolete apitool
exemption is removed.

## 2. atelet, ateom: the same enum on the wire

- Both internal protos replace `SnapshotScope` with a `SnapshotFidelity`
mirroring the public one; request fields become `fidelity` and
`desired_fidelity`. Every hop maps level by level, ROOTFS included, and
nothing is promoted: an unexpected value goes out as UNSPECIFIED for the
next hop to reject.
- Defense in depth for ROOTFS: atelet rejects it in checkpoint, restore,
and paused-upload validation, and both ateom runtimes reject it on
`CheckpointWorkload` and `RestoreWorkload` through a shared
`resources.ValidateSnapshotFidelity`.
- A golden snapshot must record MEMORY (UNSPECIFIED is no longer
tolerated).
- The metric attribute `ate.snapshot.scope` (`full`/`data`/`unknown`)
becomes `ate.snapshot.fidelity` (`volumes`/`rootfs`/`memory`/`unknown`)
in the OTel registry, `ateattr`, the ateom phase logs, and the snapshot
manifest atelet writes and reads back when narrowing a paused upload.
Dashboards keyed on the old label need updating.

## Behavior changes

- **Breaking API change.** Templates that set `onCommit` or use
`SNAPSHOT_CONTENT_SCOPE_*` values are rejected. Per the pre-v1.0 policy
there is no compatibility shim, and previously stored templates and
snapshots are not expected to exist.
- Snapshot manifests written with the old `scope` key read back with no
fidelity, which the upload path already refuses with a clear message.
- Unchanged: what MEMORY and VOLUMES capture and restore, the
golden-actor MEMORY override, and the paused-upload narrowing from
MEMORY to VOLUMES.

## Relationship to agent-substrate#2289

agent-substrate#2289 restructures the snapshot status into `Snapshot` and
`SnapshotStorage`. If it lands first, the `fidelity` field in this PR
moves onto `SnapshotStorage` (per copy), which is where it belongs
long-term; the enum, the template field, and the data-plane change are
the same either way.

## Follow-ups (step III)

`required_fidelity` with `required <= preferred` validation and the
CRASHED rule, `ResumeActorResponse.resumed_fidelity`, and ROOTFS support
in the micro-VM runtime.

## Verification

- `go test` for ateapi (unit and functional), kubectl-ate, atelet,
ateom-microvm, ateom-gvisor, ateomphaselog, ateattr, resources, and
apitool pass; new validator at 100% coverage.
- golangci-lint, kube-api-linter, python-protos, boilerplate verifiers
pass; codegen, proto-fmt, licenses, go.mod show no drift.
- Not run locally: the Weaver metrics-registry check and shellcheck,
which need Docker; relying on CI for those. e2e suites compile but were
not run against a cluster.

- [x] Breaking change
- [x] Not backward compatible

🤖 Generated with [Claude Code](https://claude.com/claude-code)
cloudonly pushed a commit to cloudonly/substrate that referenced this pull request Oct 11, 2026
…r micro-VM templates (agent-substrate#2448)

Step III of agent-substrate#2102, tracked in agent-substrate#2447. Step I (agent-substrate#2309) removed `on_pause`;
step II (agent-substrate#2364) introduced the `SnapshotFidelity` ladder with ROOTFS
defined but rejected everywhere.

**ROOTFS becomes real on micro-VM.** A ROOTFS snapshot holds the durable
volumes plus each container's root-filesystem changes, without guest
memory: the actor cold-boots on resume but keeps what it wrote to its
rootfs. gVisor keeps rootfs changes inside its memory checkpoint and
cannot serve it yet; it now declares that explicitly and a gVisor
template that asks for ROOTFS is rejected at admission.

Four commits, meant to be read in order. Each builds and passes tests on
its own, and ROOTFS stays unreachable from the public API until the
third.

## 1. atelet: forward ROOTFS and stop narrowing paused uploads

Atelet accepts every level of the ladder and forwards it; which levels a
runtime serves is the runtime's knowledge. The paused-upload path still
converted a MEMORY pause into a VOLUMES upload, a leftover from the days
of separate pause and commit scopes. Pause and suspend now capture one
fidelity, a paused actor's template cannot change, and golden actors are
never paused, so captured and requested fidelity are always equal; the
conversion is replaced by an equality check that refuses any mismatch as
drift.

## 2. ateom-microvm: capture and restore ROOTFS

Every piece already existed: MEMORY checkpoints tar the host-side rootfs
upper per container alongside the Cloud Hypervisor snapshot, and MEMORY
restores untar those uppers before assembling the overlays. ROOTFS is
the same checkpoint without the VM snapshot, and the upper restore
followed by the VOLUMES cold-boot path with a new `keepRootfsUpper` boot
parameter so the cold boot does not wipe the restored upper (safe across
its retry). Only a MEMORY restore resumes the guest's CPU counters;
ROOTFS counts from zero like VOLUMES.

`resources.ValidateSnapshotFidelity` now takes the runtime's supported
set: micro-VM passes all three, gVisor passes VOLUMES and MEMORY and
rejects ROOTFS before taking the actor lock, naming the levels it
serves.

## 3. ateapi: allow ROOTFS on micro-VM templates

The field-level rejection moves to the template-level cross-field hook:
ROOTFS requires `sandbox_config.sandbox_class: SANDBOX_CLASS_MICROVM`.
An actor can only be repointed to a template with an identical sandbox
config, so a ROOTFS template never ends up under a gVisor actor. The
control plane needs no logic change: a ROOTFS snapshot restores at
ROOTFS, and a repointed actor drops to VOLUMES as it does for MEMORY,
since a rootfs upper is only valid over the image it was written on.

## 4. e2e: prove it on the counter demo

The counter app gains a third counter in a file on its root filesystem,
printed in the response, so a ROOTFS snapshot is distinguishable from
VOLUMES. All lifecycle rows state their rootfs expectation (MEMORY and
ROOTFS keep it, VOLUMES resets it). New micro-VM-only ROOTFS rows cover
pause, suspend, suspend from PAUSED, and the template repoint and
revert. On the gVisor CI lane the ROOTFS rows skip; the new rootfs
assertions on MEMORY and VOLUMES rows run there as a side check.

## Behavior changes

- **API**: `SNAPSHOT_FIDELITY_ROOTFS` is accepted on micro-VM templates
and rejected on gVisor ones with a message naming the requirement. No
field or wire changes; the `preferred_fidelity` field-level custom
validation tag is removed.
- **Atelet**: a paused-upload whose desired fidelity differs from the
captured one is now `FailedPrecondition`; previously a MEMORY capture
could be downgraded to VOLUMES. Unreachable in practice, as explained in
commit 1.
- Unchanged: MEMORY and VOLUMES capture and restore, the golden-actor
MEMORY override, metric labels (`rootfs` was already a member).
- No backward compatibility is kept, per the pre-v1.0 policy.

## Not in this PR

`required_fidelity`, `ResumeActorResponse.resumed_fidelity`, the resume
fallback rule, gVisor ROOTFS (pending the gVisor team's separable rootfs
capture), and restoring a MEMORY snapshot as a ROOTFS cold boot (the
artifacts already allow it; atelet would need the rootfs file subset in
the manifest to skip the memory image).

## Verification

- Unit tests for atelet, ateom-microvm, ateom-gvisor, resources,
ateattr, ateomphaselog, ateapi (validation, controlapi, functional),
kubectl-ate, and apitool pass; the ateom tests are linux-tagged and were
type-checked for linux locally.
- golangci-lint, kube-api-linter, python-protos, boilerplate pass;
codegen shows no drift.
- The micro-VM counter demo was run manually on a GKE cluster with a
nested-virtualization node pool and behaved as expected. The ROOTFS e2e
rows need a micro-VM lane; the gVisor lane runs the rest.

- [x] Breaking change
- [x] Not backward compatible

🤖 Generated with [Claude Code](https://claude.com/claude-code)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants