Skip to content

Restore a lost bot workspace into a replacement sandbox #449

Description

@linear-code

Problem

When a bot's remote sandbox is permanently lost, the bot loses its files and browser state. There is no backup of a remote workspace that lives apart from the remote machine, so nothing can be restored.

On main today:

  • Create, sleep, wake, and reattach exist. createRemoteBotWorkspace (apps/server/src/provider/botWorkspace.ts) opens a saved identity, BotWorkspaceLifecycle (apps/server/src/provider/workspace/) reads and writes it, and AkeruSessionResources keeps it in bot-workspaces/<workspaceId>/provider.json, so a workspace reattaches after a server restart.
  • For most providers, a sandbox that cannot be reattached fails with "missing or unavailable" and is not replaced. Ascii Box is the exception: when Ascii confirms a 404 for the saved VM, createRemoteBotWorkspace creates a new empty VM on the next use, without asking and without restoring anything. docs/user/sandboxes.md documents this.
  • There is no shared error type that separates a missing instance from an outage, an auth failure, or a rate limit. Each adapter's inspect maps provider states on its own.
  • AkeruRemoteSession (workspace/BotWorkspaceTypes.ts) exposes commands, lifecycle, a browser endpoint, and optional computer access. BotWorkspaceFilesystem provides file access through commands. Nothing provides a provider-independent checkpoint or restore. Ascii's native stop and resume snapshots only resume the same VM. CheckpointStore.restoreCheckpoint is for git checkpoints, not sandboxes.
  • The general user archive from LEO-300 does not back up sandboxes.

Goal

If a remote sandbox is permanently lost, the user can restore the bot's durable files and supported browser profile from the last good checkpoint into a replacement sandbox they explicitly approved. The Akeru workspace identity stays the same. Only the provider's native instance ID changes.

Recovery is never automatic. AKR-87's rule against silent replacement stays in force, and this issue extends it with user-approved recovery instead of weakening it.

Checkpoints

  • Versioned checkpoints are stored under the owning Akeru environment, apart from the remote machine. They include files, the metadata needed to restore permissions, and supported browser-profile state.
  • Browser credentials are sensitive backup material, never general export data.
  • Checkpoints are taken at safe lifecycle boundaries. Browser and profile writers are paused first, and every user of a shared workspace is coordinated.
  • Promotion is atomic, with checksums and bounded retention. An interrupted export never replaces a good checkpoint.

Detecting loss

  • A genuinely missing instance is told apart from a temporary outage, an auth failure, and a rate limit. Only a confirmed missing instance offers recovery.
  • Before creating a replacement, the user sees the time of the last good checkpoint and what work may be lost.

Restoring

  • Restore goes into staging, is validated, and only then switches the native identity atomically.
  • A failed restore or a server restart leaves the last good data and a recoverable state. Operations from the old instance are fenced off.
  • Durable files are restored, not running processes. The last output of each lost process is kept and the process is clearly marked as lost. Restarting a command needs an explicit user action and must not repeat external effects.

Providers

  • Each remaining provider gets an explicit capability decision, and the file and profile formats and guarantees are documented per provider. Local data is never touched by remote replacement.

Acceptance criteria

  • Create files and a deterministic browser-session fixture, take a checkpoint, destroy only the isolated test sandbox, restore into a replacement, and confirm the files, the profile state, and the unchanged Akeru workspace identity
  • A temporary outage or auth failure never creates another sandbox
  • A corrupt or partial checkpoint, an interrupted restore, a repeated recovery request, concurrent bot use, and a server restart each keep the last good data and never leave two active instances
  • Lost processes are reported honestly with their last output and a restart action, and nothing claims a process kept running
  • The bot inbox and Settings show the same recovery state and actionable failures on every client
  • Docs state retention, how sensitive browser state is handled, what can be recovered, and the limits

How to verify

Focused, deterministic lifecycle and restore tests that wait on typed receipts and worker drains, never sleeps. Real provider tests use disposable workspaces with explicit credentials, and the report names each provider actually tested and each one still blocked by missing credentials. The primary agent then checks the recovery flow in isolated web and Electron clients at desktop and narrow widths, checks the mobile recovery state where it changes, and covers local and remote connections plus the reconnect, sleep, and wake flows around recovery. Run targeted lint and typechecks only. Never write to live Akeru data.

Out of scope

  • A Docker provider, managed hosting, whole-OS disk snapshots, or migrating a workspace between providers
  • Rerunning lost commands automatically
  • Teaching routines (AKR-80) and the general user archive (LEO-300)

Open questions

  • Ascii Box already creates an empty replacement on its own after a confirmed 404. Should that change to offer recovery instead, or does it go away with AKR-127's provider cut?

Context

Selected work: Leo picked this as feature 4 of his 2026-09-07 comparison. Missing live sandbox credentials only limit the real-provider tests. The local, deterministic checkpoint, restore, and failure work does not depend on them. Use disposable fixtures and never write to live Akeru state.

Ownership: this issue owns checkpoint storage, the restore lifecycle, and how recovery is shown. AKR-87 owns create, sleep, wake, and reattach, so reuse its workspace identity and lifecycle and do not duplicate reattachment. AKR-106 owns the remote desktop and its transport. Coordinate edits to shared workspace and session files with its owner, and start with isolated checkpoint and restore modules and tests while that is agreed.

Providers: main has Local, E2B, Daytona, Vercel Sandbox, Upstash Box, Ascii Box, Railway, and Tenki (CLOUD_SANDBOX_PROVIDERS in packages/contracts/src/settings/sandbox.ts). AKR-127 is reducing this to Local, E2B, Vercel, and a Cloudflare bridge, and AKR-133 adds hosted Akeru Cloud sandboxes. Make the capability decisions for the providers that remain when this work starts.

Where to start: apps/server/src/provider/botWorkspace.ts, botWorkspaceFilesystem.ts, botWorkspacePool.ts, AkeruSessionResources.ts, apps/server/src/provider/workspace/, apps/server/src/bot-inbox/, apps/server/src/persistence/, packages/contracts/src/settings/sandbox.ts, and docs/user/sandboxes.md.

Related: AKR-87 (reattach sandboxes after a restart), AKR-106 (watch and take control of a bot's remote desktop), LEO-300 (export and restore all user-owned state), LEO-182 (canceled: export and restore sandbox state).

Created with Claude Opus 5.5 in Claude Code.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions