Skip to content

GPU sandboxes (RTX PRO 6000) reclaimed after 15-45 min of active use, far below documented 12h lifetime #11058

Description

@yethikrishna

What I am seeing

Since this morning (2026-10-01 ~12:30 UTC through 20:20 UTC), every GPU sandbox I provisioned on molab.marimo.io has been reclaimed after roughly 15-45 minutes, while it was actively working. 15 consecutive sandboxes (RTX PRO 6000 Blackwell, 4 vCPU / 32 GiB), all gone long before any limit.

Timeline (UTC, from my watchdog log)

Provisioned Dead Lifetime Last known state before death
13:01 13:22 ~21 min mid shard download
13:29 ~13:40s ~10+ min conversion running
14:18 14:42 ~24 min downloading FP8 shards
14:42 15:22 ~40 min mid-run
15:29 15:55 ~26 min mid conversion (blk.28)
15:55 16:22 ~27 min mid conversion (blk.1)
16:24 16:42 ~18 min mid conversion
16:46 17:02 ~16 min conversion exit 0, 19.6 GB GGUF on disk
17:07 17:30 ~23 min mid conversion (blk.34)
17:54 18:22 ~28 min conversion running (blk.1)
18:26 18:57 ~31 min convert DONE, server loading model
18:57 19:22 ~25 min mid-run
19:30 20:17 ~47 min conversion exit 0, 19.6 GB GGUF on disk

Symptoms

  • Liveness probe (simple print('ok') through the marimo kernel API) times out (2x100s); retrying after re-resolving the session still fails.
  • Direct requests return HTTP 410 Gone — the sandbox is permanently gone, not just the kernel.
  • Nothing was idle: several deaths happened mid GPU conversion and even after a successful 19.6 GB model conversion, with all local work lost (sandbox /tmp is not recoverable).

What the docs say

docs/guides/molab.md: "Notebooks can run for as long as 12 hours before molab shuts them down. Notebooks that are idle for more than 90 minutes are automatically shut down."

My notebooks were neither near 12 hours nor idle for anywhere near 90 minutes. This looks like GPU-capacity preemption, but there is no documented reclaim policy, no warning, and no status page indication.

Ask

  1. Is rapid reclaim of active GPU sandboxes expected behavior right now (capacity pressure), or is something broken?
  2. If it is expected, could the docs state the real policy — and ideally could reclaims carry a warning/notice so in-flight work can checkpoint?

Happy to share logs. Thanks for the great product — just trying to understand the ground rules so I can plan long jobs.

Activity

  1. self-assigned this
    on Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions