Queue, monitor and bootstrap Google Cloud TPU VMs from a terminal UI — or drive the same operations from an AI assistant over MCP.
Built for TRC / preemptible TPU workflows, where capacity appears and vanishes without warning and you want to grab it the moment it shows up.
┌ TPU Manager ─────────────────────────────────────────────────────────────┐
│ ⭐ Name Zone Latency Accel Mode State │
│ ⭐ spot-v6e64-ew4a-1 europe-west4-a ~140ms v6e-64 spot ACTIVE │
│ spot-v6e32-ew4a-1 europe-west4-a ~140ms v6e-32 spot WAITING… │
│ ondemand-v4-8-uc2-1 us-central2-b ~240ms v4-8 ond. NOT_QUEUED│
├──────────────────────────────────────────────────────────────────────────┤
│ 14:22:31 spot-v6e64-ew4a-1: PROVISIONING → ACTIVE │
│ 14:22:33 tmux sync: opened 1 new TPU window │
└──────────────────────────────────────────────────────────────────────────┘
- Fixed slot table. Each row is one reservable quota slot with its own node name, so several TPUs can be held at once. Rows are never added or removed — you queue and cancel them.
- Bucket race. Star several slots, hit
B, and it queues them all in parallel; the first to reachPROVISIONINGwins and the rest are cancelled automatically. This is how you actually catch scarce spot capacity. - tmux SSH mirror. Every
ACTIVETPU gets a live SSH window in a per-zone tmux session. Windows are never closed behind your back. - One-key VM bootstrap. Copies and runs the setup scripts over SSH with agent forwarding, then optionally joins the node to your Tailscale tailnet.
- MCP server. The same 12 operations exposed as tools, so an assistant can check state, queue capacity and open shells for you.
- Push notifications via ntfy.sh when a TPU goes live.
| Python | 3.11+ (uses tomllib) |
gcloud |
authenticated, with TPU API access to your project |
tmux |
optional — only for the SSH mirror |
xclip / wl-copy |
optional — clipboard fallback for s |
mcp>=1.2 |
optional — only for the MCP server (SDK 1.x and 2.x both work) |
An SSH agent holding your GitHub key is needed if you want the non-interactive VM bootstrap to clone private repos:
eval "$(ssh-agent -s)" && ssh-add ~/.ssh/githubStart the app from that same shell so it inherits SSH_AUTH_SOCK.
git clone https://github.com/<you>/tpu-manager.git
cd tpu-manager
pip install -e . # TUI only
pip install -e '.[mcp]' # plus the MCP serverNothing about your GCP setup is hardcoded. Write a starter config and edit it:
python -m tpu_manager.config --init # ~/.config/tpu-manager/config.toml
python -m tpu_manager.config --show # print what actually resolvedproject = "my-gcp-project"
network = "tpu-network"
ntfy_topic = "" # empty disables notifications
poll_interval = 30
poll_concurrency = 8 # max concurrent gcloud calls
auto_setup = false # run VM bootstrap automatically on ACTIVE
tmux_sync = true
[subnets]
"europe-west4-a" = "tpu-subnet-europe-west4"
[zone_codes]
"europe-west4-a" = "ew4a"
[[slots]]
name = "ondemand-v4-8-uc2-1"
zone = "us-central2-b"
accelerator = "v4-8"
mode = "on-demand"
# Generated spot rows: one row per machine that fits your v6e quota.
v6e_zones = [["europe-west4-a", "ew4a"]]
v6e_counts = [["v6e-64", 1], ["v6e-8", 8]]Any key can be overridden by an environment variable: TPU_MANAGER_POLL_INTERVAL=10,
TPU_MANAGER_AUTO_SETUP=1, and so on.
ntfy topics are public to anyone who guesses the name. Leave
ntfy_topicempty unless you pick something unguessable.
tpu-manager # or: python -m tpu_manager| Key | Action |
|---|---|
c / a |
queue selected slot / every NOT_QUEUED slot |
b / B |
toggle slot in the bucket / run the bucket race |
d |
cancel queue or delete VM — press twice |
x |
close the slot's tmux window — press twice |
o |
open a tmux SSH window |
s |
copy the SSH command to the clipboard |
S |
copy and run the VM bootstrap script |
t / T |
save a Tailscale auth key / push Tailscale to the VM |
i |
show IPs and create time |
r / q |
force refresh / quit |
tpu-mcp # or: python -m tpu_manager --mcpRegister it with Claude Code:
claude mcp add tpu-manager -- tpu-mcpOr in an MCP client config:
{
"mcpServers": {
"tpu-manager": { "command": "tpu-mcp" }
}
}| Tool | Kind | Purpose |
|---|---|---|
list_slots |
read-only | all slots and their states |
slot_info |
read-only | IPs, create time, SSH command |
ssh_command |
read-only | the SSH command, without connecting |
tmux_status |
read-only | managed windows, orphans flagged |
queue_slot / queue_slots |
mutating | request capacity |
setup_vm |
mutating | run the bootstrap over SSH |
tailscale_setup |
mutating | join the node to your tailnet |
tmux_sync / tmux_open |
mutating | open SSH windows |
cancel_slot |
destructive | delete the TPU — needs confirm=true |
tmux_close |
destructive | kill a tmux window — needs confirm=true |
Destructive tools carry destructive_hint=true annotations, so a client can
prompt before running them.
The rule, stated once and enforced everywhere:
- Each
ACTIVETPU gets one tmux window named after the slot, in sessiontpu-<zone-code>. - Reconciliation only creates. It never kills a window — not when the TPU
leaves
ACTIVE, not when a duplicate exists, not when the app exits. - A window whose slot is no longer
ACTIVEis reported as orphaned and left running. - Windows are set
remain-on-exit on, so a dropped SSH leaves a readable pane instead of the window vanishing. - The only way to close one is explicit and confirmed:
xtwice in the TUI, ortmux_close(name, confirm=true)over MCP. - Deleting a TPU leaves its window alone and warns you; pass
close_tmux=trueto do both at once.
~/.local/state/tpu-manager/state.json keeps, between runs:
- the resolved OS Login username (the
gcloudlookup is slow — caching it is what keeps startup snappy), - your bucket selection,
- the cursor position,
- which slots finished setup, and which were already announced.
Delete the file to reset; a corrupt file degrades to "first run" rather than breaking anything.
Earlier versions froze whenever a gcloud call was in flight. The current
design keeps the UI live:
- Every key handler returns immediately and dispatches to a Textual worker.
An
async def action_*is awaited by the message pump, so a slow refresh — or the bucket race, which loops until a winner appears — used to lock the whole UI. - Long jobs run in an exclusive worker group, so a second keypress cannot start a duplicate.
- Polling is bounded by a semaphore (
poll_concurrency). Unbounded, ~40 slots forked ~40gcloudprocesses per tick. - Blocking calls — the ntfy HTTP post, clipboard helpers, the OS Login lookup — run in threads, never on the event loop.
- The table repaints on a timer, so a burst of state transitions is one redraw.
src/tpu_manager/
config.py config resolution, slot table, zones, naming
state.py the small JSON store described above
core.py TPUManager — every gcloud/tmux side effect, no UI
tui.py Textual app
mcp_server.py MCP stdio server
scripts/ bootstrap scripts scp'd to each VM
tests/ no gcloud, no tmux, no network required
pip install -e '.[dev,mcp]'
pytest
ruff check .See docs/vm-setup.md for what the bootstrap scripts do and how to debug a failed setup.
MIT — see LICENSE.