Repository navigation
cloudhypervisor: VM creation errors with "Address in use" #999
Description
Activity
This issue is stale because it has been open 180 days with no activity.
- addedlifecycle/staleDenotes an issue or PR has remained open with no activity and has become stale.Denotes an issue or PR has remained open with no activity and has become stale.
on Jul 2, 2026 I'm running into a similar issue using Cloud Hypervisor running running on multiple libvirt VMs on my local machine. Claude has a theory that it's due to a race condition that's more likely to appear in nested virtualization scenarios where it may take longer for Cloud Hypervisor to bind to a socket and become ready:
1. pkg/planner/actuator.go:55-77 (executePlan) loops: call plan.Create(ctx) to get steps → run them → immediately loop back and call plan.Create(ctx) again, with no delay between iterations, until no steps remain. 2. core/steps/microvm/create.go:33-46 (ShouldDo) gates purely on vmSvc.State(ctx, id) == MicroVMStatePending. 3. infrastructure/microvm/cloudhypervisor/create.go:24-49 (Create) calls p.startCloudHypervisor(...) — which does cmd.Start() (or DetachedStart) and returns immediately, before the spawned cloud-hypervisor binary has actually created/bound its API socket or begun serving — then writes the PID file and returns nil. 4. infrastructure/microvm/cloudhypervisor/provider.go:69-132 (State) only reports MicroVMStateRunning once it can (a) find the sock file and (b) get a successful Info() response The race: executePlan re-invokes ShouldDo essentially instantly after Do returns. If cloud-hypervisor hasn't finished coming up yet (a few tens of ms), State() still says Pending (or errors transiently) → ShouldDo returns true again → Do/Create() fires a second time for the same VM. That second Create() call goes through ensureState() (create.go:186-217), which unconditionally deletes any existing socket file at vmState.SockPath() — with no check for whether a live process already owns it — then spawns a second cloud-hypervisor process. That process collides with the first on the still-held disk write-lock ("AlreadyLocked") or the socket ("Address in use"), fails immediately, while the first, actually-healthy process is left running but orphaned — its socket path was just deleted out from under it, so State() can never successfully query it again, and flintlockd retries forever, exactly matching what we observed on the live host. Minimal fix: in Create(), after starting the process, poll for the socket to exist and respond to Info() (with a short timeout) before returning — so a losing second call to ShouldDo() reliably sees Running, not Pending/error. Belt-and-suspenders: make ensureState() refuse to delete the socket if the PID file points to a still-live process.I'm not familiar enough with the codebase to know if this is correct or hallucinated, but I'm just going to try working around the issue by switching from Cloud Hypervisor to Firecracker. Hopefully this context can be useful to somebody else!
- Flintlock 0.9.0
- Host running Linux 6.18.36-1 Manjaro
- Libvirt guest running Linux 7.0.0-27 Ubuntu
- removedlifecycle/staleDenotes an issue or PR has remained open with no activity and has become stale.Denotes an issue or PR has remained open with no activity and has become stale.
on Jul 9, 2026 Ran into this on flintlock
main(729de8a, == releasedflintlockd v0.9.0) with Cloud Hypervisor v53.0 under nested KVM, while adding a CH path to an acceptance suite. Managed to root-cause it — sharing in case it helps land a fix.Symptom
With
provider: cloudhypervisor, VM creates intermittently (~1 in 3, worse under nested virt/load) never reachCREATED.cloudhypervisor.stderrfills with hundreds of:Error: Cloud Hypervisor exited with the following chain of errors: 0: Failed to start the VMM thread 1: API socket "…/cloudhypervisor.sock" is already in use by another running instanceThe guest that does win the race boots and networks fine — so it's a control-plane reconcile bug, not guest/boot.
Root cause
-
State()reports CH"Created"asPending. Ininfrastructure/microvm/cloudhypervisor/provider.go,State()callschClient.Infoand does:case cloudhypervisor.VMStateRunning: // "Running" return ports.MicroVMStateRunning, nil case cloudhypervisor.VMStateCreated: // "Created" return ports.MicroVMStatePending, nil // <-- here
CH sits in
"Created"between the API socket binding and the guest finishing boot (long under nested virt). -
The create step's guard is state-only, so it re-fires.
core/steps/microvm/create.goShouldDo()returnsstate == MicroVMStatePending;Do()callsCreateagain → a second cloud-hypervisor on the same--api-socket→Address already in use/ tapResource busy. Reconcile re-queues on the error, so it repeats rapidly. -
ensureState()deletes a live process's socket → permanent orphan. Increate.go,ensureState()unlinks the socket whenever the file exists:if sockExists, _ := afero.Exists(p.fs, vmState.SockPath()); sockExists { p.fs.Remove(vmState.SockPath()) }
When the second
Createfires while the first, healthy CH is still bound, this removes the socket out from under the running process.State()can no longerInfo()it →Unknown/error → retries exhaustMaximumRetry→FailedState, neverCREATED. Firecracker is unaffected because itsState()treats "pidfile present + process alive" asRunning, so the create guard goes false once the VM is up.
Minimal fix (either helps; together is defense-in-depth)
provider.goState(): treat CH"Created"asMicroVMStateRunning(process up + socket bound) so the create step stops re-firing.create.goensureState(): beforeRemove(SockPath()), read the pidfile and skip removal (no-op theCreate) when that PID is a live process — so a stray re-create can't orphan a running VM.
Neither can terminate a running guest. Happy to open a PR if that's useful.
-
- added a commit that references this issue
on Jul 19, 2026
Hey, apologies for the delay and Happy New Year! ✨ Here's the issue that I promised to create. Let me know if you need me to add any further info! Attaching relevant log files here:
cloudhypervisor-stderr.log
cloudhypervisor-stdout.log
cloudhypervisor.log
What happened:
Creating a VM using latest cloudhypervisor-static results in cloudhypervisor erroring out with (from
/var/lib/flintlock/.../cloudhypervisor.stderr):What did you expect to happen:
VM creation not to fail.
How to reproduce it:
I'm running flintlockd like so:
And sending the CreateMicroVM request using
fl(built from main):Environment:
/etc/os-release): Debian 11 (bullseye)