Skip to content

cloudhypervisor: VM creation errors with "Address in use" #999

Description

@icyphox

Hey, apologies for the delay and Happy New Year! ✨ Here's the issue that I promised to create. Let me know if you need me to add any further info! Attaching relevant log files here:

cloudhypervisor-stderr.log
cloudhypervisor-stdout.log
cloudhypervisor.log

What happened:

Creating a VM using latest cloudhypervisor-static results in cloudhypervisor erroring out with (from /var/lib/flintlock/.../cloudhypervisor.stderr):

Failed to start the VMM thread: Error creation API server's socket Os { code: 98, kind: AddrInUse, message: "Address in use" }
Error booting VM: VmBoot(DeviceManager(CreateVirtioNet(OpenTap(TapOpen(ConfigureTap(Os { code: 16, kind: ResourceBusy, message: "Resource busy" }))))))

What did you expect to happen:

VM creation not to fail.

How to reproduce it:

I'm running flintlockd like so:

root@flintlock-test:/home/icy# ./flintlockd_amd64 run --cloudhypervisor-bin ./cloud-hypervisor-static --containerd-socket /run/containerd-dev/containerd.sock --bridge-name flbr-tenant1 --default-provider cloudhypervisor --insecure -v5

And sending the CreateMicroVM request using fl (built from main):

./fl microvm create --name "test-ch-1" --host localhost:9090 --network-interface eth1:tap --kernel-image ghcr.io/liquidmetal-dev/cloudhypervisor-kernel:6.2 --root-image ghcr.io/liquidmetal-dev/ubuntu:22.04 --kernel-filename vmlinux.bin

Environment:

  • flintlock version: 0.7.0
  • containerd version: 1.7.24
  • OS (e.g. from /etc/os-release): Debian 11 (bullseye)
  • Running on a GCP VM with nested virtualization enabled

Activity

  1. github-actions commented on Jul 2, 2026

    @github-actions
    Contributor

    This issue is stale because it has been open 180 days with no activity.

  2. added
    lifecycle/staleDenotes an issue or PR has remained open with no activity and has become stale.
    on Jul 2, 2026
  3. Derpthemeus commented on Jul 8, 2026

    @Derpthemeus

    I'm running into a similar issue using Cloud Hypervisor running running on multiple libvirt VMs on my local machine. Claude has a theory that it's due to a race condition that's more likely to appear in nested virtualization scenarios where it may take longer for Cloud Hypervisor to bind to a socket and become ready:

     1. pkg/planner/actuator.go:55-77 (executePlan) loops: call plan.Create(ctx) to get steps → run them → immediately loop back and call plan.Create(ctx) again, with no delay between
      iterations, until no steps remain.
      2. core/steps/microvm/create.go:33-46 (ShouldDo) gates purely on vmSvc.State(ctx, id) == MicroVMStatePending.
      3. infrastructure/microvm/cloudhypervisor/create.go:24-49 (Create) calls p.startCloudHypervisor(...) — which does cmd.Start() (or DetachedStart) and returns immediately, before the
      spawned cloud-hypervisor binary has actually created/bound its API socket or begun serving — then writes the PID file and returns nil.
      4. infrastructure/microvm/cloudhypervisor/provider.go:69-132 (State) only reports MicroVMStateRunning once it can (a) find the sock file and (b) get a successful Info() response
      The race: executePlan re-invokes ShouldDo essentially instantly after Do returns. If cloud-hypervisor hasn't finished coming up yet (a few tens of ms), State() still says Pending
      (or errors transiently) → ShouldDo returns true again → Do/Create() fires a second time for the same VM.
    
      That second Create() call goes through ensureState() (create.go:186-217), which unconditionally deletes any existing socket file at vmState.SockPath() — with no check for whether a
      live process already owns it — then spawns a second cloud-hypervisor process. That process collides with the first on the still-held disk write-lock ("AlreadyLocked") or the
      socket ("Address in use"), fails immediately, while the first, actually-healthy process is left running but orphaned — its socket path was just deleted out from under it, so
      State() can never successfully query it again, and flintlockd retries forever, exactly matching what we observed on the live host.
    
      Minimal fix: in Create(), after starting the process, poll for the socket to exist and respond to Info() (with a short timeout) before returning — so a losing second call to
      ShouldDo() reliably sees Running, not Pending/error. Belt-and-suspenders: make ensureState() refuse to delete the socket if the PID file points to a still-live process.
    

    I'm not familiar enough with the codebase to know if this is correct or hallucinated, but I'm just going to try working around the issue by switching from Cloud Hypervisor to Firecracker. Hopefully this context can be useful to somebody else!

    • Flintlock 0.9.0
    • Host running Linux 6.18.36-1 Manjaro
    • Libvirt guest running Linux 7.0.0-27 Ubuntu
  4. removed
    lifecycle/staleDenotes an issue or PR has remained open with no activity and has become stale.
    on Jul 9, 2026
  5. richardcase commented on Jul 18, 2026

    @richardcase
    Member

    Ran into this on flintlock main (729de8a, == released flintlockd v0.9.0) with Cloud Hypervisor v53.0 under nested KVM, while adding a CH path to an acceptance suite. Managed to root-cause it — sharing in case it helps land a fix.

    Symptom

    With provider: cloudhypervisor, VM creates intermittently (~1 in 3, worse under nested virt/load) never reach CREATED. cloudhypervisor.stderr fills with hundreds of:

    Error: Cloud Hypervisor exited with the following chain of errors:
      0: Failed to start the VMM thread
      1: API socket "…/cloudhypervisor.sock" is already in use by another running instance
    

    The guest that does win the race boots and networks fine — so it's a control-plane reconcile bug, not guest/boot.

    Root cause

    1. State() reports CH "Created" as Pending. In infrastructure/microvm/cloudhypervisor/provider.go, State() calls chClient.Info and does:

      case cloudhypervisor.VMStateRunning:  // "Running"
          return ports.MicroVMStateRunning, nil
      case cloudhypervisor.VMStateCreated:  // "Created"
          return ports.MicroVMStatePending, nil   // <-- here

      CH sits in "Created" between the API socket binding and the guest finishing boot (long under nested virt).

    2. The create step's guard is state-only, so it re-fires. core/steps/microvm/create.go ShouldDo() returns state == MicroVMStatePending; Do() calls Create again → a second cloud-hypervisor on the same --api-socket → Address already in use / tap Resource busy. Reconcile re-queues on the error, so it repeats rapidly.

    3. ensureState() deletes a live process's socket → permanent orphan. In create.go, ensureState() unlinks the socket whenever the file exists:

      if sockExists, _ := afero.Exists(p.fs, vmState.SockPath()); sockExists {
          p.fs.Remove(vmState.SockPath())
      }

      When the second Create fires while the first, healthy CH is still bound, this removes the socket out from under the running process. State() can no longer Info() it → Unknown/error → retries exhaust MaximumRetry → FailedState, never CREATED. Firecracker is unaffected because its State() treats "pidfile present + process alive" as Running, so the create guard goes false once the VM is up.

    Minimal fix (either helps; together is defense-in-depth)

    • provider.go State(): treat CH "Created" as MicroVMStateRunning (process up + socket bound) so the create step stops re-firing.
    • create.go ensureState(): before Remove(SockPath()), read the pidfile and skip removal (no-op the Create) when that PID is a live process — so a stray re-create can't orphan a running VM.

    Neither can terminate a running guest. Happy to open a PR if that's useful.

  6. added a commit that references this issue on Jul 19, 2026
    c2e137d
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    kind/bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions