Provisioner.Provision (internal/reconciler/provision.go) passes the reconciler's context to CreateMicroVM. If that context is cancelled while the call is in flight, the client gets Canceled and never learns the uid, but flintlock may already have accepted the request and gone on to create the microVM. No vms row is written, so nothing ever deletes it.
The reconciler's context is cancelled whenever its pool is updated or deleted, and at poolmgrd shutdown. Since #112, DeletePool drains a pool by stopping the reconciler first, so deleting a pool while it is still provisioning can leave a microVM running on a flintlock host with no record of it. That is the symptom #112 set out to remove.
The related window, where CreateMicroVM has returned and the store write then fails on the cancelled context, is fixed on the #112 branch: the orphan delete now runs on a detached context. This issue is the part that fix cannot reach, because there is no uid to delete.
Possible fixes
- The microVM's
id and namespace are chosen by Provision before the call (<pool>-<8 hex>), so on a cancelled or otherwise ambiguous CreateMicroVM error it could list the host's microVMs in that namespace, find the one with that id, and delete it.
- Or run
CreateMicroVM on a context detached from the reconciler's cancellation (bounded by a timeout), so the call always returns a uid that the existing cleanup can use.
- A periodic sweep comparing each host's microVMs against the store would also catch this, along with any other orphan.
Provisioner.Provision(internal/reconciler/provision.go) passes the reconciler's context toCreateMicroVM. If that context is cancelled while the call is in flight, the client getsCanceledand never learns the uid, but flintlock may already have accepted the request and gone on to create the microVM. Novmsrow is written, so nothing ever deletes it.The reconciler's context is cancelled whenever its pool is updated or deleted, and at
poolmgrdshutdown. Since #112,DeletePooldrains a pool by stopping the reconciler first, so deleting a pool while it is still provisioning can leave a microVM running on a flintlock host with no record of it. That is the symptom #112 set out to remove.The related window, where
CreateMicroVMhas returned and the store write then fails on the cancelled context, is fixed on the #112 branch: the orphan delete now runs on a detached context. This issue is the part that fix cannot reach, because there is no uid to delete.Possible fixes
idand namespace are chosen byProvisionbefore the call (<pool>-<8 hex>), so on a cancelled or otherwise ambiguousCreateMicroVMerror it could list the host's microVMs in that namespace, find the one with that id, and delete it.CreateMicroVMon a context detached from the reconciler's cancellation (bounded by a timeout), so the call always returns a uid that the existing cleanup can use.