Hello,
On CloudStack 4.23.0.0 with CLVM (lvmlockd + sanlock), a hard host power-off
does not fail the VM over cleanly.
The surviving host can see the VM logical volume in LVM, but the device node is
missing (lv_attr -wi-------, no /dev/VG/LV). Start then refreshes the libvirt
storage pool about every 3 seconds and fails with:
Could not find volume ... volume exists in LVM but device node not
accessible
com.cloud.hypervisor.kvm.storage.ClvmStorageAdaptor.handleMissingDeviceNode
only runs lvchange -asy for volumes whose name starts with "template-". A
normal ROOT volume never calls
LibvirtComputingResource.activateClvmVolumeExclusive (lvchange -aey).
StartCommand waits up to 1800 seconds. The VM stays Starting, and the other
host never requests the sanlock lease.
The lock must not be wiped. When the device node is missing, the surviving host
should call activateClvmVolumeExclusive (lvchange -aey, 300s). Sanlock grants
the lease after the dead host's lease expires, and the VM can start on the
other node.
We hit this on a two-host KVM cluster: detection of the dead host was fast, but
start looped on the pool refresh until it timed out. After calling lvchange
-aey on the surviving host, the same VM started there.
Regards,
Umit Eyigun
Trtek Yazilim A.S.
[email protected]