Pearl1594 opened a new issue, #14218:
URL: https://github.com/apache/cloudstack/issues/14218

   ### problem
   
   Deleting a disk-only VM snapshot on a multi-disk VM can leave CloudStack's 
database permanently inconsistent with the actual state of primary storage if 
one disk's merge completes successfully while another disk's merge in the same 
operation times out. The disk that succeeded has its underlying file correctly 
committed and removed on the hypervisor, but the corresponding database update 
is never applied - because the whole operation is treated as a single 
all-or-nothing unit. The affected volume is left pointing at a file that no 
longer exists, and the VM cannot subsequently be started.
   
   ### versions
   
   4.21.0.0+
   KVM hypervisor, primary storage of type `Filesystem` / `NetworkFilesystem` / 
`SharedMountPoint`, VM with 2+ disks, disk-only VM snapshot (no memory).
   
   ### The steps to reproduce the bug
   
   1. Create a VM with at least two disks on KVM/`SharedMountPoint` (or 
`Filesystem`/`NetworkFilesystem`) primary storage - one small disk, one large 
disk with enough real delta data that a commit takes noticeably longer than 
`qcow2.delta.merge.timeout` (or 1 hour, if using the running-VM/non-events 
path).
   2. Take a disk-only VM snapshot.
   3. Write enough data to the large disk that its subsequent commit will 
exceed the timeout.
   4. Delete the VM snapshot.
   5. Observe the small disk's merge completes and its delta file is deleted, 
while the large disk's commit is killed by the timeout.
   
   
   
   
   ### What to do about it?
   
   - The operation is fully atomic: no disk's file is deleted/merged unless 
*all* disks in the snapshot are confirmed to have completed successfully, and 
the operation cleanly fails/rolls back to a well-defined error state if any 
disk times out; or
   - Per-disk completion is persisted incrementally as each disk finishes, so a 
disk that successfully merged is never left with a database record pointing at 
a deleted file, regardless of what happens to other disks in the same VM 
snapshot.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to