bhouse-nexthop opened a new issue, #14232:
URL: https://github.com/apache/cloudstack/issues/14232

   ##### ISSUE TYPE
    * Bug Report
   
   ##### COMPONENT NAME
   ~~~
   engine-orchestration, destroyVirtualMachine, KVM
   ~~~
   
   ##### CLOUDSTACK VERSION
   ~~~
   4.22
   ~~~
   
   ##### CONFIGURATION
   `vm.destroy.forcestop=true`. KVM hosts. Any rolling restart of the agents or 
the management servers while instances are being destroyed.
   
   ##### OS / ENVIRONMENT
   KVM / libvirt.
   
   ##### SUMMARY
   
   With `vm.destroy.forcestop=true`, destroying an instance whose host is 
briefly disconnected releases the instance's NICs, IP addresses and volumes 
without stopping it. The domain keeps running on the host with no record in 
CloudStack. Its IP is handed to the next instance, so two live instances answer 
for the same address, and its root volume stays in `Destroy` because the delete 
fails while the domain holds the image.
   
   Same end state as #14206, different trigger. #14206 is a stale power report 
during start; this one is the destroy path itself.
   
   ##### STEPS TO REPRODUCE
   
   1. Set `vm.destroy.forcestop=true`.
   2. Restart `cloudstack-agent` on a KVM host (or a management server, which 
disconnects the agents connected to it while they rebalance).
   3. While the host is `Disconnected`, destroy a running instance on it with 
`expunge=true`.
   
   Any client that destroys instances routinely hits this on every rolling 
upgrade.
   
   ##### EXPECTED RESULTS
   
   The destroy fails, the instance stays `Running`, and the caller retries once 
the host is back. Or the destroy waits for the host.
   
   ##### ACTUAL RESULTS
   
   From a real occurrence on 4.22, during a rolling package upgrade:
   
   ~~~
   13:39:02 WARN  Unable to stop VM instance 
{"id":3645123,...,"state":"Stopping"} due to [AgentUnavailableException:
                  Resource [Host:110] is unreachable: Host 110: Host with 
specified id is not in the right state: Disconnected]
   13:39:02 WARN  Unable to actually stop VM instance {"id":3645123,...} but 
continue with release because it's a force stop
   13:39:02 DEBUG VM instance {"id":3645123,...} is stopped on the host.  
Proceeding to release resource held.
   13:39:02 DEBUG Successfully released network resources for the VM ...
   13:39:15 DEBUG Expunged VM instance {"id":3645123,...}
   13:41:58 WARN  Host reports 9 instance(s) that do not exist in CloudStack 
DB, they are running unmanaged. host: hv104, instances: [i-625-3645123-VM, ...]
   ~~~
   
   In one rolling upgrade, 31 instances across 6 hosts were left running this 
way, and new instances were given their addresses within hours.
   
   ##### CAUSE
   
   `UserVmManagerImpl.destroyVm(DestroyVMCmd)`, 
`VirtualMachineManagerImpl.destroy()` and 
`VirtualMachineManagerImpl.advanceExpunge()` all stop the instance with 
`cleanUpEvenIfUnableToStop = vm.destroy.forcestop`. In `advanceStop()`, a 
forced stop that gets `AgentUnavailableException` or 
`OperationTimedoutException` releases the resources and marks the instance 
stopped:
   
   ~~~java
   } catch (AgentUnavailableException | OperationTimedoutException e) {
       logger.warn("Unable to stop {} due to [{}].", ...);
   } finally {
       if (!stopped) {
           if (!cleanUpEvenIfUnableToStop) { ... throw ... }
           else { logger.warn("Unable to actually stop {} but continue with 
release because it's a force stop", vm); ... }
       }
   }
   releaseVmResources(profile, cleanUpEvenIfUnableToStop);
   ~~~
   
   A forced stop means "the host cannot tell us, treat the instance as 
stopped". That is right when the host is gone (`Down`, `Removed`, `Error`). It 
is wrong when the host is `Disconnected`, `Connecting`, `Alert` or 
`Rebalancing`: those hosts are expected back with their domains still running.
   
   `vm.destroy.forcestop` is a global setting applied to every destroy, not a 
statement by the caller that it knows the host is gone. It should not bypass 
that distinction.
   
   ##### IMPACT
   
   - an instance runs unmanaged and invisible
   - its IP is reassigned, giving an address conflict between two live instances
   - its root volume stays in `Destroy` and the storage cleanup fails on it 
every run
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to