Pearl1594 opened a new issue, #14177:
URL: https://github.com/apache/cloudstack/issues/14177

   ### problem
   
   When a VM's domain is detected missing (`PowerReportMissing`) past the 
graceful period (`vm.op.wait.interval`), CloudStack releases its NIC 
(`broadcast_uri`/`isolation_uri` -> NULL, `nics_count`) and marks the VM 
`Stopped`. If the VM's power state later resyncs back to `Running` (domain 
restored, HA restart, etc.), the NIC is never re-reserved and `nics_count` is 
never restored.
   Separately, `op_networks.nics_count` never counts the network's own 
VirtualRouter NIC, so it under-counts from network creation. Combined, a single 
NIC release event can drive `nics_count` to 0 while the network still has live 
NICs (including on a `Running` VM). 
   `NetworkGarbageCollector` trusts `nics_count==0` (plus a check of 
CloudStack's own DB-tracked "no non-Stopped instances") without verifying 
against the actual `nics` table or hypervisor state, and proceeds to stop the 
VR and release the VLAN back to the dynamic allocation pool - while it may 
still be bridged to a live VM. In our environment, GC's cleanup step also 
removes the host-level bridge, making the VM unrecoverable via a normal restart.
   
   ### versions
   
   tested on 4.22 (Maybe be observed on older versions too)
   
   ### The steps to reproduce the bug
   
   1. Deploy a VM on an isolated network - note the nics_count = 1 (instead of 
2 for VM and VR)
   2. Disable HA on the VM (as HA masks the issue)
   3. Perform `virsh destroy <vm-name>` to simulate domain loss - backup the 
dumpxml `virsh dumpxml <vm-name> > backup.xml`
   4. After the graceful period, CloudStack releases the NIC and decrements 
`nics_count` to 0.
   5. Restore the domain via `virsh create <dumped-xml>` . VM resyncs to 
`Running`, but NIC stays unreserved, counter stays 0
   6. Lower `network.gc.interval`/`network.gc.wait` to accelerate the 
scavenger; restart management server.
   7. `virsh destroy` again to bring the VM's tracked state back to `Stopped` 
   8. GC fires: stops the VR, releases the VLAN, network transitions to 
`Allocated`  `broadcast_uri=NULL`, despite 2 live NICs still present in `nics`.
   ...
   
   
   ### What to do about it?
   
   - On VM power-state resync to `Running` after a missing-VM event, reconcile 
the NIC (`broadcast_uri`/`isolation_uri`/`state=Reserved`) and restore 
`nics_count`.
   - Harden `NetworkGarbageCollector` to verify live NIC count from the `nics` 
table (`WHERE network_id=? AND removed IS NULL`) before tearing down a network, 
instead of trusting the cached counter alone
   - Include the VirtualRouter's own NIC in `nics_count` from network creation


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to