prashanthr2 commented on issue #13971:
URL: https://github.com/apache/cloudstack/issues/13971#issuecomment-5439397520

   @gvaishya 
   
   Thanks for the report. This does not looks like a bug, ideally should be a 
discussions thread. A few things to separate out.
   
   - host-passthrough is unsafe for this cluster. Your hosts mix Skylake-SP 
(6148, 6154) with Cascade Lake-SP (6254). host-passthrough passes the exact 
host CPU to the guest, so a VM's visible CPUID changes depending on which host 
it starts on. Pin to the lowest common denominator, identically on all hosts:
   
   > guest.cpu.mode=custom
   
   > guest.cpu.model=Skylake-Server
   
   **NOTE**: Before changing, confirm the exact model string with virsh 
domcapabilities on a 6148 host (you may need -IBRS or -noTSX-IBRS). Needs an 
agent restart, and existing VMs need a stop/start (not a guest reboot).
   
   
   - guest.cpu.features=vmx , is this intentional? This emits <feature 
policy='require' name='vmx'/>, exposing nested virtualisation to every guest on 
that host. AFAIK for VDI there's usually no reason for this. This is the first 
thing I'd remove and check
   
   
   
   - Also, Check the 1 CPU x 2.4 GHz offering for CPU cap:
   `SELECT id, name, cpu, speed, limit_cpu_use FROM cloud.service_offering 
WHERE uuid = <offering_id_from_UI>;`
   If limit_cpu_use = 1, a CFS <quota> is applied (virsh dumpxml <vm> | grep 
-A4 cputune). 2.4 GHz is ~77% of a core on a 3.1 GHz host but ~100% on a 2.4 
GHz host, so the same offering throttles differently by placement. For VDI that 
presents as hangs/timeouts users report as crashes.
   
   
   To identify the actual cause, for one crash with a timestamp ( 
**collect/check below if the above suggested doesn't resolve the issue)**
   
   **1. Do crashes correlate with one host? If so this is a faulty host, not a 
mixed-CPU issue** --> This is the first thing I will check to isolate the issue 
further
   2. /var/log/libvirt/qemu/i-<acct>-<vmid>-VM.log , a QEMU abort or 
reason=crashed points at guest/QEMU; nothing in the log means the host killed 
the process
   3. Host dmesg -T / journalctl -k . OOM killer reaping qemu-kvm, or MCE 
entries
   4. Guest-side: Windows bugcheck code, or Linux kernel panic
   
   Could you paste grep -v '^#' /etc/cloudstack/agent/agent.properties | grep 
-v '^$' from two hosts (one 6254, one 6148)? That confirms the setting is 
actually applied and consistent.
   
   **Summary**: fix the CPU mode regardless, but with no migration occurring 
I'd rank vmx and CPU-cap throttling as the likelier causes. The libvirt/QEMU 
log plus guest bugcheck will help isolate it further.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to