> Hi Vaibhav,
> 
> I have tested this v2 patch and it is still giving me the issue I reported,
> so here is my analysis:
> 
> a) Without applying the patch :
> 
> 1) Start the guest and run stress-ng as below for sometime
> localhost:~ # stress-ng --cpu 4 --vm 2 --vm-bytes 1G --hdd 2 --hdd-bytes 1G
> --sched other --timeout 3600000s
> stress-ng: info:  [1464] setting to a 41 days, 16 hours, 0 secs run per
> stressor
> stress-ng: info:  [1464] dispatching hogs: 4 cpu, 2 vm, 2 hdd
> 
> 
> 
> 2) Start the migration from H1 to H2:
> 
> ltc-lp7:~ # virsh migrate --live --domain sles16_anu
> qemu+ssh://10.xx.xx.xx/system --verbose --undefinesource --persistent
> --auto-converge --postcopy
> ([email protected]) Password:
> Migration: [100.00 %]
> 
> 3) Migration got completed but guest is not getting recovered from
> continuous softlockups
> 
> [ 1336.003836][    C1] watchdog: BUG: soft lockup - CPU#1 stuck for 977s!
> [htxd_monitor:1337]
> [ 1336.006834][    C4] watchdog: BUG: soft lockup - CPU#4 stuck for 1002s!
> [rcu_exp_par_gp_:19]
> [ 1346.015839][    C0] BUG: workqueue lockup - pool cpus=1 node=0 flags=0x0
> nice=0 stuck for 1090s!
> [ 1346.016355][    C0] BUG: workqueue lockup - pool cpus=3 node=0 flags=0x0
> nice=0 stuck for 1107s!
> [ 1346.016874][    C0] BUG: workqueue lockup - pool cpus=7 node=0 flags=0x0
> nice=0 stuck for 1093s!
> [ 1356.007835][    C6] watchdog: BUG: soft lockup - CPU#6 stuck for 912s!
> [systemd:1353]
> [ 1356.008835][    C7] watchdog: BUG: soft lockup - CPU#7 stuck for 998s!
> [systemd-journal:570]
> [ 1360.003836][    C1] watchdog: BUG: soft lockup - CPU#1 stuck for 999s!
> [htxd_monitor:1337]
> [ 1360.006834][    C4] watchdog: BUG: soft lockup - CPU#4 stuck for 1024s!
> [rcu_exp_par_gp_:19]
> [ 1368.933835][    C4] rcu: INFO: rcu_preempt self-detected stall on CPU
> [ 1368.933973][    C4] rcu:     4-....: (1129830 ticks this GP)
> idle=afc4/1/0x4000000000000002 softirq=3694/428556 fqs=259639
> [ 1368.934106][    C4] rcu:              hardirqs   softirqs  csw/system
> [ 1368.934188][    C4] rcu:      number:        1     444039 0
> [ 1368.934271][    C4] rcu:     cputime:        3          8 1096165   ==>
> 1110021(ms)
> [ 1368.934373][    C4] rcu:     (t=1140022 jiffies g=6177 q=1684 ncpus=8)
> [ 1376.224839][    C0] BUG: workqueue lockup - pool cpus=1 node=0 flags=0x0
> nice=0 stuck for 1120s!
> [ 1376.225307][    C0] BUG: workqueue lockup - pool cpus=3 node=0 flags=0x0
> nice=0 stuck for 1138s!
> [ 1376.225428][    C0] BUG: workqueue lockup - pool cpus=5 node=0 flags=0x0
> nice=0 stuck for 715s!
> [ 1376.225548][    C0] BUG: workqueue lockup - pool cpus=6 node=0 flags=0x0
> nice=0 stuck for 1027s!
> [ 1376.225667][    C0] BUG: workqueue lockup - pool cpus=7 node=0 flags=0x0
> nice=0 stuck for 1123s!
> [ 1444.006835][    C4] watchdog: BUG: soft lockup - CPU#4 stuck for 1100s!
> [rcu_exp_par_gp_:19]
> 
> 
> b) Even after applying the patch also it is giving same softlockup issue as
> mentioned above:
> Though I have enough vcpus and memory on the guest (16 vcpus , 13Gi of
> memory)  and ample amount of memory and cpus
> present on host still these softlockups are happening after applying the
> patch too. I tried reducing stress also on the guest
> but still this issue is seen.
> 
> stress-ng --cpu 4 --vm 2 --vm-bytes 1G --hdd 2 --hdd-bytes 1G --sched other
> --timeout 3600000s
> 
> If you are planning to send next version of this patch,
> Please do add my reported-by:
> Reported-by: Anushree Mathur <[email protected]>
> 

Thanks for testing this Anushree. Upon further debugging, I found that
this issue is likely due to a bug in QEMU, for which I've sent a fix[1].

@maddy: Please don't pull in this patch for now, will ping here if this
is needed.

[1] : lore.kernel.org/qemu-devel/[email protected]/

Thanks,
Gautam

Reply via email to