On 12/08/26 9:16 AM, Vaibhav Jain wrote:
On nestedv2 the L1 converts a pending doorbell into guest DPDES state at
the top of kvmhv_vcpu_entry_nestedv2() and immediately forgets about it:

        if (vcpu->arch.doorbell_request) {
                vcpu->arch.doorbell_request = 0;
                kvmppc_set_dpdes(vcpu, 1);
        }

Clearing 'doorbell_request' at this point assumes that handing DPDES to the
L0 is equivalent to the L2 having taken the doorbell. That is not true, and
the doorbell can be lost in two ways:

   - The block runs before the lazy_irq_pending() check, so the doorbell is
     consumed even on the path that returns 0 without ever calling
     H_GUEST_RUN_VCPU.

   - DPDES stays pending in the L2 until it is actually delivered. The L2
     may exit for an unrelated reason (hcall, page fault, HDEC) with the
     doorbell still set, typically because it was running with MSR[EE]=0.
     Nothing reloads DPDES afterwards, so the L1 never learns this.

Once 'doorbell_request' has been cleared, the L1 has no record of the
pending doorbell. kvmppc_doorbell_pending() returns false, so
kvmppc_read_dpdes() reports the target thread as idle when a sibling vCPU
emulates 'mfspr DPDES', and the vCPU can be treated as having no work
pending and blocked. From the L2's point of view the doorbell is silently
lost, which shows up as an SMT guest hanging on a doorbell-based IPI.

Fix this by converting 'doorbell_request' to an atomic to track the L2's
DPDES rather than being consumed by entry. This also enables new doorbells
to be queued up on a vcpu while and active doorbell request is being
handled.

Since kvm_vcpu_arch 'doorbell_request' is also shared across APIv1 and
Bare-Metal KVM infrastructure on PPC64, the patch also updates these
paths to use kernel's 'atomic_t' helpers.

With this following changes to kvmhv_vcpu_entry_nestedv2() code flow are
done:

   - inject DPDES after the early-return paths and before
     kvmhv_nestedv2_flush_vcpu() serializes it into the vcpu run input
     buffer, and no longer clear 'doorbell_request' there,

   - after H_GUEST_RUN_VCPU, reload DPDES from the L0. The run output only
     carries the state the L0 chose to return and the 'valids' bitmap is
     zeroed on exit, so an explicit kvmhv_nestedv2_cached_reload() is
     needed to see the L2's current value,

   - if DPDES is still set the doorbell was not delivered, so keep
     'doorbell_request' pending so that it is re-injected on the next
     entry; otherwise clear it.

This keeps a pending doorbell visible to the L1 for as long as the L2 has
not consumed it, so vCPU wakeup and DPDES emulation on sibling vCPUs stay
consistent with the L2's actual state.

Testing
=======
This patch has been tested with continuous migration for a APIv2 KVM guest
running stress-ng workload with lots of IPIs flying around, top of a
PowerVM-LPAR and seems to work. Without the patch the migration loop used
to eventually endup in KVM guest reporting lockups, which arent seen
anymore with this patch.

I haven't tested it yet with Bare-Metal nor APIv1 based PPC64 KVM
guests. Will be able to share their results (Hopefully) in next version of
this patch.

Fixes: 54ec2bd9e017 ("KVM: PPC: Book3S HV nestedv2: Fix doorbell emulation")
Signed-off-by: Vaibhav Jain <[email protected]>
Assisted-by: Claude:Opus-5

---
Change-log:
Since V1:
https://lore.kernel.org/linuxppc-dev/[email protected]/T/#u
* Address review comments around a typo [Shashiko]
* Address a race condition which Shashiko had pointed to [Shashiko]
* Updated the patch description and added test results
* Refactored 'doorbell_request' to an 'atomic_t'
* Updated code to use atomic_t helpers
---
  arch/powerpc/include/asm/kvm_host.h   |  2 +-
  arch/powerpc/kvm/book3s_hv.c          | 39 ++++++++++++++++-----------
  arch/powerpc/kvm/book3s_hv_builtin.c  |  3 +--
  arch/powerpc/kvm/book3s_hv_nested.c   |  9 ++++---
  arch/powerpc/kvm/book3s_hv_p9_entry.c |  7 +++--
  5 files changed, 34 insertions(+), 26 deletions(-)

diff --git a/arch/powerpc/include/asm/kvm_host.h 
b/arch/powerpc/include/asm/kvm_host.h
index 2d139c807577..03fac7422aa2 100644
--- a/arch/powerpc/include/asm/kvm_host.h
+++ b/arch/powerpc/include/asm/kvm_host.h
@@ -775,7 +775,7 @@ struct kvm_vcpu_arch {
        unsigned long pending_exceptions;
        u8 ceded;
        u8 prodded;
-       u8 doorbell_request;
+       atomic_t doorbell_request;
        u8 irq_pending; /* Used by XIVE to signal pending guest irqs */
        unsigned long last_inst;
diff --git a/arch/powerpc/kvm/book3s_hv.c b/arch/powerpc/kvm/book3s_hv.c
index 61dbeea317f3..fb80c13dbb29 100644
--- a/arch/powerpc/kvm/book3s_hv.c
+++ b/arch/powerpc/kvm/book3s_hv.c
@@ -15,6 +15,7 @@
   * by Alexander Graf <[email protected]>.
   */
+#include "asm/guest-state-buffer.h"
  #include <linux/kvm_host.h>
  #include <linux/kernel.h>
  #include <linux/err.h>
@@ -881,7 +882,7 @@ static bool kvmppc_doorbell_pending(struct kvm_vcpu *vcpu)
        int thr;
        struct kvmppc_vcore *vc;
- if (vcpu->arch.doorbell_request)
+       if (atomic_read(&vcpu->arch.doorbell_request))
                return true;
        if (cpu_has_feature(CPU_FTR_ARCH_300))
                return false;
@@ -1557,17 +1558,16 @@ static int kvmppc_emulate_doorbell_instr(struct 
kvm_vcpu *vcpu)
                tvcpu = kvmppc_find_vcpu(kvm, vcpu->vcpu_id - thr + arg);
                if (!tvcpu)
                        break;
-               if (!tvcpu->arch.doorbell_request) {
-                       tvcpu->arch.doorbell_request = 1;
+               if (atomic_inc_return(&tvcpu->arch.doorbell_request) >= 1)
                        kvmppc_fast_vcpu_kick_hv(tvcpu);
-               }
                break;
        case OP_31_XOP_MSGCLRP:
                arg = kvmppc_get_gpr(vcpu, rb);
                if (((arg >> 27) & 0x1f) != PPC_DBELL_SERVER)
                        break;
-               vcpu->arch.vcore->dpdes = 0;
-               vcpu->arch.doorbell_request = 0;
+
+               if (atomic_dec_return(&vcpu->arch.doorbell_request) <= 0)
+                       vcpu->arch.vcore->dpdes = 0;
                break;
        case OP_31_XOP_MFSPR:
                switch (get_sprn(inst)) {
@@ -2313,7 +2313,7 @@ static int kvmppc_get_one_reg_hv(struct kvm_vcpu *vcpu, 
u64 id,
                 * On POWER8, doorbell_request is 0.
                 */
                if (cpu_has_feature(CPU_FTR_ARCH_300))
-                       *val = get_reg_val(id, vcpu->arch.doorbell_request);
+                       *val = get_reg_val(id, 
atomic_read(&vcpu->arch.doorbell_request));
                else
                        *val = get_reg_val(id, vcpu->arch.vcore->dpdes);
                break;
@@ -2565,7 +2565,7 @@ static int kvmppc_set_one_reg_hv(struct kvm_vcpu *vcpu, 
u64 id,
                break;
        case KVM_REG_PPC_DPDES:
                if (cpu_has_feature(CPU_FTR_ARCH_300))
-                       vcpu->arch.doorbell_request = set_reg_val(id, *val) & 1;
+                       atomic_inc(&vcpu->arch.doorbell_request);
                else
                        vcpu->arch.vcore->dpdes = set_reg_val(id, *val);
                break;
@@ -4253,10 +4253,6 @@ static int kvmhv_vcpu_entry_nestedv2(struct kvm_vcpu 
*vcpu, u64 time_limit,
        int trap;
        long rc;
- if (vcpu->arch.doorbell_request) {
-               vcpu->arch.doorbell_request = 0;
-               kvmppc_set_dpdes(vcpu, 1);
-       }
io = &vcpu->arch.nestedv2_io; @@ -4265,6 +4261,10 @@ static int kvmhv_vcpu_entry_nestedv2(struct kvm_vcpu *vcpu, u64 time_limit,
        if (lazy_irq_pending())
                return 0;
+ /* Set DPDES if any doorbell is requested */
+       if (atomic_read(&vcpu->arch.doorbell_request) > 0)
+               kvmppc_set_dpdes(vcpu, 1);
+
        rc = kvmhv_nestedv2_flush_vcpu(vcpu, time_limit);
        if (rc < 0)
                return -EINVAL;
@@ -4298,6 +4298,17 @@ static int kvmhv_vcpu_entry_nestedv2(struct kvm_vcpu 
*vcpu, u64 time_limit,
timer_rearm_host_dec(*tb); + /* Check if privileged door bell was requested and handled */
+       if (atomic_read(&vcpu->arch.doorbell_request) > 0) {
+               /* In case PHYP doesn't return updated dpdes in output gsb */
+               if (vcpu->arch.vcore->dpdes)
+                       kvmhv_nestedv2_cached_reload(vcpu,
+                                                    KVMPPC_GSID_DPDES);
+               /* if dpdes was handled then reduce the doorbell count */
+               if (!vcpu->arch.vcore->dpdes)
+                       atomic_dec(&vcpu->arch.doorbell_request);
+       }
+
        /* Record context switch and guest_run_time data */
        if (kvmhv_get_l2_counters_status())
                do_trace_nested_cs_time(vcpu);
@@ -4356,9 +4367,7 @@ static int kvmhv_vcpu_entry_p9_nested(struct kvm_vcpu 
*vcpu, u64 time_limit, uns
         * enables us to receive doorbells when H_ENTER_NESTED is
         * in progress for this vCPU
         */
-
-       if (vcpu->arch.doorbell_request)
-               vcpu->arch.doorbell_request = 0;
+       atomic_set(&vcpu->arch.doorbell_request, 0);
/*
         * When setting DEC, we must always deal with irq_work_raise
diff --git a/arch/powerpc/kvm/book3s_hv_builtin.c 
b/arch/powerpc/kvm/book3s_hv_builtin.c
index fa0e3a22cac0..76bda278cff9 100644
--- a/arch/powerpc/kvm/book3s_hv_builtin.c
+++ b/arch/powerpc/kvm/book3s_hv_builtin.c
@@ -594,11 +594,10 @@ void kvmppc_guest_entry_inject_int(struct kvm_vcpu *vcpu)
                }
        }
- if (vcpu->arch.doorbell_request) {
+       if (atomic_dec_if_positive(&vcpu->arch.doorbell_request) >= 0) {
                mtspr(SPRN_DPDES, 1);
                vcpu->arch.vcore->dpdes = 1;
                smp_wmb();
-               vcpu->arch.doorbell_request = 0;
        }
  }
diff --git a/arch/powerpc/kvm/book3s_hv_nested.c b/arch/powerpc/kvm/book3s_hv_nested.c
index 22e616662255..cc81089b2bcd 100644
--- a/arch/powerpc/kvm/book3s_hv_nested.c
+++ b/arch/powerpc/kvm/book3s_hv_nested.c
@@ -32,7 +32,7 @@ void kvmhv_save_hv_regs(struct kvm_vcpu *vcpu, struct 
hv_guest_state *hr)
        struct kvmppc_vcore *vc = vcpu->arch.vcore;
hr->pcr = vc->pcr | PCR_MASK;
-       hr->dpdes = vcpu->arch.doorbell_request;
+       hr->dpdes = atomic_read(&vcpu->arch.doorbell_request);
        hr->hfscr = vcpu->arch.hfscr;
        hr->tb_offset = vc->tb_offset;
        hr->dawr0 = vcpu->arch.dawr0;
@@ -105,7 +105,7 @@ static void save_hv_return_state(struct kvm_vcpu *vcpu,
  {
        struct kvmppc_vcore *vc = vcpu->arch.vcore;
- hr->dpdes = vcpu->arch.doorbell_request;
+       hr->dpdes = atomic_read(&vcpu->arch.doorbell_request);
        hr->purr = vcpu->arch.purr;
        hr->spurr = vcpu->arch.spurr;
        hr->ic = vcpu->arch.ic;
@@ -143,7 +143,7 @@ static void restore_hv_regs(struct kvm_vcpu *vcpu, const 
struct hv_guest_state *
        struct kvmppc_vcore *vc = vcpu->arch.vcore;
vc->pcr = hr->pcr | PCR_MASK;
-       vcpu->arch.doorbell_request = hr->dpdes;
+       atomic_set(&vcpu->arch.doorbell_request, hr->dpdes);
        vcpu->arch.hfscr = hr->hfscr;
        vcpu->arch.dawr0 = hr->dawr0;
        vcpu->arch.dawrx0 = hr->dawrx0;
@@ -176,7 +176,8 @@ void kvmhv_restore_hv_return_state(struct kvm_vcpu *vcpu,
         *   a) Sent after H_ENTER_NESTED was called on this vCPU 
(arch.doorbell_request would be 1)
         *   b) Doorbell was not handled and L2 exited for some other reason 
(hr->dpdes would be 1)
         */
-       vcpu->arch.doorbell_request = vcpu->arch.doorbell_request | hr->dpdes;
+       if (hr->dpdes)
+               atomic_inc(&vcpu->arch.doorbell_request);
        vcpu->arch.hfscr = hr->hfscr;
        vcpu->arch.purr = hr->purr;
        vcpu->arch.spurr = hr->spurr;
diff --git a/arch/powerpc/kvm/book3s_hv_p9_entry.c 
b/arch/powerpc/kvm/book3s_hv_p9_entry.c
index 34bc0a8a1288..fc44d8aa7a6a 100644
--- a/arch/powerpc/kvm/book3s_hv_p9_entry.c
+++ b/arch/powerpc/kvm/book3s_hv_p9_entry.c
@@ -610,10 +610,9 @@ int kvmhv_vcpu_entry_p9(struct kvm_vcpu *vcpu, u64 
time_limit, unsigned long lpc
if (vc->pcr)
                mtspr(SPRN_PCR, vc->pcr | PCR_MASK);
-       if (vcpu->arch.doorbell_request) {
-               vcpu->arch.doorbell_request = 0;
+
+       if (atomic_dec_if_positive(&vcpu->arch.doorbell_request) >= 0)
                mtspr(SPRN_DPDES, 1);
-       }
if (dawr_enabled()) {
                if (vcpu->arch.dawr0 != host_dawr0)
@@ -838,7 +837,7 @@ int kvmhv_vcpu_entry_p9(struct kvm_vcpu *vcpu, u64 
time_limit, unsigned long lpc
dpdes = mfspr(SPRN_DPDES);
        if (dpdes)
-               vcpu->arch.doorbell_request = 1;
+               atomic_inc(&vcpu->arch.doorbell_request);
vc->vtb = mfspr(SPRN_VTB);
Hi Vaibhav,

I have tested this v2 patch and it is still giving me the issue I reported, so here is my analysis:

a) Without applying the patch :

1) Start the guest and run stress-ng as below for sometime
localhost:~ # stress-ng --cpu 4 --vm 2 --vm-bytes 1G --hdd 2 --hdd-bytes 1G --sched other --timeout 3600000s stress-ng: info:  [1464] setting to a 41 days, 16 hours, 0 secs run per stressor
stress-ng: info:  [1464] dispatching hogs: 4 cpu, 2 vm, 2 hdd



2) Start the migration from H1 to H2:

ltc-lp7:~ # virsh migrate --live --domain sles16_anu qemu+ssh://10.xx.xx.xx/system --verbose --undefinesource --persistent --auto-converge --postcopy
([email protected]) Password:
Migration: [100.00 %]

3) Migration got completed but guest is not getting recovered from continuous softlockups

[ 1336.003836][    C1] watchdog: BUG: soft lockup - CPU#1 stuck for 977s! [htxd_monitor:1337] [ 1336.006834][    C4] watchdog: BUG: soft lockup - CPU#4 stuck for 1002s! [rcu_exp_par_gp_:19] [ 1346.015839][    C0] BUG: workqueue lockup - pool cpus=1 node=0 flags=0x0 nice=0 stuck for 1090s! [ 1346.016355][    C0] BUG: workqueue lockup - pool cpus=3 node=0 flags=0x0 nice=0 stuck for 1107s! [ 1346.016874][    C0] BUG: workqueue lockup - pool cpus=7 node=0 flags=0x0 nice=0 stuck for 1093s! [ 1356.007835][    C6] watchdog: BUG: soft lockup - CPU#6 stuck for 912s! [systemd:1353] [ 1356.008835][    C7] watchdog: BUG: soft lockup - CPU#7 stuck for 998s! [systemd-journal:570] [ 1360.003836][    C1] watchdog: BUG: soft lockup - CPU#1 stuck for 999s! [htxd_monitor:1337] [ 1360.006834][    C4] watchdog: BUG: soft lockup - CPU#4 stuck for 1024s! [rcu_exp_par_gp_:19]
[ 1368.933835][    C4] rcu: INFO: rcu_preempt self-detected stall on CPU
[ 1368.933973][    C4] rcu:     4-....: (1129830 ticks this GP) idle=afc4/1/0x4000000000000002 softirq=3694/428556 fqs=259639
[ 1368.934106][    C4] rcu:              hardirqs   softirqs  csw/system
[ 1368.934188][    C4] rcu:      number:        1     444039 0
[ 1368.934271][    C4] rcu:     cputime:        3          8 1096165   ==> 1110021(ms)
[ 1368.934373][    C4] rcu:     (t=1140022 jiffies g=6177 q=1684 ncpus=8)
[ 1376.224839][    C0] BUG: workqueue lockup - pool cpus=1 node=0 flags=0x0 nice=0 stuck for 1120s! [ 1376.225307][    C0] BUG: workqueue lockup - pool cpus=3 node=0 flags=0x0 nice=0 stuck for 1138s! [ 1376.225428][    C0] BUG: workqueue lockup - pool cpus=5 node=0 flags=0x0 nice=0 stuck for 715s! [ 1376.225548][    C0] BUG: workqueue lockup - pool cpus=6 node=0 flags=0x0 nice=0 stuck for 1027s! [ 1376.225667][    C0] BUG: workqueue lockup - pool cpus=7 node=0 flags=0x0 nice=0 stuck for 1123s! [ 1444.006835][    C4] watchdog: BUG: soft lockup - CPU#4 stuck for 1100s! [rcu_exp_par_gp_:19]


b) Even after applying the patch also it is giving same softlockup issue as mentioned above: Though I have enough vcpus and memory on the guest (16 vcpus , 13Gi of memory)  and ample amount of memory and cpus present on host still these softlockups are happening after applying the patch too. I tried reducing stress also on the guest
but still this issue is seen.

stress-ng --cpu 4 --vm 2 --vm-bytes 1G --hdd 2 --hdd-bytes 1G --sched other --timeout 3600000s

If you are planning to send next version of this patch,
Please do add my reported-by:
Reported-by: Anushree Mathur <[email protected]>

Reply via email to