kvmppc_hv_find_lock_hpte() requires virtual-mode callers to run with
preemption disabled, because it can return with HPTE_V_HVLOCK still held
until the caller later unlocks the HPTE.  Existing virtual-mode callers
in book3s_64_mmu_hv.c already follow that rule.

Two virtual-mode call paths do not:

1. kvmppc_handle_exit_hv() calls kvmppc_hpte_hv_fault() for hash-mode
   data-side and instruction-side faults after guest exit, with
   preemption already enabled.  kvmppc_hpte_hv_fault() calls
   kvmppc_hv_find_lock_hpte() directly.

2. kvmppc_pseries_do_hcall() executes virtual-mode handlers for HPT
   hcalls via kvmppc_pseries_do_hpt_hcall() with preemption enabled.
   The handlers for H_ENTER, H_REMOVE, H_READ, H_CLEAR_MOD, H_CLEAR_REF,
   H_PROTECT, and H_BULK_REMOVE spin on try_lock_hpte() or lock_rmap(),
   and H_ENTER also reaches kvmppc_do_h_enter(), which uses
   arch_spin_lock() on kvm->mmu_lock.  That raw lock choice is
   intentional because kvmppc_do_h_enter() is also used by real-mode
   callers, so the correct fix is to establish the proper preemption
   context at the virtual-mode caller boundary.

If a vCPU thread is preempted while holding HPTE_V_HVLOCK, any other
thread on the same CPU trying to acquire the same bit-lock can spin
indefinitely.  The same applies to lock_rmap() users and to the H_ENTER
path while holding the raw mmu_lock.

Fix this by adding preempt_disable()/preempt_enable() pairs around the
two kvmppc_hpte_hv_fault() call sites in kvmppc_handle_exit_hv(), and
around the kvmppc_pseries_do_hpt_hcall() invocation in
kvmppc_pseries_do_hcall().  Targeting only the HPT hcall helper avoids
wrapping non-HPT hcalls (such as H_RTAS, H_CONFER, H_REGISTER_VPA) that
perform page faults, blocking memory allocations, or context switches.

Fixes: 6165d5dd99db ("KVM: PPC: Book3S HV: add virtual mode handlers for HPT 
hcalls and page faults")
Cc: [email protected] # v5.14+
Signed-off-by: Amit Machhiwal <[email protected]>
---
 arch/powerpc/kvm/book3s_hv.c | 6 ++++++
 1 file changed, 6 insertions(+)

diff --git a/arch/powerpc/kvm/book3s_hv.c b/arch/powerpc/kvm/book3s_hv.c
index aa51968e206a..52f72f30baf0 100644
--- a/arch/powerpc/kvm/book3s_hv.c
+++ b/arch/powerpc/kvm/book3s_hv.c
@@ -1212,9 +1212,11 @@ int kvmppc_pseries_do_hcall(struct kvm_vcpu *vcpu)
        case H_CLEAR_REF:
        case H_PROTECT:
        case H_BULK_REMOVE:
+               preempt_disable();
                idx = srcu_read_lock(&kvm->srcu);
                ret = kvmppc_pseries_do_hpt_hcall(vcpu, req);
                srcu_read_unlock(&kvm->srcu, idx);
+               preempt_enable();
                if (ret == H_TOO_HARD)
                        return RESUME_HOST;
                break;
@@ -1834,8 +1836,10 @@ static int kvmppc_handle_exit_hv(struct kvm_vcpu *vcpu,
                else
                        vsid = vcpu->arch.fault_gpa;
 
+               preempt_disable();
                err = kvmppc_hpte_hv_fault(vcpu, vcpu->arch.fault_dar,
                                vsid, vcpu->arch.fault_dsisr, true);
+               preempt_enable();
                if (err == 0) {
                        r = RESUME_GUEST;
                } else if (err == -1 || err == -2) {
@@ -1881,8 +1885,10 @@ static int kvmppc_handle_exit_hv(struct kvm_vcpu *vcpu,
                else
                        vsid = vcpu->arch.fault_gpa;
 
+               preempt_disable();
                err = kvmppc_hpte_hv_fault(vcpu, vcpu->arch.fault_dar,
                                vsid, vcpu->arch.fault_dsisr, false);
+               preempt_enable();
                if (err == 0) {
                        r = RESUME_GUEST;
                } else if (err == -1) {
-- 
2.54.0 (Apple Git-157)


Reply via email to