kvmppc_hv_find_lock_hpte() requires virtual-mode callers to run with
preemption disabled, because it can return with HPTE_V_HVLOCK still held
until the caller later unlocks the HPTE. Existing virtual-mode callers
in book3s_64_mmu_hv.c already follow that rule.
Two virtual-mode call paths do not:
1. kvmppc_handle_exit_hv() calls kvmppc_hpte_hv_fault() for hash-mode
data-side and instruction-side faults after guest exit, with
preemption already enabled. kvmppc_hpte_hv_fault() calls
kvmppc_hv_find_lock_hpte() directly.
2. kvmppc_pseries_do_hcall() executes virtual-mode handlers for HPT
hcalls via kvmppc_pseries_do_hpt_hcall() with preemption enabled.
The handlers for H_ENTER, H_REMOVE, H_READ, H_CLEAR_MOD, H_CLEAR_REF,
H_PROTECT, and H_BULK_REMOVE spin on try_lock_hpte() or lock_rmap(),
and H_ENTER also reaches kvmppc_do_h_enter(), which uses
arch_spin_lock() on kvm->mmu_lock. That raw lock choice is
intentional because kvmppc_do_h_enter() is also used by real-mode
callers, so the correct fix is to establish the proper preemption
context at the virtual-mode caller boundary.
If a vCPU thread is preempted while holding HPTE_V_HVLOCK, any other
thread on the same CPU trying to acquire the same bit-lock can spin
indefinitely. The same applies to lock_rmap() users and to the H_ENTER
path while holding the raw mmu_lock.
Fix this by adding preempt_disable()/preempt_enable() pairs around the
two kvmppc_hpte_hv_fault() call sites in kvmppc_handle_exit_hv(), and
around the kvmppc_pseries_do_hpt_hcall() invocation in
kvmppc_pseries_do_hcall(). Targeting only the HPT hcall helper avoids
wrapping non-HPT hcalls (such as H_RTAS, H_CONFER, H_REGISTER_VPA) that
perform page faults, blocking memory allocations, or context switches.
Fixes: 6165d5dd99db ("KVM: PPC: Book3S HV: add virtual mode handlers for HPT
hcalls and page faults")
Cc: [email protected] # v5.14+
Signed-off-by: Amit Machhiwal <[email protected]>
---
arch/powerpc/kvm/book3s_hv.c | 6 ++++++
1 file changed, 6 insertions(+)
diff --git a/arch/powerpc/kvm/book3s_hv.c b/arch/powerpc/kvm/book3s_hv.c
index aa51968e206a..52f72f30baf0 100644
--- a/arch/powerpc/kvm/book3s_hv.c
+++ b/arch/powerpc/kvm/book3s_hv.c
@@ -1212,9 +1212,11 @@ int kvmppc_pseries_do_hcall(struct kvm_vcpu *vcpu)
case H_CLEAR_REF:
case H_PROTECT:
case H_BULK_REMOVE:
+ preempt_disable();
idx = srcu_read_lock(&kvm->srcu);
ret = kvmppc_pseries_do_hpt_hcall(vcpu, req);
srcu_read_unlock(&kvm->srcu, idx);
+ preempt_enable();
if (ret == H_TOO_HARD)
return RESUME_HOST;
break;
@@ -1834,8 +1836,10 @@ static int kvmppc_handle_exit_hv(struct kvm_vcpu *vcpu,
else
vsid = vcpu->arch.fault_gpa;
+ preempt_disable();
err = kvmppc_hpte_hv_fault(vcpu, vcpu->arch.fault_dar,
vsid, vcpu->arch.fault_dsisr, true);
+ preempt_enable();
if (err == 0) {
r = RESUME_GUEST;
} else if (err == -1 || err == -2) {
@@ -1881,8 +1885,10 @@ static int kvmppc_handle_exit_hv(struct kvm_vcpu *vcpu,
else
vsid = vcpu->arch.fault_gpa;
+ preempt_disable();
err = kvmppc_hpte_hv_fault(vcpu, vcpu->arch.fault_dar,
vsid, vcpu->arch.fault_dsisr, false);
+ preempt_enable();
if (err == 0) {
r = RESUME_GUEST;
} else if (err == -1) {
--
2.54.0 (Apple Git-157)