* [PATCH v3 1/3] KVM: PPC: Book3S HV: Add SRCU protection for virtual-mode HPT hcalls
2026-10-06 12:24 [PATCH v3 0/3] KVM: PPC: Fixes for Book3S HV HPT locking and paired-single decoding Amit Machhiwal
@ 2026-10-06 12:24 ` Amit Machhiwal
2026-10-06 12:24 ` [PATCH v3 2/3] KVM: PPC: Book3S HV: Add preempt_disable() around virtual-mode HPTE bit-lock users Amit Machhiwal
2026-10-06 12:24 ` [PATCH v3 3/3] KVM: PPC: Fix get_d_signext() 12-bit displacement for paired-single D-form Amit Machhiwal
2 siblings, 0 replies; 6+ messages in thread
From: Amit Machhiwal @ 2026-10-06 12:24 UTC (permalink / raw)
To: Madhavan Srinivasan, linuxppc-dev
Cc: Amit Machhiwal, Nicholas Piggin, Michael Ellerman,
Christophe Leroy (CS GROUP), Ritesh Harjani (IBM),
Shrikanth Hegde, kvm-ppc, kvm, linux-kernel, Gautam Menghani,
Harsh Prateek Bora, R Nageswara Sastry, Alexander Graf,
linux-hardening, stable, Avi Kivity
The virtual-mode HPT hcall handlers (H_ENTER, H_REMOVE, H_BULK_REMOVE,
H_CLEAR_REF, H_CLEAR_MOD) call into book3s_hv_rm_mmu.c, which accesses
memslots via kvm_memslots_raw(). kvm_memslots_raw() uses
rcu_dereference_raw_check() to bypass SRCU lockdep annotation checking.
This is safe in real mode because the entire guest entry/exit is wrapped
in srcu_read_lock/unlock inside kvmppc_run_core() and
kvmhv_run_single_vcpu().
However, since commit 6165d5dd99db ("KVM: PPC: Book3S HV: add virtual
mode handlers for HPT hcalls and page faults"), these same handlers are
also executed in virtual mode via kvmppc_pseries_do_hcall(), which runs
after SRCU has already been released on guest exit.
A concurrent KVM_SET_USER_MEMORY_REGION deletion or move can therefore
call synchronize_srcu_expedited() — which does not wait for this thread —
and then kfree(slot) and vfree(slot->arch.rmap) while the hcall handler
still holds a raw pointer to the memslot. lock_rmap() then writes to
freed vmalloc memory, leading to use-after-free and memory corruption.
Because kvm_memslots_raw() suppresses lockdep checks, this race is
entirely silent.
Fix this by factoring out the 7 HPT hcall handlers (H_REMOVE, H_ENTER,
H_READ, H_CLEAR_MOD, H_CLEAR_REF, H_PROTECT, H_BULK_REMOVE) into a
helper function, kvmppc_pseries_do_hpt_hcall(), and wrapping its call
site in srcu_read_lock(&kvm->srcu) / srcu_read_unlock(&kvm->srcu, idx).
Targeting only the HPT hcalls avoids wrapping non-HPT hcalls that sleep
(such as H_CONFER, H_REGISTER_VPA, H_PAGE_INIT) or handlers that already
acquire SRCU internally (such as H_RTAS).
Fixes: 6165d5dd99db ("KVM: PPC: Book3S HV: add virtual mode handlers for HPT hcalls and page faults")
Cc: stable@vger.kernel.org # v5.14+
Suggested-by: Ritesh Harjani (IBM) <ritesh.list@gmail.com>
Reviewed-by: Ritesh Harjani (IBM) <ritesh.list@gmail.com>
Signed-off-by: Amit Machhiwal <amachhiw@linux.ibm.com>
---
arch/powerpc/kvm/book3s_hv.c | 70 ++++++++++++++++++------------------
1 file changed, 35 insertions(+), 35 deletions(-)
diff --git a/arch/powerpc/kvm/book3s_hv.c b/arch/powerpc/kvm/book3s_hv.c
index de05d721edc4..46dd550115a4 100644
--- a/arch/powerpc/kvm/book3s_hv.c
+++ b/arch/powerpc/kvm/book3s_hv.c
@@ -1159,6 +1159,38 @@ static long kvmppc_h_rpt_invalidate(struct kvm_vcpu *vcpu,
return H_SUCCESS;
}
+static long kvmppc_pseries_do_hpt_hcall(struct kvm_vcpu *vcpu, unsigned long req)
+{
+ switch (req) {
+ case H_REMOVE:
+ return kvmppc_h_remove(vcpu, kvmppc_get_gpr(vcpu, 4),
+ kvmppc_get_gpr(vcpu, 5),
+ kvmppc_get_gpr(vcpu, 6));
+ case H_ENTER:
+ return kvmppc_h_enter(vcpu, kvmppc_get_gpr(vcpu, 4),
+ kvmppc_get_gpr(vcpu, 5),
+ kvmppc_get_gpr(vcpu, 6),
+ kvmppc_get_gpr(vcpu, 7));
+ case H_READ:
+ return kvmppc_h_read(vcpu, kvmppc_get_gpr(vcpu, 4),
+ kvmppc_get_gpr(vcpu, 5));
+ case H_CLEAR_MOD:
+ return kvmppc_h_clear_mod(vcpu, kvmppc_get_gpr(vcpu, 4),
+ kvmppc_get_gpr(vcpu, 5));
+ case H_CLEAR_REF:
+ return kvmppc_h_clear_ref(vcpu, kvmppc_get_gpr(vcpu, 4),
+ kvmppc_get_gpr(vcpu, 5));
+ case H_PROTECT:
+ return kvmppc_h_protect(vcpu, kvmppc_get_gpr(vcpu, 4),
+ kvmppc_get_gpr(vcpu, 5),
+ kvmppc_get_gpr(vcpu, 6));
+ case H_BULK_REMOVE:
+ return kvmppc_h_bulk_remove(vcpu);
+ }
+
+ return H_FUNCTION;
+}
+
int kvmppc_pseries_do_hcall(struct kvm_vcpu *vcpu)
{
struct kvm *kvm = vcpu->kvm;
@@ -1174,47 +1206,15 @@ int kvmppc_pseries_do_hcall(struct kvm_vcpu *vcpu)
switch (req) {
case H_REMOVE:
- ret = kvmppc_h_remove(vcpu, kvmppc_get_gpr(vcpu, 4),
- kvmppc_get_gpr(vcpu, 5),
- kvmppc_get_gpr(vcpu, 6));
- if (ret == H_TOO_HARD)
- return RESUME_HOST;
- break;
case H_ENTER:
- ret = kvmppc_h_enter(vcpu, kvmppc_get_gpr(vcpu, 4),
- kvmppc_get_gpr(vcpu, 5),
- kvmppc_get_gpr(vcpu, 6),
- kvmppc_get_gpr(vcpu, 7));
- if (ret == H_TOO_HARD)
- return RESUME_HOST;
- break;
case H_READ:
- ret = kvmppc_h_read(vcpu, kvmppc_get_gpr(vcpu, 4),
- kvmppc_get_gpr(vcpu, 5));
- if (ret == H_TOO_HARD)
- return RESUME_HOST;
- break;
case H_CLEAR_MOD:
- ret = kvmppc_h_clear_mod(vcpu, kvmppc_get_gpr(vcpu, 4),
- kvmppc_get_gpr(vcpu, 5));
- if (ret == H_TOO_HARD)
- return RESUME_HOST;
- break;
case H_CLEAR_REF:
- ret = kvmppc_h_clear_ref(vcpu, kvmppc_get_gpr(vcpu, 4),
- kvmppc_get_gpr(vcpu, 5));
- if (ret == H_TOO_HARD)
- return RESUME_HOST;
- break;
case H_PROTECT:
- ret = kvmppc_h_protect(vcpu, kvmppc_get_gpr(vcpu, 4),
- kvmppc_get_gpr(vcpu, 5),
- kvmppc_get_gpr(vcpu, 6));
- if (ret == H_TOO_HARD)
- return RESUME_HOST;
- break;
case H_BULK_REMOVE:
- ret = kvmppc_h_bulk_remove(vcpu);
+ idx = srcu_read_lock(&kvm->srcu);
+ ret = kvmppc_pseries_do_hpt_hcall(vcpu, req);
+ srcu_read_unlock(&kvm->srcu, idx);
if (ret == H_TOO_HARD)
return RESUME_HOST;
break;
--
2.54.0 (Apple Git-157)
^ permalink raw reply [flat|nested] 6+ messages in thread* [PATCH v3 2/3] KVM: PPC: Book3S HV: Add preempt_disable() around virtual-mode HPTE bit-lock users
2026-10-06 12:24 [PATCH v3 0/3] KVM: PPC: Fixes for Book3S HV HPT locking and paired-single decoding Amit Machhiwal
2026-10-06 12:24 ` [PATCH v3 1/3] KVM: PPC: Book3S HV: Add SRCU protection for virtual-mode HPT hcalls Amit Machhiwal
@ 2026-10-06 12:24 ` Amit Machhiwal
2026-10-07 6:18 ` Ritesh Harjani
2026-10-06 12:24 ` [PATCH v3 3/3] KVM: PPC: Fix get_d_signext() 12-bit displacement for paired-single D-form Amit Machhiwal
2 siblings, 1 reply; 6+ messages in thread
From: Amit Machhiwal @ 2026-10-06 12:24 UTC (permalink / raw)
To: Madhavan Srinivasan, linuxppc-dev
Cc: Amit Machhiwal, Nicholas Piggin, Michael Ellerman,
Christophe Leroy (CS GROUP), Ritesh Harjani (IBM),
Shrikanth Hegde, kvm-ppc, kvm, linux-kernel, Gautam Menghani,
Harsh Prateek Bora, R Nageswara Sastry, Alexander Graf,
linux-hardening, stable, Avi Kivity
kvmppc_hv_find_lock_hpte() requires virtual-mode callers to run with
preemption disabled, because it can return with HPTE_V_HVLOCK still held
until the caller later unlocks the HPTE. Existing virtual-mode callers
in book3s_64_mmu_hv.c already follow that rule, but several paths do
not.
kvmppc_handle_exit_hv() calls kvmppc_hpte_hv_fault() for hash-mode
data-side and instruction-side faults after guest exit with preemption
enabled. kvmppc_pseries_do_hcall() executes virtual-mode HPT hcall
handlers via kvmppc_pseries_do_hpt_hcall() with preemption enabled; the
handlers for H_ENTER, H_REMOVE, H_READ, H_CLEAR_MOD, H_CLEAR_REF,
H_PROTECT, and H_BULK_REMOVE all spin on try_lock_hpte() or lock_rmap().
H_ENTER also reaches kvmppc_do_h_enter(), which uses arch_spin_lock() on
kvm->mmu_lock. That raw lock choice is intentional because
kvmppc_do_h_enter() is also called from real-mode paths, so the correct
fix is to establish the proper preemption context at the virtual-mode
caller boundary.
On the host side, kvm_unmap_rmapp(), kvm_age_rmapp(),
kvm_test_clear_dirty_npages(), and resize_hpt_rehash_hpte() also acquire
HPTE_V_HVLOCK via try_lock_hpte() in process context with preemption
enabled, serving MMU notifier callbacks, dirty-log harvesting, and HPT
resize respectively.
If any of these threads is preempted while holding HPTE_V_HVLOCK, any
other thread on the same CPU spinning on the same bit-lock can never
make progress, as the lock owner cannot be rescheduled to release it.
This is particularly acute when the spinning thread has preemption
disabled: it will never yield, causing a permanent CPU hang.
Fix this by adding preempt_disable()/preempt_enable() pairs around the
two kvmppc_hpte_hv_fault() call sites in kvmppc_handle_exit_hv() and
around the kvmppc_pseries_do_hpt_hcall() invocation in
kvmppc_pseries_do_hcall().
For kvm_unmap_rmapp() and kvm_age_rmapp(), place preempt_disable() before
lock_rmap() so that both the rmap chain lock and the subsequent
HPTE_V_HVLOCK bit-lock are held under a single non-preemptible window.
There is an ABBA ordering constraint between the two locks: the rmap chain
lock must be dropped before spinning on the HPTE bit-lock (documented in
the comment above the try_lock_hpte() call in kvm_unmap_rmapp()). To
preserve this, preempt_enable() is called after unlock_rmap() on the
failed try_lock_hpte() retry path and on any early-exit path, before the
cpu_relax() spin, so the HPTE lock owner can be scheduled.
For kvm_test_clear_dirty_npages(), remove the per-iteration
preempt_disable()/preempt_enable() pairs: this function has a single call
site, kvmppc_hv_get_dirty_log_hpt(), which already holds preempt_disable()
across the entire loop, making the inner guards redundant.
For resize_hpt_rehash_hpte(), place preempt_disable() before the
unconditional try_lock_hpte() spin loop and preempt_enable() after
unlock_hpte() at the single exit point. This function is called from
kvm_vm_ioctl_resize_hpt_commit(), which first quiesces all vCPUs by
clearing kvm->arch.mmu_ready and calling on_each_cpu() to flush any vCPU
currently running in guest mode back to host. With all vCPUs out of the
guest, no vCPU thread can hold HPTE_V_HVLOCK; any remaining lock holder
(an MMU notifier callback or dirty-log walker) runs on a separate CPU
and is not preempted, so the spin always makes forward progress.
Fixes: 6165d5dd99db ("KVM: PPC: Book3S HV: add virtual mode handlers for HPT hcalls and page faults")
Cc: stable@vger.kernel.org # v5.14+
Reviewed-by: Shrikanth Hegde <sshegde@linux.ibm.com>
Signed-off-by: Amit Machhiwal <amachhiw@linux.ibm.com>
---
Changes in v3:
- In kvm_unmap_rmapp() and kvm_age_rmapp(), moved preempt_disable() before
lock_rmap() so both rmap and HPTE locks are held under a single
non-preemptible section; added preempt_enable() on early-exit and retry paths.
- Removed redundant per-iteration preempt_disable()/preempt_enable() from
kvm_test_clear_dirty_npages() as kvmppc_hv_get_dirty_log_hpt() already holds
it.
- Updated commit log to explain the locking design per function.
- Picked up Reviewed-by tag from Shrikanth Hegde.
Changes in v2:
- Extended preempt_disable()/preempt_enable() to also cover four
host-side virtual-mode HPTE bit-lock users in book3s_64_mmu_hv.c.
- Added warning comment above kvmppc_pseries_do_hpt_hcall().
- Dropped Reviewed-by as the patch was materially extended.
arch/powerpc/kvm/book3s_64_mmu_hv.c | 10 ++++++++++
arch/powerpc/kvm/book3s_hv.c | 10 ++++++++++
2 files changed, 20 insertions(+)
diff --git a/arch/powerpc/kvm/book3s_64_mmu_hv.c b/arch/powerpc/kvm/book3s_64_mmu_hv.c
index 2ccb3d138f46..908495f2b001 100644
--- a/arch/powerpc/kvm/book3s_64_mmu_hv.c
+++ b/arch/powerpc/kvm/book3s_64_mmu_hv.c
@@ -810,9 +810,11 @@ static void kvm_unmap_rmapp(struct kvm *kvm, struct kvm_memory_slot *memslot,
rmapp = &memslot->arch.rmap[gfn - memslot->base_gfn];
for (;;) {
+ preempt_disable();
lock_rmap(rmapp);
if (!(*rmapp & KVMPPC_RMAP_PRESENT)) {
unlock_rmap(rmapp);
+ preempt_enable();
break;
}
@@ -826,6 +828,7 @@ static void kvm_unmap_rmapp(struct kvm *kvm, struct kvm_memory_slot *memslot,
if (!try_lock_hpte(hptep, HPTE_V_HVLOCK)) {
/* unlock rmap before spinning on the HPTE lock */
unlock_rmap(rmapp);
+ preempt_enable();
while (be64_to_cpu(hptep[0]) & HPTE_V_HVLOCK)
cpu_relax();
continue;
@@ -834,6 +837,7 @@ static void kvm_unmap_rmapp(struct kvm *kvm, struct kvm_memory_slot *memslot,
kvmppc_unmap_hpte(kvm, i, memslot, rmapp, gfn);
unlock_rmap(rmapp);
__unlock_hpte(hptep, be64_to_cpu(hptep[0]));
+ preempt_enable();
}
}
@@ -890,6 +894,7 @@ static bool kvm_age_rmapp(struct kvm *kvm, struct kvm_memory_slot *memslot,
rmapp = &memslot->arch.rmap[gfn - memslot->base_gfn];
retry:
+ preempt_disable();
lock_rmap(rmapp);
if (*rmapp & KVMPPC_RMAP_REFERENCED) {
*rmapp &= ~KVMPPC_RMAP_REFERENCED;
@@ -897,6 +902,7 @@ static bool kvm_age_rmapp(struct kvm *kvm, struct kvm_memory_slot *memslot,
}
if (!(*rmapp & KVMPPC_RMAP_PRESENT)) {
unlock_rmap(rmapp);
+ preempt_enable();
return ret;
}
@@ -912,6 +918,7 @@ static bool kvm_age_rmapp(struct kvm *kvm, struct kvm_memory_slot *memslot,
if (!try_lock_hpte(hptep, HPTE_V_HVLOCK)) {
/* unlock rmap before spinning on the HPTE lock */
unlock_rmap(rmapp);
+ preempt_enable();
while (be64_to_cpu(hptep[0]) & HPTE_V_HVLOCK)
cpu_relax();
goto retry;
@@ -931,6 +938,7 @@ static bool kvm_age_rmapp(struct kvm *kvm, struct kvm_memory_slot *memslot,
} while ((i = j) != head);
unlock_rmap(rmapp);
+ preempt_enable();
return ret;
}
@@ -1219,6 +1227,7 @@ static unsigned long resize_hpt_rehash_hpte(struct kvm_resize_hpt *resize,
if (!(vpte & HPTE_V_VALID) && !(vpte & HPTE_V_ABSENT))
return 0; /* nothing to do */
+ preempt_disable();
while (!try_lock_hpte(hptep, HPTE_V_HVLOCK))
cpu_relax();
@@ -1346,6 +1355,7 @@ static unsigned long resize_hpt_rehash_hpte(struct kvm_resize_hpt *resize,
out:
unlock_hpte(hptep, vpte);
+ preempt_enable();
return ret;
}
diff --git a/arch/powerpc/kvm/book3s_hv.c b/arch/powerpc/kvm/book3s_hv.c
index 46dd550115a4..56083b415a29 100644
--- a/arch/powerpc/kvm/book3s_hv.c
+++ b/arch/powerpc/kvm/book3s_hv.c
@@ -1159,6 +1159,10 @@ static long kvmppc_h_rpt_invalidate(struct kvm_vcpu *vcpu,
return H_SUCCESS;
}
+/*
+ * Must be called with preemption disabled. The HPT hcall handlers spin
+ * on HPTE bit-locks and cannot make any blocking/sleeping calls.
+ */
static long kvmppc_pseries_do_hpt_hcall(struct kvm_vcpu *vcpu, unsigned long req)
{
switch (req) {
@@ -1212,9 +1216,11 @@ int kvmppc_pseries_do_hcall(struct kvm_vcpu *vcpu)
case H_CLEAR_REF:
case H_PROTECT:
case H_BULK_REMOVE:
+ preempt_disable();
idx = srcu_read_lock(&kvm->srcu);
ret = kvmppc_pseries_do_hpt_hcall(vcpu, req);
srcu_read_unlock(&kvm->srcu, idx);
+ preempt_enable();
if (ret == H_TOO_HARD)
return RESUME_HOST;
break;
@@ -1834,8 +1840,10 @@ static int kvmppc_handle_exit_hv(struct kvm_vcpu *vcpu,
else
vsid = vcpu->arch.fault_gpa;
+ preempt_disable();
err = kvmppc_hpte_hv_fault(vcpu, vcpu->arch.fault_dar,
vsid, vcpu->arch.fault_dsisr, true);
+ preempt_enable();
if (err == 0) {
r = RESUME_GUEST;
} else if (err == -1 || err == -2) {
@@ -1881,8 +1889,10 @@ static int kvmppc_handle_exit_hv(struct kvm_vcpu *vcpu,
else
vsid = vcpu->arch.fault_gpa;
+ preempt_disable();
err = kvmppc_hpte_hv_fault(vcpu, vcpu->arch.fault_dar,
vsid, vcpu->arch.fault_dsisr, false);
+ preempt_enable();
if (err == 0) {
r = RESUME_GUEST;
} else if (err == -1) {
--
2.54.0 (Apple Git-157)
^ permalink raw reply [flat|nested] 6+ messages in thread* Re: [PATCH v3 2/3] KVM: PPC: Book3S HV: Add preempt_disable() around virtual-mode HPTE bit-lock users
2026-10-06 12:24 ` [PATCH v3 2/3] KVM: PPC: Book3S HV: Add preempt_disable() around virtual-mode HPTE bit-lock users Amit Machhiwal
@ 2026-10-07 6:18 ` Ritesh Harjani
2026-10-07 11:45 ` Amit Machhiwal
0 siblings, 1 reply; 6+ messages in thread
From: Ritesh Harjani @ 2026-10-07 6:18 UTC (permalink / raw)
To: Amit Machhiwal, Madhavan Srinivasan, linuxppc-dev
Cc: Amit Machhiwal, Nicholas Piggin, Michael Ellerman,
Christophe Leroy (CS GROUP),
Shrikanth Hegde, kvm-ppc, kvm, linux-kernel, Gautam Menghani,
Harsh Prateek Bora, R Nageswara Sastry, Alexander Graf,
linux-hardening, stable, Avi Kivity
Amit Machhiwal <amachhiw@linux.ibm.com> writes:
> kvmppc_hv_find_lock_hpte() requires virtual-mode callers to run with
> preemption disabled, because it can return with HPTE_V_HVLOCK still held
> until the caller later unlocks the HPTE. Existing virtual-mode callers
> in book3s_64_mmu_hv.c already follow that rule, but several paths do
> not.
>
> kvmppc_handle_exit_hv() calls kvmppc_hpte_hv_fault() for hash-mode
> data-side and instruction-side faults after guest exit with preemption
> enabled. kvmppc_pseries_do_hcall() executes virtual-mode HPT hcall
> handlers via kvmppc_pseries_do_hpt_hcall() with preemption enabled; the
> handlers for H_ENTER, H_REMOVE, H_READ, H_CLEAR_MOD, H_CLEAR_REF,
> H_PROTECT, and H_BULK_REMOVE all spin on try_lock_hpte() or lock_rmap().
> H_ENTER also reaches kvmppc_do_h_enter(), which uses arch_spin_lock() on
> kvm->mmu_lock. That raw lock choice is intentional because
> kvmppc_do_h_enter() is also called from real-mode paths, so the correct
> fix is to establish the proper preemption context at the virtual-mode
> caller boundary.
>
> On the host side, kvm_unmap_rmapp(), kvm_age_rmapp(),
> kvm_test_clear_dirty_npages(), and resize_hpt_rehash_hpte() also acquire
> HPTE_V_HVLOCK via try_lock_hpte() in process context with preemption
> enabled, serving MMU notifier callbacks, dirty-log harvesting, and HPT
> resize respectively.
>
I was going over all the callers of lock_rmap() and try_lock_hpte() on,
and I see that we might have missed kvm_htab_write() path...
... after spending sometime looks like we need this diff for
kvm_htab_write() path as well, since it calls kvmppc_do_h_remove() which
calls try_lock_hpte() and lock_rmap(), although the race window is much
narrower and maybe very hard to hit.
But still, could you kindly look into this and if needed please take it
forward too. Please note that this is not tested, so hoping that you
could take care of that too.
Thanks!
-ritesh
diff --git a/arch/powerpc/kvm/book3s_64_mmu_hv.c b/arch/powerpc/kvm/book3s_64_mmu_hv.c
index 908495f2b001b..916db7ceb51a9 100644
--- a/arch/powerpc/kvm/book3s_64_mmu_hv.c
+++ b/arch/powerpc/kvm/book3s_64_mmu_hv.c
@@ -47,6 +47,8 @@
static long kvmppc_virtmode_do_h_enter(struct kvm *kvm, unsigned long flags,
long pte_index, unsigned long pteh,
unsigned long ptel, unsigned long *pte_idx_ret);
+static void kvmppc_virtmode_do_h_remove(struct kvm *kvm,
+ unsigned long pte_index, unsigned long *hpret);
struct kvm_resize_hpt {
/* These fields read-only after init */
@@ -308,6 +310,19 @@ static long kvmppc_virtmode_do_h_enter(struct kvm *kvm, unsigned long flags,
}
+/*
+ * Virtual-mode H_REMOVE. kvmppc_do_h_remove() is also called from real
+ * mode, where preempt_disable() is not usable, so the guard stays here.
+ * The helper takes HPTE_V_HVLOCK and the rmap bit and does not sleep.
+ */
+static void kvmppc_virtmode_do_h_remove(struct kvm *kvm,
+ unsigned long pte_index, unsigned long *hpret)
+{
+ preempt_disable();
+ kvmppc_do_h_remove(kvm, 0, pte_index, 0, hpret);
+ preempt_enable();
+}
+
static struct kvmppc_slb *kvmppc_mmu_book3s_hv_find_slbe(struct kvm_vcpu *vcpu,
gva_t eaddr)
{
@@ -1878,7 +1893,7 @@ static ssize_t kvm_htab_write(struct file *file, const char __user *buf,
nb += HPTE_SIZE;
if (be64_to_cpu(hptp[0]) & (HPTE_V_VALID | HPTE_V_ABSENT))
- kvmppc_do_h_remove(kvm, 0, i, 0, tmp);
+ kvmppc_virtmode_do_h_remove(kvm, i, tmp);
err = -EIO;
ret = kvmppc_virtmode_do_h_enter(kvm, H_EXACT, i, v, r,
tmp);
@@ -1907,7 +1922,7 @@ static ssize_t kvm_htab_write(struct file *file, const char __user *buf,
for (j = 0; j < hdr.n_invalid; ++j) {
if (be64_to_cpu(hptp[0]) & (HPTE_V_VALID | HPTE_V_ABSENT))
- kvmppc_do_h_remove(kvm, 0, i, 0, tmp);
+ kvmppc_virtmode_do_h_remove(kvm, i, tmp);
++i;
hptp += 2;
}
^ permalink raw reply [flat|nested] 6+ messages in thread* Re: [PATCH v3 2/3] KVM: PPC: Book3S HV: Add preempt_disable() around virtual-mode HPTE bit-lock users
2026-10-07 6:18 ` Ritesh Harjani
@ 2026-10-07 11:45 ` Amit Machhiwal
0 siblings, 0 replies; 6+ messages in thread
From: Amit Machhiwal @ 2026-10-07 11:45 UTC (permalink / raw)
To: Ritesh Harjani
Cc: Amit Machhiwal, Madhavan Srinivasan, linuxppc-dev,
Nicholas Piggin, Michael Ellerman, Christophe Leroy (CS GROUP),
Shrikanth Hegde, kvm-ppc, kvm, linux-kernel, Gautam Menghani,
Harsh Prateek Bora, R Nageswara Sastry, Alexander Graf,
linux-hardening, stable, Avi Kivity
Hi Ritesh,
Thanks for reviewing this patch. Please find my response below.
On 2026/10/07 11:48 AM, Ritesh Harjani wrote:
> Amit Machhiwal <amachhiw@linux.ibm.com> writes:
>
> > kvmppc_hv_find_lock_hpte() requires virtual-mode callers to run with
> > preemption disabled, because it can return with HPTE_V_HVLOCK still held
> > until the caller later unlocks the HPTE. Existing virtual-mode callers
> > in book3s_64_mmu_hv.c already follow that rule, but several paths do
> > not.
> >
> > kvmppc_handle_exit_hv() calls kvmppc_hpte_hv_fault() for hash-mode
> > data-side and instruction-side faults after guest exit with preemption
> > enabled. kvmppc_pseries_do_hcall() executes virtual-mode HPT hcall
> > handlers via kvmppc_pseries_do_hpt_hcall() with preemption enabled; the
> > handlers for H_ENTER, H_REMOVE, H_READ, H_CLEAR_MOD, H_CLEAR_REF,
> > H_PROTECT, and H_BULK_REMOVE all spin on try_lock_hpte() or lock_rmap().
> > H_ENTER also reaches kvmppc_do_h_enter(), which uses arch_spin_lock() on
> > kvm->mmu_lock. That raw lock choice is intentional because
> > kvmppc_do_h_enter() is also called from real-mode paths, so the correct
> > fix is to establish the proper preemption context at the virtual-mode
> > caller boundary.
> >
> > On the host side, kvm_unmap_rmapp(), kvm_age_rmapp(),
> > kvm_test_clear_dirty_npages(), and resize_hpt_rehash_hpte() also acquire
> > HPTE_V_HVLOCK via try_lock_hpte() in process context with preemption
> > enabled, serving MMU notifier callbacks, dirty-log harvesting, and HPT
> > resize respectively.
> >
>
> I was going over all the callers of lock_rmap() and try_lock_hpte() on,
> and I see that we might have missed kvm_htab_write() path...
>
> ... after spending sometime looks like we need this diff for
> kvm_htab_write() path as well, since it calls kvmppc_do_h_remove() which
> calls try_lock_hpte() and lock_rmap(), although the race window is much
> narrower and maybe very hard to hit.
>
> But still, could you kindly look into this and if needed please take it
> forward too. Please note that this is not tested, so hoping that you
> could take care of that too.
Thanks for catching this — you're right, that was a genuine miss. I will fix it
in v4 by introducing kvmppc_virtmode_do_h_remove() as you suggested.
One small divergence from your proposed diff: I retained the full flags and avpn
parameters rather than baking in 0, 0. The flags field carries the PAPR-defined
H_AVPN and H_ANDCOND conditional guards, and a future virtual-mode caller may
legitimately need them. Baking the constants in would silently prevent that,
while passing them through costs nothing. This mirrors exactly how
kvmppc_virtmode_do_h_enter() is structured.
I'm in the middle of making the changes and testing the same. The v4 will be
out soon.
Thanks,
Amit
^ permalink raw reply [flat|nested] 6+ messages in thread
* [PATCH v3 3/3] KVM: PPC: Fix get_d_signext() 12-bit displacement for paired-single D-form
2026-10-06 12:24 [PATCH v3 0/3] KVM: PPC: Fixes for Book3S HV HPT locking and paired-single decoding Amit Machhiwal
2026-10-06 12:24 ` [PATCH v3 1/3] KVM: PPC: Book3S HV: Add SRCU protection for virtual-mode HPT hcalls Amit Machhiwal
2026-10-06 12:24 ` [PATCH v3 2/3] KVM: PPC: Book3S HV: Add preempt_disable() around virtual-mode HPTE bit-lock users Amit Machhiwal
@ 2026-10-06 12:24 ` Amit Machhiwal
2 siblings, 0 replies; 6+ messages in thread
From: Amit Machhiwal @ 2026-10-06 12:24 UTC (permalink / raw)
To: Madhavan Srinivasan, linuxppc-dev
Cc: Amit Machhiwal, Nicholas Piggin, Michael Ellerman,
Christophe Leroy (CS GROUP), Ritesh Harjani (IBM),
Shrikanth Hegde, kvm-ppc, kvm, linux-kernel, Gautam Menghani,
Harsh Prateek Bora, R Nageswara Sastry, Alexander Graf,
linux-hardening, stable, Avi Kivity
get_d_signext() has two compounding bugs since its introduction in 2010:
1. The extraction mask 0x8FF silently drops bits 8-10 of the 12-bit D
field (PPC ISA bits 21-23), corrupting any displacement that has any
of those bits set.
2. The sign-magnitude idiom "-(d & 0x7ff)" is wrong for two's-complement:
for D=0xFFC (encoding of -4) it returns -252 instead of -4.
Together these errors produce a wrong effective address for psq_l, psq_lu,
psq_st, and psq_stu whenever the displacement is negative or is a positive
value >= 0x100 with bits 8-10 set — essentially any real-world paired-
single stack-relative access.
The D field occupies the bottom 12 bits of the instruction word (confirmed
by the adjacent W and I extractions via inst_get_field(inst,16,16) and
inst_get_field(inst,17,19)). Replace the open-coded logic with the standard
sign_extend32(inst & 0xfff, 11), which correctly performs two's-complement
sign extension from 12 bits to 32 bits.
Fixes: 831317b605e7 ("KVM: PPC: Implement Paired Single emulation")
Cc: stable@vger.kernel.org
Reviewed-by: Shrikanth Hegde <sshegde@linux.ibm.com>
Signed-off-by: Amit Machhiwal <amachhiw@linux.ibm.com>
---
Changes in v3:
- Picked up Reviewed-by tag from Shrikanth Hegde.
arch/powerpc/kvm/book3s_paired_singles.c | 7 +------
1 file changed, 1 insertion(+), 6 deletions(-)
diff --git a/arch/powerpc/kvm/book3s_paired_singles.c b/arch/powerpc/kvm/book3s_paired_singles.c
index bc39c76c9d9f..532f96293de0 100644
--- a/arch/powerpc/kvm/book3s_paired_singles.c
+++ b/arch/powerpc/kvm/book3s_paired_singles.c
@@ -479,12 +479,7 @@ static bool kvmppc_inst_is_paired_single(struct kvm_vcpu *vcpu, u32 inst)
static int get_d_signext(u32 inst)
{
- int d = inst & 0x8ff;
-
- if (d & 0x800)
- return -(d & 0x7ff);
-
- return (d & 0x7ff);
+ return sign_extend32(inst & 0xfff, 11);
}
static int kvmppc_ps_three_in(struct kvm_vcpu *vcpu, bool rc,
--
2.54.0 (Apple Git-157)
^ permalink raw reply [flat|nested] 6+ messages in thread