* [PATCH 0/6] LoongArch: KVM: irqchip and steal time fixes
@ 2026-09-27 7:52 Tao Cui
2026-09-27 7:52 ` [PATCH 1/6] LoongArch: KVM: Clear device pointer in irqchip destroy callbacks Tao Cui
` (5 more replies)
0 siblings, 6 replies; 15+ messages in thread
From: Tao Cui @ 2026-09-27 7:52 UTC (permalink / raw)
To: maobibo, gaosong, zhaotianrui
Cc: loongarch, kvm, linux-kernel, chenhuacai, kernel,
nagachaithanya9911, cui.tao, Tao Cui
From: Tao Cui <cuitao@kylinos.cn>
Hi,
Six fixes for the LoongArch KVM irqchip and paravirtual steal time
code:
- Patch 1 fixes a host use-after-free: when KVM_CREATE_DEVICE succeeds
but the following fd allocation fails (e.g. under RLIMIT_NOFILE),
ops->destroy() frees the irqchip while kvm->arch.* still points to
it, so later interrupt injection walks into freed memory. Verified
on Loongson-3A6000: with the fd budget exhausted, KVM_IRQ_LINE
returns success and pch_pic_set_irq() executes on the freed object
(kprobe); with the fix the same sequence returns -ENXIO.
- Patch 2 makes the irqfd injection entry points tolerate a NULL
irqchip device: they dispatch through the irq routing table without
the irqchip_in_kernel() gate that protects KVM_IRQ_LINE, so a
routing entry that outlives the device dereferences it.
- Patch 3 loads kvm->arch.dmsintc once in the MSI injection path:
both pch_msi_set_irq() and dmsintc_set_irq() re-read the pointer
between the non-NULL check and the following dereferences, so a
concurrent device removal is observed between them.
- Patch 4 fixes steal time accounting: the PVTIME attribute
initializes its baseline from the control thread's run_delay while
the accumulation happens in the vcpu thread, so the first delta can
wrap and the guest reads a steal time near 2^64. This hits the
default QEMU topology (control thread sets the attribute, vcpu
threads run KVM_RUN) once the guest enables steal time. Verified on
Loongson-3A6000: a guest steal value of 2^64 - 7.58s before the fix,
a small value after it.
- Patch 5 aligns kvm_pch_pic_create() with kvm_eiointc_create() by
propagating the real error code; the kvm_ipi_create() counterpart is
being fixed separately by "Return the actual error code in
kvm_ipi_create()" (Chaithanya Lagisetty).
- Patch 6 rejects a repeated PCH-PIC CTRL_INIT: every call registers
the device on the MMIO bus at the new address while destroy removes
only one range, so stale ranges silently swallow MMIO accesses
(measured with kprobes: three accepted inits, one unregister).
All patches carry Fixes tags.
Tao Cui (6):
LoongArch: KVM: Clear device pointer in irqchip destroy callbacks
LoongArch: KVM: Guard against NULL irqchip in irq injection
LoongArch: KVM: Load dmsintc pointer once in pch_msi_set_irq
LoongArch: KVM: Rebase steal time counter in vcpu context
LoongArch: KVM: Propagate real error code in kvm_pch_pic_create
LoongArch: KVM: Reject repeated PCH-PIC CTRL_INIT
arch/loongarch/kvm/intc/dmsintc.c | 12 ++++++++++--
arch/loongarch/kvm/intc/eiointc.c | 6 +++++-
arch/loongarch/kvm/intc/ipi.c | 1 +
arch/loongarch/kvm/intc/pch_pic.c | 23 +++++++++++++++++------
arch/loongarch/kvm/vcpu.c | 25 +++++++++++++++++--------
5 files changed, 50 insertions(+), 17 deletions(-)
--
2.43.0
^ permalink raw reply [flat|nested] 15+ messages in thread
* [PATCH 1/6] LoongArch: KVM: Clear device pointer in irqchip destroy callbacks
2026-09-27 7:52 [PATCH 0/6] LoongArch: KVM: irqchip and steal time fixes Tao Cui
@ 2026-09-27 7:52 ` Tao Cui
2026-09-28 4:07 ` Bibo Mao
2026-09-27 7:52 ` [PATCH 2/6] LoongArch: KVM: Guard against NULL irqchip in irq injection Tao Cui
` (4 subsequent siblings)
5 siblings, 1 reply; 15+ messages in thread
From: Tao Cui @ 2026-09-27 7:52 UTC (permalink / raw)
To: maobibo, gaosong, zhaotianrui
Cc: loongarch, kvm, linux-kernel, chenhuacai, kernel,
nagachaithanya9911, cui.tao, Tao Cui
From: Tao Cui <cuitao@kylinos.cn>
The destroy callbacks of the four irqchip devices free the device
structure without clearing kvm->arch.{ipi,eiointc,pch_pic,dmsintc},
leaving a dangling pointer. ops->destroy() is not only called at VM
teardown but also when KVM_CREATE_DEVICE succeeds and the following
anon_inode_getfd() fails (e.g. under RLIMIT_NOFILE); the VM stays
alive, kvm_arch_irqchip_in_kernel() still reports true, and interrupt
injection dereferences the freed device.
Clear the pointer before freeing.
Fixes: c532de5a67a7 ("LoongArch: KVM: Add IPI device support")
Fixes: 2e8b9df82631 ("LoongArch: KVM: Add EIOINTC device support")
Fixes: e785dfacf7e7 ("LoongArch: KVM: Add PCHPIC device support")
Fixes: 229132c309d6 ("LoongArch: KVM: Add DMSINTC device support")
Signed-off-by: Tao Cui <cuitao@kylinos.cn>
---
arch/loongarch/kvm/intc/dmsintc.c | 5 ++++-
arch/loongarch/kvm/intc/eiointc.c | 1 +
arch/loongarch/kvm/intc/ipi.c | 1 +
arch/loongarch/kvm/intc/pch_pic.c | 1 +
4 files changed, 7 insertions(+), 1 deletion(-)
diff --git a/arch/loongarch/kvm/intc/dmsintc.c b/arch/loongarch/kvm/intc/dmsintc.c
index 89f980d867be..63072595c01b 100644
--- a/arch/loongarch/kvm/intc/dmsintc.c
+++ b/arch/loongarch/kvm/intc/dmsintc.c
@@ -161,11 +161,14 @@ static int kvm_dmsintc_create(struct kvm_device *dev, u32 type)
static void kvm_dmsintc_destroy(struct kvm_device *dev)
{
+ struct loongarch_dmsintc *s;
if (!dev || !dev->kvm || !dev->kvm->arch.dmsintc)
return;
- kfree(dev->kvm->arch.dmsintc);
+ s = dev->kvm->arch.dmsintc;
+ dev->kvm->arch.dmsintc = NULL;
+ kfree(s);
kfree(dev);
}
diff --git a/arch/loongarch/kvm/intc/eiointc.c b/arch/loongarch/kvm/intc/eiointc.c
index 80f78e07c74a..fe0a1918f26f 100644
--- a/arch/loongarch/kvm/intc/eiointc.c
+++ b/arch/loongarch/kvm/intc/eiointc.c
@@ -675,6 +675,7 @@ static void kvm_eiointc_destroy(struct kvm_device *dev)
kvm = dev->kvm;
eiointc = kvm->arch.eiointc;
+ kvm->arch.eiointc = NULL;
mutex_lock(&kvm->slots_lock);
kvm_io_bus_unregister_dev(kvm, KVM_IOCSR_BUS, &eiointc->device);
kvm_io_bus_unregister_dev(kvm, KVM_IOCSR_BUS, &eiointc->device_vext);
diff --git a/arch/loongarch/kvm/intc/ipi.c b/arch/loongarch/kvm/intc/ipi.c
index 7b333a4a0430..6ee90c10827a 100644
--- a/arch/loongarch/kvm/intc/ipi.c
+++ b/arch/loongarch/kvm/intc/ipi.c
@@ -444,6 +444,7 @@ static void kvm_ipi_destroy(struct kvm_device *dev)
kvm = dev->kvm;
ipi = kvm->arch.ipi;
+ kvm->arch.ipi = NULL;
mutex_lock(&kvm->slots_lock);
kvm_io_bus_unregister_dev(kvm, KVM_IOCSR_BUS, &ipi->device);
mutex_unlock(&kvm->slots_lock);
diff --git a/arch/loongarch/kvm/intc/pch_pic.c b/arch/loongarch/kvm/intc/pch_pic.c
index 2b63b0c2c7ce..62c09b5f3937 100644
--- a/arch/loongarch/kvm/intc/pch_pic.c
+++ b/arch/loongarch/kvm/intc/pch_pic.c
@@ -483,6 +483,7 @@ static void kvm_pch_pic_destroy(struct kvm_device *dev)
kvm = dev->kvm;
s = kvm->arch.pch_pic;
+ kvm->arch.pch_pic = NULL;
/* unregister pch pic device and free it's memory */
mutex_lock(&kvm->slots_lock);
kvm_io_bus_unregister_dev(kvm, KVM_MMIO_BUS, &s->device);
--
2.43.0
^ permalink raw reply [flat|nested] 15+ messages in thread
* [PATCH 2/6] LoongArch: KVM: Guard against NULL irqchip in irq injection
2026-09-27 7:52 [PATCH 0/6] LoongArch: KVM: irqchip and steal time fixes Tao Cui
2026-09-27 7:52 ` [PATCH 1/6] LoongArch: KVM: Clear device pointer in irqchip destroy callbacks Tao Cui
@ 2026-09-27 7:52 ` Tao Cui
2026-09-28 4:14 ` Bibo Mao
2026-09-27 7:52 ` [PATCH 3/6] LoongArch: KVM: Load dmsintc pointer once in pch_msi_set_irq Tao Cui
` (3 subsequent siblings)
5 siblings, 1 reply; 15+ messages in thread
From: Tao Cui @ 2026-09-27 7:52 UTC (permalink / raw)
To: maobibo, gaosong, zhaotianrui
Cc: loongarch, kvm, linux-kernel, chenhuacai, kernel,
nagachaithanya9911, cui.tao, Tao Cui
From: Tao Cui <cuitao@kylinos.cn>
The irqfd injection path dispatches through the irq routing table
without the kvm_arch_irqchip_in_kernel() gate that protects
KVM_IRQ_LINE: kvm_set_pic_irq() and kvm_arch_set_irq_inatomic() call
pch_pic_set_irq(kvm->arch.pch_pic, ...) and pch_msi_set_irq() calls
eiointc_set_irq(kvm->arch.eiointc, ...) directly, so a routing entry
that outlives the corresponding in-kernel irqchip dereferences a NULL
(or, before the previous patch, a freed) device.
Make pch_pic_set_irq(), eiointc_set_irq() and dmsintc_set_irq()
tolerate a NULL device.
Fixes: 1928254c5ccb ("LoongArch: KVM: Add irqfd support")
Signed-off-by: Tao Cui <cuitao@kylinos.cn>
---
arch/loongarch/kvm/intc/dmsintc.c | 3 +++
arch/loongarch/kvm/intc/eiointc.c | 5 ++++-
arch/loongarch/kvm/intc/pch_pic.c | 3 +++
3 files changed, 10 insertions(+), 1 deletion(-)
diff --git a/arch/loongarch/kvm/intc/dmsintc.c b/arch/loongarch/kvm/intc/dmsintc.c
index 63072595c01b..e27f448bb54f 100644
--- a/arch/loongarch/kvm/intc/dmsintc.c
+++ b/arch/loongarch/kvm/intc/dmsintc.c
@@ -70,6 +70,9 @@ int dmsintc_set_irq(struct kvm *kvm, u64 addr, int data, int level)
unsigned int irq, cpu;
struct kvm_vcpu *vcpu;
+ if (!kvm->arch.dmsintc)
+ return -EINVAL;
+
irq = (addr >> AVEC_IRQ_SHIFT) & AVEC_IRQ_MASK;
cpu = (addr >> AVEC_CPU_SHIFT) & kvm->arch.dmsintc->cpu_mask;
if (cpu >= KVM_MAX_VCPUS)
diff --git a/arch/loongarch/kvm/intc/eiointc.c b/arch/loongarch/kvm/intc/eiointc.c
index fe0a1918f26f..c68f9033638a 100644
--- a/arch/loongarch/kvm/intc/eiointc.c
+++ b/arch/loongarch/kvm/intc/eiointc.c
@@ -112,8 +112,11 @@ static inline void eiointc_update_sw_coremap(struct loongarch_eiointc *s,
void eiointc_set_irq(struct loongarch_eiointc *s, int irq, int level)
{
unsigned long flags;
- unsigned long *isr = (unsigned long *)s->isr;
+ unsigned long *isr;
+ if (!s)
+ return;
+ isr = (unsigned long *)s->isr;
spin_lock_irqsave(&s->lock, flags);
level ? __set_bit(irq, isr) : __clear_bit(irq, isr);
eiointc_update_irq(s, irq, level);
diff --git a/arch/loongarch/kvm/intc/pch_pic.c b/arch/loongarch/kvm/intc/pch_pic.c
index 62c09b5f3937..88666425d800 100644
--- a/arch/loongarch/kvm/intc/pch_pic.c
+++ b/arch/loongarch/kvm/intc/pch_pic.c
@@ -48,6 +48,9 @@ void pch_pic_set_irq(struct loongarch_pch_pic *s, int irq, int level)
{
u64 mask = BIT(irq);
+ if (!s)
+ return;
+
spin_lock(&s->lock);
if (level)
s->irr |= mask; /* set irr */
--
2.43.0
^ permalink raw reply [flat|nested] 15+ messages in thread
* [PATCH 3/6] LoongArch: KVM: Load dmsintc pointer once in pch_msi_set_irq
2026-09-27 7:52 [PATCH 0/6] LoongArch: KVM: irqchip and steal time fixes Tao Cui
2026-09-27 7:52 ` [PATCH 1/6] LoongArch: KVM: Clear device pointer in irqchip destroy callbacks Tao Cui
2026-09-27 7:52 ` [PATCH 2/6] LoongArch: KVM: Guard against NULL irqchip in irq injection Tao Cui
@ 2026-09-27 7:52 ` Tao Cui
2026-09-27 7:52 ` [PATCH 4/6] LoongArch: KVM: Rebase steal time counter in vcpu context Tao Cui
` (2 subsequent siblings)
5 siblings, 0 replies; 15+ messages in thread
From: Tao Cui @ 2026-09-27 7:52 UTC (permalink / raw)
To: maobibo, gaosong, zhaotianrui
Cc: loongarch, kvm, linux-kernel, chenhuacai, kernel,
nagachaithanya9911, cui.tao, Tao Cui
From: Tao Cui <cuitao@kylinos.cn>
pch_msi_set_irq() reads kvm->arch.dmsintc several times: the non-NULL
check and the address-window comparison each reload the pointer, so a
concurrent device removal can be observed between them and the
following dereference hits a freed object.
Load the pointer once at the top of pch_msi_set_irq(); also use the
local in dmsintc_set_irq() for its own cpu_mask read.
Fixes: 03de5eecb0f0 ("LoongArch: KVM: Add DMSINTC inject msi to vCPU")
Signed-off-by: Tao Cui <cuitao@kylinos.cn>
---
arch/loongarch/kvm/intc/dmsintc.c | 6 ++++--
arch/loongarch/kvm/intc/pch_pic.c | 7 ++++---
2 files changed, 8 insertions(+), 5 deletions(-)
diff --git a/arch/loongarch/kvm/intc/dmsintc.c b/arch/loongarch/kvm/intc/dmsintc.c
index e27f448bb54f..d42567eba3ed 100644
--- a/arch/loongarch/kvm/intc/dmsintc.c
+++ b/arch/loongarch/kvm/intc/dmsintc.c
@@ -70,11 +70,13 @@ int dmsintc_set_irq(struct kvm *kvm, u64 addr, int data, int level)
unsigned int irq, cpu;
struct kvm_vcpu *vcpu;
- if (!kvm->arch.dmsintc)
+ struct loongarch_dmsintc *s = kvm->arch.dmsintc;
+
+ if (!s)
return -EINVAL;
irq = (addr >> AVEC_IRQ_SHIFT) & AVEC_IRQ_MASK;
- cpu = (addr >> AVEC_CPU_SHIFT) & kvm->arch.dmsintc->cpu_mask;
+ cpu = (addr >> AVEC_CPU_SHIFT) & s->cpu_mask;
if (cpu >= KVM_MAX_VCPUS)
return -EINVAL;
vcpu = kvm_get_vcpu_by_cpuid(kvm, cpu);
diff --git a/arch/loongarch/kvm/intc/pch_pic.c b/arch/loongarch/kvm/intc/pch_pic.c
index 88666425d800..e33d6693c37c 100644
--- a/arch/loongarch/kvm/intc/pch_pic.c
+++ b/arch/loongarch/kvm/intc/pch_pic.c
@@ -74,10 +74,11 @@ void pch_pic_set_irq(struct loongarch_pch_pic *s, int irq, int level)
int pch_msi_set_irq(struct kvm *kvm, struct kvm_kernel_irq_routing_entry *e, int level)
{
u64 msg_addr = (((u64)e->msi.address_hi) << 32) | e->msi.address_lo;
+ struct loongarch_dmsintc *dmsintc = kvm->arch.dmsintc;
- if (cpu_has_msgint && kvm->arch.dmsintc &&
- msg_addr >= kvm->arch.dmsintc->msg_addr_base &&
- msg_addr < (kvm->arch.dmsintc->msg_addr_base + kvm->arch.dmsintc->msg_addr_size)) {
+ if (cpu_has_msgint && dmsintc &&
+ msg_addr >= dmsintc->msg_addr_base &&
+ msg_addr < (dmsintc->msg_addr_base + dmsintc->msg_addr_size)) {
return dmsintc_set_irq(kvm, msg_addr, e->msi.data, level);
}
--
2.43.0
^ permalink raw reply [flat|nested] 15+ messages in thread
* [PATCH 4/6] LoongArch: KVM: Rebase steal time counter in vcpu context
2026-09-27 7:52 [PATCH 0/6] LoongArch: KVM: irqchip and steal time fixes Tao Cui
` (2 preceding siblings ...)
2026-09-27 7:52 ` [PATCH 3/6] LoongArch: KVM: Load dmsintc pointer once in pch_msi_set_irq Tao Cui
@ 2026-09-27 7:52 ` Tao Cui
2026-09-28 7:23 ` Bibo Mao
2026-09-27 7:52 ` [PATCH 5/6] LoongArch: KVM: Propagate real error code in kvm_pch_pic_create Tao Cui
2026-09-27 7:52 ` [PATCH 6/6] LoongArch: KVM: Reject repeated PCH-PIC CTRL_INIT Tao Cui
5 siblings, 1 reply; 15+ messages in thread
From: Tao Cui @ 2026-09-27 7:52 UTC (permalink / raw)
To: maobibo, gaosong, zhaotianrui
Cc: loongarch, kvm, linux-kernel, chenhuacai, kernel,
nagachaithanya9911, cui.tao, Tao Cui
From: Tao Cui <cuitao@kylinos.cn>
KVM_SET_DEVICE_ATTR(PVTIME GPA) initializes st.last_steal from the
ioctl thread's run_delay, but kvm_update_stolen_time() accumulates the
vcpu thread's run_delay; the first delta can be negative and wraps in
u64, so the guest reads a steal time close to 2^64.
Drop the initialization from the attr path and lazily rebase the
counter on the first steal update, which runs in vcpu context like the
hypercall path. Re-registering a GPA clears the rebase sentinel first
and publishes the address with a write barrier; the reader side loads
the address and sentinel with READ_ONCE and a read barrier so a
concurrent re-registration is not observed half-applied.
Fixes: b4ba157044ea ("LoongArch: KVM: Add PV steal time support in host side")
Signed-off-by: Tao Cui <cuitao@kylinos.cn>
---
arch/loongarch/kvm/vcpu.c | 25 +++++++++++++++++--------
1 file changed, 17 insertions(+), 8 deletions(-)
diff --git a/arch/loongarch/kvm/vcpu.c b/arch/loongarch/kvm/vcpu.c
index 8e028be3f0a9..d931a5a802f6 100644
--- a/arch/loongarch/kvm/vcpu.c
+++ b/arch/loongarch/kvm/vcpu.c
@@ -154,12 +154,13 @@ static void kvm_update_stolen_time(struct kvm_vcpu *vcpu)
u32 version;
u64 steal;
gpa_t gpa;
+ u64 last_steal;
struct kvm_memslots *slots;
struct kvm_steal_time __user *st;
struct gfn_to_hva_cache *ghc;
ghc = &vcpu->arch.st.cache;
- gpa = vcpu->arch.st.guest_addr;
+ gpa = READ_ONCE(vcpu->arch.st.guest_addr);
if (!(gpa & KVM_STEAL_PHYS_VALID))
return;
@@ -187,8 +188,14 @@ static void kvm_update_stolen_time(struct kvm_vcpu *vcpu)
smp_wmb();
unsafe_get_user(steal, &st->steal, out);
- steal += current->sched_info.run_delay - vcpu->arch.st.last_steal;
- vcpu->arch.st.last_steal = current->sched_info.run_delay;
+ /* acquire pairs with the smp_wmb in the attr path */
+ smp_rmb();
+ last_steal = READ_ONCE(vcpu->arch.st.last_steal);
+ if (!last_steal)
+ /* first update in vcpu context: rebase the counter */
+ last_steal = current->sched_info.run_delay;
+ steal += current->sched_info.run_delay - last_steal;
+ WRITE_ONCE(vcpu->arch.st.last_steal, current->sched_info.run_delay);
unsafe_put_user(steal, &st->steal, out);
smp_wmb();
@@ -1122,7 +1129,7 @@ static int kvm_loongarch_pvtime_get_attr(struct kvm_vcpu *vcpu,
|| attr->attr != KVM_LOONGARCH_VCPU_PVTIME_GPA)
return -ENXIO;
- gpa = vcpu->arch.st.guest_addr;
+ gpa = READ_ONCE(vcpu->arch.st.guest_addr);
if (put_user(gpa, user))
return -EFAULT;
@@ -1197,7 +1204,7 @@ static int kvm_loongarch_pvtime_set_attr(struct kvm_vcpu *vcpu,
return -EINVAL;
if (!(gpa & KVM_STEAL_PHYS_VALID)) {
- vcpu->arch.st.guest_addr = gpa;
+ WRITE_ONCE(vcpu->arch.st.guest_addr, gpa);
return 0;
}
@@ -1208,8 +1215,10 @@ static int kvm_loongarch_pvtime_set_attr(struct kvm_vcpu *vcpu,
srcu_read_unlock(&kvm->srcu, idx);
if (!ret) {
- vcpu->arch.st.guest_addr = gpa;
- vcpu->arch.st.last_steal = current->sched_info.run_delay;
+ WRITE_ONCE(vcpu->arch.st.last_steal, 0);
+ /* publish the new address only after clearing the rebase sentinel */
+ smp_wmb();
+ WRITE_ONCE(vcpu->arch.st.guest_addr, gpa);
kvm_make_request(KVM_REQ_STEAL_UPDATE, vcpu);
}
@@ -1800,7 +1809,7 @@ static void kvm_vcpu_set_pv_preempted(struct kvm_vcpu *vcpu)
struct kvm_memslots *slots;
struct kvm_steal_time __user *st;
- gpa = vcpu->arch.st.guest_addr;
+ gpa = READ_ONCE(vcpu->arch.st.guest_addr);
if (!(gpa & KVM_STEAL_PHYS_VALID))
return;
--
2.43.0
^ permalink raw reply [flat|nested] 15+ messages in thread
* [PATCH 5/6] LoongArch: KVM: Propagate real error code in kvm_pch_pic_create
2026-09-27 7:52 [PATCH 0/6] LoongArch: KVM: irqchip and steal time fixes Tao Cui
` (3 preceding siblings ...)
2026-09-27 7:52 ` [PATCH 4/6] LoongArch: KVM: Rebase steal time counter in vcpu context Tao Cui
@ 2026-09-27 7:52 ` Tao Cui
2026-09-28 7:25 ` Bibo Mao
2026-09-27 7:52 ` [PATCH 6/6] LoongArch: KVM: Reject repeated PCH-PIC CTRL_INIT Tao Cui
5 siblings, 1 reply; 15+ messages in thread
From: Tao Cui @ 2026-09-27 7:52 UTC (permalink / raw)
To: maobibo, gaosong, zhaotianrui
Cc: loongarch, kvm, linux-kernel, chenhuacai, kernel,
nagachaithanya9911, cui.tao, Tao Cui
From: Tao Cui <cuitao@kylinos.cn>
kvm_pch_pic_create() replaces the real error of
kvm_setup_default_irq_routing() with a fixed -ENOMEM. Return the
actual code, as kvm_eiointc_create() already does; the kvm_ipi_create()
case is fixed separately by "Return the actual error code in
kvm_ipi_create()" (Chaithanya Lagisetty).
Fixes: 1928254c5ccb ("LoongArch: KVM: Add irqfd support")
Signed-off-by: Tao Cui <cuitao@kylinos.cn>
---
arch/loongarch/kvm/intc/pch_pic.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/arch/loongarch/kvm/intc/pch_pic.c b/arch/loongarch/kvm/intc/pch_pic.c
index e33d6693c37c..136211f154d4 100644
--- a/arch/loongarch/kvm/intc/pch_pic.c
+++ b/arch/loongarch/kvm/intc/pch_pic.c
@@ -448,7 +448,7 @@ static int kvm_pch_pic_create(struct kvm_device *dev, u32 type)
ret = kvm_setup_default_irq_routing(kvm);
if (ret)
- return -ENOMEM;
+ return ret;
s = kzalloc_obj(struct loongarch_pch_pic);
if (!s)
--
2.43.0
^ permalink raw reply [flat|nested] 15+ messages in thread
* [PATCH 6/6] LoongArch: KVM: Reject repeated PCH-PIC CTRL_INIT
2026-09-27 7:52 [PATCH 0/6] LoongArch: KVM: irqchip and steal time fixes Tao Cui
` (4 preceding siblings ...)
2026-09-27 7:52 ` [PATCH 5/6] LoongArch: KVM: Propagate real error code in kvm_pch_pic_create Tao Cui
@ 2026-09-27 7:52 ` Tao Cui
2026-09-28 8:01 ` Bibo Mao
5 siblings, 1 reply; 15+ messages in thread
From: Tao Cui @ 2026-09-27 7:52 UTC (permalink / raw)
To: maobibo, gaosong, zhaotianrui
Cc: loongarch, kvm, linux-kernel, chenhuacai, kernel,
nagachaithanya9911, cui.tao, Tao Cui
From: Tao Cui <cuitao@kylinos.cn>
KVM_DEV_LOONGARCH_PCH_PIC_CTRL_INIT has no guard against repeated
invocation: every call overwrites pch_pic_base and registers the same
kvm_io_device on the MMIO bus at the new address, while
kvm_pch_pic_destroy() unregisters only one bus range. After a repeated
init, MMIO to the stale ranges computes its register offset against the
new base and silently reads 0 / drops writes, and the leftover bus
entries persist until the VM is destroyed.
Reject a repeated init with -EBUSY, a null address with -EINVAL, and
only set pch_pic_base after the bus registration succeeds so a failed
init does not leave the device half-initialized. The registration
error is propagated instead of being replaced with -EFAULT.
Fixes: d206d9514873 ("LoongArch: KVM: Add PCHPIC user mode read and write functions")
Signed-off-by: Tao Cui <cuitao@kylinos.cn>
---
arch/loongarch/kvm/intc/pch_pic.c | 10 ++++++++--
1 file changed, 8 insertions(+), 2 deletions(-)
diff --git a/arch/loongarch/kvm/intc/pch_pic.c b/arch/loongarch/kvm/intc/pch_pic.c
index 136211f154d4..2baf9ecf0c8a 100644
--- a/arch/loongarch/kvm/intc/pch_pic.c
+++ b/arch/loongarch/kvm/intc/pch_pic.c
@@ -285,7 +285,10 @@ static int kvm_pch_pic_init(struct kvm_device *dev, u64 addr)
struct kvm_io_device *device;
struct loongarch_pch_pic *s = dev->kvm->arch.pch_pic;
- s->pch_pic_base = addr;
+ if (!addr)
+ return -EINVAL;
+ if (s->pch_pic_base)
+ return -EBUSY;
device = &s->device;
/* init device by pch pic writing and reading ops */
kvm_iodevice_init(device, &kvm_pch_pic_ops);
@@ -293,8 +296,11 @@ static int kvm_pch_pic_init(struct kvm_device *dev, u64 addr)
/* register pch pic device */
ret = kvm_io_bus_register_dev(kvm, KVM_MMIO_BUS, addr, PCH_PIC_SIZE, device);
mutex_unlock(&kvm->slots_lock);
+ if (ret < 0)
+ return ret;
- return (ret < 0) ? -EFAULT : 0;
+ s->pch_pic_base = addr;
+ return 0;
}
/* used by user space to get or set pch pic registers */
--
2.43.0
^ permalink raw reply [flat|nested] 15+ messages in thread
* Re: [PATCH 1/6] LoongArch: KVM: Clear device pointer in irqchip destroy callbacks
2026-09-27 7:52 ` [PATCH 1/6] LoongArch: KVM: Clear device pointer in irqchip destroy callbacks Tao Cui
@ 2026-09-28 4:07 ` Bibo Mao
0 siblings, 0 replies; 15+ messages in thread
From: Bibo Mao @ 2026-09-28 4:07 UTC (permalink / raw)
To: Tao Cui, gaosong, zhaotianrui
Cc: loongarch, kvm, linux-kernel, chenhuacai, kernel,
nagachaithanya9911, Tao Cui
On 2026/9/27 下午3:52, Tao Cui wrote:
> From: Tao Cui <cuitao@kylinos.cn>
>
> The destroy callbacks of the four irqchip devices free the device
> structure without clearing kvm->arch.{ipi,eiointc,pch_pic,dmsintc},
> leaving a dangling pointer. ops->destroy() is not only called at VM
> teardown but also when KVM_CREATE_DEVICE succeeds and the following
> anon_inode_getfd() fails (e.g. under RLIMIT_NOFILE); the VM stays
> alive, kvm_arch_irqchip_in_kernel() still reports true, and interrupt
> injection dereferences the freed device.
>
> Clear the pointer before freeing.
>
> Fixes: c532de5a67a7 ("LoongArch: KVM: Add IPI device support")
> Fixes: 2e8b9df82631 ("LoongArch: KVM: Add EIOINTC device support")
> Fixes: e785dfacf7e7 ("LoongArch: KVM: Add PCHPIC device support")
> Fixes: 229132c309d6 ("LoongArch: KVM: Add DMSINTC device support")
> Signed-off-by: Tao Cui <cuitao@kylinos.cn>
> ---
> arch/loongarch/kvm/intc/dmsintc.c | 5 ++++-
> arch/loongarch/kvm/intc/eiointc.c | 1 +
> arch/loongarch/kvm/intc/ipi.c | 1 +
> arch/loongarch/kvm/intc/pch_pic.c | 1 +
> 4 files changed, 7 insertions(+), 1 deletion(-)
>
> diff --git a/arch/loongarch/kvm/intc/dmsintc.c b/arch/loongarch/kvm/intc/dmsintc.c
> index 89f980d867be..63072595c01b 100644
> --- a/arch/loongarch/kvm/intc/dmsintc.c
> +++ b/arch/loongarch/kvm/intc/dmsintc.c
> @@ -161,11 +161,14 @@ static int kvm_dmsintc_create(struct kvm_device *dev, u32 type)
>
> static void kvm_dmsintc_destroy(struct kvm_device *dev)
> {
> + struct loongarch_dmsintc *s;
>
> if (!dev || !dev->kvm || !dev->kvm->arch.dmsintc)
> return;
>
> - kfree(dev->kvm->arch.dmsintc);
> + s = dev->kvm->arch.dmsintc;
> + dev->kvm->arch.dmsintc = NULL;
> + kfree(s);
one line patch seems better here such as:
+ dev->kvm->arch.dmsintc = NULL;
Regards
Bibo Mao
> kfree(dev);
> }
>
> diff --git a/arch/loongarch/kvm/intc/eiointc.c b/arch/loongarch/kvm/intc/eiointc.c
> index 80f78e07c74a..fe0a1918f26f 100644
> --- a/arch/loongarch/kvm/intc/eiointc.c
> +++ b/arch/loongarch/kvm/intc/eiointc.c
> @@ -675,6 +675,7 @@ static void kvm_eiointc_destroy(struct kvm_device *dev)
>
> kvm = dev->kvm;
> eiointc = kvm->arch.eiointc;
> + kvm->arch.eiointc = NULL;
> mutex_lock(&kvm->slots_lock);
> kvm_io_bus_unregister_dev(kvm, KVM_IOCSR_BUS, &eiointc->device);
> kvm_io_bus_unregister_dev(kvm, KVM_IOCSR_BUS, &eiointc->device_vext);
> diff --git a/arch/loongarch/kvm/intc/ipi.c b/arch/loongarch/kvm/intc/ipi.c
> index 7b333a4a0430..6ee90c10827a 100644
> --- a/arch/loongarch/kvm/intc/ipi.c
> +++ b/arch/loongarch/kvm/intc/ipi.c
> @@ -444,6 +444,7 @@ static void kvm_ipi_destroy(struct kvm_device *dev)
>
> kvm = dev->kvm;
> ipi = kvm->arch.ipi;
> + kvm->arch.ipi = NULL;
> mutex_lock(&kvm->slots_lock);
> kvm_io_bus_unregister_dev(kvm, KVM_IOCSR_BUS, &ipi->device);
> mutex_unlock(&kvm->slots_lock);
> diff --git a/arch/loongarch/kvm/intc/pch_pic.c b/arch/loongarch/kvm/intc/pch_pic.c
> index 2b63b0c2c7ce..62c09b5f3937 100644
> --- a/arch/loongarch/kvm/intc/pch_pic.c
> +++ b/arch/loongarch/kvm/intc/pch_pic.c
> @@ -483,6 +483,7 @@ static void kvm_pch_pic_destroy(struct kvm_device *dev)
>
> kvm = dev->kvm;
> s = kvm->arch.pch_pic;
> + kvm->arch.pch_pic = NULL;
> /* unregister pch pic device and free it's memory */
> mutex_lock(&kvm->slots_lock);
> kvm_io_bus_unregister_dev(kvm, KVM_MMIO_BUS, &s->device);
>
^ permalink raw reply [flat|nested] 15+ messages in thread
* Re: [PATCH 2/6] LoongArch: KVM: Guard against NULL irqchip in irq injection
2026-09-27 7:52 ` [PATCH 2/6] LoongArch: KVM: Guard against NULL irqchip in irq injection Tao Cui
@ 2026-09-28 4:14 ` Bibo Mao
2026-09-28 9:24 ` Tao Cui
0 siblings, 1 reply; 15+ messages in thread
From: Bibo Mao @ 2026-09-28 4:14 UTC (permalink / raw)
To: Tao Cui, gaosong, zhaotianrui
Cc: loongarch, kvm, linux-kernel, chenhuacai, kernel,
nagachaithanya9911, Tao Cui
On 2026/9/27 下午3:52, Tao Cui wrote:
> From: Tao Cui <cuitao@kylinos.cn>
>
> The irqfd injection path dispatches through the irq routing table
> without the kvm_arch_irqchip_in_kernel() gate that protects
> KVM_IRQ_LINE: kvm_set_pic_irq() and kvm_arch_set_irq_inatomic() call
> pch_pic_set_irq(kvm->arch.pch_pic, ...) and pch_msi_set_irq() calls
> eiointc_set_irq(kvm->arch.eiointc, ...) directly, so a routing entry
> that outlives the corresponding in-kernel irqchip dereferences a NULL
> (or, before the previous patch, a freed) device.
>
> Make pch_pic_set_irq(), eiointc_set_irq() and dmsintc_set_irq()
> tolerate a NULL device.
>
> Fixes: 1928254c5ccb ("LoongArch: KVM: Add irqfd support")
> Signed-off-by: Tao Cui <cuitao@kylinos.cn>
> ---
> arch/loongarch/kvm/intc/dmsintc.c | 3 +++
> arch/loongarch/kvm/intc/eiointc.c | 5 ++++-
> arch/loongarch/kvm/intc/pch_pic.c | 3 +++
> 3 files changed, 10 insertions(+), 1 deletion(-)
>
> diff --git a/arch/loongarch/kvm/intc/dmsintc.c b/arch/loongarch/kvm/intc/dmsintc.c
> index 63072595c01b..e27f448bb54f 100644
> --- a/arch/loongarch/kvm/intc/dmsintc.c
> +++ b/arch/loongarch/kvm/intc/dmsintc.c
> @@ -70,6 +70,9 @@ int dmsintc_set_irq(struct kvm *kvm, u64 addr, int data, int level)
> unsigned int irq, cpu;
> struct kvm_vcpu *vcpu;
>
> + if (!kvm->arch.dmsintc)
> + return -EINVAL;
> +
There is kvm->arch.dmsintc checking in its caller function
pch_msi_set_irq().
> irq = (addr >> AVEC_IRQ_SHIFT) & AVEC_IRQ_MASK;
> cpu = (addr >> AVEC_CPU_SHIFT) & kvm->arch.dmsintc->cpu_mask;
> if (cpu >= KVM_MAX_VCPUS)
> diff --git a/arch/loongarch/kvm/intc/eiointc.c b/arch/loongarch/kvm/intc/eiointc.c
> index fe0a1918f26f..c68f9033638a 100644
> --- a/arch/loongarch/kvm/intc/eiointc.c
> +++ b/arch/loongarch/kvm/intc/eiointc.c
> @@ -112,8 +112,11 @@ static inline void eiointc_update_sw_coremap(struct loongarch_eiointc *s,
> void eiointc_set_irq(struct loongarch_eiointc *s, int irq, int level)
> {
> unsigned long flags;
> - unsigned long *isr = (unsigned long *)s->isr;
> + unsigned long *isr;
>
> + if (!s)
> + return;
When is it possible that parameter s is NULL? There is
kvm_arch_irqchip_in_kernel() checking in kvm_send_userspace_msi().
> + isr = (unsigned long *)s->isr;
> spin_lock_irqsave(&s->lock, flags);
> level ? __set_bit(irq, isr) : __clear_bit(irq, isr);
> eiointc_update_irq(s, irq, level);
> diff --git a/arch/loongarch/kvm/intc/pch_pic.c b/arch/loongarch/kvm/intc/pch_pic.c
> index 62c09b5f3937..88666425d800 100644
> --- a/arch/loongarch/kvm/intc/pch_pic.c
> +++ b/arch/loongarch/kvm/intc/pch_pic.c
> @@ -48,6 +48,9 @@ void pch_pic_set_irq(struct loongarch_pch_pic *s, int irq, int level)
> {
> u64 mask = BIT(irq);
>
> + if (!s)
> + return;
> +
When is it possible that parameter s is NULL? I think that when in
kernel irqchip is removed, there is no way to inject line IRQ in kernel
side.
Regards
Bibo Mao
> spin_lock(&s->lock);
> if (level)
> s->irr |= mask; /* set irr */
>
^ permalink raw reply [flat|nested] 15+ messages in thread
* Re: [PATCH 4/6] LoongArch: KVM: Rebase steal time counter in vcpu context
2026-09-27 7:52 ` [PATCH 4/6] LoongArch: KVM: Rebase steal time counter in vcpu context Tao Cui
@ 2026-09-28 7:23 ` Bibo Mao
2026-09-28 9:28 ` Tao Cui
0 siblings, 1 reply; 15+ messages in thread
From: Bibo Mao @ 2026-09-28 7:23 UTC (permalink / raw)
To: Tao Cui, gaosong, zhaotianrui
Cc: loongarch, kvm, linux-kernel, chenhuacai, kernel,
nagachaithanya9911, Tao Cui
On 2026/9/27 下午3:52, Tao Cui wrote:
> From: Tao Cui <cuitao@kylinos.cn>
>
> KVM_SET_DEVICE_ATTR(PVTIME GPA) initializes st.last_steal from the
> ioctl thread's run_delay, but kvm_update_stolen_time() accumulates the
> vcpu thread's run_delay; the first delta can be negative and wraps in
I think that ioctl thread is the vCPU thread itself in KVM mode, there
will be many potential problems when one thread set registers of vCPU
while vCPU is running.
Regards
Bibo Mao
> u64, so the guest reads a steal time close to 2^64.
>
> Drop the initialization from the attr path and lazily rebase the
> counter on the first steal update, which runs in vcpu context like the
> hypercall path. Re-registering a GPA clears the rebase sentinel first
> and publishes the address with a write barrier; the reader side loads
> the address and sentinel with READ_ONCE and a read barrier so a
> concurrent re-registration is not observed half-applied.
>
> Fixes: b4ba157044ea ("LoongArch: KVM: Add PV steal time support in host side")
> Signed-off-by: Tao Cui <cuitao@kylinos.cn>
> ---
> arch/loongarch/kvm/vcpu.c | 25 +++++++++++++++++--------
> 1 file changed, 17 insertions(+), 8 deletions(-)
>
> diff --git a/arch/loongarch/kvm/vcpu.c b/arch/loongarch/kvm/vcpu.c
> index 8e028be3f0a9..d931a5a802f6 100644
> --- a/arch/loongarch/kvm/vcpu.c
> +++ b/arch/loongarch/kvm/vcpu.c
> @@ -154,12 +154,13 @@ static void kvm_update_stolen_time(struct kvm_vcpu *vcpu)
> u32 version;
> u64 steal;
> gpa_t gpa;
> + u64 last_steal;
> struct kvm_memslots *slots;
> struct kvm_steal_time __user *st;
> struct gfn_to_hva_cache *ghc;
>
> ghc = &vcpu->arch.st.cache;
> - gpa = vcpu->arch.st.guest_addr;
> + gpa = READ_ONCE(vcpu->arch.st.guest_addr);
> if (!(gpa & KVM_STEAL_PHYS_VALID))
> return;
>
> @@ -187,8 +188,14 @@ static void kvm_update_stolen_time(struct kvm_vcpu *vcpu)
> smp_wmb();
>
> unsafe_get_user(steal, &st->steal, out);
> - steal += current->sched_info.run_delay - vcpu->arch.st.last_steal;
> - vcpu->arch.st.last_steal = current->sched_info.run_delay;
> + /* acquire pairs with the smp_wmb in the attr path */
> + smp_rmb();
> + last_steal = READ_ONCE(vcpu->arch.st.last_steal);
> + if (!last_steal)
> + /* first update in vcpu context: rebase the counter */
> + last_steal = current->sched_info.run_delay;
> + steal += current->sched_info.run_delay - last_steal;
> + WRITE_ONCE(vcpu->arch.st.last_steal, current->sched_info.run_delay);
> unsafe_put_user(steal, &st->steal, out);
>
> smp_wmb();
> @@ -1122,7 +1129,7 @@ static int kvm_loongarch_pvtime_get_attr(struct kvm_vcpu *vcpu,
> || attr->attr != KVM_LOONGARCH_VCPU_PVTIME_GPA)
> return -ENXIO;
>
> - gpa = vcpu->arch.st.guest_addr;
> + gpa = READ_ONCE(vcpu->arch.st.guest_addr);
> if (put_user(gpa, user))
> return -EFAULT;
>
> @@ -1197,7 +1204,7 @@ static int kvm_loongarch_pvtime_set_attr(struct kvm_vcpu *vcpu,
> return -EINVAL;
>
> if (!(gpa & KVM_STEAL_PHYS_VALID)) {
> - vcpu->arch.st.guest_addr = gpa;
> + WRITE_ONCE(vcpu->arch.st.guest_addr, gpa);
> return 0;
> }
>
> @@ -1208,8 +1215,10 @@ static int kvm_loongarch_pvtime_set_attr(struct kvm_vcpu *vcpu,
> srcu_read_unlock(&kvm->srcu, idx);
>
> if (!ret) {
> - vcpu->arch.st.guest_addr = gpa;
> - vcpu->arch.st.last_steal = current->sched_info.run_delay;
> + WRITE_ONCE(vcpu->arch.st.last_steal, 0);
> + /* publish the new address only after clearing the rebase sentinel */
> + smp_wmb();
> + WRITE_ONCE(vcpu->arch.st.guest_addr, gpa);
> kvm_make_request(KVM_REQ_STEAL_UPDATE, vcpu);
> }
>
> @@ -1800,7 +1809,7 @@ static void kvm_vcpu_set_pv_preempted(struct kvm_vcpu *vcpu)
> struct kvm_memslots *slots;
> struct kvm_steal_time __user *st;
>
> - gpa = vcpu->arch.st.guest_addr;
> + gpa = READ_ONCE(vcpu->arch.st.guest_addr);
> if (!(gpa & KVM_STEAL_PHYS_VALID))
> return;
>
>
^ permalink raw reply [flat|nested] 15+ messages in thread
* Re: [PATCH 5/6] LoongArch: KVM: Propagate real error code in kvm_pch_pic_create
2026-09-27 7:52 ` [PATCH 5/6] LoongArch: KVM: Propagate real error code in kvm_pch_pic_create Tao Cui
@ 2026-09-28 7:25 ` Bibo Mao
0 siblings, 0 replies; 15+ messages in thread
From: Bibo Mao @ 2026-09-28 7:25 UTC (permalink / raw)
To: Tao Cui, gaosong, zhaotianrui
Cc: loongarch, kvm, linux-kernel, chenhuacai, kernel,
nagachaithanya9911, Tao Cui
On 2026/9/27 下午3:52, Tao Cui wrote:
> From: Tao Cui <cuitao@kylinos.cn>
>
> kvm_pch_pic_create() replaces the real error of
> kvm_setup_default_irq_routing() with a fixed -ENOMEM. Return the
> actual code, as kvm_eiointc_create() already does; the kvm_ipi_create()
> case is fixed separately by "Return the actual error code in
> kvm_ipi_create()" (Chaithanya Lagisetty).
>
> Fixes: 1928254c5ccb ("LoongArch: KVM: Add irqfd support")
> Signed-off-by: Tao Cui <cuitao@kylinos.cn>
> ---
> arch/loongarch/kvm/intc/pch_pic.c | 2 +-
> 1 file changed, 1 insertion(+), 1 deletion(-)
>
> diff --git a/arch/loongarch/kvm/intc/pch_pic.c b/arch/loongarch/kvm/intc/pch_pic.c
> index e33d6693c37c..136211f154d4 100644
> --- a/arch/loongarch/kvm/intc/pch_pic.c
> +++ b/arch/loongarch/kvm/intc/pch_pic.c
> @@ -448,7 +448,7 @@ static int kvm_pch_pic_create(struct kvm_device *dev, u32 type)
>
> ret = kvm_setup_default_irq_routing(kvm);
> if (ret)
> - return -ENOMEM;
> + return ret;
>
> s = kzalloc_obj(struct loongarch_pch_pic);
> if (!s)
>
Reviewed-by: Bibo Mao <maobibo@loongson.cn>
^ permalink raw reply [flat|nested] 15+ messages in thread
* Re: [PATCH 6/6] LoongArch: KVM: Reject repeated PCH-PIC CTRL_INIT
2026-09-27 7:52 ` [PATCH 6/6] LoongArch: KVM: Reject repeated PCH-PIC CTRL_INIT Tao Cui
@ 2026-09-28 8:01 ` Bibo Mao
2026-09-28 9:31 ` Tao Cui
0 siblings, 1 reply; 15+ messages in thread
From: Bibo Mao @ 2026-09-28 8:01 UTC (permalink / raw)
To: Tao Cui, gaosong, zhaotianrui
Cc: loongarch, kvm, linux-kernel, chenhuacai, kernel,
nagachaithanya9911, Tao Cui
On 2026/9/27 下午3:52, Tao Cui wrote:
> From: Tao Cui <cuitao@kylinos.cn>
>
> KVM_DEV_LOONGARCH_PCH_PIC_CTRL_INIT has no guard against repeated
> invocation: every call overwrites pch_pic_base and registers the same
> kvm_io_device on the MMIO bus at the new address, while
> kvm_pch_pic_destroy() unregisters only one bus range. After a repeated
> init, MMIO to the stale ranges computes its register offset against the
> new base and silently reads 0 / drops writes, and the leftover bus
> entries persist until the VM is destroyed.
>
> Reject a repeated init with -EBUSY, a null address with -EINVAL, and
> only set pch_pic_base after the bus registration succeeds so a failed
> init does not leave the device half-initialized. The registration
> error is propagated instead of being replaced with -EFAULT.
>
> Fixes: d206d9514873 ("LoongArch: KVM: Add PCHPIC user mode read and write functions")
> Signed-off-by: Tao Cui <cuitao@kylinos.cn>
> ---
> arch/loongarch/kvm/intc/pch_pic.c | 10 ++++++++--
> 1 file changed, 8 insertions(+), 2 deletions(-)
>
> diff --git a/arch/loongarch/kvm/intc/pch_pic.c b/arch/loongarch/kvm/intc/pch_pic.c
> index 136211f154d4..2baf9ecf0c8a 100644
> --- a/arch/loongarch/kvm/intc/pch_pic.c
> +++ b/arch/loongarch/kvm/intc/pch_pic.c
> @@ -285,7 +285,10 @@ static int kvm_pch_pic_init(struct kvm_device *dev, u64 addr)
> struct kvm_io_device *device;
> struct loongarch_pch_pic *s = dev->kvm->arch.pch_pic;
>
> - s->pch_pic_base = addr;
> + if (!addr)
> + return -EINVAL;
To detect repeated init, I think it will be better add a BOOL type
variable such ready/has_init for this. In theory device with physical
address 0 is possible.
> + if (s->pch_pic_base)
> + return -EBUSY;
-EEXIST or 0 if it is created already?
Regards
Bibo Mao
> device = &s->device;
> /* init device by pch pic writing and reading ops */
> kvm_iodevice_init(device, &kvm_pch_pic_ops);
> @@ -293,8 +296,11 @@ static int kvm_pch_pic_init(struct kvm_device *dev, u64 addr)
> /* register pch pic device */
> ret = kvm_io_bus_register_dev(kvm, KVM_MMIO_BUS, addr, PCH_PIC_SIZE, device);
> mutex_unlock(&kvm->slots_lock);
> + if (ret < 0)
> + return ret;
>
> - return (ret < 0) ? -EFAULT : 0;
> + s->pch_pic_base = addr;
> + return 0;
> }
>
> /* used by user space to get or set pch pic registers */
>
^ permalink raw reply [flat|nested] 15+ messages in thread
* Re: [PATCH 2/6] LoongArch: KVM: Guard against NULL irqchip in irq injection
2026-09-28 4:14 ` Bibo Mao
@ 2026-09-28 9:24 ` Tao Cui
0 siblings, 0 replies; 15+ messages in thread
From: Tao Cui @ 2026-09-28 9:24 UTC (permalink / raw)
To: Bibo Mao, gaosong, zhaotianrui
Cc: cui.tao, loongarch, kvm, linux-kernel, chenhuacai, kernel,
nagachaithanya9911, Tao Cui
在 2026/9/28 12:14, Bibo Mao 写道:
>
>
> On 2026/9/27 下午3:52, Tao Cui wrote:
>> From: Tao Cui <cuitao@kylinos.cn>
>>
>> The irqfd injection path dispatches through the irq routing table
>> without the kvm_arch_irqchip_in_kernel() gate that protects
>> KVM_IRQ_LINE: kvm_set_pic_irq() and kvm_arch_set_irq_inatomic() call
>> pch_pic_set_irq(kvm->arch.pch_pic, ...) and pch_msi_set_irq() calls
>> eiointc_set_irq(kvm->arch.eiointc, ...) directly, so a routing entry
>> that outlives the corresponding in-kernel irqchip dereferences a NULL
>> (or, before the previous patch, a freed) device.
>>
>> Make pch_pic_set_irq(), eiointc_set_irq() and dmsintc_set_irq()
>> tolerate a NULL device.
>>
>> Fixes: 1928254c5ccb ("LoongArch: KVM: Add irqfd support")
>> Signed-off-by: Tao Cui <cuitao@kylinos.cn>
>> ---
>> arch/loongarch/kvm/intc/dmsintc.c | 3 +++
>> arch/loongarch/kvm/intc/eiointc.c | 5 ++++-
>> arch/loongarch/kvm/intc/pch_pic.c | 3 +++
>> 3 files changed, 10 insertions(+), 1 deletion(-)
>>
>> diff --git a/arch/loongarch/kvm/intc/dmsintc.c b/arch/loongarch/kvm/intc/dmsintc.c
>> index 63072595c01b..e27f448bb54f 100644
>> --- a/arch/loongarch/kvm/intc/dmsintc.c
>> +++ b/arch/loongarch/kvm/intc/dmsintc.c
>> @@ -70,6 +70,9 @@ int dmsintc_set_irq(struct kvm *kvm, u64 addr, int data, int level)
>> unsigned int irq, cpu;
>> struct kvm_vcpu *vcpu;
>> + if (!kvm->arch.dmsintc)
>> + return -EINVAL;
>> +
> There is kvm->arch.dmsintc checking in its caller function pch_msi_set_irq().
>
>> irq = (addr >> AVEC_IRQ_SHIFT) & AVEC_IRQ_MASK;
>> cpu = (addr >> AVEC_CPU_SHIFT) & kvm->arch.dmsintc->cpu_mask;
>> if (cpu >= KVM_MAX_VCPUS)
>> diff --git a/arch/loongarch/kvm/intc/eiointc.c b/arch/loongarch/kvm/intc/eiointc.c
>> index fe0a1918f26f..c68f9033638a 100644
>> --- a/arch/loongarch/kvm/intc/eiointc.c
>> +++ b/arch/loongarch/kvm/intc/eiointc.c
>> @@ -112,8 +112,11 @@ static inline void eiointc_update_sw_coremap(struct loongarch_eiointc *s,
>> void eiointc_set_irq(struct loongarch_eiointc *s, int irq, int level)
>> {
>> unsigned long flags;
>> - unsigned long *isr = (unsigned long *)s->isr;
>> + unsigned long *isr;
>> + if (!s)
>> + return;
> When is it possible that parameter s is NULL? There is kvm_arch_irqchip_in_kernel() checking in kvm_send_userspace_msi().
>
>> + isr = (unsigned long *)s->isr;
>> spin_lock_irqsave(&s->lock, flags);
>> level ? __set_bit(irq, isr) : __clear_bit(irq, isr);
>> eiointc_update_irq(s, irq, level);
>> diff --git a/arch/loongarch/kvm/intc/pch_pic.c b/arch/loongarch/kvm/intc/pch_pic.c
>> index 62c09b5f3937..88666425d800 100644
>> --- a/arch/loongarch/kvm/intc/pch_pic.c
>> +++ b/arch/loongarch/kvm/intc/pch_pic.c
>> @@ -48,6 +48,9 @@ void pch_pic_set_irq(struct loongarch_pch_pic *s, int irq, int level)
>> {
>> u64 mask = BIT(irq);
>> + if (!s)
>> + return;
>> +
> When is it possible that parameter s is NULL? I think that when in kernel irqchip is removed, there is no way to inject line IRQ in kernel side.
>
Thanks for looking into this.
You're right. I rechecked the lifetime rules after your comments, and
my reasoning for this patch was flawed.
I had assumed an irqfd could remain active after the corresponding
in-kernel irqchip had been removed. In fact, that state is not
reachable. The irqchip devices are not destroyed during the lifetime
of a running VM, and during VM teardown irqfds are released before
the irqchip devices are destroyed. For DMSINTC, the rollback path is
already guarded by `pch_msi_set_irq()`.
So these NULL checks are not protecting a real execution path. I'll
drop this patch in the next revision.
Thanks,
Tao
> Regards
> Bibo Mao
>> spin_lock(&s->lock);
>> if (level)
>> s->irr |= mask; /* set irr */
>>
>
^ permalink raw reply [flat|nested] 15+ messages in thread
* Re: [PATCH 4/6] LoongArch: KVM: Rebase steal time counter in vcpu context
2026-09-28 7:23 ` Bibo Mao
@ 2026-09-28 9:28 ` Tao Cui
0 siblings, 0 replies; 15+ messages in thread
From: Tao Cui @ 2026-09-28 9:28 UTC (permalink / raw)
To: Bibo Mao, gaosong, zhaotianrui
Cc: cui.tao, loongarch, kvm, linux-kernel, chenhuacai, kernel,
nagachaithanya9911, Tao Cui
在 2026/9/28 15:23, Bibo Mao 写道:
>
>
> On 2026/9/27 下午3:52, Tao Cui wrote:
>> From: Tao Cui <cuitao@kylinos.cn>
>>
>> KVM_SET_DEVICE_ATTR(PVTIME GPA) initializes st.last_steal from the
>> ioctl thread's run_delay, but kvm_update_stolen_time() accumulates the
>> vcpu thread's run_delay; the first delta can be negative and wraps in
> I think that ioctl thread is the vCPU thread itself in KVM mode, there will be many potential problems when one thread set registers of vCPU while vCPU is running.
>
You're right. I assumed KVM_SET_DEVICE_ATTR() could be issued
from a thread different from the vCPU thread. Looking at the expected
userspace flow again, the pvtime GPA is registered by the guest kernel
via hypercall (which runs in vCPU context), or set by the vCPU thread
itself before entering KVM_RUN. In both cases `last_steal` and the
subsequent steal updates use the same task's `run_delay`, so the
negative delta I described cannot occur.
I'll drop this patch as well.
Thanks,
Tao> Regards
> Bibo Mao
>> u64, so the guest reads a steal time close to 2^64.
>>
>> Drop the initialization from the attr path and lazily rebase the
>> counter on the first steal update, which runs in vcpu context like the
>> hypercall path. Re-registering a GPA clears the rebase sentinel first
>> and publishes the address with a write barrier; the reader side loads
>> the address and sentinel with READ_ONCE and a read barrier so a
>> concurrent re-registration is not observed half-applied.
>>
>> Fixes: b4ba157044ea ("LoongArch: KVM: Add PV steal time support in host side")
>> Signed-off-by: Tao Cui <cuitao@kylinos.cn>
>> ---
>> arch/loongarch/kvm/vcpu.c | 25 +++++++++++++++++--------
>> 1 file changed, 17 insertions(+), 8 deletions(-)
>>
>> diff --git a/arch/loongarch/kvm/vcpu.c b/arch/loongarch/kvm/vcpu.c
>> index 8e028be3f0a9..d931a5a802f6 100644
>> --- a/arch/loongarch/kvm/vcpu.c
>> +++ b/arch/loongarch/kvm/vcpu.c
>> @@ -154,12 +154,13 @@ static void kvm_update_stolen_time(struct kvm_vcpu *vcpu)
>> u32 version;
>> u64 steal;
>> gpa_t gpa;
>> + u64 last_steal;
>> struct kvm_memslots *slots;
>> struct kvm_steal_time __user *st;
>> struct gfn_to_hva_cache *ghc;
>> ghc = &vcpu->arch.st.cache;
>> - gpa = vcpu->arch.st.guest_addr;
>> + gpa = READ_ONCE(vcpu->arch.st.guest_addr);
>> if (!(gpa & KVM_STEAL_PHYS_VALID))
>> return;
>> @@ -187,8 +188,14 @@ static void kvm_update_stolen_time(struct kvm_vcpu *vcpu)
>> smp_wmb();
>> unsafe_get_user(steal, &st->steal, out);
>> - steal += current->sched_info.run_delay - vcpu->arch.st.last_steal;
>> - vcpu->arch.st.last_steal = current->sched_info.run_delay;
>> + /* acquire pairs with the smp_wmb in the attr path */
>> + smp_rmb();
>> + last_steal = READ_ONCE(vcpu->arch.st.last_steal);
>> + if (!last_steal)
>> + /* first update in vcpu context: rebase the counter */
>> + last_steal = current->sched_info.run_delay;
>> + steal += current->sched_info.run_delay - last_steal;
>> + WRITE_ONCE(vcpu->arch.st.last_steal, current->sched_info.run_delay);
>> unsafe_put_user(steal, &st->steal, out);
>> smp_wmb();
>> @@ -1122,7 +1129,7 @@ static int kvm_loongarch_pvtime_get_attr(struct kvm_vcpu *vcpu,
>> || attr->attr != KVM_LOONGARCH_VCPU_PVTIME_GPA)
>> return -ENXIO;
>> - gpa = vcpu->arch.st.guest_addr;
>> + gpa = READ_ONCE(vcpu->arch.st.guest_addr);
>> if (put_user(gpa, user))
>> return -EFAULT;
>> @@ -1197,7 +1204,7 @@ static int kvm_loongarch_pvtime_set_attr(struct kvm_vcpu *vcpu,
>> return -EINVAL;
>> if (!(gpa & KVM_STEAL_PHYS_VALID)) {
>> - vcpu->arch.st.guest_addr = gpa;
>> + WRITE_ONCE(vcpu->arch.st.guest_addr, gpa);
>> return 0;
>> }
>> @@ -1208,8 +1215,10 @@ static int kvm_loongarch_pvtime_set_attr(struct kvm_vcpu *vcpu,
>> srcu_read_unlock(&kvm->srcu, idx);
>> if (!ret) {
>> - vcpu->arch.st.guest_addr = gpa;
>> - vcpu->arch.st.last_steal = current->sched_info.run_delay;
>> + WRITE_ONCE(vcpu->arch.st.last_steal, 0);
>> + /* publish the new address only after clearing the rebase sentinel */
>> + smp_wmb();
>> + WRITE_ONCE(vcpu->arch.st.guest_addr, gpa);
>> kvm_make_request(KVM_REQ_STEAL_UPDATE, vcpu);
>> }
>> @@ -1800,7 +1809,7 @@ static void kvm_vcpu_set_pv_preempted(struct kvm_vcpu *vcpu)
>> struct kvm_memslots *slots;
>> struct kvm_steal_time __user *st;
>> - gpa = vcpu->arch.st.guest_addr;
>> + gpa = READ_ONCE(vcpu->arch.st.guest_addr);
>> if (!(gpa & KVM_STEAL_PHYS_VALID))
>> return;
>>
>
^ permalink raw reply [flat|nested] 15+ messages in thread
* Re: [PATCH 6/6] LoongArch: KVM: Reject repeated PCH-PIC CTRL_INIT
2026-09-28 8:01 ` Bibo Mao
@ 2026-09-28 9:31 ` Tao Cui
0 siblings, 0 replies; 15+ messages in thread
From: Tao Cui @ 2026-09-28 9:31 UTC (permalink / raw)
To: Bibo Mao, gaosong, zhaotianrui
Cc: cui.tao, loongarch, kvm, linux-kernel, chenhuacai, kernel,
nagachaithanya9911, Tao Cui
在 2026/9/28 16:01, Bibo Mao 写道:
>
>
> On 2026/9/27 下午3:52, Tao Cui wrote:
>> From: Tao Cui <cuitao@kylinos.cn>
>>
>> KVM_DEV_LOONGARCH_PCH_PIC_CTRL_INIT has no guard against repeated
>> invocation: every call overwrites pch_pic_base and registers the same
>> kvm_io_device on the MMIO bus at the new address, while
>> kvm_pch_pic_destroy() unregisters only one bus range. After a repeated
>> init, MMIO to the stale ranges computes its register offset against the
>> new base and silently reads 0 / drops writes, and the leftover bus
>> entries persist until the VM is destroyed.
>>
>> Reject a repeated init with -EBUSY, a null address with -EINVAL, and
>> only set pch_pic_base after the bus registration succeeds so a failed
>> init does not leave the device half-initialized. The registration
>> error is propagated instead of being replaced with -EFAULT.
>>
>> Fixes: d206d9514873 ("LoongArch: KVM: Add PCHPIC user mode read and write functions")
>> Signed-off-by: Tao Cui <cuitao@kylinos.cn>
>> ---
>> arch/loongarch/kvm/intc/pch_pic.c | 10 ++++++++--
>> 1 file changed, 8 insertions(+), 2 deletions(-)
>>
>> diff --git a/arch/loongarch/kvm/intc/pch_pic.c b/arch/loongarch/kvm/intc/pch_pic.c
>> index 136211f154d4..2baf9ecf0c8a 100644
>> --- a/arch/loongarch/kvm/intc/pch_pic.c
>> +++ b/arch/loongarch/kvm/intc/pch_pic.c
>> @@ -285,7 +285,10 @@ static int kvm_pch_pic_init(struct kvm_device *dev, u64 addr)
>> struct kvm_io_device *device;
>> struct loongarch_pch_pic *s = dev->kvm->arch.pch_pic;
>> - s->pch_pic_base = addr;
>> + if (!addr)
>> + return -EINVAL;
> To detect repeated init, I think it will be better add a BOOL type variable such ready/has_init for this. In theory device with physical address 0 is possible.
>
Thanks for the review.
Good point. Using `pch_pic_base == 0` as the initialization state mixes the device state with the configured address, a separate boolean is clearer.
>> + if (s->pch_pic_base)
>> + return -EBUSY;
> -EEXIST or 0 if it is created already?
>
I also agree that `-EEXIST` is a better fit than `-EBUSY` for repeated initialization. I'll update the patch accordingly.
Thanks,
Tao> Regards
> Bibo Mao
>> device = &s->device;
>> /* init device by pch pic writing and reading ops */
>> kvm_iodevice_init(device, &kvm_pch_pic_ops);
>> @@ -293,8 +296,11 @@ static int kvm_pch_pic_init(struct kvm_device *dev, u64 addr)
>> /* register pch pic device */
>> ret = kvm_io_bus_register_dev(kvm, KVM_MMIO_BUS, addr, PCH_PIC_SIZE, device);
>> mutex_unlock(&kvm->slots_lock);
>> + if (ret < 0)
>> + return ret;
>> - return (ret < 0) ? -EFAULT : 0;
>> + s->pch_pic_base = addr;
>> + return 0;
>> }
>> /* used by user space to get or set pch pic registers */
>>
>
^ permalink raw reply [flat|nested] 15+ messages in thread
end of thread, other threads:[~2026-09-28 9:32 UTC | newest]
Thread overview: 15+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-27 7:52 [PATCH 0/6] LoongArch: KVM: irqchip and steal time fixes Tao Cui
2026-09-27 7:52 ` [PATCH 1/6] LoongArch: KVM: Clear device pointer in irqchip destroy callbacks Tao Cui
2026-09-28 4:07 ` Bibo Mao
2026-09-27 7:52 ` [PATCH 2/6] LoongArch: KVM: Guard against NULL irqchip in irq injection Tao Cui
2026-09-28 4:14 ` Bibo Mao
2026-09-28 9:24 ` Tao Cui
2026-09-27 7:52 ` [PATCH 3/6] LoongArch: KVM: Load dmsintc pointer once in pch_msi_set_irq Tao Cui
2026-09-27 7:52 ` [PATCH 4/6] LoongArch: KVM: Rebase steal time counter in vcpu context Tao Cui
2026-09-28 7:23 ` Bibo Mao
2026-09-28 9:28 ` Tao Cui
2026-09-27 7:52 ` [PATCH 5/6] LoongArch: KVM: Propagate real error code in kvm_pch_pic_create Tao Cui
2026-09-28 7:25 ` Bibo Mao
2026-09-27 7:52 ` [PATCH 6/6] LoongArch: KVM: Reject repeated PCH-PIC CTRL_INIT Tao Cui
2026-09-28 8:01 ` Bibo Mao
2026-09-28 9:31 ` Tao Cui
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®