* Re: [PATCH] KVM: selftests: arm64: Add test for cross-vCPU LPI disable race [not found] <20260922102959.41FFA1F000FF@smtp.kernel.org> @ 2026-09-22 10:53 ` Yuchao Zhang 2026-09-22 10:53 ` [PATCH v2] " Yuchao Zhang 1 sibling, 0 replies; 3+ messages in thread From: Yuchao Zhang @ 2026-09-22 10:53 UTC (permalink / raw) To: Marc Zyngier Cc: Oliver Upton, Fuad Tabba, James Morse, Suzuki K Poulose, Zenghui Yu, Catalin Marinas, Will Deacon, sashiko-reviews, kvmarm, linux-arm-kernel, linux-kernel, Yuchao Zhang Thanks for the Sashiko review. All three points are legitimate; v2 addresses them: 1. INVALL/SYNC on unmapped collection: Agreed - INVALL for collection 1 was a command error that stalls the virtual ITS queue. v2 only sends INVALL and SYNC for TARGET_VCPU_ID (the only mapped collection and the only vCPU receiving ITS commands); unmapped collections and untouched vCPUs are skipped. 2. configure_lpis() overflow via -e: Agreed. v2 bounds-checks the -e argument in main() (must fit in one 64K ITT page) and adds an explicit assertion in configure_lpis() that nr_lpis <= SZ_64K to guarantee the property table is never overrun. 3. Silently ignored KVM_SIGNAL_MSI failures: Agreed in spirit. Injection failures are expected during the brief window where the disable path has invalidated the ITS caches and the guest has not yet remapped them. v2 counts successful injections (checking KVM_SIGNAL_MSI return value > 0) and asserts at the end that at least one succeeded, so a permanently broken environment can no longer yield a false pass. Additionally, v2 re-establishes the ITS mappings after every EnableLPIs toggle. Without this, the disable path's cache invalidation kills all mappings after the first iteration and the MSI flood goes silent - now the overflow window exists on every iteration, matching the re-initialisation sequence a real guest performs on re-enable. Tested status is unchanged: no KVM-capable hardware available, so compile- and TCG-plumbing-tested only; no Tested-by. pw-bot: cr Thanks, Yuchao ^ permalink raw reply [flat|nested] 3+ messages in thread
* [PATCH v2] KVM: selftests: arm64: Add test for cross-vCPU LPI disable race [not found] <20260922102959.41FFA1F000FF@smtp.kernel.org> 2026-09-22 10:53 ` [PATCH] KVM: selftests: arm64: Add test for cross-vCPU LPI disable race Yuchao Zhang @ 2026-09-22 10:53 ` Yuchao Zhang 1 sibling, 0 replies; 3+ messages in thread From: Yuchao Zhang @ 2026-09-22 10:53 UTC (permalink / raw) To: Marc Zyngier Cc: Oliver Upton, Fuad Tabba, James Morse, Suzuki K Poulose, Zenghui Yu, Catalin Marinas, Will Deacon, sashiko-reviews, kvmarm, linux-arm-kernel, linux-kernel, Yuchao Zhang Add a selftest that validates the behavior of remote LPI disabling while the target vCPU has in-flight/overflowing LPIs. The test configures an ITS with multiple LPIs targeting vCPU 0, which receives a continuous stream of MSIs forcing its List Registers to overflow into the ap_list. Concurrently, vCPU 1 repeatedly toggles GICR_CTLR.EnableLPIs on vCPU 0's redistributor. On unpatched kernels, this race can lead to a use-after-free or host kernel panic in vgic_fold_lr_state() due to a dangling last_lr_irq pointer. With the fix in place (stopping the VM and holding a refcount on last_lr_irq), the test runs to completion without errors. Signed-off-by: Yuchao Zhang <ndaugoing@gmail.com> --- v2: - Only send INVALL and SYNC for the mapped collection / target vCPU; commands targeting unmapped collections or untouched vCPUs are unnecessary and can stall the virtual ITS queue (Sashiko review). - Re-establish the ITS mappings after every EnableLPIs toggle, since the disable path invalidates the ITS caches and the MSI flood would otherwise go silent after the first iteration. - Bounds-check the -e argument and assert nr_lpis <= SZ_64K in configure_lpis() so it cannot overflow the 64K ITT or LPI prop tables (Sashiko review). - Correctly check KVM_SIGNAL_MSI return value (>0 on success) and count successful injections, asserting at the end that at least one succeeded (Sashiko review). tools/testing/selftests/kvm/Makefile.kvm | 1 + .../selftests/kvm/arm64/vgic_lpi_disable.c | 430 ++++++++++++++++++ 2 files changed, 431 insertions(+) create mode 100644 tools/testing/selftests/kvm/arm64/vgic_lpi_disable.c diff --git a/tools/testing/selftests/kvm/Makefile.kvm b/tools/testing/selftests/kvm/Makefile.kvm index 96bab7002d39..cf0baec3f6c7 100644 --- a/tools/testing/selftests/kvm/Makefile.kvm +++ b/tools/testing/selftests/kvm/Makefile.kvm @@ -190,6 +190,7 @@ TEST_GEN_PROGS_arm64 += arm64/vcpu_width_config TEST_GEN_PROGS_arm64 += arm64/vgic_init TEST_GEN_PROGS_arm64 += arm64/vgic_irq TEST_GEN_PROGS_arm64 += arm64/vgic_lpi_stress +TEST_GEN_PROGS_arm64 += arm64/vgic_lpi_disable TEST_GEN_PROGS_arm64 += arm64/vgic_v5 TEST_GEN_PROGS_arm64 += arm64/vpmu_counter_access TEST_GEN_PROGS_arm64 += arm64/no-vgic diff --git a/tools/testing/selftests/kvm/arm64/vgic_lpi_disable.c b/tools/testing/selftests/kvm/arm64/vgic_lpi_disable.c new file mode 100644 index 000000000000..d3480369d760 --- /dev/null +++ b/tools/testing/selftests/kvm/arm64/vgic_lpi_disable.c @@ -0,0 +1,430 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * vgic_lpi_disable - Test cross-vCPU LPI disabling race condition + * + * Copyright (c) 2026 Yuchao Zhang <ndaugoing@gmail.com> + * + * This test verifies that disabling LPIs from a remote vCPU while the + * target vCPU has in-flight/overflowing LPIs does not lead to use-after-free + * or kernel panic. + */ + +#include <linux/sizes.h> +#include <pthread.h> +#include <stdatomic.h> +#include <sys/sysinfo.h> + +#include "kvm_util.h" +#include "delay.h" +#include "gic.h" +#include "gic_v3.h" +#include "gic_v3_its.h" +#include "processor.h" +#include "ucall.h" +#include "vgic.h" + +#define TEST_MEMSLOT_INDEX 1 +#define GIC_LPI_OFFSET 8192 + +#define TARGET_VCPU_ID 0 +#define DISABLER_VCPU_ID 1 + +static size_t nr_iterations = 200; +static gpa_t gpa_base; +static atomic_int lpis_injected; + +static struct kvm_vm *vm; +static struct kvm_vcpu **vcpus; +static int its_fd; + +static struct test_data { + bool request_vcpus_stop; + u32 nr_cpus; + u32 nr_devices; + u32 nr_event_ids; + + gpa_t device_table; + gpa_t collection_table; + gpa_t cmdq_base; + void *cmdq_base_va; + gpa_t itt_tables; + + gpa_t lpi_prop_table; + gpa_t lpi_pend_tables; +} test_data = { + .nr_cpus = 2, + .nr_devices = 1, + .nr_event_ids = 64, +}; + +static void guest_irq_handler(struct ex_regs *regs) +{ + u32 intid = gic_get_and_ack_irq(); + + if (intid == IAR_SPURIOUS) + return; + + GUEST_ASSERT(intid >= GIC_LPI_OFFSET); + gic_set_eoi(intid); +} + +static void guest_setup_its_mappings(void) +{ + u32 device_id, event_id, intid = GIC_LPI_OFFSET; + u32 nr_events = test_data.nr_event_ids; + u32 nr_devices = test_data.nr_devices; + + /* Map collection 0 to TARGET_VCPU_ID */ + its_send_mapc_cmd(test_data.cmdq_base_va, TARGET_VCPU_ID, TARGET_VCPU_ID, true); + + /* Map all LPIs to TARGET_VCPU_ID to force LR overflow */ + for (device_id = 0; device_id < nr_devices; device_id++) { + gpa_t itt_base = test_data.itt_tables + (device_id * SZ_64K); + + its_send_mapd_cmd(test_data.cmdq_base_va, device_id, + itt_base, SZ_64K, true); + + for (event_id = 0; event_id < nr_events; event_id++) { + its_send_mapti_cmd(test_data.cmdq_base_va, device_id, + event_id, TARGET_VCPU_ID, intid++); + } + } +} + +static void guest_invalidate_rdists(void) +{ + /* + * Only collection TARGET_VCPU_ID is mapped; INVALL for an + * unmapped collection is a command error that stalls the + * virtual ITS command queue. + */ + its_send_invall_cmd(test_data.cmdq_base_va, TARGET_VCPU_ID); +} + +static void guest_setup_gic(void) +{ + static atomic_int nr_cpus_ready; + u32 cpuid = guest_get_vcpuid(); + + gic_init(GIC_V3, test_data.nr_cpus); + gic_rdist_enable_lpis(test_data.lpi_prop_table, SZ_64K, + test_data.lpi_pend_tables + (cpuid * SZ_64K)); + + atomic_fetch_add(&nr_cpus_ready, 1); + + if (cpuid > 0) + return; + + while (atomic_load(&nr_cpus_ready) < test_data.nr_cpus) + cpu_relax(); + + its_init(test_data.collection_table, SZ_64K, + test_data.device_table, SZ_64K, + test_data.cmdq_base, SZ_64K); + + guest_setup_its_mappings(); + guest_invalidate_rdists(); + + /* SYNC to ensure ITS setup is complete */ + its_send_sync_cmd(test_data.cmdq_base_va, TARGET_VCPU_ID); +} + +static inline void *test_gicr_base_cpu(u32 cpu) +{ + return (void *)(GICR_BASE_GPA + cpu * SZ_64K * 2); +} + +static void test_gicv3_gicr_wait_for_rwp(u32 cpu) +{ + unsigned int count = 100000; + + while (readl(test_gicr_base_cpu(cpu) + GICR_CTLR) & GICR_CTLR_RWP) { + GUEST_ASSERT(count--); + udelay(10); + } +} + +static void guest_code(size_t nr_lpis) +{ + u32 cpuid = guest_get_vcpuid(); + + guest_setup_gic(); + + if (cpuid == TARGET_VCPU_ID) { + local_irq_enable(); + GUEST_SYNC(0); + + while (!READ_ONCE(test_data.request_vcpus_stop)) + cpu_relax(); + } else { + GUEST_SYNC(0); + + for (size_t i = 0; i < nr_iterations; i++) { + /* Remotely disable LPIs on target vCPU */ + writel(0, test_gicr_base_cpu(TARGET_VCPU_ID) + GICR_CTLR); + test_gicv3_gicr_wait_for_rwp(TARGET_VCPU_ID); + + for (int d = 0; d < 50; d++) + cpu_relax(); + + /* Remotely re-enable LPIs on target vCPU */ + writel(GICR_CTLR_ENABLE_LPIS, + test_gicr_base_cpu(TARGET_VCPU_ID) + GICR_CTLR); + test_gicv3_gicr_wait_for_rwp(TARGET_VCPU_ID); + + /* + * The disable path invalidates the ITS caches, so + * re-establish the mappings for the next round of + * injections. + */ + guest_setup_its_mappings(); + guest_invalidate_rdists(); + its_send_sync_cmd(test_data.cmdq_base_va, + TARGET_VCPU_ID); + } + + WRITE_ONCE(test_data.request_vcpus_stop, true); + } + + GUEST_DONE(); +} + +static void setup_memslot(void) +{ + size_t pages; + size_t sz; + + sz = (3 + test_data.nr_devices) * SZ_64K; + sz += (1 + test_data.nr_cpus) * SZ_64K; + + pages = sz / vm->page_size; + gpa_base = ((vm_compute_max_gfn(vm) + 1) * vm->page_size) - sz; + vm_userspace_mem_region_add(vm, VM_MEM_SRC_ANONYMOUS, gpa_base, + TEST_MEMSLOT_INDEX, pages, 0); +} + +#define LPI_PROP_DEFAULT_PRIO 0xa0 + +static void configure_lpis(void) +{ + size_t nr_lpis = test_data.nr_devices * test_data.nr_event_ids; + u8 *tbl = addr_gpa2hva(vm, test_data.lpi_prop_table); + size_t i; + + TEST_ASSERT(nr_lpis <= SZ_64K, + "nr_lpis (%zu) exceeds 64K prop table", nr_lpis); + + for (i = 0; i < nr_lpis; i++) { + tbl[i] = LPI_PROP_DEFAULT_PRIO | + LPI_PROP_GROUP1 | + LPI_PROP_ENABLED; + } +} + +static void setup_test_data(void) +{ + size_t pages_per_64k = vm_calc_num_guest_pages(vm->mode, SZ_64K); + u32 nr_devices = test_data.nr_devices; + u32 nr_cpus = test_data.nr_cpus; + gpa_t cmdq_base; + + test_data.device_table = vm_phy_pages_alloc(vm, pages_per_64k, + gpa_base, + TEST_MEMSLOT_INDEX); + + test_data.collection_table = vm_phy_pages_alloc(vm, pages_per_64k, + gpa_base, + TEST_MEMSLOT_INDEX); + + cmdq_base = vm_phy_pages_alloc(vm, pages_per_64k, gpa_base, + TEST_MEMSLOT_INDEX); + virt_map(vm, cmdq_base, cmdq_base, pages_per_64k); + test_data.cmdq_base = cmdq_base; + test_data.cmdq_base_va = (void *)cmdq_base; + + test_data.itt_tables = vm_phy_pages_alloc(vm, pages_per_64k * nr_devices, + gpa_base, TEST_MEMSLOT_INDEX); + + test_data.lpi_prop_table = vm_phy_pages_alloc(vm, pages_per_64k, + gpa_base, TEST_MEMSLOT_INDEX); + configure_lpis(); + + test_data.lpi_pend_tables = vm_phy_pages_alloc(vm, pages_per_64k * nr_cpus, + gpa_base, TEST_MEMSLOT_INDEX); + + sync_global_to_guest(vm, test_data); +} + +static void setup_gic(void) +{ + its_fd = vgic_its_setup(vm); +} + +static void signal_lpi(u32 device_id, u32 event_id) +{ + gpa_t db_addr = GITS_BASE_GPA + GITS_TRANSLATER; + + struct kvm_msi msi = { + .address_lo = db_addr, + .address_hi = db_addr >> 32, + .data = event_id, + .devid = device_id, + .flags = KVM_MSI_VALID_DEVID, + }; + + int ret = __vm_ioctl(vm, KVM_SIGNAL_MSI, &msi); + + /* + * KVM_SIGNAL_MSI returns 1 on success (MSI delivered to target vCPU), + * and 0 if the MSI was blocked (e.g. while LPIs are disabled or being + * remapped). Count successes so a permanently broken environment cannot + * yield a false pass. + */ + if (ret > 0) + atomic_fetch_add(&lpis_injected, 1); +} + +static pthread_barrier_t test_setup_barrier; + +static atomic_bool stop_lpi_thread; + +static void *lpi_worker_thread(void *data) +{ + u32 device_id = (size_t)data; + u32 event_id; + + pthread_barrier_wait(&test_setup_barrier); + + while (!atomic_load(&stop_lpi_thread)) { + for (event_id = 0; event_id < test_data.nr_event_ids; event_id++) + signal_lpi(device_id, event_id); + usleep(100); + } + + return NULL; +} + +static void *vcpu_worker_thread(void *data) +{ + struct kvm_vcpu *vcpu = data; + struct ucall uc; + + while (true) { + vcpu_run(vcpu); + + switch (get_ucall(vcpu, &uc)) { + case UCALL_SYNC: + pthread_barrier_wait(&test_setup_barrier); + continue; + case UCALL_DONE: + if (vcpu == vcpus[DISABLER_VCPU_ID]) { + atomic_store(&stop_lpi_thread, true); + write_guest_global(vm, test_data.request_vcpus_stop, true); + } + return NULL; + case UCALL_ABORT: + REPORT_GUEST_ASSERT(uc); + break; + default: + TEST_FAIL("Unknown ucall: %lu", uc.cmd); + } + } + + return NULL; +} + +static void run_test(void) +{ + pthread_t *vcpu_threads; + pthread_t lpi_thread; + u32 i; + + pthread_barrier_init(&test_setup_barrier, NULL, test_data.nr_cpus + 1); + + vcpu_threads = malloc(sizeof(pthread_t) * test_data.nr_cpus); + TEST_ASSERT(vcpu_threads, "Failed to allocate vcpu_threads"); + + for (i = 0; i < test_data.nr_cpus; i++) + pthread_create(&vcpu_threads[i], NULL, vcpu_worker_thread, vcpus[i]); + + pthread_create(&lpi_thread, NULL, lpi_worker_thread, (void *)(size_t)0); + + pthread_join(lpi_thread, NULL); + for (i = 0; i < test_data.nr_cpus; i++) + pthread_join(vcpu_threads[i], NULL); + + TEST_ASSERT(atomic_load(&lpis_injected) > 0, + "no LPI was ever injected; broken test environment?"); + + free(vcpu_threads); +} + +static void setup_vm(void) +{ + int i; + + vm = vm_create_with_vcpus(test_data.nr_cpus, guest_code, vcpus); + + vm_init_descriptor_tables(vm); + for (i = 0; i < test_data.nr_cpus; i++) + vcpu_init_descriptor_tables(vcpus[i]); + + vm_install_exception_handler(vm, VECTOR_IRQ_CURRENT, guest_irq_handler); + + setup_memslot(); + setup_gic(); + setup_test_data(); +} + +static void destroy_vm(void) +{ + close(its_fd); + kvm_vm_free(vm); +} + +static void help(const char *name) +{ + pr_info("Usage: %s [-i iterations] [-e event_ids]\n", name); + pr_info(" -i: number of iterations to toggle GICR_CTLR.EnableLPIs (default %lu)\n", + nr_iterations); + pr_info(" -e: number of event IDs/LPIs to inject (default %u)\n", + test_data.nr_event_ids); +} + +int main(int argc, char **argv) +{ + int opt; + + TEST_REQUIRE(kvm_supports_vgic_v3()); + + while ((opt = getopt(argc, argv, "i:e:h")) != -1) { + switch (opt) { + case 'i': + nr_iterations = atoi_positive("iterations", optarg); + break; + case 'e': + test_data.nr_event_ids = atoi_positive("event_ids", optarg); + TEST_ASSERT(test_data.nr_event_ids <= SZ_64K / 8, + "event_ids must fit in one 64K ITT page"); + break; + case 'h': + default: + help(argv[0]); + exit(opt == 'h' ? 0 : 1); + } + } + + vcpus = malloc(sizeof(struct kvm_vcpu *) * test_data.nr_cpus); + TEST_ASSERT(vcpus, "Failed to allocate vcpus array"); + + setup_vm(); + run_test(); + destroy_vm(); + + free(vcpus); + + pr_info("Completed %lu iterations of remote LPI disable successfully\n", + nr_iterations); + + return 0; +} -- 2.53.0 ^ permalink raw reply [flat|nested] 3+ messages in thread
* Re: [PATCH v3 0/1] KVM: arm64: vgic: Drop last_lr_irq and serialize overflow EOI replay @ 2026-09-20 23:47 Marc Zyngier 2026-09-22 10:16 ` [PATCH] KVM: selftests: arm64: Add test for cross-vCPU LPI disable race Yuchao Zhang 0 siblings, 1 reply; 3+ messages in thread From: Marc Zyngier @ 2026-09-20 23:47 UTC (permalink / raw) To: Yuchao Zhang Cc: Oliver Upton, Fuad Tabba, James Morse, Suzuki K Poulose, Zenghui Yu, Catalin Marinas, Will Deacon, kvmarm, linux-arm-kernel, linux-kernel On Sun, 20 Sep 2026 02:49:10 +0100, Yuchao Zhang <ndaugoing@gmail.com> wrote: > > Hi Marc, Oliver, Fuad, and KVM/arm64 maintainers, > > Following up on the discussion around the remote LPI disable vs LR fold > race [1], this series addresses the issue at its root: the last_lr_irq > cursor introduced in commit 6da5e537f5af ("KVM: arm64: vgic: Pick EOIcount > deactivations from AP-list tail"). > > Problem: > vgic_v3_fold_lr_state() / vgic_v2_fold_lr_state() walk the overflow tail > of the ap_list starting from *host_data_ptr(last_lr_irq) without holding > ap_list_lock. Caching this raw pointer across the entire guest execution > leaves it vulnerable to concurrent modification: when a remote vCPU > disables LPIs via GICR_CTLR, vgic_flush_pending_lpis() unlinks the node > with list_del() and drops its reference, leaving last_lr_irq pointing to > a poisoned or freed object. When the vCPU exits, the walk dereferences > corrupted memory, causing a kernel panic or UAF. Why can't this be solved by simply taking a reference on the object pointed to by last_lr_irq? > > Oliver and Marc suggested taking ap_list_lock in > vgic_v3_fold_lr_state() [2][3]. I tried that approach first, but it > runs into the following lock-order problems: > 1. kvm_notify_acked_irq() grabs regular spinlocks and can re-enter > vgic_queue_irq_unlock() (which takes ap_list_lock), causing deadlock. > 2. vgic_put_irq() is a no-op for SPI/PPI, but for LPIs it calls > refcount_dec_and_lock_irqsave() which acquires dist->lpi_xa.xa_lock > when dropping the last reference. That lock sits above ap_list_lock > in the lock ordering, so calling vgic_put_irq() under ap_list_lock > causes lock inversion. Which is why we have vgic_put_irq_norelease() and vgic_release_deleted_lpis(), which allow deferring the release until we're in a suitable context. But that's beside the point. > Addressing these under a global fold lock requires deferring all EOI'ed > SPI notifications to a stack bitmap and deferring LPI releases with Which stack bitmap? > vgic_put_irq_norelease(), penalizing the fast path for all exits even > though folding hardware LRs does not touch ap_list at all. It also still > requires pinning and clearing the per-CPU last_lr_irq pointer. An uncontended atomic access is hardly an overhead, is it? Where is the overhead? And I don't understand what you're saying about the LRs not affecting the ap_list... They *always* do. > > Changes since v2: > - Replaced the skip-unlink approach of v2 with dropping the last_lr_irq > cursor entirely and serializing only the overflow EOI replay under > ap_list_lock (per Oliver and Marc's suggestion [2][3]). > > The cleaner approach in this patch: > 1. Drop the fragile last_lr_irq per-CPU cursor entirely. > 2. The common fast path (folding hardware LRs) runs natively without > ap_list_lock. We record the INTIDs of the used LRs in a small stack > array (VGIC_V3_MAX_LRS / VGIC_V2_MAX_LRS entries). Why is v2 even under consideration? LPIs are strictly v3 (ignoring v5 here), and non-LPIs are statically allocated, meaning they can't vanish under your feet. > 3. If eoicount == 0 (the vast majority of guest exits), clear > cpuif->used_lrs = 0 and return immediately without taking ap_list_lock. > 4. If unlikely(eoicount > 0), acquire ap_list_lock only to scan the ap_list > and pin (via vgic_get_irq_ref) up to eoicount active interrupts that > were not in hardware LRs. The scan is a linear walk over at most 16/64 > LR INTIDs per candidate, on the rare eoicount > 0 path - bounded and What makes you think this is acceptable? It really isn't. The point is that this is not limited to 16 entries. That's the whole point of EOIcount, which spans up to 31 simultaneously active priorities. I have no idea what you describe is achieving, TBH. And looking at the patch, I see a quadratic behaviour, which doesn't strike me as low overhead... > acceptable. lr_intids[] only records INTIDs from the used_lrs range; > the extraction mask mirrors vgic_fold_lr() exactly > (ICH_LR_VIRTUAL_ID_MASK for GICv3, GICH_LR_VIRTUALID for GICv2), > so no stale or invalid slot can produce a false match. > 5. Drop ap_list_lock immediately, and then replay their deactivations > outside the lock, naturally eliminating both eventfd re-entrancy and > lpi_xa lock inversions without changing any function signatures. > vgic_fold_lr() has no error path, so cpuif->used_lrs = 0 is always > reached after a complete fold, with no risk of partial cleanup. > > Note on EOIcount hardware limits: > ICH_HCR_EL2.EOIcount (GICv3) and GICH_HCR.EOICount (GICv2) are both > 5-bit fields, giving a maximum value of 31. The targets[32] stack array > and min_t(u32, eoicount, ARRAY_SIZE(targets)) bound together ensure no > overflow even if hardware writes an unexpected value. > > Note on EOIcount source: > For GICv3, eoicount is read from cpuif->vgic_hcr, which is populated by > __vgic_v3_save_state right before ICH_HCR_EL2 is cleared in hardware. > For GICv2, vgic_v2_save_state reads GICH_HCR via MMIO into the same > cpuif->vgic_hcr field (when LRENPIE is set) before writing 0 to GICH_HCR. > In both cases the software copy is the only valid source; reading the > hardware register after save would return 0. > > Note on scope: > This series fixes the use-after-free in the ap_list traversal caused by > last_lr_irq. It does not address the separate concern raised by Oliver > in [2] about a pending LPI still sitting in an LR when RWP=0 becomes > visible to another vCPU; that may require a stronger approach (e.g. > halting the VM) and I am happy to follow up separately. This has all the hallmarks of an AI gone wild. The problem is correctly described (last_lr_irq doesn't hold a reference on the irq), but the proposed solution is completely ignoring it, and implements... something else. I've hacked something at [1], which: - changes the behaviour of last_lr_irq to only be non-NULL when the LRs are full. - take a reference on the irq flagged as last_lr_irq, and drop this reference in vgic_prune_ap_list(), contributing to the LPIs being freed once the ap_list_lock is dropped. - stop the world when a vcpu disable LPIs. We could do slightly better, but it isn't worth the hassle for something that *never* happens. I only boot tested a small nested guest (L1 + L2) with lockdep on my laptop, and nothing caught fire. Please give it a go. You also seem to have a reproducer for this, it'd be good if you could turn it into a selftest. Thanks, M. [1] https://web.git.kernel.org/pub/scm/linux/kernel/git/maz/arm-platforms.git/log/?h=kvm-arm64/vgic-last_lr_irq-fixes -- Jazz isn't dead. It just smells funny. ^ permalink raw reply [flat|nested] 3+ messages in thread
* [PATCH] KVM: selftests: arm64: Add test for cross-vCPU LPI disable race 2026-09-20 23:47 [PATCH v3 0/1] KVM: arm64: vgic: Drop last_lr_irq and serialize overflow EOI replay Marc Zyngier @ 2026-09-22 10:16 ` Yuchao Zhang 0 siblings, 0 replies; 3+ messages in thread From: Yuchao Zhang @ 2026-09-22 10:16 UTC (permalink / raw) To: Marc Zyngier Cc: Oliver Upton, Fuad Tabba, James Morse, Suzuki K Poulose, Zenghui Yu, Catalin Marinas, Will Deacon, kvmarm, linux-arm-kernel, linux-kernel, Yuchao Zhang Add a selftest that validates the behavior of remote LPI disabling while the target vCPU has in-flight/overflowing LPIs. The test configures an ITS with multiple LPIs targeting vCPU 0, which receives a continuous stream of MSIs forcing its List Registers to overflow into the ap_list. Concurrently, vCPU 1 repeatedly toggles GICR_CTLR.EnableLPIs on vCPU 0's redistributor. On unpatched kernels, this race can lead to a use-after-free or host kernel panic in vgic_fold_lr_state() due to a dangling last_lr_irq pointer. With the fix in place (stopping the VM and holding a refcount on last_lr_irq), the test runs to completion without errors. Signed-off-by: Yuchao Zhang <ndaugoing@gmail.com> --- tools/testing/selftests/kvm/Makefile.kvm | 1 + .../selftests/kvm/arm64/vgic_lpi_disable.c | 401 ++++++++++++++++++ 2 files changed, 402 insertions(+) create mode 100644 tools/testing/selftests/kvm/arm64/vgic_lpi_disable.c diff --git a/tools/testing/selftests/kvm/Makefile.kvm b/tools/testing/selftests/kvm/Makefile.kvm index 96bab7002d39..cf0baec3f6c7 100644 --- a/tools/testing/selftests/kvm/Makefile.kvm +++ b/tools/testing/selftests/kvm/Makefile.kvm @@ -190,6 +190,7 @@ TEST_GEN_PROGS_arm64 += arm64/vcpu_width_config TEST_GEN_PROGS_arm64 += arm64/vgic_init TEST_GEN_PROGS_arm64 += arm64/vgic_irq TEST_GEN_PROGS_arm64 += arm64/vgic_lpi_stress +TEST_GEN_PROGS_arm64 += arm64/vgic_lpi_disable TEST_GEN_PROGS_arm64 += arm64/vgic_v5 TEST_GEN_PROGS_arm64 += arm64/vpmu_counter_access TEST_GEN_PROGS_arm64 += arm64/no-vgic diff --git a/tools/testing/selftests/kvm/arm64/vgic_lpi_disable.c b/tools/testing/selftests/kvm/arm64/vgic_lpi_disable.c new file mode 100644 index 000000000000..078134919228 --- /dev/null +++ b/tools/testing/selftests/kvm/arm64/vgic_lpi_disable.c @@ -0,0 +1,401 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * vgic_lpi_disable - Test cross-vCPU LPI disabling race condition + * + * Copyright (c) 2026 Yuchao Zhang <ndaugoing@gmail.com> + * + * This test verifies that disabling LPIs from a remote vCPU while the + * target vCPU has in-flight/overflowing LPIs does not lead to use-after-free + * or kernel panic. + */ + +#include <linux/sizes.h> +#include <pthread.h> +#include <stdatomic.h> +#include <sys/sysinfo.h> + +#include "kvm_util.h" +#include "delay.h" +#include "gic.h" +#include "gic_v3.h" +#include "gic_v3_its.h" +#include "processor.h" +#include "ucall.h" +#include "vgic.h" + +#define TEST_MEMSLOT_INDEX 1 +#define GIC_LPI_OFFSET 8192 + +#define TARGET_VCPU_ID 0 +#define DISABLER_VCPU_ID 1 + +static size_t nr_iterations = 200; +static gpa_t gpa_base; + +static struct kvm_vm *vm; +static struct kvm_vcpu **vcpus; +static int its_fd; + +static struct test_data { + bool request_vcpus_stop; + u32 nr_cpus; + u32 nr_devices; + u32 nr_event_ids; + + gpa_t device_table; + gpa_t collection_table; + gpa_t cmdq_base; + void *cmdq_base_va; + gpa_t itt_tables; + + gpa_t lpi_prop_table; + gpa_t lpi_pend_tables; +} test_data = { + .nr_cpus = 2, + .nr_devices = 1, + .nr_event_ids = 64, +}; + +static void guest_irq_handler(struct ex_regs *regs) +{ + u32 intid = gic_get_and_ack_irq(); + + if (intid == IAR_SPURIOUS) + return; + + GUEST_ASSERT(intid >= GIC_LPI_OFFSET); + gic_set_eoi(intid); +} + +static void guest_setup_its_mappings(void) +{ + u32 device_id, event_id, intid = GIC_LPI_OFFSET; + u32 nr_events = test_data.nr_event_ids; + u32 nr_devices = test_data.nr_devices; + + /* Map collection 0 to TARGET_VCPU_ID */ + its_send_mapc_cmd(test_data.cmdq_base_va, TARGET_VCPU_ID, TARGET_VCPU_ID, true); + + /* Map all LPIs to TARGET_VCPU_ID to force LR overflow */ + for (device_id = 0; device_id < nr_devices; device_id++) { + gpa_t itt_base = test_data.itt_tables + (device_id * SZ_64K); + + its_send_mapd_cmd(test_data.cmdq_base_va, device_id, + itt_base, SZ_64K, true); + + for (event_id = 0; event_id < nr_events; event_id++) { + its_send_mapti_cmd(test_data.cmdq_base_va, device_id, + event_id, TARGET_VCPU_ID, intid++); + } + } +} + +static void guest_invalidate_all_rdists(void) +{ + int i; + + for (i = 0; i < test_data.nr_cpus; i++) + its_send_invall_cmd(test_data.cmdq_base_va, i); +} + +static void guest_setup_gic(void) +{ + static atomic_int nr_cpus_ready; + u32 cpuid = guest_get_vcpuid(); + + gic_init(GIC_V3, test_data.nr_cpus); + gic_rdist_enable_lpis(test_data.lpi_prop_table, SZ_64K, + test_data.lpi_pend_tables + (cpuid * SZ_64K)); + + atomic_fetch_add(&nr_cpus_ready, 1); + + if (cpuid > 0) + return; + + while (atomic_load(&nr_cpus_ready) < test_data.nr_cpus) + cpu_relax(); + + its_init(test_data.collection_table, SZ_64K, + test_data.device_table, SZ_64K, + test_data.cmdq_base, SZ_64K); + + guest_setup_its_mappings(); + guest_invalidate_all_rdists(); + + /* SYNC to ensure ITS setup is complete */ + for (cpuid = 0; cpuid < test_data.nr_cpus; cpuid++) + its_send_sync_cmd(test_data.cmdq_base_va, cpuid); +} + +static inline void *test_gicr_base_cpu(u32 cpu) +{ + return (void *)(GICR_BASE_GPA + cpu * SZ_64K * 2); +} + +static void test_gicv3_gicr_wait_for_rwp(u32 cpu) +{ + unsigned int count = 100000; + + while (readl(test_gicr_base_cpu(cpu) + GICR_CTLR) & GICR_CTLR_RWP) { + GUEST_ASSERT(count--); + udelay(10); + } +} + +static void guest_code(size_t nr_lpis) +{ + u32 cpuid = guest_get_vcpuid(); + + guest_setup_gic(); + + if (cpuid == TARGET_VCPU_ID) { + local_irq_enable(); + GUEST_SYNC(0); + + while (!READ_ONCE(test_data.request_vcpus_stop)) + cpu_relax(); + } else { + GUEST_SYNC(0); + + for (size_t i = 0; i < nr_iterations; i++) { + /* Remotely disable LPIs on target vCPU */ + writel(0, test_gicr_base_cpu(TARGET_VCPU_ID) + GICR_CTLR); + test_gicv3_gicr_wait_for_rwp(TARGET_VCPU_ID); + + for (int d = 0; d < 50; d++) + cpu_relax(); + + /* Remotely re-enable LPIs on target vCPU */ + writel(GICR_CTLR_ENABLE_LPIS, + test_gicr_base_cpu(TARGET_VCPU_ID) + GICR_CTLR); + test_gicv3_gicr_wait_for_rwp(TARGET_VCPU_ID); + } + + WRITE_ONCE(test_data.request_vcpus_stop, true); + } + + GUEST_DONE(); +} + +static void setup_memslot(void) +{ + size_t pages; + size_t sz; + + sz = (3 + test_data.nr_devices) * SZ_64K; + sz += (1 + test_data.nr_cpus) * SZ_64K; + + pages = sz / vm->page_size; + gpa_base = ((vm_compute_max_gfn(vm) + 1) * vm->page_size) - sz; + vm_userspace_mem_region_add(vm, VM_MEM_SRC_ANONYMOUS, gpa_base, + TEST_MEMSLOT_INDEX, pages, 0); +} + +#define LPI_PROP_DEFAULT_PRIO 0xa0 + +static void configure_lpis(void) +{ + size_t nr_lpis = test_data.nr_devices * test_data.nr_event_ids; + u8 *tbl = addr_gpa2hva(vm, test_data.lpi_prop_table); + size_t i; + + for (i = 0; i < nr_lpis; i++) { + tbl[i] = LPI_PROP_DEFAULT_PRIO | + LPI_PROP_GROUP1 | + LPI_PROP_ENABLED; + } +} + +static void setup_test_data(void) +{ + size_t pages_per_64k = vm_calc_num_guest_pages(vm->mode, SZ_64K); + u32 nr_devices = test_data.nr_devices; + u32 nr_cpus = test_data.nr_cpus; + gpa_t cmdq_base; + + test_data.device_table = vm_phy_pages_alloc(vm, pages_per_64k, + gpa_base, + TEST_MEMSLOT_INDEX); + + test_data.collection_table = vm_phy_pages_alloc(vm, pages_per_64k, + gpa_base, + TEST_MEMSLOT_INDEX); + + cmdq_base = vm_phy_pages_alloc(vm, pages_per_64k, gpa_base, + TEST_MEMSLOT_INDEX); + virt_map(vm, cmdq_base, cmdq_base, pages_per_64k); + test_data.cmdq_base = cmdq_base; + test_data.cmdq_base_va = (void *)cmdq_base; + + test_data.itt_tables = vm_phy_pages_alloc(vm, pages_per_64k * nr_devices, + gpa_base, TEST_MEMSLOT_INDEX); + + test_data.lpi_prop_table = vm_phy_pages_alloc(vm, pages_per_64k, + gpa_base, TEST_MEMSLOT_INDEX); + configure_lpis(); + + test_data.lpi_pend_tables = vm_phy_pages_alloc(vm, pages_per_64k * nr_cpus, + gpa_base, TEST_MEMSLOT_INDEX); + + sync_global_to_guest(vm, test_data); +} + +static void setup_gic(void) +{ + its_fd = vgic_its_setup(vm); +} + +static void signal_lpi(u32 device_id, u32 event_id) +{ + gpa_t db_addr = GITS_BASE_GPA + GITS_TRANSLATER; + + struct kvm_msi msi = { + .address_lo = db_addr, + .address_hi = db_addr >> 32, + .data = event_id, + .devid = device_id, + .flags = KVM_MSI_VALID_DEVID, + }; + + __vm_ioctl(vm, KVM_SIGNAL_MSI, &msi); +} + +static pthread_barrier_t test_setup_barrier; + +static atomic_bool stop_lpi_thread; + +static void *lpi_worker_thread(void *data) +{ + u32 device_id = (size_t)data; + u32 event_id; + + pthread_barrier_wait(&test_setup_barrier); + + while (!atomic_load(&stop_lpi_thread)) { + for (event_id = 0; event_id < test_data.nr_event_ids; event_id++) + signal_lpi(device_id, event_id); + usleep(100); + } + + return NULL; +} + +static void *vcpu_worker_thread(void *data) +{ + struct kvm_vcpu *vcpu = data; + struct ucall uc; + + while (true) { + vcpu_run(vcpu); + + switch (get_ucall(vcpu, &uc)) { + case UCALL_SYNC: + pthread_barrier_wait(&test_setup_barrier); + continue; + case UCALL_DONE: + if (vcpu == vcpus[DISABLER_VCPU_ID]) { + atomic_store(&stop_lpi_thread, true); + write_guest_global(vm, test_data.request_vcpus_stop, true); + } + return NULL; + case UCALL_ABORT: + REPORT_GUEST_ASSERT(uc); + break; + default: + TEST_FAIL("Unknown ucall: %lu", uc.cmd); + } + } + + return NULL; +} + +static void run_test(void) +{ + pthread_t *vcpu_threads; + pthread_t lpi_thread; + u32 i; + + pthread_barrier_init(&test_setup_barrier, NULL, test_data.nr_cpus + 1); + + vcpu_threads = malloc(sizeof(pthread_t) * test_data.nr_cpus); + TEST_ASSERT(vcpu_threads, "Failed to allocate vcpu_threads"); + + for (i = 0; i < test_data.nr_cpus; i++) + pthread_create(&vcpu_threads[i], NULL, vcpu_worker_thread, vcpus[i]); + + pthread_create(&lpi_thread, NULL, lpi_worker_thread, (void *)(size_t)0); + + pthread_join(lpi_thread, NULL); + for (i = 0; i < test_data.nr_cpus; i++) + pthread_join(vcpu_threads[i], NULL); + + free(vcpu_threads); +} + +static void setup_vm(void) +{ + int i; + + vm = vm_create_with_vcpus(test_data.nr_cpus, guest_code, vcpus); + + vm_init_descriptor_tables(vm); + for (i = 0; i < test_data.nr_cpus; i++) + vcpu_init_descriptor_tables(vcpus[i]); + + vm_install_exception_handler(vm, VECTOR_IRQ_CURRENT, guest_irq_handler); + + setup_memslot(); + setup_gic(); + setup_test_data(); +} + +static void destroy_vm(void) +{ + close(its_fd); + kvm_vm_free(vm); +} + +static void help(const char *name) +{ + pr_info("Usage: %s [-i iterations] [-e event_ids]\n", name); + pr_info(" -i: number of iterations to toggle GICR_CTLR.EnableLPIs (default %lu)\n", + nr_iterations); + pr_info(" -e: number of event IDs/LPIs to inject (default %u)\n", + test_data.nr_event_ids); +} + +int main(int argc, char **argv) +{ + int opt; + + TEST_REQUIRE(kvm_supports_vgic_v3()); + + while ((opt = getopt(argc, argv, "i:e:h")) != -1) { + switch (opt) { + case 'i': + nr_iterations = atoi_positive("iterations", optarg); + break; + case 'e': + test_data.nr_event_ids = atoi_positive("event_ids", optarg); + break; + case 'h': + default: + help(argv[0]); + exit(opt == 'h' ? 0 : 1); + } + } + + vcpus = malloc(sizeof(struct kvm_vcpu *) * test_data.nr_cpus); + TEST_ASSERT(vcpus, "Failed to allocate vcpus array"); + + setup_vm(); + run_test(); + destroy_vm(); + + free(vcpus); + + pr_info("Completed %lu iterations of remote LPI disable successfully\n", + nr_iterations); + + return 0; +} -- 2.53.0 ^ permalink raw reply [flat|nested] 3+ messages in thread
end of thread, other threads:[~2026-09-22 10:53 UTC | newest]
Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
[not found] <20260922102959.41FFA1F000FF@smtp.kernel.org>
2026-09-22 10:53 ` [PATCH] KVM: selftests: arm64: Add test for cross-vCPU LPI disable race Yuchao Zhang
2026-09-22 10:53 ` [PATCH v2] " Yuchao Zhang
2026-09-20 23:47 [PATCH v3 0/1] KVM: arm64: vgic: Drop last_lr_irq and serialize overflow EOI replay Marc Zyngier
2026-09-22 10:16 ` [PATCH] KVM: selftests: arm64: Add test for cross-vCPU LPI disable race Yuchao Zhang
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®