Summary ======= The corruption you are seeing is consistent with a known use-after-free in the x86 change_page_attr (CPA) code that is present in 6.18.44 and was only fixed in 6.18.52 and 6.18.53. cpa_collapse_large_pages() rebuilds a leaf PMD out of its 4K PTEs and then frees the old PTE table, with no lock held against the lockless page table walk that __change_page_attr() performs before it stores through the PTE pointer it cached. The stale 8-byte store of a kernel PTE value lands in whatever the buddy allocator has since handed that page out for. On a KVM host the two sides of this race are both hot. set_memory_rox() is the only caller that passes CPA_COLLAPSE, and it runs on every module load (execmem_restore_rox()), every ftrace trampoline creation (arch/x86/kernel/ftrace.c:423) and every new 2M BPF program pack. The other side is any lockless walk of the same execmem tables: set_memory_nx()/set_memory_rw() from execmem_force_rw() on module load and trampoline allocation, and vmalloc_to_page() inside __text_poke() for every patch of module text, kprobe slot, trampoline or BPF pack. Module text, kprobe slots and ftrace trampolines share the same 2M ROX cache pages, so the collapser and the victim land in the same PMD by construction. A libvirt host does all of this constantly: module autoload (tun, vhost_net, br_netfilter, ebtable_*, xt_*), per-VM seccomp filters, systemd cgroup BPF, perf and BPF probes, ftrace and static call patching on every module load. This matches your good/bad window exactly. The collapse feature was added in v6.15 and is not in 5.15. It also matches the KASAN result. free_page_is_bad() is gated on is_check_pages_enabled(), which is off unless CONFIG_DEBUG_VM is set, and KASAN changes the allocation pattern enough that the freed page tends not to be reused in the race window. The upstream reporter only reproduced it by injecting a delay at the CPA page table lookup. Recommendation: run 6.18.53 or later. 6.18.54-rc1, which you say you are about to test, has the complete series. 6.18.52 has only the cpa_lock patch, which does not close the window. Mainline (v7.3-rc5) has nothing further pending for arch/x86/mm/pat/set_memory.c. Kernel version ============== 6.18.44lb9.01 (AlmaLinux 9 build of stable 6.18.44, stable commit 1efe5d048a39). First seen on 6.18.31, last crash on 6.18.44. 5.15.x was fine. Machine ======= ASUSTeK RS720A-E12-RS12 / K14PP-D24, AMD EPYC, BIOS 2305 11/21/2025. KVM host, AlmaLinux 9, Intel ice and Mellanox NICs. Taint "G E", unsigned module only. Stack trace =========== Oops: general protection fault, probably for non-canonical address 0xfffffff0c930038: 0000 [#1] SMP NOPTI CPU: 23 UID: 189 PID: 7538 Comm: pacemaker-contr RIP: 0010:__d_lookup+0x4a/0xc0 RAX: 0fffffff0c930020 RBX: 0fffffff0c930020 RCX: 0000000000000000 RDX: ff2e6dbe0d9b6000 RSI: ff7532a13699fe60 RDI: ff2e6d1e4e630d80 RBP: 000000000b654440 Call Trace: d_lookup+0x27/0x50 lookup_dcache+0x1f/0x80 lookup_one_qstr_excl+0x1e/0xe0 filename_create+0xc4/0x160 do_mkdirat+0x5a/0x190 __x64_sys_mkdir+0x42/0x60 do_syscall_64+0x64/0xbf0 entry_SYSCALL_64_after_hwframe+0x76/0x7e Other messages you reported: Inconsistency detected by ld.so: ../sysdeps/x86_64/dl-machine.h: 548: elf_machine_rela_relative: Assertion `ELFW(R_TYPE) (reloc->r_info) == R_X86_64_RELATIVE' failed! What the oops registers say =========================== The Code bytes decode to the hash chain walk in fs/dcache.c:__d_lookup(): mov (%rbx),%rax ; load bucket->first mov %rax,%rbx and $-2,%rbx ; strip the hlist_bl lock bit cmp $1,%rax ja body loop: mov (%rbx),%rbx ; node->next test %rbx,%rbx je out body: cmp %ebp,0x18(%rbx) ; <-- faulting insn, d_name.hash_len jne loop RAX equals RBX and RAX is only ever written by the initial bucket load, so this is the first loop iteration. The corrupt word is the hlist_bl_head.first field of the dentry_hashtable bucket itself, not a dentry's d_hash.next. RDX 0xff2e6dbe0d9b6000 is the dentry_hashtable base, and the runtime constant d_hash_shift is patched to 7, so: bucket = 0xff2e6dbe0d9b6000 + (0x0b654440 >> 7) * 8 = 0xff2e6dbe0e51b440 That table is a 256MB alloc_large_system_hash() allocation from memblock. It is allocated at boot, is PG_reserved and is never freed. So this is a stray write to a fixed physical page, not a use-after-free of a recycled object. The corrupt value 0x0fffffff0c930020 is an exact x86-64 non-present swap PTE: __swp_type() is 1 and __swp_offset() is 0x79b67f, about 30.4GiB into swap device 1. The only low bit set is bit 5, _PAGE_ACCESSED. Every bit the swap layout constrains is as it should be: P, PSE and the G/PROT_NONE alias are clear, the software bits 1-3 are clear, the type is an ordinary swap type and the inverted offset gives the long run of ones in bits 32-58. The one thing the layout does not account for is bit 5. __swp_entry() leaves bits 0-8 clear and no helper sets bit 5 on a non-present PTE; arch/x86/include/asm/pgtable_64.h treats bits 5 and 6 as don't-care only because of the Intel Knights Landing erratum (00839ee3b299, _PAGE_KNL_ERRATUM_MASK), which does not apply to EPYC. So either something other than a Linux swap PTE happens to fit this layout, or the word was a swap PTE that acquired a stray bit. I cannot tell which from one word, which is why the vmcore page dump requested below matters: 511 neighbouring PTE-shaped words would settle it. Suspect commit ============== commit 41d88484c71cd4f659348da41b7b5b3dbd3be1f6 Author: Kirill A. Shutemov x86/mm/pat: restore large ROX pages after fragmentation Link: https://lore.kernel.org/r/20250126074733.1384926-4-rppt@kernel.org This added runtime collapse of split kernel large pages, driven from cpa_flush(), including freeing the PTE table that the collapsed PMD replaces. It is the Fixes: target of every fix listed below. It is in v6.15 and later, and is not in 5.15. > diff --git a/arch/x86/mm/pat/set_memory.c b/arch/x86/mm/pat/set_memory.c > --- a/arch/x86/mm/pat/set_memory.c > +++ b/arch/x86/mm/pat/set_memory.c > @@ -394,6 +408,40 @@ static void __cpa_flush_tlb(void *data) > +static void cpa_collapse_large_pages(struct cpa_data *cpa) > +{ > + unsigned long start, addr, end; > + struct ptdesc *ptdesc, *tmp; > + LIST_HEAD(pgtables); > + int collapsed = 0; > + int i; [ ... range iteration ... ] > + if (!collapsed) > + return; > + > + flush_tlb_all(); > + > + list_for_each_entry_safe(ptdesc, tmp, &pgtables, pt_list) { > + list_del(&ptdesc->pt_list); > + __free_page(ptdesc_page(ptdesc)); > + } > +} In 6.18.44 that last loop is pagetable_free(ptdesc), which is still an immediate free. There is no RCU grace period and no other deferral. The flush_tlb_all() above it only makes the hardware forget the old translation; it does nothing about a CPU that is sitting inside __change_page_attr() holding a pointer into that table. > @@ -402,7 +450,7 @@ static void cpa_flush(struct cpa_data *cpa, int cache) > if (cache && !static_cpu_has(X86_FEATURE_CLFLUSH)) { > cpa_flush_all(cache); > - return; > + goto collapse_large_pages; > } [ ... ] > @@ -427,6 +475,10 @@ static void cpa_flush(struct cpa_data *cpa, int cache) > mb(); > + > +collapse_large_pages: > + if (cpa->flags & CPA_COLLAPSE) > + cpa_collapse_large_pages(cpa); > } The collapse is hooked into cpa_flush(), which __change_page_attr_set_clr() calls after it has already dropped cpa_lock. So cpa_lock does not serialise the collapse against anything. > @@ -1196,6 +1248,161 @@ static int split_large_page(struct cpa_data *cpa, pte_t *kpte, > +static int collapse_pmd_page(pmd_t *pmd, unsigned long addr, > + struct list_head *pgtables) > +{ [ ... uniformity checks over all 512 PTEs ... ] > + old_pmd = *pmd; > + > + /* Success: set up a large page */ > + pgprot = pgprot_4k_2_large(pte_pgprot(first)); > + pgprot_val(pgprot) |= _PAGE_PSE; > + _pmd = pfn_pmd(pfn, pgprot); > + set_pmd(pmd, _pmd); > + > + /* Queue the page table to be freed after TLB flush */ > + list_add(&page_ptdesc(pmd_page(old_pmd))->pt_list, pgtables); collapse_large_pages(), the caller, takes pgd_lock around this. The lockless CPA walker never takes pgd_lock, so pgd_lock does not help either. The other side, in 6.18.44: arch/x86/mm/pat/set_memory.c:__change_page_attr() { address = __cpa_addr(cpa, cpa->curpage); repeat: kpte = _lookup_address_cpa(cpa, address, &level, &nx, &rw); ... old_pte = *kpte; ... if (level == PG_LEVEL_4K) { ... new_pte = pfn_pte(pfn, new_prot); ... if (pte_val(old_pte) != pte_val(new_pte)) { set_pte_atomic(kpte, new_pte); <-- stale cpa->flags |= CPA_FLUSHTLB; } _lookup_address_cpa() reaches lookup_address_in_pgd_attr(), which is a plain lockless walk. Between the walk and the set_pte_atomic() the caller can be preempted or take an interrupt; this runs with interrupts on. That store is the only unsafe instruction in the function. The large-page split branch further down is safe because __split_large_page() revalidates under pgd_lock. The freed table is not even a tracked page table page in 6.18.44: arch/x86/mm/pat/set_memory.c:split_large_page() { if (!debug_pagealloc_enabled()) spin_unlock(&cpa_lock); base = alloc_pages(GFP_KERNEL, 0); A bare alloc_pages(), so it goes straight back to the per-CPU free list and can be reallocated immediately. Race timeline ============= CPU A (text_poke -> CPU B (module_enable_rox / execmem_make_temp_rw -> execmem_restore_rox / set_memory_nx/rw) bpf_jit_binary_lock_ro -> set_memory_rox, CPA_COLLAPSE) ----- ----- __change_page_attr() kpte = _lookup_address_cpa() old_pte = *kpte new_pte = pfn_pte(...) preempted / interrupted __change_page_attr_set_clr() drops cpa_lock cpa_flush() cpa_collapse_large_pages() collapse_pmd_page(): all 512 PTEs uniform, set_pmd() installs a leaf, old PTE table queued flush_tlb_all() pagetable_free() -> immediate __free_pages() (any CPU) page is reallocated: .so page cache folio, QEMU guest RAM, a user PMD/PTE table, slab set_pte_atomic(kpte, new_pte) stores a PTE-shaped word into the reallocated page Where the swap PTE comes from (inferred continuation) ------------------------------------------------------ The race above writes a present kernel PTE, never a swap entry. To reach the dentry hash table with a swap-PTE-shaped word the following has to happen next. Each step is verified in the 6.18.44 code; the sequence as a whole is inferred, not proven for this oops. 1. The freed PTE table is reallocated as a QEMU page table. 2. The stale set_pte_atomic() lands in it. The injected entry is a translation into an execmem text page (case A/B below). 3. Case B: GUP-slow follows that entry and KVM maps the text page into the guest; a later zap_present_folio_ptes() does an unbalanced folio_put() and frees the still-live text page. 4. That text page is reallocated as another page table while text_poke()/the BPF JIT keep writing instruction bytes into it through the ROX mapping. Instruction bytes are now PMD entries with arbitrary pfns; some pass pmd_bad(). 5. Reclaim: try_to_unmap_one() -> page_vma_mapped_walk() -> pte_offset_map_lock() reads such a PMD, computes __va(garbage pfn) as the PTE table, and set_pte_at() stores a swap PTE there. If that pfn is the dentry_hashtable page, one bucket head becomes 0x0fffffff0c930020-like. 6. Days later __d_lookup() hashes into that bucket and faults. Step 5 is the only writer of an ordinary swap type in the mm, and the only step that can touch memory the allocator never owned. This is not speculation about the code. The same interleaving was reported upstream with a KASAN reproducer, in the commit that first tried to address it: commit 1aac65f3e651 ("x86/mm/pat: Take cpa_lock around large-page collapse") BUG: KASAN: use-after-free in __change_page_attr+0x7cc/0x7e0 Write of size 8 at addr ffff888181139718 by task modprobe ... The buggy address belongs to the physical page: pfn:0x181139 ... page_type: f2(table) Fixes: 41d88484c71c ("x86/mm/pat: restore large ROX pages after fragmentation") Signed-off-by: Denis V. Lunev Which stable releases carry the fixes ===================================== None of these are in 6.18.44. The fixes that matter are the init_mm mmap lock pair, a1c7570cedd0 and d5d8b8662e6e: the collapse runs under the init_mm write lock and the whole attribute change, including the lockless walk and the store through the cached pointer, runs under the read lock. That excludes both walker-vs-collapse and collapse-vs-collapse. 1587d3394e25 covers the third walker: __text_poke() resolves the pages it patches with vmalloc_to_page(), a lockless walk of the same execmem tables, and now takes the init_mm read lock around it. Without that, a collapse under a concurrent text_poke() returns NULL (the BUG_ON at arch/x86/kernel/alternative.c:2576 that openSUSE hit) or, if the freed table has already been reused, a wrong page that text_poke() then writes instruction bytes into. 9e4a3ec3411b makes the split tables real kernel page tables so their freeing is deferred. The earlier cpa_lock patch, 1aac65f3e651, is superseded by the write lock and adds nothing once those are applied. 6.18.52 591b6fac9df3 x86/mm/pat: Take cpa_lock around large-page collapse (upstream 1aac65f3e651) 6.18.53 35820cf8dd52 x86/mm/pat: Acquire init_mm write lock on collapse to avoid UAF (upstream a1c7570cedd0) e21a9ea81426 x86/mm/pat: Acquire init_mm read lock on attribute changes to avoid UAF (upstream d5d8b8662e6e) e164f4a25e23 x86/mm/pat: Convert split_large_page() to use ptdescs 5029589bb773 x86/mm/pat: Don't gate cpa_lock on debug_pagealloc_enabled() 84e0cd79d57f x86/mm/pat: Allocate split page tables as kernel page tables (upstream 9e4a3ec3411b) 281e6f536f2f x86/alternatives: Exclude text poking against change_page_attr() (upstream 1587d3394e25) 74a2626044de x86/mm: Fix and document DEBUG_PAGEALLOC (upstream 7da514d819a0) 6.18.52 does not fix this: it only takes cpa_lock around the collapse, and the walker never holds cpa_lock across its walk-then-store window. The init_mm mmap lock pair that closes that window is in 6.18.53, which is the first stable release with the complete set. Use 6.18.53 or later. There is no clean interim mitigation on 6.18.44. CPA_COLLAPSE cannot be disabled by a boot parameter. How this reaches the symptoms you saw ===================================== Symptoms 1 and 2 follow directly. The stale store writes a PTE-shaped qword into a page that has been reallocated. If that page is a page cache folio for a mapped .so, eight bytes of its relocation table are replaced and ld.so trips the R_X86_64_RELATIVE assertion while the file on disk is intact. If it is an anonymous page that QEMU has just populated as guest RAM on a migration destination, the guest sees eight corrupt bytes. Both of these match "right after migration": the destination host is populating gigabytes of guest RAM and allocating page tables at maximum rate, which is exactly when a just-freed page gets reused inside the race window. Symptom 3, this oops, needs one more step, because the dentry hash table is memblock memory that is never freed and therefore cannot be the directly reallocated page. The escalation is that the victim page is itself a page table. Each of the following links is verified in the 6.18.44 source, but I want to be clear that the end-to-end chain for this particular oops is plausible rather than proven from a single vmcore. Case A, the victim is a user PMD table. pmd_bad() is true for the injected value, so the first user-mode touch faults and mm/pgtable-generic.c:___pte_offset_map() heals it via pmd_clear_bad() with a "bad pmd" pr_err. But a supervisor-side copy_to_user() is permitted through a U=0 level. A read() or recvmsg() into guest RAM, or vhost-net, or kvm_write_guest(), then has the hardware walker read a qword of module text as a PTE. If that qword happens to have P and RW set, the copied data is written to an arbitrary physical address. This route leaves no taint. Case B, the victim is a user PTE table. x86 has no pte_bad(). GUP-fast rejects a U=0 entry, but GUP-slow does not: mm/gup.c:follow_page_pte() and mm/memory.c:__vm_normal_page() check neither _PAGE_USER nor PageReserved, so a live execmem page is returned and KVM maps kernel module text into a guest. A later mm/memory.c:zap_present_folio_ptes() does folio_remove_rmap_ptes() and folio_put() on a page that was never rmapped, dropping the refcount to zero, printing "BUG: Bad page map" and releasing live ROX text into the buddy allocator while text_poke() and the BPF JIT keep writing instruction bytes into it. Recycled as a user PMD table, instruction bytes are page table entries with arbitrary pfns that pass pmd_bad(), pte_offset_map() computes __va() of an arbitrary physical page, and try_to_unmap_one() writes a genuine host swap PTE into it. That last step is what the corrupt word looks like: a real host swap PTE. The stray bit 5 is unexplained on this route too. One caveat that argues against case B on this specific host: the taint is "G E" with no "B". "BUG: Bad page map" had not fired on that machine before the crash. So either the taint-free case A route or a direct stray write carries this particular chain, or the corrupting event happened on a different boot. This limits, but does not refute, the cascade. The corruption can sit in a rarely used dentry bucket for a long time before something hashes into it, which fits the 22-day uptime on this crash and your observation that the host crashes were not correlated with migration. Secondary findings ================== These came up while looking and are worth knowing about, but none of them explains this oops. 26505e1b5b54 ("KVM: SVM: make svm_flush_tlb_gva do a full asid flush if NPT enabled") is mainline only and has not been backported to 6.18.y. It is AMD-specific guest memory corruption via the Hyper-V PV TLB flush path, so it only affects Windows guests running with hv-tlbflush=on. If any of your crashing guests are Windows, this is worth backporting separately. It is independent of the CPA race. 55ddbc2ca6d5 (upstream f7491d7c81db, "x86/mm: Fix user-space data loss with MADV_FREE and THP") is a real bug in 6.18.44. pmd_modify() drops _PAGE_DIRTY, which has been the case since v6.10, and it causes data loss for MADV_FREE THP and writable file THP. It is fixed in 6.18.53. It is data loss rather than a stray write, so it does not explain the oops, but it is another reason to move off 6.18.44. Note that this one needs THP, which you have now disabled. d053eb7e09e1 (upstream 1e75a8255f11, "iommu/amd: Wait for completion instead of returning early in iommu_completion_wait()") is already in 6.18.42 and therefore in the kernel that crashed. It is ruled out for this oops, but it may be relevant to the earlier incidents you had on 6.18.31 through 6.18.41. Theories that were eliminated ============================= KVM NPT mapping at too large a level, or with the wrong base pfn. The mapping level comes from the host page tables and the pfn from GUP; KVM cannot reach memblock memory on its own. A host mm swap or migration PTE stored through a stale page table pointer. All the store sites are bounded and the pointer provenance checks out. A missed MMU notifier invalidation. Notifier ordering on the recovery path is correct, and this cannot reach never-freed memory. NIC DMA to the wrong address. Only a teardown-time page_pool use-after-free turned up, and the iommu/amd completion-wait fix is already in 6.18.42. 105d04edbec8 (upstream 33192a26cddea, mm/huge_memory huge_zero_pfn race). Real, but not present in 6.18.44, and it cannot reach memblock memory. 0a25ee42e7d1 (upstream 3d679b7cb31f, KVM x86/mmu CMPXCHG when clearing the Accessed bit in the TDP MMU). Not in 6.18.44, and benign for this. TDP MMU in-place huge page recovery. Structurally excluded: notifier zaps take mmu_lock for write, recovery takes it for read. memblock/buddy physical aliasing. This would produce "Bad page state" reports, which you have not seen. Worth confirming from the vmcore, see below. Neither THP nor NUMA balancing is involved in the CPA race. Disabling them was a reasonable precaution but it will not stop this. If you keep seeing corruption with THP off, that is consistent with the diagnosis rather than against it. What would confirm this ======================= Log greps, across all affected hosts and all boots, not just the ones that crashed: grep -i 'Bad page map' /var/log/messages* grep -i 'bad pmd' /var/log/messages* grep -i 'bad pud' /var/log/messages* grep -i 'Bad page state' /var/log/messages* grep -i 'CPA: called for zero pte' /var/log/messages* Any of these, particularly "bad pmd", is direct evidence that a freed kernel PTE table was reused as a user page table. "CPA: called for zero pte" would be the CPA walker itself tripping over a collapsed mapping. Questions: 1. swapon --show on the host, and inside the guests. Is there a swap device with index 1 and a size of at least roughly 30.5GiB? That tells us whether the corrupt word is a host swap PTE or a guest one, which distinguishes case A from case B above. 2. From the vmcore, dump the 4K page containing 0xff2e6dbe0e51b440. If the neighbouring words are also PTE-shaped, the page was being used as a page table and the diagnosis above is confirmed. If only the one word is corrupt, it was a single stray store. 3. What are CONFIG_DEBUG_VM, init_on_alloc, init_on_free and page_poison set to in the production build versus the KASAN build? free_page_is_bad() is gated on is_check_pages_enabled(), which needs CONFIG_DEBUG_VM, so the production kernel would not report the bad free even if it happened. 4. Are any of the crashing guests Windows, and is hv-tlbflush set on them? That decides whether 26505e1b5b54 matters for you. 5. Has any corruption occurred since THP was disabled? If yes, that supports the CPA race over your THP theory. 6. When you run the 6.18.54-rc1 test, please note that 6.18.52 is only a partial fix, so if you have any results from a 6.18.52 kernel they should not be treated as a clean run.