+cc Boris, Tal, Rik, Matt FYI - seems another instance of the INVLPGB bug. On Thu, Oct 01, 2026 at 09:40:58PM +0200, Nikola Ciprich wrote: > Hello again, > > the good (?) news is, in the meantime we got another crash on different machine > and I have a kdump including complete vmcore. This one was 6.18.53 > > here are some details: > > Hardware: Supermicro AS-2024US-TRT / H12DSU-iN, BIOS 3.5 (2025-09-22). Dual-socket AMD EPYC 7343 16-Core, 32 CPUs, 1024 GB RAM. > > analyzed with crash + matching vmlinux debuginfo: Thanks that's useful. We can rule out CPA at this point. > > Oops: > general protection fault, probably for non-canonical address 0xfffffff0c930038 > RIP: __d_lookup+0x4a/0xc0 > Comm: systemd PID: 1274600 CPU: 23 > Call trace: > __d_lookup > lookup_fast > walk_component > link_path_walk > path_openat > do_filp_open > do_sys_openat2 > __x64_sys_openat > do_syscall_64 > > Exception frame registers: > RAX: 0fffffff0c930020 RBX: 0fffffff0c930020 RCX: 000000000000000b > RDX: ffff985ffd223600 RSI: ffffb18962823d70 RDI: ffff989ddf09b5c0 > RBP: 000000005560450b R8: 000000007fffffff R9: fefefefefefefeff > R10: 0000000000000000 R11: 93c3d2eb02a31dda R12: ffff989ddf09b5c0 > R13: ffff989ddf09b5c0 R14: ffffb18962823d70 R15: 0000000000000000 > > Faulting instruction (cmp %ebp,0x18(%rbx)) dereferences RBX+0x18. RBX held > the non-canonical value 0x0fffffff0c930020, giving fault address > 0x0fffffff0c930038. __d_lookup was walking the d_hash (hlist_bl) bucket; > RBX was the node pointer being dereferenced. > > Observations from the vmcore: > > The faulting value 0x0fffffff0c930020 is non-canonical and is not a mapped > kernel address: > crash> kmem 0x0fffffff0c930020 > kmem: cannot determine page for fffffff0c930020 > fffffff0c930020: physical address not found in mem map > > The dentry being looked up (RDI/R12/R13 = 0xffff989ddf09b5c0) is intact and > well-formed: > name "app.slice", len 9, d_name.hash 0xCE973022 (consistent) > d_op = kernfs_dops; valid d_parent, d_inode, d_sb > d_hash.next = 0x0 (this node is the end of its bucket chain) > > The target dentry and its hash chain in the dump show no corruption; the > chain terminates cleanly. > > No page migration, compaction, or THP activity was in progress on any CPU > at panic. "bt -a" filtered for migrate*/compact*/khugepaged/kcompactd/ > kswapd/split_huge*/folio*/d_move/rename returned nothing. > > Automatic NUMA balancing was disabled at crash time (read from kernel memory): > crash> p sysctl_numa_balancing_mode > $ = 0 > > Top-level (PMD) transparent hugepage policy was "never" at crash time: > crash> p/x transparent_hugepage_flags > $ = 0x1c0 > Bits set: 6 (DEFRAG_REQ_MADV), 7 (DEFRAG_KHUGEPAGED), 8 (USE_ZERO_PAGE). > Bits 0 (TRANSPARENT_HUGEPAGE_FLAG) and 1 (REQ_MADV_FLAG) are clear, i.e. > sysfs enabled = never. Per-order mTHP controls (huge_anon_orders_*) were not > inspected for this dump, so mTHP state is not asserted here. > > No MCE/EDAC/hardware-error records are present in the kernel log for this host. > > I can provide the full vmcore and the matching vmlinux/debuginfo on request, and > run further crash queries against it. > > not sure if this is of any help? Very useful. It looks like AMD TLB invalidation issues are the leading likely cause here then - one on the host side with INVLPGB, and a separate one in KVM with INVLPGA. And it looks like a hardware bug, unfortunately. Support for this was merged in 6.15 which matches your kernel versions too. It's actually not solved yet, but there are two workarounds that can be applied here. ## Issue 1: INVLPGB See [0] for a report of the same kind of thing (segfaults like yours), and [1] for the proposed temporary workaround. TL;DR: The mitigation for this is to update your kernel command line parameters thusly: 6.18.45 and later (6.18.53, 6.18.54): tlbi=ipi <6.18.45: clearcpuid=419 If this resolves it then it confirms that this is the issue. There is also, usefully, a reproducer which should show corruption in minutes rather than weeks, see: https://gist.github.com/mfleming/ca26ad3f8d65a12d23d62fb176480fc1 https://lore.kernel.org/all/akzuETUv5XEFSRL6@matt-Precision-5490/ The fact you've seen this on a Milan machine (EPYC 7343) is new information so that could be useful for the report, so if you can reproduce _without the fix_ first that'd be very useful to know! Also then try the fix and see if it reproduces afterwards. If the reproducer doesn't work then I guess worth waiting to see if hosts reproduce over a longer time period with the tlbi=ipi issue. Also, which CPU is in the first crash host (the ASUS SP5 box)? That'd be useful to know thanks! ## Issue 2: INVLPGA This is a separate issue, as mentioned before - Red Hat have seen windows guest memory corruption on AMD with hv-tlbflush which tlbi=ipi won't touch. I had the AI generate a patch for you that applies to 6.18.53 and even 6.18.44 to make your life easier :) The patch is attached. For a kernel tree with git: git am 0001-KVM-SVM-flush_tlb_gva-full-asid-flush-6.18.y.patch Or just plain raw code: patch -p1 < 0001-KVM-SVM-flush_tlb_gva-full-asid-flush-6.18.y.patch So please test issue 1 separately with the cmdline change, but if you're still seeing issues especially with windows guests, then this KVM patch should then ALSO be applied. This is [2] which is merged upstream as commit 26505e1b5b54 ("KVM: SVM: make svm_flush_tlb_gva do a full asid flush if NPT enabled"). ## Crash + vmcore Obviously do the above and especially try the reproducer in the lab! :) But also with the vmcore, can you run these 4 commands and reply with the output please? crash> p dentry_hashtable crash> rd -64 0xffff986002783000 512 crash> search -p 0x0fffffff0c930020 crash> vtop 0xffff986002783a50 crash> search -p -m 0xfff0000000000fff ## Mitigations for production I suggest you apply both of the above for your production as it's _likely_ it will resolve the issue there. The cmdline changes are perfectly safe but you should check the generated patch, however! I suspect it'll be fine but I can't guarantee it obviously LLMs etc. (albeit it is a frontier model so at least as good as it can be ;) ## Attachments I attach the backported fix as discussed above and also the AI's full debug report FYI. -- Cheers, Lorenzo [0]:https://lore.kernel.org/all/CAAuFnRTyva1_3tsF3vrMBL+TLS1YL4EgUPT2c3O9k7A9hWUMnA@mail.gmail.com/ [1]:https://lore.kernel.org/all/20260914045823.GFaqd-7yDf6aK4JUQr@fat_crate.local/ [2]:https://lore.kernel.org/all/20260723094419.630204-1-pbonzini@redhat.com/