* hunting memory corruption bug in 6.18.x
@ 2026-09-25 8:48 Nikola Ciprich
2026-09-25 10:05 ` Lorenzo Stoakes (ARM)
` (2 more replies)
0 siblings, 3 replies; 16+ messages in thread
From: Nikola Ciprich @ 2026-09-25 8:48 UTC (permalink / raw)
To: linux-mm; +Cc: linux-kernel, akpm, david, ljs, nikola.ciprich
Hi,
I've been hunting a weird memory corruption bug for the last few weeks,
without success so far, so I'd like to report it and kindly ask for help.
We first hit it after a live VM migration between two KVM hosts:
suddenly some dynamic libraries in the host appeared to be corrupted:
Inconsistency detected by ld.so: ../sysdeps/x86_64/dl-machine.h: 548: elf_machine_rela_relative: Assertion `ELFW(R_TYPE) (reloc->r_info) == R_X86_64_RELATIVE' failed!
(Later we also hit this with libcrypto.so.3 etc.) The files on disk
were OK; the problem seemed to exist only in RAM.
I'm fairly sure this is not hardware related: there were no ECC errors,
and we have since hit this (and similar issues, more on that below) on
multiple machines.
The problems started after we moved from 5.15.x to 6.18.x kernels.
Since then I've spent a lot of time trying to reproduce it on a lab
cluster, and we were able to trigger some corruption after days of
migrating VMs back and forth. At first I suspected the Intel ice driver,
for which I found similar reports, but we saw new problems even after
backporting fixes (and also with Mellanox cards).
So far we've hit three different kinds of problems, which may or may
not be related:
- .so library corruption right after VM migration
- VM crashes (or process crashes inside VMs), possibly related to
migration (those always happened during migration)
- host crashes due to kernel structure corruption (these happened
without any VM migration)
We first hit these problems with 6.18.31; the last crash I saw was
with 6.18.44.
All affected machines use AMD EPYC CPUs and act as KVM hosts; the OS
is AlmaLinux 9.
I suspect two subsystems that have seen a lot of changes:
- transparent hugepages
- NUMA balancing
(but those are just my guesses)
As a safety measure, we've disabled THP and NUMA balancing on all hosts.
I'm aware this is still a very vague report with a lot of guessing,
but my question is: has anybody hit similar problems with 6.18 or
newer kernels?
I see a lot of patches in every stable release, but simply trying
newer kernels doesn't seem efficient here. Deploying them is also
risky, since the hosts have to be emptied by migrating VMs off them
before reboot, and that migration itself may trigger more crashes.
None of the released or queued fixes for 6.18 seem to be directly
related.
I tried running my migration tests on hosts with KASAN enabled, and
also with SLUB debugging, but was never able to reproduce the problem
with those enabled (without them, I was able to hit issues within
days).
I'll start another round of migration tests in the lab, now with
6.18.54-rc1, but I still thought it would be good to report this and
ask here.
last but not least, here's kdump from last crash (this was not related
to any VM migration, but is very similar to another few crashes
we got):
[1924553.414736] Oops: general protection fault, probably for non-canonical address 0xfffffff0c930038: 0000 [#1] SMP NOPTI
[1924553.434800] CPU: 23 UID: 189 PID: 7538 Comm: pacemaker-contr Kdump: loaded Tainted: G E 6.18.44lb9.01 #1 PREEMPT(voluntary)
[1924553.456934] Tainted: [E]=UNSIGNED_MODULE
[1924553.465551] Hardware name: ASUSTeK COMPUTER INC. RS720A-E12-RS12/K14PP-D24 Series, BIOS 2305 11/21/2025
[1924553.484152] RIP: 0010:__d_lookup+0x4a/0xc0
[1924553.492878] Code: ff 48 89 c5 c1 e8 07 48 8d 1c c2 e8 60 8f d1 ff 48 8b 03 48 89 c3 48 83 e3 fe 48 83 f8 01 77 0a eb 2f 48 8b 1b 48 85 db 74 27 <39> 6b 18 75 f3 4c 8d 63 78 4c 89 e7 e8
d5 e1 7c 00 4c 39 6b 10 74
[1924553.525191] RSP: 0018:ff7532a13699fda0 EFLAGS: 00010212
[1924553.534986] RAX: 0fffffff0c930020 RBX: 0fffffff0c930020 RCX: 0000000000000000
[1924553.546679] RDX: ff2e6dbe0d9b6000 RSI: ff7532a13699fe60 RDI: ff2e6d1e4e630d80
[1924553.558367] RBP: 000000000b654440 R08: 0000000000002403 R09: 0000000000000179
[1924553.570026] R10: 000000000000000d R11: 0000000000000000 R12: 0000000001876e5c
[1924553.581601] R13: ff2e6d1e4e630d80 R14: ff7532a13699fe60 R15: 0000000000000000
[1924553.593115] FS: 00007ff8743aaa80(0000) GS:ff2e6e5e94c45000(0000) knlGS:0000000000000000
[1924553.605576] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[1924553.615611] CR2: 00007ffcd63a5000 CR3: 00000003ee840001 CR4: 0000000000771ef0
[1924553.627022] PKRU: 55555554
[1924553.633869] Call Trace:
[1924553.640353] <TASK>
[1924553.646369] d_lookup+0x27/0x50
[1924553.653366] lookup_dcache+0x1f/0x80
[1924553.660713] lookup_one_qstr_excl+0x1e/0xe0
[1924553.668589] ? preempt_schedule_common+0x2c/0x70
[1924553.676837] filename_create+0xc4/0x160
[1924553.684209] do_mkdirat+0x5a/0x190
[1924553.691050] __x64_sys_mkdir+0x42/0x60
[1924553.698163] do_syscall_64+0x64/0xbf0
[1924553.705145] entry_SYSCALL_64_after_hwframe+0x76/0x7e
[1924553.713533] RIP: 0033:0x7ff8754ff08b
[1924553.720358] Code: 8b 05 91 bd 0f 00 41 bc ff ff ff ff 64 c7 00 16 00 00 00 e9 4f ff ff ff e8 12 f7 01 00 66 90 f3 0f 1e fa b8 53 00 00 00 0f 05 <48> 3d 01 f0 ff ff 73 01 c3 48 8b 0d 5d
bd 0f 00 f7 d8 64 89 01 48
[1924553.748552] RSP: 002b:00007ffc05e5c148 EFLAGS: 00000246 ORIG_RAX: 0000000000000053
[1924553.759401] RAX: ffffffffffffffda RBX: 00005623cc0bd513 RCX: 00007ff8754ff08b
[1924553.769760] RDX: 000000000fde421b RSI: 00000000000001c0 RDI: 00005623cc0bd4f4
[1924553.780083] RBP: f49998db0aa753ff R08: 0000000000000004 R09: 0000000000000001
[1924553.790348] R10: 00007ff87587d000 R11: 0000000000000246 R12: 8421084210842109
[1924553.800604] R13: 00005623cc0bd513 R14: 00007ff8755bd740 R15: 000000000fde421b
[1924553.810819] </TASK>
I'll be very very gratefull for any hints here..
with best regards
nikola ciprich
PS: I tried to CC maintainers of suspected subsystems, but those are just my guesses,
so I hope I won't offend anyone.
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: hunting memory corruption bug in 6.18.x
2026-09-25 8:48 hunting memory corruption bug in 6.18.x Nikola Ciprich
@ 2026-09-25 10:05 ` Lorenzo Stoakes (ARM)
2026-09-25 12:13 ` Lorenzo Stoakes (ARM)
2026-09-26 16:02 ` Luiz Capitulino
2 siblings, 0 replies; 16+ messages in thread
From: Lorenzo Stoakes (ARM) @ 2026-09-25 10:05 UTC (permalink / raw)
To: Nikola Ciprich
Cc: linux-mm, linux-kernel, akpm, david, Dave Hansen, Mike Rapoport
+cc Dave, Mike
On Fri, Sep 25, 2026 at 10:48:26AM +0200, Nikola Ciprich wrote:
> Hi,
>
> I've been hunting a weird memory corruption bug for the last few weeks,
> without success so far, so I'd like to report it and kindly ask for help.
>
> We first hit it after a live VM migration between two KVM hosts:
> suddenly some dynamic libraries in the host appeared to be corrupted:
>
> Inconsistency detected by ld.so: ../sysdeps/x86_64/dl-machine.h: 548: elf_machine_rela_relative: Assertion `ELFW(R_TYPE) (reloc->r_info) == R_X86_64_RELATIVE' failed!
>
> (Later we also hit this with libcrypto.so.3 etc.) The files on disk
> were OK; the problem seemed to exist only in RAM.
>
> I'm fairly sure this is not hardware related: there were no ECC errors,
> and we have since hit this (and similar issues, more on that below) on
> multiple machines.
>
> The problems started after we moved from 5.15.x to 6.18.x kernels.
I had an AI dig into this.
And to give a quick response before trying to wrangle/check what it said
into a coherent analysis, the TL;DR is it seems to be caused by some CPA
bugs we fixed recently in x86.
This should be fixed in 6.18.53+ could you test again with everything
re-enabled?
I'll reply again with something more detailed.
>
> Since then I've spent a lot of time trying to reproduce it on a lab
> cluster, and we were able to trigger some corruption after days of
> migrating VMs back and forth. At first I suspected the Intel ice driver,
> for which I found similar reports, but we saw new problems even after
> backporting fixes (and also with Mellanox cards).
>
> So far we've hit three different kinds of problems, which may or may
> not be related:
>
> - .so library corruption right after VM migration
> - VM crashes (or process crashes inside VMs), possibly related to
> migration (those always happened during migration)
> - host crashes due to kernel structure corruption (these happened
> without any VM migration)
>
> We first hit these problems with 6.18.31; the last crash I saw was
> with 6.18.44.
>
> All affected machines use AMD EPYC CPUs and act as KVM hosts; the OS
> is AlmaLinux 9.
>
> I suspect two subsystems that have seen a lot of changes:
>
> - transparent hugepages
> - NUMA balancing
>
> (but those are just my guesses)
>
> As a safety measure, we've disabled THP and NUMA balancing on all hosts.
>
> I'm aware this is still a very vague report with a lot of guessing,
> but my question is: has anybody hit similar problems with 6.18 or
> newer kernels?
>
> I see a lot of patches in every stable release, but simply trying
> newer kernels doesn't seem efficient here. Deploying them is also
> risky, since the hosts have to be emptied by migrating VMs off them
> before reboot, and that migration itself may trigger more crashes.
> None of the released or queued fixes for 6.18 seem to be directly
> related.
>
> I tried running my migration tests on hosts with KASAN enabled, and
> also with SLUB debugging, but was never able to reproduce the problem
> with those enabled (without them, I was able to hit issues within
> days).
>
> I'll start another round of migration tests in the lab, now with
> 6.18.54-rc1, but I still thought it would be good to report this and
> ask here.
>
> last but not least, here's kdump from last crash (this was not related
> to any VM migration, but is very similar to another few crashes
> we got):
>
> [1924553.414736] Oops: general protection fault, probably for non-canonical address 0xfffffff0c930038: 0000 [#1] SMP NOPTI
> [1924553.434800] CPU: 23 UID: 189 PID: 7538 Comm: pacemaker-contr Kdump: loaded Tainted: G E 6.18.44lb9.01 #1 PREEMPT(voluntary)
> [1924553.456934] Tainted: [E]=UNSIGNED_MODULE
> [1924553.465551] Hardware name: ASUSTeK COMPUTER INC. RS720A-E12-RS12/K14PP-D24 Series, BIOS 2305 11/21/2025
> [1924553.484152] RIP: 0010:__d_lookup+0x4a/0xc0
> [1924553.492878] Code: ff 48 89 c5 c1 e8 07 48 8d 1c c2 e8 60 8f d1 ff 48 8b 03 48 89 c3 48 83 e3 fe 48 83 f8 01 77 0a eb 2f 48 8b 1b 48 85 db 74 27 <39> 6b 18 75 f3 4c 8d 63 78 4c 89 e7 e8
> d5 e1 7c 00 4c 39 6b 10 74
> [1924553.525191] RSP: 0018:ff7532a13699fda0 EFLAGS: 00010212
> [1924553.534986] RAX: 0fffffff0c930020 RBX: 0fffffff0c930020 RCX: 0000000000000000
> [1924553.546679] RDX: ff2e6dbe0d9b6000 RSI: ff7532a13699fe60 RDI: ff2e6d1e4e630d80
> [1924553.558367] RBP: 000000000b654440 R08: 0000000000002403 R09: 0000000000000179
> [1924553.570026] R10: 000000000000000d R11: 0000000000000000 R12: 0000000001876e5c
> [1924553.581601] R13: ff2e6d1e4e630d80 R14: ff7532a13699fe60 R15: 0000000000000000
> [1924553.593115] FS: 00007ff8743aaa80(0000) GS:ff2e6e5e94c45000(0000) knlGS:0000000000000000
> [1924553.605576] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
> [1924553.615611] CR2: 00007ffcd63a5000 CR3: 00000003ee840001 CR4: 0000000000771ef0
> [1924553.627022] PKRU: 55555554
> [1924553.633869] Call Trace:
> [1924553.640353] <TASK>
> [1924553.646369] d_lookup+0x27/0x50
> [1924553.653366] lookup_dcache+0x1f/0x80
> [1924553.660713] lookup_one_qstr_excl+0x1e/0xe0
> [1924553.668589] ? preempt_schedule_common+0x2c/0x70
> [1924553.676837] filename_create+0xc4/0x160
> [1924553.684209] do_mkdirat+0x5a/0x190
> [1924553.691050] __x64_sys_mkdir+0x42/0x60
> [1924553.698163] do_syscall_64+0x64/0xbf0
> [1924553.705145] entry_SYSCALL_64_after_hwframe+0x76/0x7e
> [1924553.713533] RIP: 0033:0x7ff8754ff08b
> [1924553.720358] Code: 8b 05 91 bd 0f 00 41 bc ff ff ff ff 64 c7 00 16 00 00 00 e9 4f ff ff ff e8 12 f7 01 00 66 90 f3 0f 1e fa b8 53 00 00 00 0f 05 <48> 3d 01 f0 ff ff 73 01 c3 48 8b 0d 5d
> bd 0f 00 f7 d8 64 89 01 48
> [1924553.748552] RSP: 002b:00007ffc05e5c148 EFLAGS: 00000246 ORIG_RAX: 0000000000000053
> [1924553.759401] RAX: ffffffffffffffda RBX: 00005623cc0bd513 RCX: 00007ff8754ff08b
> [1924553.769760] RDX: 000000000fde421b RSI: 00000000000001c0 RDI: 00005623cc0bd4f4
> [1924553.780083] RBP: f49998db0aa753ff R08: 0000000000000004 R09: 0000000000000001
> [1924553.790348] R10: 00007ff87587d000 R11: 0000000000000246 R12: 8421084210842109
> [1924553.800604] R13: 00005623cc0bd513 R14: 00007ff8755bd740 R15: 000000000fde421b
> [1924553.810819] </TASK>
>
> I'll be very very gratefull for any hints here..
Hi, I had a
>
> with best regards
>
> nikola ciprich
>
> PS: I tried to CC maintainers of suspected subsystems, but those are just my guesses,
> so I hope I won't offend anyone.
>
--
Cheers, Lorenzo
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: hunting memory corruption bug in 6.18.x
2026-09-25 8:48 hunting memory corruption bug in 6.18.x Nikola Ciprich
2026-09-25 10:05 ` Lorenzo Stoakes (ARM)
@ 2026-09-25 12:13 ` Lorenzo Stoakes (ARM)
2026-09-26 5:58 ` Nikola Ciprich
2026-09-26 16:02 ` Luiz Capitulino
2 siblings, 1 reply; 16+ messages in thread
From: Lorenzo Stoakes (ARM) @ 2026-09-25 12:13 UTC (permalink / raw)
To: Nikola Ciprich
Cc: linux-mm, linux-kernel, akpm, david, Mike Rapoport, Dave Hansen,
Pedro Falcato, Kiryl Shutsemau
+cc various
Tl;DR before I dig in, I had an AI dig into the report (as they're
essentially superhuman at this kind of thing so always worth doing), and it
seems the recently fixed CPA bugs are likely to be the underlying cause
here.
EDIT: OK so I spent 2+ hrs analysing this :>))) but hopefully it's useful,
I wanted to make sure what the LLM came up with was vaguely sensible.
It's speculative, but I really think the below is the best explanation for
what you're observing.
And TL;DR is that 6.18.53 should fix it.
On Fri, Sep 25, 2026 at 10:48:26AM +0200, Nikola Ciprich wrote:
> Hi,
>
> I've been hunting a weird memory corruption bug for the last few weeks,
> without success so far, so I'd like to report it and kindly ask for help.
Thanks for the detailed report!
>
> We first hit it after a live VM migration between two KVM hosts:
> suddenly some dynamic libraries in the host appeared to be corrupted:
>
> Inconsistency detected by ld.so: ../sysdeps/x86_64/dl-machine.h: 548: elf_machine_rela_relative: Assertion `ELFW(R_TYPE) (reloc->r_info) == R_X86_64_RELATIVE' failed!
>
> (Later we also hit this with libcrypto.so.3 etc.) The files on disk
> were OK; the problem seemed to exist only in RAM.
>
> I'm fairly sure this is not hardware related: there were no ECC errors,
> and we have since hit this (and similar issues, more on that below) on
> multiple machines.
>
> The problems started after we moved from 5.15.x to 6.18.x kernels.
It seems that commit 41d88484c71c ("x86/mm/pat: restore large ROX pages after
fragmentation") is the underlying cause (landed in 6.15).
It impacts CPA or 'Change Page Attributes' which is the means by which direct
mapping page table entries are updated to reflect underlying attribute changes
for ranges, often (and the motivation behind this change) read-only executable
ranges for e.g. JITters etc.
When it does this it sees if the range being changed can be 'collapsed' into a
huge page, i.e. mapped at PMD level for instance rather than PTE level to avoid
fragmentation of the direct map.
In particular the change introduces cpa_collapse_large_pages(), which frees
kernel page tables when it does this.
And this is problematic, because it did that without properly synchronising
against concurrent readers.
And I think in particular the issue here is the one fixed by Pedro in commit
1587d3394e25 ("x86/alternatives: Exclude text poking against
change_page_attr()").
>
> Since then I've spent a lot of time trying to reproduce it on a lab
> cluster, and we were able to trigger some corruption after days of
> migrating VMs back and forth. At first I suspected the Intel ice driver,
Yeah these race bugs can be VERY painful, sorry about that!
> for which I found similar reports, but we saw new problems even after
> backporting fixes (and also with Mellanox cards).
>
> So far we've hit three different kinds of problems, which may or may
> not be related:
>
> - .so library corruption right after VM migration
> - VM crashes (or process crashes inside VMs), possibly related to
> migration (those always happened during migration)
> - host crashes due to kernel structure corruption (these happened
> without any VM migration)
So there are two sides to the race: set_memory_rox() - triggered on module
load, ftrace trampoline creation and every new BPF 2M program pack.
The other side is execmem_force_rw() -> set_memory_[nx,rw]() (concurrent
module load or ftrace trampoline allocation) __text_poke() ->
vmalloc_to_page() for patching module text, kprobe slots, trampolines or
BPF.
Both are happening a lot at KVM host bringup (module autoload, per-VM
seccomp filters, perf, BPF probes, etc.
So this aligns with the theory.
>
> We first hit these problems with 6.18.31; the last crash I saw was
> with 6.18.44.
Yeah, the fact you didn't see an issue with 5.15 matches commit
41d88484c71c ("x86/mm/pat: restore large ROX pages after fragmentation")
being the cause.
>
> All affected machines use AMD EPYC CPUs and act as KVM hosts; the OS
> is AlmaLinux 9.
>
> I suspect two subsystems that have seen a lot of changes:
>
> - transparent hugepages
> - NUMA balancing
>
> (but those are just my guesses)
>
> As a safety measure, we've disabled THP and NUMA balancing on all hosts.
Actually I think doing this doesn't actually save you at all, since the
collapse happens even without THP enabled, and NUMA balancing shouldn't
impact any of this.
>
> I'm aware this is still a very vague report with a lot of guessing,
> but my question is: has anybody hit similar problems with 6.18 or
> newer kernels?
Yeah, the description of a fix for this mentions something that seems
exactly like this bug:
https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=1aac65f3e651334259ecb2a5f5ddb81c01f02599
Though note that that patch doesn't actually solve the problem, you need
fixes from 6.18.53 to resolve the bug:
Commit a1c7570cedd0 ("x86/mm/pat: Acquire init_mm write lock on collapse to avoid UAF")
Commit d5d8b8662e6e ("x86/mm/pat: Acquire init_mm read lock on attribute changes to avoid UAF")
These are prerequisites ^^^ for the actual fix for this vvv
Commit 1587d3394e25 ("x86/alternatives: Exclude text poking against change_page_attr()")
>
> I see a lot of patches in every stable release, but simply trying
> newer kernels doesn't seem efficient here. Deploying them is also
> risky, since the hosts have to be emptied by migrating VMs off them
> before reboot, and that migration itself may trigger more crashes.
> None of the released or queued fixes for 6.18 seem to be directly
> related.
>
> I tried running my migration tests on hosts with KASAN enabled, and
> also with SLUB debugging, but was never able to reproduce the problem
> with those enabled (without them, I was able to hit issues within
> days).
Ugh, unhelpful, but makes sense as it changes race windows.
>
> I'll start another round of migration tests in the lab, now with
> 6.18.54-rc1, but I still thought it would be good to report this and
> ask here.
>
Ah yeah you're already going to be testing the fixed series then :)
Obviously if the issue re-triggers there, back to the drawing board. But I
don't think it will.
> last but not least, here's kdump from last crash (this was not related
> to any VM migration, but is very similar to another few crashes
> we got):
>
> [1924553.414736] Oops: general protection fault, probably for non-canonical address 0xfffffff0c930038: 0000 [#1] SMP NOPTI
> [1924553.434800] CPU: 23 UID: 189 PID: 7538 Comm: pacemaker-contr Kdump: loaded Tainted: G E 6.18.44lb9.01 #1 PREEMPT(voluntary)
> [1924553.456934] Tainted: [E]=UNSIGNED_MODULE
> [1924553.465551] Hardware name: ASUSTeK COMPUTER INC. RS720A-E12-RS12/K14PP-D24 Series, BIOS 2305 11/21/2025
> [1924553.484152] RIP: 0010:__d_lookup+0x4a/0xc0
So __d_lookup() is where the invalid address oops happened.
> [1924553.492878] Code: ff 48 89 c5 c1 e8 07 48 8d 1c c2 e8 60 8f d1 ff 48 8b 03 48 89 c3 48 83 e3 fe 48 83 f8 01 77 0a eb 2f 48 8b 1b 48 85 db 74 27 <39> 6b 18 75 f3 4c 8d 63 78 4c 89 e7 e8
> d5 e1 7c 00 4c 39 6b 10 74
> [1924553.525191] RSP: 0018:ff7532a13699fda0 EFLAGS: 00010212
> [1924553.534986] RAX: 0fffffff0c930020 RBX: 0fffffff0c930020 RCX: 0000000000000000
> [1924553.546679] RDX: ff2e6dbe0d9b6000 RSI: ff7532a13699fe60 RDI: ff2e6d1e4e630d80
> [1924553.558367] RBP: 000000000b654440 R08: 0000000000002403 R09: 0000000000000179
> [1924553.570026] R10: 000000000000000d R11: 0000000000000000 R12: 0000000001876e5c
> [1924553.581601] R13: ff2e6d1e4e630d80 R14: ff7532a13699fe60 R15: 0000000000000000
> [1924553.593115] FS: 00007ff8743aaa80(0000) GS:ff2e6e5e94c45000(0000) knlGS:0000000000000000
> [1924553.605576] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
> [1924553.615611] CR2: 00007ffcd63a5000 CR3: 00000003ee840001 CR4: 0000000000771ef0
> [1924553.627022] PKRU: 55555554
> [1924553.633869] Call Trace:
> [1924553.640353] <TASK>
> [1924553.646369] d_lookup+0x27/0x50
> [1924553.653366] lookup_dcache+0x1f/0x80
> [1924553.660713] lookup_one_qstr_excl+0x1e/0xe0
> [1924553.668589] ? preempt_schedule_common+0x2c/0x70
> [1924553.676837] filename_create+0xc4/0x160
> [1924553.684209] do_mkdirat+0x5a/0x190
> [1924553.691050] __x64_sys_mkdir+0x42/0x60
> [1924553.698163] do_syscall_64+0x64/0xbf0
> [1924553.705145] entry_SYSCALL_64_after_hwframe+0x76/0x7e
> [1924553.713533] RIP: 0033:0x7ff8754ff08b
> [1924553.720358] Code: 8b 05 91 bd 0f 00 41 bc ff ff ff ff 64 c7 00 16 00 00 00 e9 4f ff ff ff e8 12 f7 01 00 66 90 f3 0f 1e fa b8 53 00 00 00 0f 05 <48> 3d 01 f0 ff ff 73 01 c3 48 8b 0d 5d
> bd 0f 00 f7 d8 64 89 01 48
> [1924553.748552] RSP: 002b:00007ffc05e5c148 EFLAGS: 00000246 ORIG_RAX: 0000000000000053
> [1924553.759401] RAX: ffffffffffffffda RBX: 00005623cc0bd513 RCX: 00007ff8754ff08b
> [1924553.769760] RDX: 000000000fde421b RSI: 00000000000001c0 RDI: 00005623cc0bd4f4
> [1924553.780083] RBP: f49998db0aa753ff R08: 0000000000000004 R09: 0000000000000001
> [1924553.790348] R10: 00007ff87587d000 R11: 0000000000000246 R12: 8421084210842109
> [1924553.800604] R13: 00005623cc0bd513 R14: 00007ff8755bd740 R15: 000000000fde421b
> [1924553.810819] </TASK>
So the LLM went to town on this and it's quite interesting.
The code (6.18.44) disassembles to:
struct dentry *__d_lookup(const struct dentry *parent, const struct qstr *name)
{
...
hlist_bl_for_each_entry_rcu(dentry, node, b, d_hash) {
Which is:
#define hlist_bl_for_each_entry_rcu(tpos, pos, head, member) \
for (pos = hlist_bl_first_rcu(head); \
pos && \
({ tpos = hlist_bl_entry(pos, typeof(*tpos), member); 1; }); \
pos = rcu_dereference_raw(pos->next))
And:
static inline struct hlist_bl_node *hlist_bl_first_rcu(struct hlist_bl_head *h)
{
return (struct hlist_bl_node *)
((unsigned long)rcu_dereference_check(h->first, hlist_bl_is_locked(h)) & ~LIST_BL_LOCKMASK);
}
And:
static inline bool hlist_bl_is_locked(struct hlist_bl_head *b)
{
return bit_spin_is_locked(0, (unsigned long *)b);
}
mov (%rbx),%rax ; rax = h->first = 0x0fffffff0c930020
mov %rax,%rbx
and $-2,%rbx ; strip hlist_bl lock bit (no-op, bit 0 clear)
cmp $1,%rax ; hlist_bl_is_locked()
ja body ; non-empty, enter loop
...
if (dentry->d_name.hash != hash)
continue;
body:
cmp %ebp,0x18(%rbx) ; <-- FAULT: deref of dentry->d_name.hash
So dentry is corrupted.
The code and registers are consistent with this being the first iteration
of the loop, which you'd expect with corrupted dentry.
Looking further back in the code:
struct dentry *__d_lookup(const struct dentry *parent, const struct qstr *name)
{
...
struct hlist_bl_head *b = d_hash(hash);
And:
static inline struct hlist_bl_head *d_hash(unsigned long hashlen)
{
return runtime_const_ptr(dentry_hashtable) +
runtime_const_shift_right_32(hashlen, d_hash_shift);
}
Only the trailing ff is in the code output from the splat but the movabs is
there -> RDX: so:
movabs $0xff2e6dbe0d9b6000,%rdx ; runtime_const_ptr(dentry_hashtable)
mov %rax, %rbp
shr $0x7, %eax ; runtime_const_shift_right_32(hashlen, d_hash_shift);
lea (%rdx, %rax, 8), %rbx ; bucket = &dentry_hashtable[hash >> 7]
(The 8 is multiplying the size of the 8 byte pointers)
Note that rbp retains RAX's value = 0xb654440 (not clobbered elsewehre), so
the hlist_bl_head bucket pointer is
0xff2e6dbe0d9b6000 + (0xb654440 >> 7) * 8
So:
b = 0xff2e6dbe0e51b440
This matters, because in hlist_bl_first_rcu() this pointer is treated as a
valid struct hlist_bl_head pointer:
struct hlist_bl_head {
struct hlist_bl_node *first;
};
And the data at 0xff2e6dbe0e51b440 contains 0x0fffffff0c930020 (rbx), which
is assumed to be a valid struct hlist_bl_node * embedded in a dentry.
IOW, RDX contains the dentry_hashtable pointer:
static struct hlist_bl_head *dentry_hashtable __ro_after_init __used;
That is allocated in dcache_init_early() and never freed:
static void __init dcache_init_early(void)
{
...
dentry_hashtable =
alloc_large_system_hash("Dentry cache",
sizeof(struct hlist_bl_head),
dhash_entries,
13,
HASH_EARLY | HASH_ZERO,
&d_hash_shift,
NULL,
2,
0);
...
}
So that allocate has to be legit, somehow the data there got corrupted.
This is allocated early by memblock.
d_hash_shift = 7 = 32 - lg(entries), so entries = 25, and 2^25 entries of 8
bytes each = a 256 MiB table.
So it's legit data that got corrupted.
RDI contains the parent dentry at 0xff2e6d1e4e630d80 (the disassembly at
__d_lookup() confirms).
KASLR makes things tricky but this is most definitely a slab allocation in
the direct map and the stack (VMAP_STACK) is at 0xff7532a13699fda0 (rsp)
~70 TiB higher which sits at least 10 TiB padded above the direct map so
it's safe to say that this is in the direct map.
And the corrupted value (rbx) 0x0fffffff0c930020 is interesting - it's an
exactly a x86-64 swap softleaf value:
__swp_type() = val >> (64 - SWP_TYPE_BITS=5) = val >> 59 = 1
__swp_offset() = (~(x).val << SWP_TYPE_BITS >> SWP_OFFSET_SHIFT) = (~val << 6 >> 14)
= 0x79b67f
I.e. it's a swap softleaf entry 0x79b67f 4 KiB pages into the swap = ~30.4 GiB.
It's also the _second_ swap in the system (Nikola - if you have a 2nd swap
on the reporting system of >=~30 GiB that kinda confirms it).
Also, interstingly, bit 5 is set, which is one of the bits allowed to be
set (possibly by hardware) in the swap entry.
__swp_entry() clears bits 0-8 and only bits 1-3 are software bits in a
non-present PTE.
This is _PAGE_ACCESSED (1 << _PAGE_BIT_ACCESSED = 5) = 0x20. So it makes
sense that hardware might have set it.
In effect - every single bit is exactly how it should be for a valid swap
entry (since commit 00839ee3b299 ("x86/mm: Move swap offset/type up in PTE
to work around erratum"))..
It seems more than a coincidence :)
This speaks to some kind of memory corruption that has resulted in a store
to an arbitrary physical address that happens to be the dentry.
Now looking to the proposed CPA cause (ultimately the thing solved by
So the speculated race here is:
pfn_exec: pfn of the module text page CPU A is changing attributes on
pfn_dh: pfn of the dentry hash table page the swap PTE ends up in
CPU A: set_memory_nx() (text_poke) CPU B: set_memory_rox() (module load)
------------------------------------ -----------------------------------------
__change_page_attr()
kpte = lookup (lockless)
<preempted>
cpa_collapse_large_pages()
set_pmd(leaf)
__free_pages(old PTE table)
some process: pte_alloc() gets
that page as a user page table
UAF!!! Writing into arbitrary memory
set_pte_atomic(kpte,
pfn_pte(pfn_exec, prot)) <- lands in that page table:
entry -> pfn_exec (module text)
process faults on it, GUP follows
it, later zap frees the data page at
pfn_exec while it is still live module
text!!!
That data page is reallocated as a PMD
table.
text_poke() keeps writing
code bytes into it -> entries
with arbitrary pfns, one of them
pfn_dh
reclaim: try_to_unmap_one()
pte_offset_map() -> __va(pfn_dh)
set_pte_at(swap PTE) <- dentry
hash table
(The reason it needs to be interpreted as a PMD page table is otherwise
reclaim wouldn't be trying to write a swap PTE entry into it).
And this is exactly what commit 1587d3394e25 ("x86/alternatives: Exclude
text poking against change_page_attr()") protects against.
Yes it's out on a limb (and I spent FAR TOO LONG going through this
analysis) but there's really no other sensible explanation as to why dentry
data got corrupted like that to that exact shape.
Which goes to show how nasty this kind of data corruption issue can be.
>
> I'll be very very gratefull for any hints here..
As above :)
>
> with best regards
>
> nikola ciprich
>
> PS: I tried to CC maintainers of suspected subsystems, but those are just my guesses,
> so I hope I won't offend anyone.
>
Don't worry about that, I think we're all happy to get legit bug
reports. We just might be too busy to reply right away :)
The LLM added on some hints for confirmation of this:
Schlopp>>
Log greps, across all affected hosts and all boots, not just the ones
that crashed:
grep -i 'Bad page map' /var/log/messages*
grep -i 'bad pmd' /var/log/messages*
grep -i 'bad pud' /var/log/messages*
grep -i 'Bad page state' /var/log/messages*
grep -i 'CPA: called for zero pte' /var/log/messages*
Any of these, particularly "bad pmd", is direct evidence that a freed
kernel PTE table was reused as a user page table. "CPA: called for zero
pte" would be the CPA walker itself tripping over a collapsed mapping.
Questions:
1. swapon --show on the host, and inside the guests. Is there a swap
device with index 1 and a size of at least roughly 30.5GiB? That
tells us whether the corrupt word is a host swap PTE or a guest one,
which distinguishes case A from case B above.
2. From the vmcore, dump the 4K page containing 0xff2e6dbe0e51b440. If
the neighbouring words are also PTE-shaped, the page was being used
as a page table and the diagnosis above is confirmed. If only the
one word is corrupt, it was a single stray store.
3. What are CONFIG_DEBUG_VM, init_on_alloc, init_on_free and
page_poison set to in the production build versus the KASAN build?
free_page_is_bad() is gated on is_check_pages_enabled(), which needs
CONFIG_DEBUG_VM, so the production kernel would not report the bad
free even if it happened.
4. Are any of the crashing guests Windows, and is hv-tlbflush set on
them? That decides whether 26505e1b5b54 matters for you.
5. Has any corruption occurred since THP was disabled? If yes, that
supports the CPA race over your THP theory.
6. When you run the 6.18.54-rc1 test, please note that 6.18.52 is only
a partial fix, so if you have any results from a 6.18.52 kernel they
should not be treated as a clean run.
<< Schlopp
But really a run against 6.18.53 being OK under heavy testing for several
days should confirm it also.
If it turns out it's not this then back to the drawing board I guess! Let
us know.
--
Cheers, Lorenzo
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: hunting memory corruption bug in 6.18.x
2026-09-25 12:13 ` Lorenzo Stoakes (ARM)
@ 2026-09-26 5:58 ` Nikola Ciprich
2026-09-26 9:32 ` Lorenzo Stoakes (ARM)
0 siblings, 1 reply; 16+ messages in thread
From: Nikola Ciprich @ 2026-09-26 5:58 UTC (permalink / raw)
To: Lorenzo Stoakes (ARM)
Cc: linux-mm, linux-kernel, akpm, david, Mike Rapoport, Dave Hansen,
Pedro Falcato, Kiryl Shutsemau, Nikola Ciprich
Hello Lorenzo (and others),
thank you for your time looking into this..
(replies inline)
On Fri, Sep 25, 2026 at 01:13:52PM +0100, Lorenzo Stoakes (ARM) wrote:
> +cc various
>
> Tl;DR before I dig in, I had an AI dig into the report (as they're
> essentially superhuman at this kind of thing so always worth doing), and it
> seems the recently fixed CPA bugs are likely to be the underlying cause
> here.
>
> EDIT: OK so I spent 2+ hrs analysing this :>))) but hopefully it's useful,
> I wanted to make sure what the LLM came up with was vaguely sensible.
>
> It's speculative, but I really think the below is the best explanation for
> what you're observing.
>
> And TL;DR is that 6.18.53 should fix it.
>
> On Fri, Sep 25, 2026 at 10:48:26AM +0200, Nikola Ciprich wrote:
.. truncated ..
> >
> > The problems started after we moved from 5.15.x to 6.18.x kernels.
>
> It seems that commit 41d88484c71c ("x86/mm/pat: restore large ROX pages after
> fragmentation") is the underlying cause (landed in 6.15).
>
> It impacts CPA or 'Change Page Attributes' which is the means by which direct
> mapping page table entries are updated to reflect underlying attribute changes
> for ranges, often (and the motivation behind this change) read-only executable
> ranges for e.g. JITters etc.
>
> When it does this it sees if the range being changed can be 'collapsed' into a
> huge page, i.e. mapped at PMD level for instance rather than PTE level to avoid
> fragmentation of the direct map.
>
> In particular the change introduces cpa_collapse_large_pages(), which frees
> kernel page tables when it does this.
>
> And this is problematic, because it did that without properly synchronising
> against concurrent readers.
>
> And I think in particular the issue here is the one fixed by Pedro in commit
> 1587d3394e25 ("x86/alternatives: Exclude text poking against
> change_page_attr()").
>
> >
> > Since then I've spent a lot of time trying to reproduce it on a lab
> > cluster, and we were able to trigger some corruption after days of
> > migrating VMs back and forth. At first I suspected the Intel ice driver,
>
> Yeah these race bugs can be VERY painful, sorry about that!
>
> > for which I found similar reports, but we saw new problems even after
> > backporting fixes (and also with Mellanox cards).
> >
> > So far we've hit three different kinds of problems, which may or may
> > not be related:
> >
> > - .so library corruption right after VM migration
> > - VM crashes (or process crashes inside VMs), possibly related to
> > migration (those always happened during migration)
> > - host crashes due to kernel structure corruption (these happened
> > without any VM migration)
>
> So there are two sides to the race: set_memory_rox() - triggered on module
> load, ftrace trampoline creation and every new BPF 2M program pack.
>
> The other side is execmem_force_rw() -> set_memory_[nx,rw]() (concurrent
> module load or ftrace trampoline allocation) __text_poke() ->
> vmalloc_to_page() for patching module text, kprobe slots, trampolines or
> BPF.
>
> Both are happening a lot at KVM host bringup (module autoload, per-VM
> seccomp filters, perf, BPF probes, etc.
one note here, at least last mentioned crash (with 6.18.44) happened with
host running only windows guest, in general we're seeing those problems
mosly with windows VM hosting machines.. so maybe they're triggerng the
problem with some other, but similar mechanism?
>
> So this aligns with the theory.
>
> >
> > We first hit these problems with 6.18.31; the last crash I saw was
> > with 6.18.44.
>
> Yeah, the fact you didn't see an issue with 5.15 matches commit
> 41d88484c71c ("x86/mm/pat: restore large ROX pages after fragmentation")
> being the cause.
>
> >
> > All affected machines use AMD EPYC CPUs and act as KVM hosts; the OS
> > is AlmaLinux 9.
> >
> > I suspect two subsystems that have seen a lot of changes:
> >
> > - transparent hugepages
> > - NUMA balancing
> >
> > (but those are just my guesses)
> >
> > As a safety measure, we've disabled THP and NUMA balancing on all hosts.
>
> Actually I think doing this doesn't actually save you at all, since the
> collapse happens even without THP enabled, and NUMA balancing shouldn't
> impact any of this.
>
> >
> > I'm aware this is still a very vague report with a lot of guessing,
> > but my question is: has anybody hit similar problems with 6.18 or
> > newer kernels?
>
> Yeah, the description of a fix for this mentions something that seems
> exactly like this bug:
>
> https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=1aac65f3e651334259ecb2a5f5ddb81c01f02599
>
> Though note that that patch doesn't actually solve the problem, you need
> fixes from 6.18.53 to resolve the bug:
>
> Commit a1c7570cedd0 ("x86/mm/pat: Acquire init_mm write lock on collapse to avoid UAF")
> Commit d5d8b8662e6e ("x86/mm/pat: Acquire init_mm read lock on attribute changes to avoid UAF")
>
> These are prerequisites ^^^ for the actual fix for this vvv
>
> Commit 1587d3394e25 ("x86/alternatives: Exclude text poking against change_page_attr()")
>
> >
> > I see a lot of patches in every stable release, but simply trying
> > newer kernels doesn't seem efficient here. Deploying them is also
> > risky, since the hosts have to be emptied by migrating VMs off them
> > before reboot, and that migration itself may trigger more crashes.
> > None of the released or queued fixes for 6.18 seem to be directly
> > related.
> >
> > I tried running my migration tests on hosts with KASAN enabled, and
> > also with SLUB debugging, but was never able to reproduce the problem
> > with those enabled (without them, I was able to hit issues within
> > days).
>
> Ugh, unhelpful, but makes sense as it changes race windows.
>
> >
> > I'll start another round of migration tests in the lab, now with
> > 6.18.54-rc1, but I still thought it would be good to report this and
> > ask here.
> >
>
> Ah yeah you're already going to be testing the fixed series then :)
>
> Obviously if the issue re-triggers there, back to the drawing board. But I
> don't think it will.
>
> > last but not least, here's kdump from last crash (this was not related
> > to any VM migration, but is very similar to another few crashes
> > we got):
> >
> > [1924553.414736] Oops: general protection fault, probably for non-canonical address 0xfffffff0c930038: 0000 [#1] SMP NOPTI
> > [1924553.434800] CPU: 23 UID: 189 PID: 7538 Comm: pacemaker-contr Kdump: loaded Tainted: G E 6.18.44lb9.01 #1 PREEMPT(voluntary)
> > [1924553.456934] Tainted: [E]=UNSIGNED_MODULE
> > [1924553.465551] Hardware name: ASUSTeK COMPUTER INC. RS720A-E12-RS12/K14PP-D24 Series, BIOS 2305 11/21/2025
> > [1924553.484152] RIP: 0010:__d_lookup+0x4a/0xc0
>
> So __d_lookup() is where the invalid address oops happened.
>
> > [1924553.492878] Code: ff 48 89 c5 c1 e8 07 48 8d 1c c2 e8 60 8f d1 ff 48 8b 03 48 89 c3 48 83 e3 fe 48 83 f8 01 77 0a eb 2f 48 8b 1b 48 85 db 74 27 <39> 6b 18 75 f3 4c 8d 63 78 4c 89 e7 e8
> > d5 e1 7c 00 4c 39 6b 10 74
> > [1924553.525191] RSP: 0018:ff7532a13699fda0 EFLAGS: 00010212
> > [1924553.534986] RAX: 0fffffff0c930020 RBX: 0fffffff0c930020 RCX: 0000000000000000
> > [1924553.546679] RDX: ff2e6dbe0d9b6000 RSI: ff7532a13699fe60 RDI: ff2e6d1e4e630d80
> > [1924553.558367] RBP: 000000000b654440 R08: 0000000000002403 R09: 0000000000000179
> > [1924553.570026] R10: 000000000000000d R11: 0000000000000000 R12: 0000000001876e5c
> > [1924553.581601] R13: ff2e6d1e4e630d80 R14: ff7532a13699fe60 R15: 0000000000000000
> > [1924553.593115] FS: 00007ff8743aaa80(0000) GS:ff2e6e5e94c45000(0000) knlGS:0000000000000000
> > [1924553.605576] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
> > [1924553.615611] CR2: 00007ffcd63a5000 CR3: 00000003ee840001 CR4: 0000000000771ef0
> > [1924553.627022] PKRU: 55555554
> > [1924553.633869] Call Trace:
> > [1924553.640353] <TASK>
> > [1924553.646369] d_lookup+0x27/0x50
> > [1924553.653366] lookup_dcache+0x1f/0x80
> > [1924553.660713] lookup_one_qstr_excl+0x1e/0xe0
> > [1924553.668589] ? preempt_schedule_common+0x2c/0x70
> > [1924553.676837] filename_create+0xc4/0x160
> > [1924553.684209] do_mkdirat+0x5a/0x190
> > [1924553.691050] __x64_sys_mkdir+0x42/0x60
> > [1924553.698163] do_syscall_64+0x64/0xbf0
> > [1924553.705145] entry_SYSCALL_64_after_hwframe+0x76/0x7e
> > [1924553.713533] RIP: 0033:0x7ff8754ff08b
> > [1924553.720358] Code: 8b 05 91 bd 0f 00 41 bc ff ff ff ff 64 c7 00 16 00 00 00 e9 4f ff ff ff e8 12 f7 01 00 66 90 f3 0f 1e fa b8 53 00 00 00 0f 05 <48> 3d 01 f0 ff ff 73 01 c3 48 8b 0d 5d
> > bd 0f 00 f7 d8 64 89 01 48
> > [1924553.748552] RSP: 002b:00007ffc05e5c148 EFLAGS: 00000246 ORIG_RAX: 0000000000000053
> > [1924553.759401] RAX: ffffffffffffffda RBX: 00005623cc0bd513 RCX: 00007ff8754ff08b
> > [1924553.769760] RDX: 000000000fde421b RSI: 00000000000001c0 RDI: 00005623cc0bd4f4
> > [1924553.780083] RBP: f49998db0aa753ff R08: 0000000000000004 R09: 0000000000000001
> > [1924553.790348] R10: 00007ff87587d000 R11: 0000000000000246 R12: 8421084210842109
> > [1924553.800604] R13: 00005623cc0bd513 R14: 00007ff8755bd740 R15: 000000000fde421b
> > [1924553.810819] </TASK>
>
> So the LLM went to town on this and it's quite interesting.
>
> The code (6.18.44) disassembles to:
>
> struct dentry *__d_lookup(const struct dentry *parent, const struct qstr *name)
> {
> ...
>
> hlist_bl_for_each_entry_rcu(dentry, node, b, d_hash) {
>
> Which is:
>
> #define hlist_bl_for_each_entry_rcu(tpos, pos, head, member) \
> for (pos = hlist_bl_first_rcu(head); \
> pos && \
> ({ tpos = hlist_bl_entry(pos, typeof(*tpos), member); 1; }); \
> pos = rcu_dereference_raw(pos->next))
>
> And:
>
> static inline struct hlist_bl_node *hlist_bl_first_rcu(struct hlist_bl_head *h)
> {
> return (struct hlist_bl_node *)
> ((unsigned long)rcu_dereference_check(h->first, hlist_bl_is_locked(h)) & ~LIST_BL_LOCKMASK);
> }
>
> And:
>
> static inline bool hlist_bl_is_locked(struct hlist_bl_head *b)
> {
> return bit_spin_is_locked(0, (unsigned long *)b);
> }
>
> mov (%rbx),%rax ; rax = h->first = 0x0fffffff0c930020
>
> mov %rax,%rbx
> and $-2,%rbx ; strip hlist_bl lock bit (no-op, bit 0 clear)
> cmp $1,%rax ; hlist_bl_is_locked()
> ja body ; non-empty, enter loop
> ...
>
> if (dentry->d_name.hash != hash)
> continue;
> body:
> cmp %ebp,0x18(%rbx) ; <-- FAULT: deref of dentry->d_name.hash
>
> So dentry is corrupted.
>
> The code and registers are consistent with this being the first iteration
> of the loop, which you'd expect with corrupted dentry.
>
> Looking further back in the code:
>
> struct dentry *__d_lookup(const struct dentry *parent, const struct qstr *name)
> {
> ...
> struct hlist_bl_head *b = d_hash(hash);
>
> And:
>
> static inline struct hlist_bl_head *d_hash(unsigned long hashlen)
> {
> return runtime_const_ptr(dentry_hashtable) +
> runtime_const_shift_right_32(hashlen, d_hash_shift);
> }
>
> Only the trailing ff is in the code output from the splat but the movabs is
> there -> RDX: so:
>
> movabs $0xff2e6dbe0d9b6000,%rdx ; runtime_const_ptr(dentry_hashtable)
> mov %rax, %rbp
> shr $0x7, %eax ; runtime_const_shift_right_32(hashlen, d_hash_shift);
> lea (%rdx, %rax, 8), %rbx ; bucket = &dentry_hashtable[hash >> 7]
>
> (The 8 is multiplying the size of the 8 byte pointers)
>
> Note that rbp retains RAX's value = 0xb654440 (not clobbered elsewehre), so
> the hlist_bl_head bucket pointer is
>
> 0xff2e6dbe0d9b6000 + (0xb654440 >> 7) * 8
>
> So:
>
> b = 0xff2e6dbe0e51b440
>
> This matters, because in hlist_bl_first_rcu() this pointer is treated as a
> valid struct hlist_bl_head pointer:
>
> struct hlist_bl_head {
> struct hlist_bl_node *first;
> };
>
> And the data at 0xff2e6dbe0e51b440 contains 0x0fffffff0c930020 (rbx), which
> is assumed to be a valid struct hlist_bl_node * embedded in a dentry.
>
> IOW, RDX contains the dentry_hashtable pointer:
>
> static struct hlist_bl_head *dentry_hashtable __ro_after_init __used;
>
> That is allocated in dcache_init_early() and never freed:
>
> static void __init dcache_init_early(void)
> {
> ...
> dentry_hashtable =
> alloc_large_system_hash("Dentry cache",
> sizeof(struct hlist_bl_head),
> dhash_entries,
> 13,
> HASH_EARLY | HASH_ZERO,
> &d_hash_shift,
> NULL,
> 2,
> 0);
> ...
> }
>
> So that allocate has to be legit, somehow the data there got corrupted.
>
> This is allocated early by memblock.
>
> d_hash_shift = 7 = 32 - lg(entries), so entries = 25, and 2^25 entries of 8
> bytes each = a 256 MiB table.
>
> So it's legit data that got corrupted.
>
> RDI contains the parent dentry at 0xff2e6d1e4e630d80 (the disassembly at
> __d_lookup() confirms).
>
> KASLR makes things tricky but this is most definitely a slab allocation in
> the direct map and the stack (VMAP_STACK) is at 0xff7532a13699fda0 (rsp)
> ~70 TiB higher which sits at least 10 TiB padded above the direct map so
> it's safe to say that this is in the direct map.
>
> And the corrupted value (rbx) 0x0fffffff0c930020 is interesting - it's an
> exactly a x86-64 swap softleaf value:
>
> __swp_type() = val >> (64 - SWP_TYPE_BITS=5) = val >> 59 = 1
> __swp_offset() = (~(x).val << SWP_TYPE_BITS >> SWP_OFFSET_SHIFT) = (~val << 6 >> 14)
> = 0x79b67f
>
> I.e. it's a swap softleaf entry 0x79b67f 4 KiB pages into the swap = ~30.4 GiB.
>
> It's also the _second_ swap in the system (Nikola - if you have a 2nd swap
> on the reporting system of >=~30 GiB that kinda confirms it).
I suspect this may be a bit of a red herring...
actually there is NO swap on that machine, also there were no linux guests.. so
that might just be a coincidence? not sure if it changes anything..
>
> Also, interstingly, bit 5 is set, which is one of the bits allowed to be
> set (possibly by hardware) in the swap entry.
>
> __swp_entry() clears bits 0-8 and only bits 1-3 are software bits in a
> non-present PTE.
>
> This is _PAGE_ACCESSED (1 << _PAGE_BIT_ACCESSED = 5) = 0x20. So it makes
> sense that hardware might have set it.
>
> In effect - every single bit is exactly how it should be for a valid swap
> entry (since commit 00839ee3b299 ("x86/mm: Move swap offset/type up in PTE
> to work around erratum"))..
>
> It seems more than a coincidence :)
>
> This speaks to some kind of memory corruption that has resulted in a store
> to an arbitrary physical address that happens to be the dentry.
>
> Now looking to the proposed CPA cause (ultimately the thing solved by
>
>
> So the speculated race here is:
>
> pfn_exec: pfn of the module text page CPU A is changing attributes on
> pfn_dh: pfn of the dentry hash table page the swap PTE ends up in
>
> CPU A: set_memory_nx() (text_poke) CPU B: set_memory_rox() (module load)
> ------------------------------------ -----------------------------------------
> __change_page_attr()
> kpte = lookup (lockless)
> <preempted>
> cpa_collapse_large_pages()
> set_pmd(leaf)
> __free_pages(old PTE table)
>
> some process: pte_alloc() gets
> that page as a user page table
>
> UAF!!! Writing into arbitrary memory
> set_pte_atomic(kpte,
> pfn_pte(pfn_exec, prot)) <- lands in that page table:
> entry -> pfn_exec (module text)
>
> process faults on it, GUP follows
> it, later zap frees the data page at
> pfn_exec while it is still live module
> text!!!
>
> That data page is reallocated as a PMD
> table.
>
> text_poke() keeps writing
> code bytes into it -> entries
> with arbitrary pfns, one of them
> pfn_dh
>
> reclaim: try_to_unmap_one()
> pte_offset_map() -> __va(pfn_dh)
> set_pte_at(swap PTE) <- dentry
> hash table
>
> (The reason it needs to be interpreted as a PMD page table is otherwise
> reclaim wouldn't be trying to write a swap PTE entry into it).
>
> And this is exactly what commit 1587d3394e25 ("x86/alternatives: Exclude
> text poking against change_page_attr()") protects against.
>
> Yes it's out on a limb (and I spent FAR TOO LONG going through this
> analysis) but there's really no other sensible explanation as to why dentry
> data got corrupted like that to that exact shape.
>
> Which goes to show how nasty this kind of data corruption issue can be.
>
> >
> > I'll be very very gratefull for any hints here..
>
> As above :)
>
> >
> > with best regards
> >
> > nikola ciprich
> >
> > PS: I tried to CC maintainers of suspected subsystems, but those are just my guesses,
> > so I hope I won't offend anyone.
> >
>
> Don't worry about that, I think we're all happy to get legit bug
> reports. We just might be too busy to reply right away :)
>
> The LLM added on some hints for confirmation of this:
>
> Schlopp>>
>
> Log greps, across all affected hosts and all boots, not just the ones
> that crashed:
>
> grep -i 'Bad page map' /var/log/messages*
> grep -i 'bad pmd' /var/log/messages*
> grep -i 'bad pud' /var/log/messages*
> grep -i 'Bad page state' /var/log/messages*
> grep -i 'CPA: called for zero pte' /var/log/messages*
not a single occurance (this machine uses journal, but I checked those
and no such messages.. in general i tend to check dmesg and system logs
a lot, so I'd have already reported such messages..
>
> Any of these, particularly "bad pmd", is direct evidence that a freed
> kernel PTE table was reused as a user page table. "CPA: called for zero
> pte" would be the CPA walker itself tripping over a collapsed mapping.
>
> Questions:
>
> 1. swapon --show on the host, and inside the guests. Is there a swap
> device with index 1 and a size of at least roughly 30.5GiB? That
> tells us whether the corrupt word is a host swap PTE or a guest one,
> which distinguishes case A from case B above.
>
> 2. From the vmcore, dump the 4K page containing 0xff2e6dbe0e51b440. If
> the neighbouring words are also PTE-shaped, the page was being used
> as a page table and the diagnosis above is confirmed. If only the
> one word is corrupt, it was a single stray store.
unfortunately I don't have full vmcore from that crash, as it didn't fit
to /var/crash, backtrace I posted is from vmcore-dmesg.txt so can't confirm
that..
>
> 3. What are CONFIG_DEBUG_VM, init_on_alloc, init_on_free and
> page_poison set to in the production build versus the KASAN build?
> free_page_is_bad() is gated on is_check_pages_enabled(), which needs
> CONFIG_DEBUG_VM, so the production kernel would not report the bad
> free even if it happened.
I don't have CONFIG_DEBUG_VM enabled in production..
>
> 4. Are any of the crashing guests Windows, and is hv-tlbflush set on
> them? That decides whether 26505e1b5b54 matters for you.
yes, windows, but hv-tlbflush is enabled on a sigle VM and it runs od
different node all the time.
>
> 5. Has any corruption occurred since THP was disabled? If yes, that
> supports the CPA race over your THP theory.
not yet, but it's not happening that often, so unsure here
>
> 6. When you run the 6.18.54-rc1 test, please note that 6.18.52 is only
> a partial fix, so if you have any results from a 6.18.52 kernel they
> should not be treated as a clean run.
sure, I'll start today with 6.18.54, won't consider older tests.
>
> << Schlopp
>
> But really a run against 6.18.53 being OK under heavy testing for several
> days should confirm it also.
>
> If it turns out it's not this then back to the drawing board I guess! Let
> us know.
I surely will!
cheers, nik
>
> --
> Cheers, Lorenzo
>
--
Ing. Nikola CIPRICH
technický ředitel
+420 591 166 214
+420 777 093 799
nikola.ciprich@linuxbox.cz
www.linuxbox.cz
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: hunting memory corruption bug in 6.18.x
2026-09-26 5:58 ` Nikola Ciprich
@ 2026-09-26 9:32 ` Lorenzo Stoakes (ARM)
2026-09-28 9:05 ` Nikola Ciprich
0 siblings, 1 reply; 16+ messages in thread
From: Lorenzo Stoakes (ARM) @ 2026-09-26 9:32 UTC (permalink / raw)
To: Nikola Ciprich
Cc: linux-mm, linux-kernel, akpm, david, Mike Rapoport, Dave Hansen,
Pedro Falcato, Kiryl Shutsemau
[-- Attachment #1: Type: text/plain, Size: 8169 bytes --]
On Sat, Sep 26, 2026 at 07:58:22AM +0200, Nikola Ciprich wrote:
> Hello Lorenzo (and others),
>
> thank you for your time looking into this..
No worries!
> >
> > So there are two sides to the race: set_memory_rox() - triggered on module
> > load, ftrace trampoline creation and every new BPF 2M program pack.
> >
> > The other side is execmem_force_rw() -> set_memory_[nx,rw]() (concurrent
> > module load or ftrace trampoline allocation) __text_poke() ->
> > vmalloc_to_page() for patching module text, kprobe slots, trampolines or
> > BPF.
> >
> > Both are happening a lot at KVM host bringup (module autoload, per-VM
> > seccomp filters, perf, BPF probes, etc.
>
> one note here, at least last mentioned crash (with 6.18.44) happened with
> host running only windows guest, in general we're seeing those problems
> mosly with windows VM hosting machines.. so maybe they're triggerng the
> problem with some other, but similar mechanism?
Interesting! But indeed all this is host-side.
Though it came up with a possible finding around commit 26505e1b5b54 ("KVM:
SVM: make svm_flush_tlb_gva do a full asid flush if NPT enabled") which has
not been backported yet.
That's is a. AMD-specific guest memory corruption and b. only if
hv-tlbflush=on.
This is independent of the CPA stuff.
So if the CPA stuff turns out to be a red herring that's one worth looking
at? Are you able to run a modified kernel with this applied on top?
> > KASLR makes things tricky but this is most definitely a slab allocation in
> > the direct map and the stack (VMAP_STACK) is at 0xff7532a13699fda0 (rsp)
> > ~70 TiB higher which sits at least 10 TiB padded above the direct map so
> > it's safe to say that this is in the direct map.
> >
> > And the corrupted value (rbx) 0x0fffffff0c930020 is interesting - it's an
> > exactly a x86-64 swap softleaf value:
> >
> > __swp_type() = val >> (64 - SWP_TYPE_BITS=5) = val >> 59 = 1
> > __swp_offset() = (~(x).val << SWP_TYPE_BITS >> SWP_OFFSET_SHIFT) = (~val << 6 >> 14)
> > = 0x79b67f
> >
> > I.e. it's a swap softleaf entry 0x79b67f 4 KiB pages into the swap = ~30.4 GiB.
> >
> > It's also the _second_ swap in the system (Nikola - if you have a 2nd swap
> > on the reporting system of >=~30 GiB that kinda confirms it).
>
> I suspect this may be a bit of a red herring...
>
> actually there is NO swap on that machine, also there were no linux guests.. so
> that might just be a coincidence? not sure if it changes anything..
>
Hmm that's really really odd. But if you had swap before or VMs before this
is a long-lasting corruption that could have been sat there for days before
you triggered it.
> > The LLM added on some hints for confirmation of this:
> >
> > Schlopp>>
> >
> > Log greps, across all affected hosts and all boots, not just the ones
> > that crashed:
> >
> > grep -i 'Bad page map' /var/log/messages*
> > grep -i 'bad pmd' /var/log/messages*
> > grep -i 'bad pud' /var/log/messages*
> > grep -i 'Bad page state' /var/log/messages*
> > grep -i 'CPA: called for zero pte' /var/log/messages*
>
> not a single occurance (this machine uses journal, but I checked those
> and no such messages.. in general i tend to check dmesg and system logs
> a lot, so I'd have already reported such messages..
Yeah I don't know why it assumed you used antiquated logging..! :)
OK that's interesting.
>
> >
> > Any of these, particularly "bad pmd", is direct evidence that a freed
> > kernel PTE table was reused as a user page table. "CPA: called for zero
> > pte" would be the CPA walker itself tripping over a collapsed mapping.
> >
> > Questions:
> >
> > 1. swapon --show on the host, and inside the guests. Is there a swap
> > device with index 1 and a size of at least roughly 30.5GiB? That
> > tells us whether the corrupt word is a host swap PTE or a guest one,
> > which distinguishes case A from case B above.
> >
> > 2. From the vmcore, dump the 4K page containing 0xff2e6dbe0e51b440. If
> > the neighbouring words are also PTE-shaped, the page was being used
> > as a page table and the diagnosis above is confirmed. If only the
> > one word is corrupt, it was a single stray store.
> unfortunately I don't have full vmcore from that crash, as it didn't fit
> to /var/crash, backtrace I posted is from vmcore-dmesg.txt so can't confirm
> that..
Ah that's a pity!
>
>
> >
> > 3. What are CONFIG_DEBUG_VM, init_on_alloc, init_on_free and
> > page_poison set to in the production build versus the KASAN build?
> > free_page_is_bad() is gated on is_check_pages_enabled(), which needs
> > CONFIG_DEBUG_VM, so the production kernel would not report the bad
> > free even if it happened.
>
> I don't have CONFIG_DEBUG_VM enabled in production..
Well that explains the lack of bad reports above. I don't know why it'd
assume you'd run kernels with that (we do not recommend that for production
:)
>
> >
> > 4. Are any of the crashing guests Windows, and is hv-tlbflush set on
> > them? That decides whether 26505e1b5b54 matters for you.
> yes, windows, but hv-tlbflush is enabled on a sigle VM and it runs od
> different node all the time.
Ah but that could be enough to cause memory corruption. The reports seem to
be about guest memory corruption though.
To be clear - are you observing it in the guest or host? I gathered host
from the splat.
>
>
> >
> > 5. Has any corruption occurred since THP was disabled? If yes, that
> > supports the CPA race over your THP theory.
> not yet, but it's not happening that often, so unsure here
Yeah it seems to be a hard one to hit. If you tried a kernel with KASAN
enabled it might be flagged earlier? But that could also kill the race
window and would slow the system down a lot.
>
> >
> > 6. When you run the 6.18.54-rc1 test, please note that 6.18.52 is only
> > a partial fix, so if you have any results from a 6.18.52 kernel they
> > should not be treated as a clean run.
> sure, I'll start today with 6.18.54, won't consider older tests.
Ack, that's the best thing to do at the moment to be honest.
If you were consistently getting corruption after X days previously, 2*X
days let's say of none can give confidence it's fixed there.
>
>
> >
> > << Schlopp
> >
> > But really a run against 6.18.53 being OK under heavy testing for several
> > days should confirm it also.
> >
> > If it turns out it's not this then back to the drawing board I guess! Let
> > us know.
> I surely will!
>
> cheers, nik
Thanks! Given the nature of the bug and the fact the LLM went a little out
on a limb.
Some more stuff from the report, which I also enclose in full here FYI.
schlopp>>
26505e1b5b54 ("KVM: SVM: make svm_flush_tlb_gva do a full asid flush if
NPT enabled") is mainline only and has not been backported to 6.18.y.
It is AMD-specific guest memory corruption via the Hyper-V PV TLB flush
path, so it only affects Windows guests running with hv-tlbflush=on. If
any of your crashing guests are Windows, this is worth backporting
separately. It is independent of the CPA race.
55ddbc2ca6d5 (upstream f7491d7c81db, "x86/mm: Fix user-space data loss
with MADV_FREE and THP") is a real bug in 6.18.44. pmd_modify() drops
_PAGE_DIRTY, which has been the case since v6.10, and it causes data
loss for MADV_FREE THP and writable file THP. It is fixed in 6.18.53.
It is data loss rather than a stray write, so it does not explain the
oops, but it is another reason to move off 6.18.44. Note that this one
needs THP, which you have now disabled.
d053eb7e09e1 (upstream 1e75a8255f11, "iommu/amd: Wait for completion
instead of returning early in iommu_completion_wait()") is already in
6.18.42 and therefore in the kernel that crashed. It is ruled out for
this oops, but it may be relevant to the earlier incidents you had on
6.18.31 through 6.18.41.
<<schlopp
Let us know how the tests get on! If you trigger a bug on 54 let us know
ASAP so we can investigate alternative theories.
Thanks!
>
>
>
>
> >
> > --
> > Cheers, Lorenzo
> >
>
> --
> Ing. Nikola CIPRICH
> technický ředitel
>
> +420 591 166 214
> +420 777 093 799
> nikola.ciprich@linuxbox.cz
>
> www.linuxbox.cz
--
Cheers, Lorenzo
[-- Attachment #2: debug-report.txt --]
[-- Type: text/plain, Size: 23666 bytes --]
Summary
=======
The corruption you are seeing is consistent with a known use-after-free
in the x86 change_page_attr (CPA) code that is present in 6.18.44 and
was only fixed in 6.18.52 and 6.18.53.
cpa_collapse_large_pages() rebuilds a leaf PMD out of its 4K PTEs and
then frees the old PTE table, with no lock held against the lockless
page table walk that __change_page_attr() performs before it stores
through the PTE pointer it cached. The stale 8-byte store of a kernel
PTE value lands in whatever the buddy allocator has since handed that
page out for.
On a KVM host the two sides of this race are both hot. set_memory_rox()
is the only caller that passes CPA_COLLAPSE, and it runs on every module
load (execmem_restore_rox()), every ftrace trampoline creation
(arch/x86/kernel/ftrace.c:423) and every new 2M BPF program pack. The
other side is any lockless walk of the same execmem tables:
set_memory_nx()/set_memory_rw() from execmem_force_rw() on module load
and trampoline allocation, and vmalloc_to_page() inside __text_poke()
for every patch of module text, kprobe slot, trampoline or BPF pack.
Module text, kprobe slots and ftrace trampolines share the same 2M ROX
cache pages, so the collapser and the victim land in the same PMD by
construction. A libvirt host does all of this constantly: module
autoload (tun, vhost_net, br_netfilter, ebtable_*, xt_*), per-VM seccomp
filters, systemd cgroup BPF, perf and BPF probes, ftrace and static call
patching on every module load.
This matches your good/bad window exactly. The collapse feature was
added in v6.15 and is not in 5.15.
It also matches the KASAN result. free_page_is_bad() is gated on
is_check_pages_enabled(), which is off unless CONFIG_DEBUG_VM is set,
and KASAN changes the allocation pattern enough that the freed page
tends not to be reused in the race window. The upstream reporter only
reproduced it by injecting a delay at the CPA page table lookup.
Recommendation: run 6.18.53 or later. 6.18.54-rc1, which you say you
are about to test, has the complete series. 6.18.52 has only the
cpa_lock patch, which does not close the window. Mainline (v7.3-rc5)
has nothing further pending for arch/x86/mm/pat/set_memory.c.
Kernel version
==============
6.18.44lb9.01 (AlmaLinux 9 build of stable 6.18.44, stable commit
1efe5d048a39). First seen on 6.18.31, last crash on 6.18.44. 5.15.x
was fine.
Machine
=======
ASUSTeK RS720A-E12-RS12 / K14PP-D24, AMD EPYC, BIOS 2305 11/21/2025.
KVM host, AlmaLinux 9, Intel ice and Mellanox NICs. Taint "G E",
unsigned module only.
Stack trace
===========
Oops: general protection fault, probably for non-canonical address
0xfffffff0c930038: 0000 [#1] SMP NOPTI
CPU: 23 UID: 189 PID: 7538 Comm: pacemaker-contr
RIP: 0010:__d_lookup+0x4a/0xc0
RAX: 0fffffff0c930020 RBX: 0fffffff0c930020 RCX: 0000000000000000
RDX: ff2e6dbe0d9b6000 RSI: ff7532a13699fe60 RDI: ff2e6d1e4e630d80
RBP: 000000000b654440
Call Trace:
<TASK>
d_lookup+0x27/0x50
lookup_dcache+0x1f/0x80
lookup_one_qstr_excl+0x1e/0xe0
filename_create+0xc4/0x160
do_mkdirat+0x5a/0x190
__x64_sys_mkdir+0x42/0x60
do_syscall_64+0x64/0xbf0
entry_SYSCALL_64_after_hwframe+0x76/0x7e
</TASK>
Other messages you reported:
Inconsistency detected by ld.so: ../sysdeps/x86_64/dl-machine.h: 548:
elf_machine_rela_relative: Assertion `ELFW(R_TYPE) (reloc->r_info) ==
R_X86_64_RELATIVE' failed!
What the oops registers say
===========================
The Code bytes decode to the hash chain walk in fs/dcache.c:__d_lookup():
mov (%rbx),%rax ; load bucket->first
mov %rax,%rbx
and $-2,%rbx ; strip the hlist_bl lock bit
cmp $1,%rax
ja body
loop:
mov (%rbx),%rbx ; node->next
test %rbx,%rbx
je out
body:
cmp %ebp,0x18(%rbx) ; <-- faulting insn, d_name.hash_len
jne loop
RAX equals RBX and RAX is only ever written by the initial bucket load,
so this is the first loop iteration. The corrupt word is the
hlist_bl_head.first field of the dentry_hashtable bucket itself, not a
dentry's d_hash.next.
RDX 0xff2e6dbe0d9b6000 is the dentry_hashtable base, and the runtime
constant d_hash_shift is patched to 7, so:
bucket = 0xff2e6dbe0d9b6000 + (0x0b654440 >> 7) * 8
= 0xff2e6dbe0e51b440
That table is a 256MB alloc_large_system_hash() allocation from
memblock. It is allocated at boot, is PG_reserved and is never freed.
So this is a stray write to a fixed physical page, not a
use-after-free of a recycled object.
The corrupt value 0x0fffffff0c930020 is an exact x86-64 non-present
swap PTE: __swp_type() is 1 and __swp_offset() is 0x79b67f, about
30.4GiB into swap device 1. The only low bit set is bit 5,
_PAGE_ACCESSED.
Every bit the swap layout constrains is as it should be: P, PSE and
the G/PROT_NONE alias are clear, the software bits 1-3 are clear, the
type is an ordinary swap type and the inverted offset gives the long
run of ones in bits 32-58. The one thing the layout does not account
for is bit 5. __swp_entry() leaves bits 0-8 clear and no helper sets
bit 5 on a non-present PTE; arch/x86/include/asm/pgtable_64.h treats
bits 5 and 6 as don't-care only because of the Intel Knights Landing
erratum (00839ee3b299, _PAGE_KNL_ERRATUM_MASK), which does not apply to
EPYC. So either something other than a Linux swap PTE happens to fit
this layout, or the word was a swap PTE that acquired a stray bit. I
cannot tell which from one word, which is why the vmcore page dump
requested below matters: 511 neighbouring PTE-shaped words would settle
it.
Suspect commit
==============
commit 41d88484c71cd4f659348da41b7b5b3dbd3be1f6
Author: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
x86/mm/pat: restore large ROX pages after fragmentation
Link: https://lore.kernel.org/r/20250126074733.1384926-4-rppt@kernel.org
This added runtime collapse of split kernel large pages, driven from
cpa_flush(), including freeing the PTE table that the collapsed PMD
replaces. It is the Fixes: target of every fix listed below. It is in
v6.15 and later, and is not in 5.15.
> diff --git a/arch/x86/mm/pat/set_memory.c b/arch/x86/mm/pat/set_memory.c
> --- a/arch/x86/mm/pat/set_memory.c
> +++ b/arch/x86/mm/pat/set_memory.c
> @@ -394,6 +408,40 @@ static void __cpa_flush_tlb(void *data)
> +static void cpa_collapse_large_pages(struct cpa_data *cpa)
> +{
> + unsigned long start, addr, end;
> + struct ptdesc *ptdesc, *tmp;
> + LIST_HEAD(pgtables);
> + int collapsed = 0;
> + int i;
[ ... range iteration ... ]
> + if (!collapsed)
> + return;
> +
> + flush_tlb_all();
> +
> + list_for_each_entry_safe(ptdesc, tmp, &pgtables, pt_list) {
> + list_del(&ptdesc->pt_list);
> + __free_page(ptdesc_page(ptdesc));
> + }
> +}
In 6.18.44 that last loop is pagetable_free(ptdesc), which is still an
immediate free. There is no RCU grace period and no other deferral. The
flush_tlb_all() above it only makes the hardware forget the old
translation; it does nothing about a CPU that is sitting inside
__change_page_attr() holding a pointer into that table.
> @@ -402,7 +450,7 @@ static void cpa_flush(struct cpa_data *cpa, int cache)
> if (cache && !static_cpu_has(X86_FEATURE_CLFLUSH)) {
> cpa_flush_all(cache);
> - return;
> + goto collapse_large_pages;
> }
[ ... ]
> @@ -427,6 +475,10 @@ static void cpa_flush(struct cpa_data *cpa, int cache)
> mb();
> +
> +collapse_large_pages:
> + if (cpa->flags & CPA_COLLAPSE)
> + cpa_collapse_large_pages(cpa);
> }
The collapse is hooked into cpa_flush(), which
__change_page_attr_set_clr() calls after it has already dropped
cpa_lock. So cpa_lock does not serialise the collapse against anything.
> @@ -1196,6 +1248,161 @@ static int split_large_page(struct cpa_data *cpa, pte_t *kpte,
> +static int collapse_pmd_page(pmd_t *pmd, unsigned long addr,
> + struct list_head *pgtables)
> +{
[ ... uniformity checks over all 512 PTEs ... ]
> + old_pmd = *pmd;
> +
> + /* Success: set up a large page */
> + pgprot = pgprot_4k_2_large(pte_pgprot(first));
> + pgprot_val(pgprot) |= _PAGE_PSE;
> + _pmd = pfn_pmd(pfn, pgprot);
> + set_pmd(pmd, _pmd);
> +
> + /* Queue the page table to be freed after TLB flush */
> + list_add(&page_ptdesc(pmd_page(old_pmd))->pt_list, pgtables);
collapse_large_pages(), the caller, takes pgd_lock around this. The
lockless CPA walker never takes pgd_lock, so pgd_lock does not help
either.
The other side, in 6.18.44:
arch/x86/mm/pat/set_memory.c:__change_page_attr() {
address = __cpa_addr(cpa, cpa->curpage);
repeat:
kpte = _lookup_address_cpa(cpa, address, &level, &nx, &rw);
...
old_pte = *kpte;
...
if (level == PG_LEVEL_4K) {
...
new_pte = pfn_pte(pfn, new_prot);
...
if (pte_val(old_pte) != pte_val(new_pte)) {
set_pte_atomic(kpte, new_pte); <-- stale
cpa->flags |= CPA_FLUSHTLB;
}
_lookup_address_cpa() reaches lookup_address_in_pgd_attr(), which is a
plain lockless walk. Between the walk and the set_pte_atomic() the
caller can be preempted or take an interrupt; this runs with interrupts
on. That store is the only unsafe instruction in the function. The
large-page split branch further down is safe because __split_large_page()
revalidates under pgd_lock.
The freed table is not even a tracked page table page in 6.18.44:
arch/x86/mm/pat/set_memory.c:split_large_page() {
if (!debug_pagealloc_enabled())
spin_unlock(&cpa_lock);
base = alloc_pages(GFP_KERNEL, 0);
A bare alloc_pages(), so it goes straight back to the per-CPU free list
and can be reallocated immediately.
Race timeline
=============
CPU A (text_poke -> CPU B (module_enable_rox /
execmem_make_temp_rw -> execmem_restore_rox /
set_memory_nx/rw) bpf_jit_binary_lock_ro ->
set_memory_rox, CPA_COLLAPSE)
----- -----
__change_page_attr()
kpte = _lookup_address_cpa()
old_pte = *kpte
new_pte = pfn_pte(...)
preempted / interrupted
__change_page_attr_set_clr()
drops cpa_lock
cpa_flush()
cpa_collapse_large_pages()
collapse_pmd_page(): all 512
PTEs uniform, set_pmd() installs
a leaf, old PTE table queued
flush_tlb_all()
pagetable_free() -> immediate
__free_pages()
(any CPU) page is reallocated:
.so page cache folio, QEMU guest
RAM, a user PMD/PTE table, slab
set_pte_atomic(kpte, new_pte)
stores a PTE-shaped word into
the reallocated page
Where the swap PTE comes from (inferred continuation)
------------------------------------------------------
The race above writes a present kernel PTE, never a swap entry. To
reach the dentry hash table with a swap-PTE-shaped word the following
has to happen next. Each step is verified in the 6.18.44 code; the
sequence as a whole is inferred, not proven for this oops.
1. The freed PTE table is reallocated as a QEMU page table.
2. The stale set_pte_atomic() lands in it. The injected entry is a
translation into an execmem text page (case A/B below).
3. Case B: GUP-slow follows that entry and KVM maps the text page
into the guest; a later zap_present_folio_ptes() does an
unbalanced folio_put() and frees the still-live text page.
4. That text page is reallocated as another page table while
text_poke()/the BPF JIT keep writing instruction bytes into it
through the ROX mapping. Instruction bytes are now PMD entries
with arbitrary pfns; some pass pmd_bad().
5. Reclaim: try_to_unmap_one() -> page_vma_mapped_walk() ->
pte_offset_map_lock() reads such a PMD, computes
__va(garbage pfn) as the PTE table, and set_pte_at() stores a
swap PTE there. If that pfn is the dentry_hashtable page, one
bucket head becomes 0x0fffffff0c930020-like.
6. Days later __d_lookup() hashes into that bucket and faults.
Step 5 is the only writer of an ordinary swap type in the mm, and
the only step that can touch memory the allocator never owned.
This is not speculation about the code. The same interleaving was
reported upstream with a KASAN reproducer, in the commit that first
tried to address it:
commit 1aac65f3e651 ("x86/mm/pat: Take cpa_lock around large-page
collapse")
BUG: KASAN: use-after-free in __change_page_attr+0x7cc/0x7e0
Write of size 8 at addr ffff888181139718 by task modprobe
...
The buggy address belongs to the physical page:
pfn:0x181139 ... page_type: f2(table)
Fixes: 41d88484c71c ("x86/mm/pat: restore large ROX pages after
fragmentation")
Signed-off-by: Denis V. Lunev <den@openvz.org>
Which stable releases carry the fixes
=====================================
None of these are in 6.18.44.
The fixes that matter are the init_mm mmap lock pair, a1c7570cedd0 and
d5d8b8662e6e: the collapse runs under the init_mm write lock and the
whole attribute change, including the lockless walk and the store
through the cached pointer, runs under the read lock. That excludes
both walker-vs-collapse and collapse-vs-collapse. 1587d3394e25 covers
the third walker: __text_poke() resolves the pages it patches with
vmalloc_to_page(), a lockless walk of the same execmem tables, and now
takes the init_mm read lock around it. Without that, a collapse under
a concurrent text_poke() returns NULL (the BUG_ON at
arch/x86/kernel/alternative.c:2576 that openSUSE hit) or, if the freed
table has already been reused, a wrong page that text_poke() then
writes instruction bytes into. 9e4a3ec3411b makes the split tables
real kernel page tables so their freeing is deferred. The earlier
cpa_lock patch, 1aac65f3e651, is superseded by the write lock and adds
nothing once those are applied.
6.18.52
591b6fac9df3 x86/mm/pat: Take cpa_lock around large-page collapse
(upstream 1aac65f3e651)
6.18.53
35820cf8dd52 x86/mm/pat: Acquire init_mm write lock on collapse to
avoid UAF (upstream a1c7570cedd0)
e21a9ea81426 x86/mm/pat: Acquire init_mm read lock on attribute
changes to avoid UAF (upstream d5d8b8662e6e)
e164f4a25e23 x86/mm/pat: Convert split_large_page() to use ptdescs
5029589bb773 x86/mm/pat: Don't gate cpa_lock on
debug_pagealloc_enabled()
84e0cd79d57f x86/mm/pat: Allocate split page tables as kernel page
tables (upstream 9e4a3ec3411b)
281e6f536f2f x86/alternatives: Exclude text poking against
change_page_attr() (upstream 1587d3394e25)
74a2626044de x86/mm: Fix and document DEBUG_PAGEALLOC
(upstream 7da514d819a0)
6.18.52 does not fix this: it only takes cpa_lock around the collapse,
and the walker never holds cpa_lock across its walk-then-store window.
The init_mm mmap lock pair that closes that window is in 6.18.53, which
is the first stable release with the complete set. Use 6.18.53 or
later.
There is no clean interim mitigation on 6.18.44. CPA_COLLAPSE cannot be
disabled by a boot parameter.
How this reaches the symptoms you saw
=====================================
Symptoms 1 and 2 follow directly. The stale store writes a PTE-shaped
qword into a page that has been reallocated. If that page is a page
cache folio for a mapped .so, eight bytes of its relocation table are
replaced and ld.so trips the R_X86_64_RELATIVE assertion while the file
on disk is intact. If it is an anonymous page that QEMU has just
populated as guest RAM on a migration destination, the guest sees eight
corrupt bytes. Both of these match "right after migration": the
destination host is populating gigabytes of guest RAM and allocating
page tables at maximum rate, which is exactly when a just-freed page
gets reused inside the race window.
Symptom 3, this oops, needs one more step, because the dentry hash table
is memblock memory that is never freed and therefore cannot be the
directly reallocated page. The escalation is that the victim page is
itself a page table. Each of the following links is verified in the
6.18.44 source, but I want to be clear that the end-to-end chain for
this particular oops is plausible rather than proven from a single
vmcore.
Case A, the victim is a user PMD table. pmd_bad() is true for the
injected value, so the first user-mode touch faults and
mm/pgtable-generic.c:___pte_offset_map() heals it via pmd_clear_bad()
with a "bad pmd" pr_err. But a supervisor-side copy_to_user() is
permitted through a U=0 level. A read() or recvmsg() into guest RAM, or
vhost-net, or kvm_write_guest(), then has the hardware walker read a
qword of module text as a PTE. If that qword happens to have P and RW
set, the copied data is written to an arbitrary physical address. This
route leaves no taint.
Case B, the victim is a user PTE table. x86 has no pte_bad(). GUP-fast
rejects a U=0 entry, but GUP-slow does not:
mm/gup.c:follow_page_pte() and mm/memory.c:__vm_normal_page() check
neither _PAGE_USER nor PageReserved, so a live execmem page is returned
and KVM maps kernel module text into a guest. A later
mm/memory.c:zap_present_folio_ptes() does folio_remove_rmap_ptes() and
folio_put() on a page that was never rmapped, dropping the refcount to
zero, printing "BUG: Bad page map" and releasing live ROX text into the
buddy allocator while text_poke() and the BPF JIT keep writing
instruction bytes into it. Recycled as a user PMD table, instruction
bytes are page table entries with arbitrary pfns that pass pmd_bad(),
pte_offset_map() computes __va() of an arbitrary physical page, and
try_to_unmap_one() writes a genuine host swap PTE into it.
That last step is what the corrupt word looks like: a real host swap
PTE. The stray bit 5 is unexplained on this route too.
One caveat that argues against case B on this specific host: the taint
is "G E" with no "B". "BUG: Bad page map" had not fired on that
machine before the crash. So either the taint-free case A route or a
direct stray write carries this particular chain, or the corrupting
event happened on a different boot. This limits, but does not refute,
the cascade. The corruption can sit in a rarely used dentry bucket for
a long time before something hashes into it, which fits the 22-day
uptime on this crash and your observation that the host crashes were
not correlated with migration.
Secondary findings
==================
These came up while looking and are worth knowing about, but none of
them explains this oops.
26505e1b5b54 ("KVM: SVM: make svm_flush_tlb_gva do a full asid flush if
NPT enabled") is mainline only and has not been backported to 6.18.y.
It is AMD-specific guest memory corruption via the Hyper-V PV TLB flush
path, so it only affects Windows guests running with hv-tlbflush=on. If
any of your crashing guests are Windows, this is worth backporting
separately. It is independent of the CPA race.
55ddbc2ca6d5 (upstream f7491d7c81db, "x86/mm: Fix user-space data loss
with MADV_FREE and THP") is a real bug in 6.18.44. pmd_modify() drops
_PAGE_DIRTY, which has been the case since v6.10, and it causes data
loss for MADV_FREE THP and writable file THP. It is fixed in 6.18.53.
It is data loss rather than a stray write, so it does not explain the
oops, but it is another reason to move off 6.18.44. Note that this one
needs THP, which you have now disabled.
d053eb7e09e1 (upstream 1e75a8255f11, "iommu/amd: Wait for completion
instead of returning early in iommu_completion_wait()") is already in
6.18.42 and therefore in the kernel that crashed. It is ruled out for
this oops, but it may be relevant to the earlier incidents you had on
6.18.31 through 6.18.41.
Theories that were eliminated
=============================
KVM NPT mapping at too large a level, or with the wrong base pfn. The
mapping level comes from the host page tables and the pfn from GUP;
KVM cannot reach memblock memory on its own.
A host mm swap or migration PTE stored through a stale page table
pointer. All the store sites are bounded and the pointer provenance
checks out.
A missed MMU notifier invalidation. Notifier ordering on the recovery
path is correct, and this cannot reach never-freed memory.
NIC DMA to the wrong address. Only a teardown-time page_pool
use-after-free turned up, and the iommu/amd completion-wait fix is
already in 6.18.42.
105d04edbec8 (upstream 33192a26cddea, mm/huge_memory huge_zero_pfn
race). Real, but not present in 6.18.44, and it cannot reach memblock
memory.
0a25ee42e7d1 (upstream 3d679b7cb31f, KVM x86/mmu CMPXCHG when clearing
the Accessed bit in the TDP MMU). Not in 6.18.44, and benign for this.
TDP MMU in-place huge page recovery. Structurally excluded: notifier
zaps take mmu_lock for write, recovery takes it for read.
memblock/buddy physical aliasing. This would produce "Bad page state"
reports, which you have not seen. Worth confirming from the vmcore, see
below.
Neither THP nor NUMA balancing is involved in the CPA race. Disabling
them was a reasonable precaution but it will not stop this. If you keep
seeing corruption with THP off, that is consistent with the diagnosis
rather than against it.
What would confirm this
=======================
Log greps, across all affected hosts and all boots, not just the ones
that crashed:
grep -i 'Bad page map' /var/log/messages*
grep -i 'bad pmd' /var/log/messages*
grep -i 'bad pud' /var/log/messages*
grep -i 'Bad page state' /var/log/messages*
grep -i 'CPA: called for zero pte' /var/log/messages*
Any of these, particularly "bad pmd", is direct evidence that a freed
kernel PTE table was reused as a user page table. "CPA: called for zero
pte" would be the CPA walker itself tripping over a collapsed mapping.
Questions:
1. swapon --show on the host, and inside the guests. Is there a swap
device with index 1 and a size of at least roughly 30.5GiB? That
tells us whether the corrupt word is a host swap PTE or a guest one,
which distinguishes case A from case B above.
2. From the vmcore, dump the 4K page containing 0xff2e6dbe0e51b440. If
the neighbouring words are also PTE-shaped, the page was being used
as a page table and the diagnosis above is confirmed. If only the
one word is corrupt, it was a single stray store.
3. What are CONFIG_DEBUG_VM, init_on_alloc, init_on_free and
page_poison set to in the production build versus the KASAN build?
free_page_is_bad() is gated on is_check_pages_enabled(), which needs
CONFIG_DEBUG_VM, so the production kernel would not report the bad
free even if it happened.
4. Are any of the crashing guests Windows, and is hv-tlbflush set on
them? That decides whether 26505e1b5b54 matters for you.
5. Has any corruption occurred since THP was disabled? If yes, that
supports the CPA race over your THP theory.
6. When you run the 6.18.54-rc1 test, please note that 6.18.52 is only
a partial fix, so if you have any results from a 6.18.52 kernel they
should not be treated as a clean run.
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: hunting memory corruption bug in 6.18.x
2026-09-25 8:48 hunting memory corruption bug in 6.18.x Nikola Ciprich
2026-09-25 10:05 ` Lorenzo Stoakes (ARM)
2026-09-25 12:13 ` Lorenzo Stoakes (ARM)
@ 2026-09-26 16:02 ` Luiz Capitulino
2026-09-28 8:47 ` Nikola Ciprich
2 siblings, 1 reply; 16+ messages in thread
From: Luiz Capitulino @ 2026-09-26 16:02 UTC (permalink / raw)
To: Nikola Ciprich, linux-mm; +Cc: linux-kernel, akpm, david, ljs
On 9/25/26 4:48 AM, Nikola Ciprich wrote:
> Hi,
>
> I've been hunting a weird memory corruption bug for the last few weeks,
> without success so far, so I'd like to report it and kindly ask for help.
>
> We first hit it after a live VM migration between two KVM hosts:
> suddenly some dynamic libraries in the host appeared to be corrupted:
>
> Inconsistency detected by ld.so: ../sysdeps/x86_64/dl-machine.h: 548: elf_machine_rela_relative: Assertion `ELFW(R_TYPE) (reloc->r_info) == R_X86_64_RELATIVE' failed!
>
> (Later we also hit this with libcrypto.so.3 etc.) The files on disk
> were OK; the problem seemed to exist only in RAM.
>
> I'm fairly sure this is not hardware related: there were no ECC errors,
> and we have since hit this (and similar issues, more on that below) on
> multiple machines.
>
> The problems started after we moved from 5.15.x to 6.18.x kernels.
How long does it take to reproduce? Can you reliably distinguish good
from bad?
I know that Lorenzo jumped in and gave some good suggestions already,
but in case you still find yourself without any further options you
could consider if bisection is feasible: start with manual bisection
to identify the first bad kernel between v5.15 and v6.18 and then the
first bad -rc. You could go to git bisect from here, but it may take
several weeks depending on how long it takes to reproduce.
Another option is to try latest Linus tree to see if the issue is there.
If it's not there then it might have been fixed, in this case you could
bisect for the fix (if feasible, of course).
>
> Since then I've spent a lot of time trying to reproduce it on a lab
> cluster, and we were able to trigger some corruption after days of
> migrating VMs back and forth. At first I suspected the Intel ice driver,
> for which I found similar reports, but we saw new problems even after
> backporting fixes (and also with Mellanox cards).
>
> So far we've hit three different kinds of problems, which may or may
> not be related:
>
> - .so library corruption right after VM migration
> - VM crashes (or process crashes inside VMs), possibly related to
> migration (those always happened during migration)
> - host crashes due to kernel structure corruption (these happened
> without any VM migration)
>
> We first hit these problems with 6.18.31; the last crash I saw was
> with 6.18.44.
>
> All affected machines use AMD EPYC CPUs and act as KVM hosts; the OS
> is AlmaLinux 9.
>
> I suspect two subsystems that have seen a lot of changes:
>
> - transparent hugepages
> - NUMA balancing
>
> (but those are just my guesses)
>
> As a safety measure, we've disabled THP and NUMA balancing on all hosts.
>
> I'm aware this is still a very vague report with a lot of guessing,
> but my question is: has anybody hit similar problems with 6.18 or
> newer kernels?
>
> I see a lot of patches in every stable release, but simply trying
> newer kernels doesn't seem efficient here. Deploying them is also
> risky, since the hosts have to be emptied by migrating VMs off them
> before reboot, and that migration itself may trigger more crashes.
> None of the released or queued fixes for 6.18 seem to be directly
> related.
>
> I tried running my migration tests on hosts with KASAN enabled, and
> also with SLUB debugging, but was never able to reproduce the problem
> with those enabled (without them, I was able to hit issues within
> days).
>
> I'll start another round of migration tests in the lab, now with
> 6.18.54-rc1, but I still thought it would be good to report this and
> ask here.
>
> last but not least, here's kdump from last crash (this was not related
> to any VM migration, but is very similar to another few crashes
> we got):
>
> [1924553.414736] Oops: general protection fault, probably for non-canonical address 0xfffffff0c930038: 0000 [#1] SMP NOPTI
> [1924553.434800] CPU: 23 UID: 189 PID: 7538 Comm: pacemaker-contr Kdump: loaded Tainted: G E 6.18.44lb9.01 #1 PREEMPT(voluntary)
> [1924553.456934] Tainted: [E]=UNSIGNED_MODULE
> [1924553.465551] Hardware name: ASUSTeK COMPUTER INC. RS720A-E12-RS12/K14PP-D24 Series, BIOS 2305 11/21/2025
> [1924553.484152] RIP: 0010:__d_lookup+0x4a/0xc0
> [1924553.492878] Code: ff 48 89 c5 c1 e8 07 48 8d 1c c2 e8 60 8f d1 ff 48 8b 03 48 89 c3 48 83 e3 fe 48 83 f8 01 77 0a eb 2f 48 8b 1b 48 85 db 74 27 <39> 6b 18 75 f3 4c 8d 63 78 4c 89 e7 e8
> d5 e1 7c 00 4c 39 6b 10 74
> [1924553.525191] RSP: 0018:ff7532a13699fda0 EFLAGS: 00010212
> [1924553.534986] RAX: 0fffffff0c930020 RBX: 0fffffff0c930020 RCX: 0000000000000000
> [1924553.546679] RDX: ff2e6dbe0d9b6000 RSI: ff7532a13699fe60 RDI: ff2e6d1e4e630d80
> [1924553.558367] RBP: 000000000b654440 R08: 0000000000002403 R09: 0000000000000179
> [1924553.570026] R10: 000000000000000d R11: 0000000000000000 R12: 0000000001876e5c
> [1924553.581601] R13: ff2e6d1e4e630d80 R14: ff7532a13699fe60 R15: 0000000000000000
> [1924553.593115] FS: 00007ff8743aaa80(0000) GS:ff2e6e5e94c45000(0000) knlGS:0000000000000000
> [1924553.605576] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
> [1924553.615611] CR2: 00007ffcd63a5000 CR3: 00000003ee840001 CR4: 0000000000771ef0
> [1924553.627022] PKRU: 55555554
> [1924553.633869] Call Trace:
> [1924553.640353] <TASK>
> [1924553.646369] d_lookup+0x27/0x50
> [1924553.653366] lookup_dcache+0x1f/0x80
> [1924553.660713] lookup_one_qstr_excl+0x1e/0xe0
> [1924553.668589] ? preempt_schedule_common+0x2c/0x70
> [1924553.676837] filename_create+0xc4/0x160
> [1924553.684209] do_mkdirat+0x5a/0x190
> [1924553.691050] __x64_sys_mkdir+0x42/0x60
> [1924553.698163] do_syscall_64+0x64/0xbf0
> [1924553.705145] entry_SYSCALL_64_after_hwframe+0x76/0x7e
> [1924553.713533] RIP: 0033:0x7ff8754ff08b
> [1924553.720358] Code: 8b 05 91 bd 0f 00 41 bc ff ff ff ff 64 c7 00 16 00 00 00 e9 4f ff ff ff e8 12 f7 01 00 66 90 f3 0f 1e fa b8 53 00 00 00 0f 05 <48> 3d 01 f0 ff ff 73 01 c3 48 8b 0d 5d
> bd 0f 00 f7 d8 64 89 01 48
> [1924553.748552] RSP: 002b:00007ffc05e5c148 EFLAGS: 00000246 ORIG_RAX: 0000000000000053
> [1924553.759401] RAX: ffffffffffffffda RBX: 00005623cc0bd513 RCX: 00007ff8754ff08b
> [1924553.769760] RDX: 000000000fde421b RSI: 00000000000001c0 RDI: 00005623cc0bd4f4
> [1924553.780083] RBP: f49998db0aa753ff R08: 0000000000000004 R09: 0000000000000001
> [1924553.790348] R10: 00007ff87587d000 R11: 0000000000000246 R12: 8421084210842109
> [1924553.800604] R13: 00005623cc0bd513 R14: 00007ff8755bd740 R15: 000000000fde421b
> [1924553.810819] </TASK>
>
> I'll be very very gratefull for any hints here..
>
> with best regards
>
> nikola ciprich
>
> PS: I tried to CC maintainers of suspected subsystems, but those are just my guesses,
> so I hope I won't offend anyone.
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: hunting memory corruption bug in 6.18.x
2026-09-26 16:02 ` Luiz Capitulino
@ 2026-09-28 8:47 ` Nikola Ciprich
0 siblings, 0 replies; 16+ messages in thread
From: Nikola Ciprich @ 2026-09-28 8:47 UTC (permalink / raw)
To: Luiz Capitulino; +Cc: linux-mm, linux-kernel, akpm, david, ljs, Nikola Ciprich
Hi Luiz,
> > The problems started after we moved from 5.15.x to 6.18.x kernels.
>
> How long does it take to reproduce? Can you reliably distinguish good
> from bad?
Unfortunately it is painfully hard to reproduce and thus almost impossible
to bisect. A few times, after more than a week of successful tests, I
deployed a "fixed" kernel to production... and got another crash after
three weeks :(
If I were able to reproduce it more easily, bisecting would probably be
the first thing I'd try, but I still haven't found an easy way to trigger
it.
I tried heavily loaded guests running MSSQL being hammered by HammerDB
(one of the affected customers runs lots of Windows guests with MSSQL),
and others running kernel builds in a loop on top of a tmpfs ramdisk,
all of them being migrated back and forth.
>
> I know that Lorenzo jumped in and gave some good suggestions already,
> but in case you still find yourself without any further options you
> could consider if bisection is feasible: start with manual bisection
> to identify the first bad kernel between v5.15 and v6.18 and then the
> first bad -rc. You could go to git bisect from here, but it may take
> several weeks depending on how long it takes to reproduce.
>
> Another option is to try latest Linus tree to see if the issue is there.
> If it's not there then it might have been fixed, in this case you could
> bisect for the fix (if feasible, of course).
>
Yes, both approaches would be feasible if I were able to reproduce it
more easily :(
So for now I'm running another round of tests on 6.18.54, and I guess
I'll deploy it to the affected production clusters. That's still better
than waiting for a crash on the older release.
I'll report back once I have something new (from the lab or production).
cheers
nik
--
Ing. Nikola CIPRICH
technický ředitel
+420 591 166 214
+420 777 093 799
nikola.ciprich@linuxbox.cz
www.linuxbox.cz
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: hunting memory corruption bug in 6.18.x
2026-09-26 9:32 ` Lorenzo Stoakes (ARM)
@ 2026-09-28 9:05 ` Nikola Ciprich
2026-09-29 19:11 ` Nikola Ciprich
0 siblings, 1 reply; 16+ messages in thread
From: Nikola Ciprich @ 2026-09-28 9:05 UTC (permalink / raw)
To: Lorenzo Stoakes (ARM)
Cc: linux-mm, linux-kernel, akpm, david, Mike Rapoport, Dave Hansen,
Pedro Falcato, Kiryl Shutsemau, Nikola Ciprich
> > one note here, at least last mentioned crash (with 6.18.44) happened with
> > host running only windows guest, in general we're seeing those problems
> > mosly with windows VM hosting machines.. so maybe they're triggerng the
> > problem with some other, but similar mechanism?
>
> Interesting! But indeed all this is host-side.
>
> Though it came up with a possible finding around commit 26505e1b5b54 ("KVM:
> SVM: make svm_flush_tlb_gva do a full asid flush if NPT enabled") which has
> not been backported yet.
>
> That's is a. AMD-specific guest memory corruption and b. only if
> hv-tlbflush=on.
>
> This is independent of the CPA stuff.
>
> So if the CPA stuff turns out to be a red herring that's one worth looking
> at? Are you able to run a modified kernel with this applied on top?
hmm, this one looks quite promising.
the patch doesn't apply cleanly on top of latest 6.18, I'll have a look at it
however, I double checked, there is single guest with hv-tlbflush running on
affected cluster and it never run on recently crashed node. so unless this
fixes also some different case, it's probably not the bug we're hunting here..
(but I won't be surprised, if all this is caused with multiple independent
bugs)
>
> > > KASLR makes things tricky but this is most definitely a slab allocation in
> > > the direct map and the stack (VMAP_STACK) is at 0xff7532a13699fda0 (rsp)
> > > ~70 TiB higher which sits at least 10 TiB padded above the direct map so
> > > it's safe to say that this is in the direct map.
> > >
> > > And the corrupted value (rbx) 0x0fffffff0c930020 is interesting - it's an
> > > exactly a x86-64 swap softleaf value:
> > >
> > > __swp_type() = val >> (64 - SWP_TYPE_BITS=5) = val >> 59 = 1
> > > __swp_offset() = (~(x).val << SWP_TYPE_BITS >> SWP_OFFSET_SHIFT) = (~val << 6 >> 14)
> > > = 0x79b67f
> > >
> > > I.e. it's a swap softleaf entry 0x79b67f 4 KiB pages into the swap = ~30.4 GiB.
> > >
> > > It's also the _second_ swap in the system (Nikola - if you have a 2nd swap
> > > on the reporting system of >=~30 GiB that kinda confirms it).
> >
> > I suspect this may be a bit of a red herring...
> >
> > actually there is NO swap on that machine, also there were no linux guests.. so
> > that might just be a coincidence? not sure if it changes anything..
> >
>
> Hmm that's really really odd. But if you had swap before or VMs before this
> is a long-lasting corruption that could have been sat there for days before
> you triggered it.
>
> > > The LLM added on some hints for confirmation of this:
> > >
> > > Schlopp>>
> > >
> > > Log greps, across all affected hosts and all boots, not just the ones
> > > that crashed:
> > >
> > > grep -i 'Bad page map' /var/log/messages*
> > > grep -i 'bad pmd' /var/log/messages*
> > > grep -i 'bad pud' /var/log/messages*
> > > grep -i 'Bad page state' /var/log/messages*
> > > grep -i 'CPA: called for zero pte' /var/log/messages*
> >
> > not a single occurance (this machine uses journal, but I checked those
> > and no such messages.. in general i tend to check dmesg and system logs
> > a lot, so I'd have already reported such messages..
>
> Yeah I don't know why it assumed you used antiquated logging..! :)
>
> OK that's interesting.
>
> >
> > >
> > > Any of these, particularly "bad pmd", is direct evidence that a freed
> > > kernel PTE table was reused as a user page table. "CPA: called for zero
> > > pte" would be the CPA walker itself tripping over a collapsed mapping.
> > >
> > > Questions:
> > >
> > > 1. swapon --show on the host, and inside the guests. Is there a swap
> > > device with index 1 and a size of at least roughly 30.5GiB? That
> > > tells us whether the corrupt word is a host swap PTE or a guest one,
> > > which distinguishes case A from case B above.
> > >
> > > 2. From the vmcore, dump the 4K page containing 0xff2e6dbe0e51b440. If
> > > the neighbouring words are also PTE-shaped, the page was being used
> > > as a page table and the diagnosis above is confirmed. If only the
> > > one word is corrupt, it was a single stray store.
>
> > unfortunately I don't have full vmcore from that crash, as it didn't fit
> > to /var/crash, backtrace I posted is from vmcore-dmesg.txt so can't confirm
> > that..
>
> Ah that's a pity!
>
> >
> >
> > >
> > > 3. What are CONFIG_DEBUG_VM, init_on_alloc, init_on_free and
> > > page_poison set to in the production build versus the KASAN build?
> > > free_page_is_bad() is gated on is_check_pages_enabled(), which needs
> > > CONFIG_DEBUG_VM, so the production kernel would not report the bad
> > > free even if it happened.
> >
> > I don't have CONFIG_DEBUG_VM enabled in production..
>
> Well that explains the lack of bad reports above. I don't know why it'd
> assume you'd run kernels with that (we do not recommend that for production
> :)
>
> >
> > >
> > > 4. Are any of the crashing guests Windows, and is hv-tlbflush set on
> > > them? That decides whether 26505e1b5b54 matters for you.
> > yes, windows, but hv-tlbflush is enabled on a sigle VM and it runs od
> > different node all the time.
>
> Ah but that could be enough to cause memory corruption. The reports seem to
> be about guest memory corruption though.
>
> To be clear - are you observing it in the guest or host? I gathered host
> from the splat.
I experienced multiple host memory corruptions (broken .so libraries, etc)
and also few visible from guests (GCC crashes while doing kernel builds in a loop,
some guest panics.. and then few windows crashes (not sure about causes there,
I'm no windows expert) but I guess all this can be cause by HOST side corruption
> > > 5. Has any corruption occurred since THP was disabled? If yes, that
> > > supports the CPA race over your THP theory.
> > not yet, but it's not happening that often, so unsure here
>
> Yeah it seems to be a hard one to hit. If you tried a kernel with KASAN
> enabled it might be flagged earlier? But that could also kill the race
> window and would slow the system down a lot.
>
> >
> > >
> > > 6. When you run the 6.18.54-rc1 test, please note that 6.18.52 is only
> > > a partial fix, so if you have any results from a 6.18.52 kernel they
> > > should not be treated as a clean run.
> > sure, I'll start today with 6.18.54, won't consider older tests.
>
> Ack, that's the best thing to do at the moment to be honest.
>
> If you were consistently getting corruption after X days previously, 2*X
> days let's say of none can give confidence it's fixed there.
I'm quite unsure what is the safe period here, as I mentioned to Luiz today,
at least two times, I thought it's fixed after ~10 days of stress tests, deployed
kernel to production.. and got another crash after few weeks..
> > I surely will!
> >
> > cheers, nik
>
> Thanks! Given the nature of the bug and the fact the LLM went a little out
> on a limb.
>
> Some more stuff from the report, which I also enclose in full here FYI.
>
> schlopp>>
>
> 26505e1b5b54 ("KVM: SVM: make svm_flush_tlb_gva do a full asid flush if
> NPT enabled") is mainline only and has not been backported to 6.18.y.
> It is AMD-specific guest memory corruption via the Hyper-V PV TLB flush
> path, so it only affects Windows guests running with hv-tlbflush=on. If
> any of your crashing guests are Windows, this is worth backporting
> separately. It is independent of the CPA race.
>
> 55ddbc2ca6d5 (upstream f7491d7c81db, "x86/mm: Fix user-space data loss
> with MADV_FREE and THP") is a real bug in 6.18.44. pmd_modify() drops
> _PAGE_DIRTY, which has been the case since v6.10, and it causes data
> loss for MADV_FREE THP and writable file THP. It is fixed in 6.18.53.
> It is data loss rather than a stray write, so it does not explain the
> oops, but it is another reason to move off 6.18.44. Note that this one
> needs THP, which you have now disabled.
>
> d053eb7e09e1 (upstream 1e75a8255f11, "iommu/amd: Wait for completion
> instead of returning early in iommu_completion_wait()") is already in
> 6.18.42 and therefore in the kernel that crashed. It is ruled out for
> this oops, but it may be relevant to the earlier incidents you had on
> 6.18.31 through 6.18.41.
>
> <<schlopp
>
> Let us know how the tests get on! If you trigger a bug on 54 let us know
> ASAP so we can investigate alternative theories.
sure!
BR nik
>
> Thanks!
>
> >
> >
> >
> >
> > >
> > > --
> > > Cheers, Lorenzo
> > >
> >
> > --
> > Ing. Nikola CIPRICH
> > technický ředitel
> >
> > +420 591 166 214
> > +420 777 093 799
> > nikola.ciprich@linuxbox.cz
> >
> > www.linuxbox.cz
>
> --
> Cheers, Lorenzo
> Summary
> =======
>
> The corruption you are seeing is consistent with a known use-after-free
> in the x86 change_page_attr (CPA) code that is present in 6.18.44 and
> was only fixed in 6.18.52 and 6.18.53.
>
> cpa_collapse_large_pages() rebuilds a leaf PMD out of its 4K PTEs and
> then frees the old PTE table, with no lock held against the lockless
> page table walk that __change_page_attr() performs before it stores
> through the PTE pointer it cached. The stale 8-byte store of a kernel
> PTE value lands in whatever the buddy allocator has since handed that
> page out for.
>
> On a KVM host the two sides of this race are both hot. set_memory_rox()
> is the only caller that passes CPA_COLLAPSE, and it runs on every module
> load (execmem_restore_rox()), every ftrace trampoline creation
> (arch/x86/kernel/ftrace.c:423) and every new 2M BPF program pack. The
> other side is any lockless walk of the same execmem tables:
> set_memory_nx()/set_memory_rw() from execmem_force_rw() on module load
> and trampoline allocation, and vmalloc_to_page() inside __text_poke()
> for every patch of module text, kprobe slot, trampoline or BPF pack.
> Module text, kprobe slots and ftrace trampolines share the same 2M ROX
> cache pages, so the collapser and the victim land in the same PMD by
> construction. A libvirt host does all of this constantly: module
> autoload (tun, vhost_net, br_netfilter, ebtable_*, xt_*), per-VM seccomp
> filters, systemd cgroup BPF, perf and BPF probes, ftrace and static call
> patching on every module load.
>
> This matches your good/bad window exactly. The collapse feature was
> added in v6.15 and is not in 5.15.
>
> It also matches the KASAN result. free_page_is_bad() is gated on
> is_check_pages_enabled(), which is off unless CONFIG_DEBUG_VM is set,
> and KASAN changes the allocation pattern enough that the freed page
> tends not to be reused in the race window. The upstream reporter only
> reproduced it by injecting a delay at the CPA page table lookup.
>
> Recommendation: run 6.18.53 or later. 6.18.54-rc1, which you say you
> are about to test, has the complete series. 6.18.52 has only the
> cpa_lock patch, which does not close the window. Mainline (v7.3-rc5)
> has nothing further pending for arch/x86/mm/pat/set_memory.c.
>
>
> Kernel version
> ==============
>
> 6.18.44lb9.01 (AlmaLinux 9 build of stable 6.18.44, stable commit
> 1efe5d048a39). First seen on 6.18.31, last crash on 6.18.44. 5.15.x
> was fine.
>
>
> Machine
> =======
>
> ASUSTeK RS720A-E12-RS12 / K14PP-D24, AMD EPYC, BIOS 2305 11/21/2025.
> KVM host, AlmaLinux 9, Intel ice and Mellanox NICs. Taint "G E",
> unsigned module only.
>
>
> Stack trace
> ===========
>
> Oops: general protection fault, probably for non-canonical address
> 0xfffffff0c930038: 0000 [#1] SMP NOPTI
> CPU: 23 UID: 189 PID: 7538 Comm: pacemaker-contr
> RIP: 0010:__d_lookup+0x4a/0xc0
> RAX: 0fffffff0c930020 RBX: 0fffffff0c930020 RCX: 0000000000000000
> RDX: ff2e6dbe0d9b6000 RSI: ff7532a13699fe60 RDI: ff2e6d1e4e630d80
> RBP: 000000000b654440
> Call Trace:
> <TASK>
> d_lookup+0x27/0x50
> lookup_dcache+0x1f/0x80
> lookup_one_qstr_excl+0x1e/0xe0
> filename_create+0xc4/0x160
> do_mkdirat+0x5a/0x190
> __x64_sys_mkdir+0x42/0x60
> do_syscall_64+0x64/0xbf0
> entry_SYSCALL_64_after_hwframe+0x76/0x7e
> </TASK>
>
> Other messages you reported:
>
> Inconsistency detected by ld.so: ../sysdeps/x86_64/dl-machine.h: 548:
> elf_machine_rela_relative: Assertion `ELFW(R_TYPE) (reloc->r_info) ==
> R_X86_64_RELATIVE' failed!
>
>
> What the oops registers say
> ===========================
>
> The Code bytes decode to the hash chain walk in fs/dcache.c:__d_lookup():
>
> mov (%rbx),%rax ; load bucket->first
> mov %rax,%rbx
> and $-2,%rbx ; strip the hlist_bl lock bit
> cmp $1,%rax
> ja body
> loop:
> mov (%rbx),%rbx ; node->next
> test %rbx,%rbx
> je out
> body:
> cmp %ebp,0x18(%rbx) ; <-- faulting insn, d_name.hash_len
> jne loop
>
> RAX equals RBX and RAX is only ever written by the initial bucket load,
> so this is the first loop iteration. The corrupt word is the
> hlist_bl_head.first field of the dentry_hashtable bucket itself, not a
> dentry's d_hash.next.
>
> RDX 0xff2e6dbe0d9b6000 is the dentry_hashtable base, and the runtime
> constant d_hash_shift is patched to 7, so:
>
> bucket = 0xff2e6dbe0d9b6000 + (0x0b654440 >> 7) * 8
> = 0xff2e6dbe0e51b440
>
> That table is a 256MB alloc_large_system_hash() allocation from
> memblock. It is allocated at boot, is PG_reserved and is never freed.
> So this is a stray write to a fixed physical page, not a
> use-after-free of a recycled object.
>
> The corrupt value 0x0fffffff0c930020 is an exact x86-64 non-present
> swap PTE: __swp_type() is 1 and __swp_offset() is 0x79b67f, about
> 30.4GiB into swap device 1. The only low bit set is bit 5,
> _PAGE_ACCESSED.
>
> Every bit the swap layout constrains is as it should be: P, PSE and
> the G/PROT_NONE alias are clear, the software bits 1-3 are clear, the
> type is an ordinary swap type and the inverted offset gives the long
> run of ones in bits 32-58. The one thing the layout does not account
> for is bit 5. __swp_entry() leaves bits 0-8 clear and no helper sets
> bit 5 on a non-present PTE; arch/x86/include/asm/pgtable_64.h treats
> bits 5 and 6 as don't-care only because of the Intel Knights Landing
> erratum (00839ee3b299, _PAGE_KNL_ERRATUM_MASK), which does not apply to
> EPYC. So either something other than a Linux swap PTE happens to fit
> this layout, or the word was a swap PTE that acquired a stray bit. I
> cannot tell which from one word, which is why the vmcore page dump
> requested below matters: 511 neighbouring PTE-shaped words would settle
> it.
>
>
> Suspect commit
> ==============
>
> commit 41d88484c71cd4f659348da41b7b5b3dbd3be1f6
> Author: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
>
> x86/mm/pat: restore large ROX pages after fragmentation
>
> Link: https://lore.kernel.org/r/20250126074733.1384926-4-rppt@kernel.org
>
> This added runtime collapse of split kernel large pages, driven from
> cpa_flush(), including freeing the PTE table that the collapsed PMD
> replaces. It is the Fixes: target of every fix listed below. It is in
> v6.15 and later, and is not in 5.15.
>
> > diff --git a/arch/x86/mm/pat/set_memory.c b/arch/x86/mm/pat/set_memory.c
> > --- a/arch/x86/mm/pat/set_memory.c
> > +++ b/arch/x86/mm/pat/set_memory.c
>
> > @@ -394,6 +408,40 @@ static void __cpa_flush_tlb(void *data)
> > +static void cpa_collapse_large_pages(struct cpa_data *cpa)
> > +{
> > + unsigned long start, addr, end;
> > + struct ptdesc *ptdesc, *tmp;
> > + LIST_HEAD(pgtables);
> > + int collapsed = 0;
> > + int i;
>
> [ ... range iteration ... ]
>
> > + if (!collapsed)
> > + return;
> > +
> > + flush_tlb_all();
> > +
> > + list_for_each_entry_safe(ptdesc, tmp, &pgtables, pt_list) {
> > + list_del(&ptdesc->pt_list);
> > + __free_page(ptdesc_page(ptdesc));
> > + }
> > +}
>
> In 6.18.44 that last loop is pagetable_free(ptdesc), which is still an
> immediate free. There is no RCU grace period and no other deferral. The
> flush_tlb_all() above it only makes the hardware forget the old
> translation; it does nothing about a CPU that is sitting inside
> __change_page_attr() holding a pointer into that table.
>
> > @@ -402,7 +450,7 @@ static void cpa_flush(struct cpa_data *cpa, int cache)
> > if (cache && !static_cpu_has(X86_FEATURE_CLFLUSH)) {
> > cpa_flush_all(cache);
> > - return;
> > + goto collapse_large_pages;
> > }
>
> [ ... ]
>
> > @@ -427,6 +475,10 @@ static void cpa_flush(struct cpa_data *cpa, int cache)
> > mb();
> > +
> > +collapse_large_pages:
> > + if (cpa->flags & CPA_COLLAPSE)
> > + cpa_collapse_large_pages(cpa);
> > }
>
> The collapse is hooked into cpa_flush(), which
> __change_page_attr_set_clr() calls after it has already dropped
> cpa_lock. So cpa_lock does not serialise the collapse against anything.
>
> > @@ -1196,6 +1248,161 @@ static int split_large_page(struct cpa_data *cpa, pte_t *kpte,
> > +static int collapse_pmd_page(pmd_t *pmd, unsigned long addr,
> > + struct list_head *pgtables)
> > +{
>
> [ ... uniformity checks over all 512 PTEs ... ]
>
> > + old_pmd = *pmd;
> > +
> > + /* Success: set up a large page */
> > + pgprot = pgprot_4k_2_large(pte_pgprot(first));
> > + pgprot_val(pgprot) |= _PAGE_PSE;
> > + _pmd = pfn_pmd(pfn, pgprot);
> > + set_pmd(pmd, _pmd);
> > +
> > + /* Queue the page table to be freed after TLB flush */
> > + list_add(&page_ptdesc(pmd_page(old_pmd))->pt_list, pgtables);
>
> collapse_large_pages(), the caller, takes pgd_lock around this. The
> lockless CPA walker never takes pgd_lock, so pgd_lock does not help
> either.
>
> The other side, in 6.18.44:
>
> arch/x86/mm/pat/set_memory.c:__change_page_attr() {
> address = __cpa_addr(cpa, cpa->curpage);
> repeat:
> kpte = _lookup_address_cpa(cpa, address, &level, &nx, &rw);
> ...
> old_pte = *kpte;
> ...
> if (level == PG_LEVEL_4K) {
> ...
> new_pte = pfn_pte(pfn, new_prot);
> ...
> if (pte_val(old_pte) != pte_val(new_pte)) {
> set_pte_atomic(kpte, new_pte); <-- stale
> cpa->flags |= CPA_FLUSHTLB;
> }
>
> _lookup_address_cpa() reaches lookup_address_in_pgd_attr(), which is a
> plain lockless walk. Between the walk and the set_pte_atomic() the
> caller can be preempted or take an interrupt; this runs with interrupts
> on. That store is the only unsafe instruction in the function. The
> large-page split branch further down is safe because __split_large_page()
> revalidates under pgd_lock.
>
> The freed table is not even a tracked page table page in 6.18.44:
>
> arch/x86/mm/pat/set_memory.c:split_large_page() {
> if (!debug_pagealloc_enabled())
> spin_unlock(&cpa_lock);
> base = alloc_pages(GFP_KERNEL, 0);
>
> A bare alloc_pages(), so it goes straight back to the per-CPU free list
> and can be reallocated immediately.
>
>
> Race timeline
> =============
>
> CPU A (text_poke -> CPU B (module_enable_rox /
> execmem_make_temp_rw -> execmem_restore_rox /
> set_memory_nx/rw) bpf_jit_binary_lock_ro ->
> set_memory_rox, CPA_COLLAPSE)
> ----- -----
> __change_page_attr()
> kpte = _lookup_address_cpa()
> old_pte = *kpte
> new_pte = pfn_pte(...)
> preempted / interrupted
> __change_page_attr_set_clr()
> drops cpa_lock
> cpa_flush()
> cpa_collapse_large_pages()
> collapse_pmd_page(): all 512
> PTEs uniform, set_pmd() installs
> a leaf, old PTE table queued
> flush_tlb_all()
> pagetable_free() -> immediate
> __free_pages()
>
> (any CPU) page is reallocated:
> .so page cache folio, QEMU guest
> RAM, a user PMD/PTE table, slab
>
> set_pte_atomic(kpte, new_pte)
> stores a PTE-shaped word into
> the reallocated page
>
> Where the swap PTE comes from (inferred continuation)
> ------------------------------------------------------
>
> The race above writes a present kernel PTE, never a swap entry. To
> reach the dentry hash table with a swap-PTE-shaped word the following
> has to happen next. Each step is verified in the 6.18.44 code; the
> sequence as a whole is inferred, not proven for this oops.
>
> 1. The freed PTE table is reallocated as a QEMU page table.
>
> 2. The stale set_pte_atomic() lands in it. The injected entry is a
> translation into an execmem text page (case A/B below).
>
> 3. Case B: GUP-slow follows that entry and KVM maps the text page
> into the guest; a later zap_present_folio_ptes() does an
> unbalanced folio_put() and frees the still-live text page.
>
> 4. That text page is reallocated as another page table while
> text_poke()/the BPF JIT keep writing instruction bytes into it
> through the ROX mapping. Instruction bytes are now PMD entries
> with arbitrary pfns; some pass pmd_bad().
>
> 5. Reclaim: try_to_unmap_one() -> page_vma_mapped_walk() ->
> pte_offset_map_lock() reads such a PMD, computes
> __va(garbage pfn) as the PTE table, and set_pte_at() stores a
> swap PTE there. If that pfn is the dentry_hashtable page, one
> bucket head becomes 0x0fffffff0c930020-like.
>
> 6. Days later __d_lookup() hashes into that bucket and faults.
>
> Step 5 is the only writer of an ordinary swap type in the mm, and
> the only step that can touch memory the allocator never owned.
>
> This is not speculation about the code. The same interleaving was
> reported upstream with a KASAN reproducer, in the commit that first
> tried to address it:
>
> commit 1aac65f3e651 ("x86/mm/pat: Take cpa_lock around large-page
> collapse")
>
> BUG: KASAN: use-after-free in __change_page_attr+0x7cc/0x7e0
> Write of size 8 at addr ffff888181139718 by task modprobe
> ...
> The buggy address belongs to the physical page:
> pfn:0x181139 ... page_type: f2(table)
>
> Fixes: 41d88484c71c ("x86/mm/pat: restore large ROX pages after
> fragmentation")
> Signed-off-by: Denis V. Lunev <den@openvz.org>
>
>
> Which stable releases carry the fixes
> =====================================
>
> None of these are in 6.18.44.
>
> The fixes that matter are the init_mm mmap lock pair, a1c7570cedd0 and
> d5d8b8662e6e: the collapse runs under the init_mm write lock and the
> whole attribute change, including the lockless walk and the store
> through the cached pointer, runs under the read lock. That excludes
> both walker-vs-collapse and collapse-vs-collapse. 1587d3394e25 covers
> the third walker: __text_poke() resolves the pages it patches with
> vmalloc_to_page(), a lockless walk of the same execmem tables, and now
> takes the init_mm read lock around it. Without that, a collapse under
> a concurrent text_poke() returns NULL (the BUG_ON at
> arch/x86/kernel/alternative.c:2576 that openSUSE hit) or, if the freed
> table has already been reused, a wrong page that text_poke() then
> writes instruction bytes into. 9e4a3ec3411b makes the split tables
> real kernel page tables so their freeing is deferred. The earlier
> cpa_lock patch, 1aac65f3e651, is superseded by the write lock and adds
> nothing once those are applied.
>
> 6.18.52
> 591b6fac9df3 x86/mm/pat: Take cpa_lock around large-page collapse
> (upstream 1aac65f3e651)
>
> 6.18.53
> 35820cf8dd52 x86/mm/pat: Acquire init_mm write lock on collapse to
> avoid UAF (upstream a1c7570cedd0)
> e21a9ea81426 x86/mm/pat: Acquire init_mm read lock on attribute
> changes to avoid UAF (upstream d5d8b8662e6e)
> e164f4a25e23 x86/mm/pat: Convert split_large_page() to use ptdescs
> 5029589bb773 x86/mm/pat: Don't gate cpa_lock on
> debug_pagealloc_enabled()
> 84e0cd79d57f x86/mm/pat: Allocate split page tables as kernel page
> tables (upstream 9e4a3ec3411b)
> 281e6f536f2f x86/alternatives: Exclude text poking against
> change_page_attr() (upstream 1587d3394e25)
> 74a2626044de x86/mm: Fix and document DEBUG_PAGEALLOC
> (upstream 7da514d819a0)
>
> 6.18.52 does not fix this: it only takes cpa_lock around the collapse,
> and the walker never holds cpa_lock across its walk-then-store window.
> The init_mm mmap lock pair that closes that window is in 6.18.53, which
> is the first stable release with the complete set. Use 6.18.53 or
> later.
>
> There is no clean interim mitigation on 6.18.44. CPA_COLLAPSE cannot be
> disabled by a boot parameter.
>
>
> How this reaches the symptoms you saw
> =====================================
>
> Symptoms 1 and 2 follow directly. The stale store writes a PTE-shaped
> qword into a page that has been reallocated. If that page is a page
> cache folio for a mapped .so, eight bytes of its relocation table are
> replaced and ld.so trips the R_X86_64_RELATIVE assertion while the file
> on disk is intact. If it is an anonymous page that QEMU has just
> populated as guest RAM on a migration destination, the guest sees eight
> corrupt bytes. Both of these match "right after migration": the
> destination host is populating gigabytes of guest RAM and allocating
> page tables at maximum rate, which is exactly when a just-freed page
> gets reused inside the race window.
>
> Symptom 3, this oops, needs one more step, because the dentry hash table
> is memblock memory that is never freed and therefore cannot be the
> directly reallocated page. The escalation is that the victim page is
> itself a page table. Each of the following links is verified in the
> 6.18.44 source, but I want to be clear that the end-to-end chain for
> this particular oops is plausible rather than proven from a single
> vmcore.
>
> Case A, the victim is a user PMD table. pmd_bad() is true for the
> injected value, so the first user-mode touch faults and
> mm/pgtable-generic.c:___pte_offset_map() heals it via pmd_clear_bad()
> with a "bad pmd" pr_err. But a supervisor-side copy_to_user() is
> permitted through a U=0 level. A read() or recvmsg() into guest RAM, or
> vhost-net, or kvm_write_guest(), then has the hardware walker read a
> qword of module text as a PTE. If that qword happens to have P and RW
> set, the copied data is written to an arbitrary physical address. This
> route leaves no taint.
>
> Case B, the victim is a user PTE table. x86 has no pte_bad(). GUP-fast
> rejects a U=0 entry, but GUP-slow does not:
> mm/gup.c:follow_page_pte() and mm/memory.c:__vm_normal_page() check
> neither _PAGE_USER nor PageReserved, so a live execmem page is returned
> and KVM maps kernel module text into a guest. A later
> mm/memory.c:zap_present_folio_ptes() does folio_remove_rmap_ptes() and
> folio_put() on a page that was never rmapped, dropping the refcount to
> zero, printing "BUG: Bad page map" and releasing live ROX text into the
> buddy allocator while text_poke() and the BPF JIT keep writing
> instruction bytes into it. Recycled as a user PMD table, instruction
> bytes are page table entries with arbitrary pfns that pass pmd_bad(),
> pte_offset_map() computes __va() of an arbitrary physical page, and
> try_to_unmap_one() writes a genuine host swap PTE into it.
>
> That last step is what the corrupt word looks like: a real host swap
> PTE. The stray bit 5 is unexplained on this route too.
>
> One caveat that argues against case B on this specific host: the taint
> is "G E" with no "B". "BUG: Bad page map" had not fired on that
> machine before the crash. So either the taint-free case A route or a
> direct stray write carries this particular chain, or the corrupting
> event happened on a different boot. This limits, but does not refute,
> the cascade. The corruption can sit in a rarely used dentry bucket for
> a long time before something hashes into it, which fits the 22-day
> uptime on this crash and your observation that the host crashes were
> not correlated with migration.
>
>
> Secondary findings
> ==================
>
> These came up while looking and are worth knowing about, but none of
> them explains this oops.
>
> 26505e1b5b54 ("KVM: SVM: make svm_flush_tlb_gva do a full asid flush if
> NPT enabled") is mainline only and has not been backported to 6.18.y.
> It is AMD-specific guest memory corruption via the Hyper-V PV TLB flush
> path, so it only affects Windows guests running with hv-tlbflush=on. If
> any of your crashing guests are Windows, this is worth backporting
> separately. It is independent of the CPA race.
>
> 55ddbc2ca6d5 (upstream f7491d7c81db, "x86/mm: Fix user-space data loss
> with MADV_FREE and THP") is a real bug in 6.18.44. pmd_modify() drops
> _PAGE_DIRTY, which has been the case since v6.10, and it causes data
> loss for MADV_FREE THP and writable file THP. It is fixed in 6.18.53.
> It is data loss rather than a stray write, so it does not explain the
> oops, but it is another reason to move off 6.18.44. Note that this one
> needs THP, which you have now disabled.
>
> d053eb7e09e1 (upstream 1e75a8255f11, "iommu/amd: Wait for completion
> instead of returning early in iommu_completion_wait()") is already in
> 6.18.42 and therefore in the kernel that crashed. It is ruled out for
> this oops, but it may be relevant to the earlier incidents you had on
> 6.18.31 through 6.18.41.
>
>
> Theories that were eliminated
> =============================
>
> KVM NPT mapping at too large a level, or with the wrong base pfn. The
> mapping level comes from the host page tables and the pfn from GUP;
> KVM cannot reach memblock memory on its own.
>
> A host mm swap or migration PTE stored through a stale page table
> pointer. All the store sites are bounded and the pointer provenance
> checks out.
>
> A missed MMU notifier invalidation. Notifier ordering on the recovery
> path is correct, and this cannot reach never-freed memory.
>
> NIC DMA to the wrong address. Only a teardown-time page_pool
> use-after-free turned up, and the iommu/amd completion-wait fix is
> already in 6.18.42.
>
> 105d04edbec8 (upstream 33192a26cddea, mm/huge_memory huge_zero_pfn
> race). Real, but not present in 6.18.44, and it cannot reach memblock
> memory.
>
> 0a25ee42e7d1 (upstream 3d679b7cb31f, KVM x86/mmu CMPXCHG when clearing
> the Accessed bit in the TDP MMU). Not in 6.18.44, and benign for this.
>
> TDP MMU in-place huge page recovery. Structurally excluded: notifier
> zaps take mmu_lock for write, recovery takes it for read.
>
> memblock/buddy physical aliasing. This would produce "Bad page state"
> reports, which you have not seen. Worth confirming from the vmcore, see
> below.
>
> Neither THP nor NUMA balancing is involved in the CPA race. Disabling
> them was a reasonable precaution but it will not stop this. If you keep
> seeing corruption with THP off, that is consistent with the diagnosis
> rather than against it.
>
>
> What would confirm this
> =======================
>
> Log greps, across all affected hosts and all boots, not just the ones
> that crashed:
>
> grep -i 'Bad page map' /var/log/messages*
> grep -i 'bad pmd' /var/log/messages*
> grep -i 'bad pud' /var/log/messages*
> grep -i 'Bad page state' /var/log/messages*
> grep -i 'CPA: called for zero pte' /var/log/messages*
>
> Any of these, particularly "bad pmd", is direct evidence that a freed
> kernel PTE table was reused as a user page table. "CPA: called for zero
> pte" would be the CPA walker itself tripping over a collapsed mapping.
>
> Questions:
>
> 1. swapon --show on the host, and inside the guests. Is there a swap
> device with index 1 and a size of at least roughly 30.5GiB? That
> tells us whether the corrupt word is a host swap PTE or a guest one,
> which distinguishes case A from case B above.
>
> 2. From the vmcore, dump the 4K page containing 0xff2e6dbe0e51b440. If
> the neighbouring words are also PTE-shaped, the page was being used
> as a page table and the diagnosis above is confirmed. If only the
> one word is corrupt, it was a single stray store.
>
> 3. What are CONFIG_DEBUG_VM, init_on_alloc, init_on_free and
> page_poison set to in the production build versus the KASAN build?
> free_page_is_bad() is gated on is_check_pages_enabled(), which needs
> CONFIG_DEBUG_VM, so the production kernel would not report the bad
> free even if it happened.
>
> 4. Are any of the crashing guests Windows, and is hv-tlbflush set on
> them? That decides whether 26505e1b5b54 matters for you.
>
> 5. Has any corruption occurred since THP was disabled? If yes, that
> supports the CPA race over your THP theory.
>
> 6. When you run the 6.18.54-rc1 test, please note that 6.18.52 is only
> a partial fix, so if you have any results from a 6.18.52 kernel they
> should not be treated as a clean run.
--
Ing. Nikola CIPRICH
technický ředitel
+420 591 166 214
+420 777 093 799
nikola.ciprich@linuxbox.cz
www.linuxbox.cz
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: hunting memory corruption bug in 6.18.x
2026-09-28 9:05 ` Nikola Ciprich
@ 2026-09-29 19:11 ` Nikola Ciprich
2026-09-30 9:07 ` Lorenzo Stoakes (ARM)
0 siblings, 1 reply; 16+ messages in thread
From: Nikola Ciprich @ 2026-09-29 19:11 UTC (permalink / raw)
To: Lorenzo Stoakes (ARM)
Cc: linux-mm, linux-kernel, akpm, david, Mike Rapoport, Dave Hansen,
Pedro Falcato, Kiryl Shutsemau, luizcap, Nikola Ciprich
Hi Lorenzo, Luiz,
while the lab cluster running 6.18.54 happily migrates VMs (one running MSSQL
with hammerDB, another doing kernel builds in ramdisk in a loop) there and
back without any issue, my colleague tried upgrading one production cluster to
the same kernel and immediately after migrating few VMs back got system in almost
unusable state, everything started crashing:
Sep 29 18:42:33 nrbphav4a kernel: qemu-system-x86[19014]: segfault at 0 ip 00007f9a3ee7e61e sp 00007ffd96294380 error 6 in libglib-2.0.so.0.6800.4[a461e,7f9a3edf7000+91000] likely on CPU 10>
Sep 29 18:42:33 nrbphav4a kernel: qemu-system-x86[18946]: segfault at 0 ip 00007f713a90a61e sp 00007ffc58f46c60 error 6
Sep 29 18:42:33 nrbphav4a kernel: Code: 29 9c 24 80 00 00 00 0f 29 a4 24 90 00 00 00 0f 29 ac 24 a0 00 00 00 0f 29 b4 ff ff 00 00 00 0f 29 bc ff ff 00 00 00 64 00 00 <00> 00 28 00 00 00 48 >
Sep 29 18:42:33 nrbphav4a kernel: qemu-system-x86[18736]: segfault at 0 ip 00007fa0be9f461e sp 00007fffb246df50 error 6
Sep 29 18:42:33 nrbphav4a kernel: in libglib-2.0.so.0.6800.4[a461e,7f713a883000+91000]
Sep 29 18:42:33 nrbphav4a kernel: pacemaker-execd[17696]: segfault at 0 ip 00007f3c6ed6a61e sp 00007fffa95873d0 error 6 in libglib-2.0.so.0.6800.4[a461e,7f3c6ece3000+91000] likely on CPU 20>
Sep 29 18:42:33 nrbphav4a kernel: Code: 29 9c 24 80 00 00 00 0f 29 a4 24 90 00 00 00 0f 29 ac 24 a0 00 00 00 0f 29 b4 ff ff 00 00 00 0f 29 bc ff ff 00 00 00 64 00 00 <00> 00 28 00 00 00 48 >
Sep 29 18:42:33 nrbphav4a kernel: in libglib-2.0.so.0.6800.4[a461e,7fa0be96d000+91000]
Sep 29 18:42:33 nrbphav4a kernel: likely on CPU 22 (core 6, socket 1)
Sep 29 18:42:33 nrbphav4a kernel: Code: 29 9c 24 80 00 00 00 0f 29 a4 24 90 00 00 00 0f 29 ac 24 a0 00 00 00 0f 29 b4 ff ff 00 00 00 0f 29 bc ff ff 00 00 00 64 00 00 <00> 00 28 00 00 00 48 >
Sep 29 18:42:33 nrbphav4a kernel: qemu-system-x86[19275]: segfault at 0 ip 00007fcc8201d61e sp 00007ffd9cf1b660 error 6
Sep 29 18:42:33 nrbphav4a kernel: likely on CPU 28 (core 12, socket 1)
Sep 29 18:42:33 nrbphav4a kernel: in libglib-2.0.so.0.6800.4[a461e,7fcc81f96000+91000] likely on CPU 19 (core 3, socket 1)
Sep 29 18:42:33 nrbphav4a kernel: Code: 29 9c 24 80 00 00 00 0f 29 a4 24 90 00 00 00 0f 29 ac 24 a0 00 00 00 0f 29 b4 ff ff 00 00 00 0f 29 bc ff ff 00 00 00 64 00 00 <00> 00 28 00 00 00 48 >
Sep 29 18:42:33 nrbphav4a kernel: pacemakerd[17690]: segfault at 0 ip 00007f752776a61e sp 00007fff3cbcc8f0 error 6
Sep 29 18:42:33 nrbphav4a kernel: pacemaker-fence[17695]: segfault at 0 ip 00007f9d9496a61e sp 00007ffe50dd0210 error 6 in libglib-2.0.so.0.6800.4[a461e,7f9d948e3000+91000] likely on CPU 21>
Sep 29 18:42:33 nrbphav4a kernel: Code: 29 9c 24 80 00 00 00 0f 29 a4 24 90 00 00 00 0f 29 ac 24 a0 00 00 00 0f 29 b4 ff ff 00 00 00 0f 29 bc ff ff 00 00 00 64 00 00 <00> 00 28 00 00 00 48 >
Sep 29 18:42:33 nrbphav4a kernel: pacemaker-contr[17699]: segfault at 0 ip 00007faeef96a61e sp 00007ffd04022700 error 6
Sep 29 18:42:33 nrbphav4a kernel:
Sep 29 18:42:33 nrbphav4a kernel: Code: 29 9c 24 80 00 00 00 0f 29 a4 24 90 00 00 00 0f 29 ac 24 a0 00 00 00 0f 29 b4 ff ff 00 00 00 0f 29 bc ff ff 00 00 00 64 00 00 <00> 00 28 00 00 00 48 >
Sep 29 18:42:33 nrbphav4a kernel: pacemaker-sched[17698]: segfault at 0 ip 00007ff29f96a61e sp 00007ffd6b0b2010 error 6 in libglib-2.0.so.0.6800.4[a461e,7ff29f8e3000+91000] likely on CPU 13>
Sep 29 18:42:33 nrbphav4a kernel: Code: 29 9c 24 80 00 00 00 0f 29 a4 24 90 00 00 00 0f 29 ac 24 a0 00 00 00 0f 29 b4 ff ff 00 00 00 0f 29 bc ff ff 00 00 00 64 00 00 <00> 00 28 00 00 00 48 >
Sep 29 18:42:33 nrbphav4a kernel: pacemaker-attrd[17697]: segfault at 0 ip 00007f2d8176a61e sp 00007fff8ac49030 error 6 in libglib-2.0.so.0.6800.4[a461e,7f2d816e3000+91000] likely on CPU 13>
Sep 29 18:42:33 nrbphav4a kernel: Code: 29 9c 24 80 00 00 00 0f 29 a4 24 90 00 00 00 0f 29 ac 24 a0 00 00 00 0f 29 b4 ff ff 00 00 00 0f 29 bc ff ff 00 00 00 64 00 00 <00> 00 28 00 00 00 48 >
Sep 29 18:42:33 nrbphav4a kernel: in libglib-2.0.so.0.6800.4[a461e,7f75276e3000+91000]
Sep 29 18:42:33 nrbphav4a kernel: in libglib-2.0.so.0.6800.4[a461e,7faeef8e3000+91000] likely on CPU 22 (core 6, socket 1)
Sep 29 18:42:33 nrbphav4a kernel: Code: 29 9c 24 80 00 00 00 0f 29 a4 24 90 00 00 00 0f 29 ac 24 a0 00 00 00 0f 29 b4 ff ff 00 00 00 0f 29 bc ff ff 00 00 00 64 00 00 <00> 00 28 00 00 00 48 >
Sep 29 18:42:33 nrbphav4a kernel: likely on CPU 20 (core 4, socket 1)
Sep 29 18:42:33 nrbphav4a kernel: Code: 29 9c 24 80 00 00 00 0f 29 a4 24 90 00 00 00 0f 29 ac 24 a0 00 00 00 0f 29 b4 ff ff 00 00 00 0f 29 bc ff ff 00 00 00 64 00 00 <00> 00 28 00 00 00 48 >
Sep 29 18:42:41 nrbphav4a kernel: show_signal_msg: 7 callbacks suppressed
Sep 29 18:42:41 nrbphav4a kernel: NetworkManager[19482]: segfault at 0 ip 00007f26c920161e sp 00007ffe553f7960 error 6 in libglib-2.0.so.0.6800.4[a461e,7f26c917a000+91000] likely on CPU 23 >
Sep 29 18:42:41 nrbphav4a kernel: Code: 29 9c 24 80 00 00 00 0f 29 a4 24 90 00 00 00 0f 29 ac 24 a0 00 00 00 0f 29 b4 ff ff 00 00 00 0f 29 bc ff ff 00 00 00 64 00 00 <00> 00 28 00 00 00 48 >
Sep 29 18:42:41 nrbphav4a kernel: irqbalance[2085]: segfault at 0 ip 00007f642283e61e sp 00007fff55260dd0 error 6 in libglib-2.0.so.0.6800.4[a461e,7f64227b7000+91000] likely on CPU 20 (core>
Sep 29 18:42:41 nrbphav4a kernel: Code: 29 9c 24 80 00 00 00 0f 29 a4 24 90 00 00 00 0f 29 ac 24 a0 00 00 00 0f 29 b4 ff ff 00 00 00 0f 29 bc ff ff 00 00 00 64 00 00 <00> 00 28 00 00 00 48 >
Sep 29 18:42:41 nrbphav4a kernel: pacemakerd[19486]: segfault at 0 ip 00007ff395b6a61e sp 00007ffc9165b240 error 6 in libglib-2.0.so.0.6800.4[a461e,7ff395ae3000+91000] likely on CPU 8 (core>
Sep 29 18:42:41 nrbphav4a kernel: Code: 29 9c 24 80 00 00 00 0f 29 a4 24 90 00 00 00 0f 29 ac 24 a0 00 00 00 0f 29 b4 ff ff 00 00 00 0f 29 bc ff ff 00 00 00 64 00 00 <00> 00 28 00 00 00 48 >
Sep 29 18:42:43 nrbphav4a kernel: NetworkManager[19513]: segfault at 0 ip 00007f8089d6a61e sp 00007ffe9d566e60 error 6 in libglib-2.0.so.0.6800.4[a461e,7f8089ce3000+91000] likely on CPU 7 (>
Sep 29 18:42:43 nrbphav4a kernel: Code: 29 9c 24 80 00 00 00 0f 29 a4 24 90 00 00 00 0f 29 ac 24 a0 00 00 00 0f 29 b4 ff ff 00 00 00 0f 29 bc ff ff 00 00 00 64 00 00 <00> 00 28 00 00 00 48 >
Sep 29 18:42:43 nrbphav4a kernel: libvirtd[19514]: segfault at 0 ip 00007f35f97fe61e sp 00007ffdecdd9dc0 error 6 in libglib-2.0.so.0.6800.4[a461e,7f35f9777000+91000] likely on CPU 3 (core 3>
Sep 29 18:42:43 nrbphav4a kernel: Code: 29 9c 24 80 00 00 00 0f 29 a4 24 90 00 00 00 0f 29 ac 24 a0 00 00 00 0f 29 b4 ff ff 00 00 00 0f 29 bc ff ff 00 00 00 64 00 00 <00> 00 28 00 00 00 48 >
Sep 29 18:42:47 nrbphav4a kernel: pacemakerd[19563]: segfault at 0 ip 00007f7e64b6a61e sp 00007ffc27ea2590 error 6 in libglib-2.0.so.0.6800.4[a461e,7f7e64ae3000+91000] likely on CPU 28 (cor>
Sep 29 18:42:47 nrbphav4a kernel: Code: 29 9c 24 80 00 00 00 0f 29 a4 24 90 00 00 00 0f 29 ac 24 a0 00 00 00 0f 29 b4 ff ff 00 00 00 0f 29 bc ff ff 00 00 00 64 00 00 <00> 00 28 00 00 00 48 >
Sep 29 18:42:47 nrbphav4a kernel: virtlogd[18719]: segfault at 0 ip 00007f1f3796a61e sp 00007ffd6c2586f0 error 6 in libglib-2.0.so.0.6800.4[a461e,7f1f378e3000+91000] likely on CPU 18 (core >
Sep 29 18:42:47 nrbphav4a kernel: Code: 29 9c 24 80 00 00 00 0f 29 a4 24 90 00 00 00 0f 29 ac 24 a0 00 00 00 0f 29 b4 ff ff 00 00 00 0f 29 bc ff ff 00 00 00 64 00 00 <00> 00 28 00 00 00 48 >
so he reverted back to 6.18.15
then he got crash of another node (to which he was migrating VMs from the first updated node), that one
was still running 6.18.15
pretty unfortunate planned work outcome :-/
but it confirms that both 6.18.15 and unfortunately also 6.18.54 are affected
maybe we'll be able to dedicate one or two nodes of this cluster to some tests,
I'll check that tomorrow
BR
nik
On Mon, Sep 28, 2026 at 11:05:15AM +0200, Nikola Ciprich wrote:
> > > one note here, at least last mentioned crash (with 6.18.44) happened with
> > > host running only windows guest, in general we're seeing those problems
> > > mosly with windows VM hosting machines.. so maybe they're triggerng the
> > > problem with some other, but similar mechanism?
> >
> > Interesting! But indeed all this is host-side.
> >
> > Though it came up with a possible finding around commit 26505e1b5b54 ("KVM:
> > SVM: make svm_flush_tlb_gva do a full asid flush if NPT enabled") which has
> > not been backported yet.
> >
> > That's is a. AMD-specific guest memory corruption and b. only if
> > hv-tlbflush=on.
> >
> > This is independent of the CPA stuff.
> >
> > So if the CPA stuff turns out to be a red herring that's one worth looking
> > at? Are you able to run a modified kernel with this applied on top?
>
> hmm, this one looks quite promising.
>
> the patch doesn't apply cleanly on top of latest 6.18, I'll have a look at it
>
> however, I double checked, there is single guest with hv-tlbflush running on
> affected cluster and it never run on recently crashed node. so unless this
> fixes also some different case, it's probably not the bug we're hunting here..
>
> (but I won't be surprised, if all this is caused with multiple independent
> bugs)
>
> >
> > > > KASLR makes things tricky but this is most definitely a slab allocation in
> > > > the direct map and the stack (VMAP_STACK) is at 0xff7532a13699fda0 (rsp)
> > > > ~70 TiB higher which sits at least 10 TiB padded above the direct map so
> > > > it's safe to say that this is in the direct map.
> > > >
> > > > And the corrupted value (rbx) 0x0fffffff0c930020 is interesting - it's an
> > > > exactly a x86-64 swap softleaf value:
> > > >
> > > > __swp_type() = val >> (64 - SWP_TYPE_BITS=5) = val >> 59 = 1
> > > > __swp_offset() = (~(x).val << SWP_TYPE_BITS >> SWP_OFFSET_SHIFT) = (~val << 6 >> 14)
> > > > = 0x79b67f
> > > >
> > > > I.e. it's a swap softleaf entry 0x79b67f 4 KiB pages into the swap = ~30.4 GiB.
> > > >
> > > > It's also the _second_ swap in the system (Nikola - if you have a 2nd swap
> > > > on the reporting system of >=~30 GiB that kinda confirms it).
> > >
> > > I suspect this may be a bit of a red herring...
> > >
> > > actually there is NO swap on that machine, also there were no linux guests.. so
> > > that might just be a coincidence? not sure if it changes anything..
> > >
> >
> > Hmm that's really really odd. But if you had swap before or VMs before this
> > is a long-lasting corruption that could have been sat there for days before
> > you triggered it.
> >
> > > > The LLM added on some hints for confirmation of this:
> > > >
> > > > Schlopp>>
> > > >
> > > > Log greps, across all affected hosts and all boots, not just the ones
> > > > that crashed:
> > > >
> > > > grep -i 'Bad page map' /var/log/messages*
> > > > grep -i 'bad pmd' /var/log/messages*
> > > > grep -i 'bad pud' /var/log/messages*
> > > > grep -i 'Bad page state' /var/log/messages*
> > > > grep -i 'CPA: called for zero pte' /var/log/messages*
> > >
> > > not a single occurance (this machine uses journal, but I checked those
> > > and no such messages.. in general i tend to check dmesg and system logs
> > > a lot, so I'd have already reported such messages..
> >
> > Yeah I don't know why it assumed you used antiquated logging..! :)
> >
> > OK that's interesting.
> >
> > >
> > > >
> > > > Any of these, particularly "bad pmd", is direct evidence that a freed
> > > > kernel PTE table was reused as a user page table. "CPA: called for zero
> > > > pte" would be the CPA walker itself tripping over a collapsed mapping.
> > > >
> > > > Questions:
> > > >
> > > > 1. swapon --show on the host, and inside the guests. Is there a swap
> > > > device with index 1 and a size of at least roughly 30.5GiB? That
> > > > tells us whether the corrupt word is a host swap PTE or a guest one,
> > > > which distinguishes case A from case B above.
> > > >
> > > > 2. From the vmcore, dump the 4K page containing 0xff2e6dbe0e51b440. If
> > > > the neighbouring words are also PTE-shaped, the page was being used
> > > > as a page table and the diagnosis above is confirmed. If only the
> > > > one word is corrupt, it was a single stray store.
> >
> > > unfortunately I don't have full vmcore from that crash, as it didn't fit
> > > to /var/crash, backtrace I posted is from vmcore-dmesg.txt so can't confirm
> > > that..
> >
> > Ah that's a pity!
> >
> > >
> > >
> > > >
> > > > 3. What are CONFIG_DEBUG_VM, init_on_alloc, init_on_free and
> > > > page_poison set to in the production build versus the KASAN build?
> > > > free_page_is_bad() is gated on is_check_pages_enabled(), which needs
> > > > CONFIG_DEBUG_VM, so the production kernel would not report the bad
> > > > free even if it happened.
> > >
> > > I don't have CONFIG_DEBUG_VM enabled in production..
> >
> > Well that explains the lack of bad reports above. I don't know why it'd
> > assume you'd run kernels with that (we do not recommend that for production
> > :)
> >
> > >
> > > >
> > > > 4. Are any of the crashing guests Windows, and is hv-tlbflush set on
> > > > them? That decides whether 26505e1b5b54 matters for you.
> > > yes, windows, but hv-tlbflush is enabled on a sigle VM and it runs od
> > > different node all the time.
> >
> > Ah but that could be enough to cause memory corruption. The reports seem to
> > be about guest memory corruption though.
> >
> > To be clear - are you observing it in the guest or host? I gathered host
> > from the splat.
>
> I experienced multiple host memory corruptions (broken .so libraries, etc)
> and also few visible from guests (GCC crashes while doing kernel builds in a loop,
> some guest panics.. and then few windows crashes (not sure about causes there,
> I'm no windows expert) but I guess all this can be cause by HOST side corruption
>
> > > > 5. Has any corruption occurred since THP was disabled? If yes, that
> > > > supports the CPA race over your THP theory.
> > > not yet, but it's not happening that often, so unsure here
> >
> > Yeah it seems to be a hard one to hit. If you tried a kernel with KASAN
> > enabled it might be flagged earlier? But that could also kill the race
> > window and would slow the system down a lot.
> >
> > >
> > > >
> > > > 6. When you run the 6.18.54-rc1 test, please note that 6.18.52 is only
> > > > a partial fix, so if you have any results from a 6.18.52 kernel they
> > > > should not be treated as a clean run.
> > > sure, I'll start today with 6.18.54, won't consider older tests.
> >
> > Ack, that's the best thing to do at the moment to be honest.
> >
> > If you were consistently getting corruption after X days previously, 2*X
> > days let's say of none can give confidence it's fixed there.
> I'm quite unsure what is the safe period here, as I mentioned to Luiz today,
> at least two times, I thought it's fixed after ~10 days of stress tests, deployed
> kernel to production.. and got another crash after few weeks..
>
>
> > > I surely will!
> > >
> > > cheers, nik
> >
> > Thanks! Given the nature of the bug and the fact the LLM went a little out
> > on a limb.
> >
> > Some more stuff from the report, which I also enclose in full here FYI.
> >
> > schlopp>>
> >
> > 26505e1b5b54 ("KVM: SVM: make svm_flush_tlb_gva do a full asid flush if
> > NPT enabled") is mainline only and has not been backported to 6.18.y.
> > It is AMD-specific guest memory corruption via the Hyper-V PV TLB flush
> > path, so it only affects Windows guests running with hv-tlbflush=on. If
> > any of your crashing guests are Windows, this is worth backporting
> > separately. It is independent of the CPA race.
> >
> > 55ddbc2ca6d5 (upstream f7491d7c81db, "x86/mm: Fix user-space data loss
> > with MADV_FREE and THP") is a real bug in 6.18.44. pmd_modify() drops
> > _PAGE_DIRTY, which has been the case since v6.10, and it causes data
> > loss for MADV_FREE THP and writable file THP. It is fixed in 6.18.53.
> > It is data loss rather than a stray write, so it does not explain the
> > oops, but it is another reason to move off 6.18.44. Note that this one
> > needs THP, which you have now disabled.
> >
> > d053eb7e09e1 (upstream 1e75a8255f11, "iommu/amd: Wait for completion
> > instead of returning early in iommu_completion_wait()") is already in
> > 6.18.42 and therefore in the kernel that crashed. It is ruled out for
> > this oops, but it may be relevant to the earlier incidents you had on
> > 6.18.31 through 6.18.41.
> >
> > <<schlopp
> >
> > Let us know how the tests get on! If you trigger a bug on 54 let us know
> > ASAP so we can investigate alternative theories.
>
> sure!
>
> BR nik
>
>
> >
> > Thanks!
> >
> > >
> > >
> > >
> > >
> > > >
> > > > --
> > > > Cheers, Lorenzo
> > > >
> > >
> > > --
> > > Ing. Nikola CIPRICH
> > > technický ředitel
> > >
> > > +420 591 166 214
> > > +420 777 093 799
> > > nikola.ciprich@linuxbox.cz
> > >
> > > www.linuxbox.cz
> >
> > --
> > Cheers, Lorenzo
>
> > Summary
> > =======
> >
> > The corruption you are seeing is consistent with a known use-after-free
> > in the x86 change_page_attr (CPA) code that is present in 6.18.44 and
> > was only fixed in 6.18.52 and 6.18.53.
> >
> > cpa_collapse_large_pages() rebuilds a leaf PMD out of its 4K PTEs and
> > then frees the old PTE table, with no lock held against the lockless
> > page table walk that __change_page_attr() performs before it stores
> > through the PTE pointer it cached. The stale 8-byte store of a kernel
> > PTE value lands in whatever the buddy allocator has since handed that
> > page out for.
> >
> > On a KVM host the two sides of this race are both hot. set_memory_rox()
> > is the only caller that passes CPA_COLLAPSE, and it runs on every module
> > load (execmem_restore_rox()), every ftrace trampoline creation
> > (arch/x86/kernel/ftrace.c:423) and every new 2M BPF program pack. The
> > other side is any lockless walk of the same execmem tables:
> > set_memory_nx()/set_memory_rw() from execmem_force_rw() on module load
> > and trampoline allocation, and vmalloc_to_page() inside __text_poke()
> > for every patch of module text, kprobe slot, trampoline or BPF pack.
> > Module text, kprobe slots and ftrace trampolines share the same 2M ROX
> > cache pages, so the collapser and the victim land in the same PMD by
> > construction. A libvirt host does all of this constantly: module
> > autoload (tun, vhost_net, br_netfilter, ebtable_*, xt_*), per-VM seccomp
> > filters, systemd cgroup BPF, perf and BPF probes, ftrace and static call
> > patching on every module load.
> >
> > This matches your good/bad window exactly. The collapse feature was
> > added in v6.15 and is not in 5.15.
> >
> > It also matches the KASAN result. free_page_is_bad() is gated on
> > is_check_pages_enabled(), which is off unless CONFIG_DEBUG_VM is set,
> > and KASAN changes the allocation pattern enough that the freed page
> > tends not to be reused in the race window. The upstream reporter only
> > reproduced it by injecting a delay at the CPA page table lookup.
> >
> > Recommendation: run 6.18.53 or later. 6.18.54-rc1, which you say you
> > are about to test, has the complete series. 6.18.52 has only the
> > cpa_lock patch, which does not close the window. Mainline (v7.3-rc5)
> > has nothing further pending for arch/x86/mm/pat/set_memory.c.
> >
> >
> > Kernel version
> > ==============
> >
> > 6.18.44lb9.01 (AlmaLinux 9 build of stable 6.18.44, stable commit
> > 1efe5d048a39). First seen on 6.18.31, last crash on 6.18.44. 5.15.x
> > was fine.
> >
> >
> > Machine
> > =======
> >
> > ASUSTeK RS720A-E12-RS12 / K14PP-D24, AMD EPYC, BIOS 2305 11/21/2025.
> > KVM host, AlmaLinux 9, Intel ice and Mellanox NICs. Taint "G E",
> > unsigned module only.
> >
> >
> > Stack trace
> > ===========
> >
> > Oops: general protection fault, probably for non-canonical address
> > 0xfffffff0c930038: 0000 [#1] SMP NOPTI
> > CPU: 23 UID: 189 PID: 7538 Comm: pacemaker-contr
> > RIP: 0010:__d_lookup+0x4a/0xc0
> > RAX: 0fffffff0c930020 RBX: 0fffffff0c930020 RCX: 0000000000000000
> > RDX: ff2e6dbe0d9b6000 RSI: ff7532a13699fe60 RDI: ff2e6d1e4e630d80
> > RBP: 000000000b654440
> > Call Trace:
> > <TASK>
> > d_lookup+0x27/0x50
> > lookup_dcache+0x1f/0x80
> > lookup_one_qstr_excl+0x1e/0xe0
> > filename_create+0xc4/0x160
> > do_mkdirat+0x5a/0x190
> > __x64_sys_mkdir+0x42/0x60
> > do_syscall_64+0x64/0xbf0
> > entry_SYSCALL_64_after_hwframe+0x76/0x7e
> > </TASK>
> >
> > Other messages you reported:
> >
> > Inconsistency detected by ld.so: ../sysdeps/x86_64/dl-machine.h: 548:
> > elf_machine_rela_relative: Assertion `ELFW(R_TYPE) (reloc->r_info) ==
> > R_X86_64_RELATIVE' failed!
> >
> >
> > What the oops registers say
> > ===========================
> >
> > The Code bytes decode to the hash chain walk in fs/dcache.c:__d_lookup():
> >
> > mov (%rbx),%rax ; load bucket->first
> > mov %rax,%rbx
> > and $-2,%rbx ; strip the hlist_bl lock bit
> > cmp $1,%rax
> > ja body
> > loop:
> > mov (%rbx),%rbx ; node->next
> > test %rbx,%rbx
> > je out
> > body:
> > cmp %ebp,0x18(%rbx) ; <-- faulting insn, d_name.hash_len
> > jne loop
> >
> > RAX equals RBX and RAX is only ever written by the initial bucket load,
> > so this is the first loop iteration. The corrupt word is the
> > hlist_bl_head.first field of the dentry_hashtable bucket itself, not a
> > dentry's d_hash.next.
> >
> > RDX 0xff2e6dbe0d9b6000 is the dentry_hashtable base, and the runtime
> > constant d_hash_shift is patched to 7, so:
> >
> > bucket = 0xff2e6dbe0d9b6000 + (0x0b654440 >> 7) * 8
> > = 0xff2e6dbe0e51b440
> >
> > That table is a 256MB alloc_large_system_hash() allocation from
> > memblock. It is allocated at boot, is PG_reserved and is never freed.
> > So this is a stray write to a fixed physical page, not a
> > use-after-free of a recycled object.
> >
> > The corrupt value 0x0fffffff0c930020 is an exact x86-64 non-present
> > swap PTE: __swp_type() is 1 and __swp_offset() is 0x79b67f, about
> > 30.4GiB into swap device 1. The only low bit set is bit 5,
> > _PAGE_ACCESSED.
> >
> > Every bit the swap layout constrains is as it should be: P, PSE and
> > the G/PROT_NONE alias are clear, the software bits 1-3 are clear, the
> > type is an ordinary swap type and the inverted offset gives the long
> > run of ones in bits 32-58. The one thing the layout does not account
> > for is bit 5. __swp_entry() leaves bits 0-8 clear and no helper sets
> > bit 5 on a non-present PTE; arch/x86/include/asm/pgtable_64.h treats
> > bits 5 and 6 as don't-care only because of the Intel Knights Landing
> > erratum (00839ee3b299, _PAGE_KNL_ERRATUM_MASK), which does not apply to
> > EPYC. So either something other than a Linux swap PTE happens to fit
> > this layout, or the word was a swap PTE that acquired a stray bit. I
> > cannot tell which from one word, which is why the vmcore page dump
> > requested below matters: 511 neighbouring PTE-shaped words would settle
> > it.
> >
> >
> > Suspect commit
> > ==============
> >
> > commit 41d88484c71cd4f659348da41b7b5b3dbd3be1f6
> > Author: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
> >
> > x86/mm/pat: restore large ROX pages after fragmentation
> >
> > Link: https://lore.kernel.org/r/20250126074733.1384926-4-rppt@kernel.org
> >
> > This added runtime collapse of split kernel large pages, driven from
> > cpa_flush(), including freeing the PTE table that the collapsed PMD
> > replaces. It is the Fixes: target of every fix listed below. It is in
> > v6.15 and later, and is not in 5.15.
> >
> > > diff --git a/arch/x86/mm/pat/set_memory.c b/arch/x86/mm/pat/set_memory.c
> > > --- a/arch/x86/mm/pat/set_memory.c
> > > +++ b/arch/x86/mm/pat/set_memory.c
> >
> > > @@ -394,6 +408,40 @@ static void __cpa_flush_tlb(void *data)
> > > +static void cpa_collapse_large_pages(struct cpa_data *cpa)
> > > +{
> > > + unsigned long start, addr, end;
> > > + struct ptdesc *ptdesc, *tmp;
> > > + LIST_HEAD(pgtables);
> > > + int collapsed = 0;
> > > + int i;
> >
> > [ ... range iteration ... ]
> >
> > > + if (!collapsed)
> > > + return;
> > > +
> > > + flush_tlb_all();
> > > +
> > > + list_for_each_entry_safe(ptdesc, tmp, &pgtables, pt_list) {
> > > + list_del(&ptdesc->pt_list);
> > > + __free_page(ptdesc_page(ptdesc));
> > > + }
> > > +}
> >
> > In 6.18.44 that last loop is pagetable_free(ptdesc), which is still an
> > immediate free. There is no RCU grace period and no other deferral. The
> > flush_tlb_all() above it only makes the hardware forget the old
> > translation; it does nothing about a CPU that is sitting inside
> > __change_page_attr() holding a pointer into that table.
> >
> > > @@ -402,7 +450,7 @@ static void cpa_flush(struct cpa_data *cpa, int cache)
> > > if (cache && !static_cpu_has(X86_FEATURE_CLFLUSH)) {
> > > cpa_flush_all(cache);
> > > - return;
> > > + goto collapse_large_pages;
> > > }
> >
> > [ ... ]
> >
> > > @@ -427,6 +475,10 @@ static void cpa_flush(struct cpa_data *cpa, int cache)
> > > mb();
> > > +
> > > +collapse_large_pages:
> > > + if (cpa->flags & CPA_COLLAPSE)
> > > + cpa_collapse_large_pages(cpa);
> > > }
> >
> > The collapse is hooked into cpa_flush(), which
> > __change_page_attr_set_clr() calls after it has already dropped
> > cpa_lock. So cpa_lock does not serialise the collapse against anything.
> >
> > > @@ -1196,6 +1248,161 @@ static int split_large_page(struct cpa_data *cpa, pte_t *kpte,
> > > +static int collapse_pmd_page(pmd_t *pmd, unsigned long addr,
> > > + struct list_head *pgtables)
> > > +{
> >
> > [ ... uniformity checks over all 512 PTEs ... ]
> >
> > > + old_pmd = *pmd;
> > > +
> > > + /* Success: set up a large page */
> > > + pgprot = pgprot_4k_2_large(pte_pgprot(first));
> > > + pgprot_val(pgprot) |= _PAGE_PSE;
> > > + _pmd = pfn_pmd(pfn, pgprot);
> > > + set_pmd(pmd, _pmd);
> > > +
> > > + /* Queue the page table to be freed after TLB flush */
> > > + list_add(&page_ptdesc(pmd_page(old_pmd))->pt_list, pgtables);
> >
> > collapse_large_pages(), the caller, takes pgd_lock around this. The
> > lockless CPA walker never takes pgd_lock, so pgd_lock does not help
> > either.
> >
> > The other side, in 6.18.44:
> >
> > arch/x86/mm/pat/set_memory.c:__change_page_attr() {
> > address = __cpa_addr(cpa, cpa->curpage);
> > repeat:
> > kpte = _lookup_address_cpa(cpa, address, &level, &nx, &rw);
> > ...
> > old_pte = *kpte;
> > ...
> > if (level == PG_LEVEL_4K) {
> > ...
> > new_pte = pfn_pte(pfn, new_prot);
> > ...
> > if (pte_val(old_pte) != pte_val(new_pte)) {
> > set_pte_atomic(kpte, new_pte); <-- stale
> > cpa->flags |= CPA_FLUSHTLB;
> > }
> >
> > _lookup_address_cpa() reaches lookup_address_in_pgd_attr(), which is a
> > plain lockless walk. Between the walk and the set_pte_atomic() the
> > caller can be preempted or take an interrupt; this runs with interrupts
> > on. That store is the only unsafe instruction in the function. The
> > large-page split branch further down is safe because __split_large_page()
> > revalidates under pgd_lock.
> >
> > The freed table is not even a tracked page table page in 6.18.44:
> >
> > arch/x86/mm/pat/set_memory.c:split_large_page() {
> > if (!debug_pagealloc_enabled())
> > spin_unlock(&cpa_lock);
> > base = alloc_pages(GFP_KERNEL, 0);
> >
> > A bare alloc_pages(), so it goes straight back to the per-CPU free list
> > and can be reallocated immediately.
> >
> >
> > Race timeline
> > =============
> >
> > CPU A (text_poke -> CPU B (module_enable_rox /
> > execmem_make_temp_rw -> execmem_restore_rox /
> > set_memory_nx/rw) bpf_jit_binary_lock_ro ->
> > set_memory_rox, CPA_COLLAPSE)
> > ----- -----
> > __change_page_attr()
> > kpte = _lookup_address_cpa()
> > old_pte = *kpte
> > new_pte = pfn_pte(...)
> > preempted / interrupted
> > __change_page_attr_set_clr()
> > drops cpa_lock
> > cpa_flush()
> > cpa_collapse_large_pages()
> > collapse_pmd_page(): all 512
> > PTEs uniform, set_pmd() installs
> > a leaf, old PTE table queued
> > flush_tlb_all()
> > pagetable_free() -> immediate
> > __free_pages()
> >
> > (any CPU) page is reallocated:
> > .so page cache folio, QEMU guest
> > RAM, a user PMD/PTE table, slab
> >
> > set_pte_atomic(kpte, new_pte)
> > stores a PTE-shaped word into
> > the reallocated page
> >
> > Where the swap PTE comes from (inferred continuation)
> > ------------------------------------------------------
> >
> > The race above writes a present kernel PTE, never a swap entry. To
> > reach the dentry hash table with a swap-PTE-shaped word the following
> > has to happen next. Each step is verified in the 6.18.44 code; the
> > sequence as a whole is inferred, not proven for this oops.
> >
> > 1. The freed PTE table is reallocated as a QEMU page table.
> >
> > 2. The stale set_pte_atomic() lands in it. The injected entry is a
> > translation into an execmem text page (case A/B below).
> >
> > 3. Case B: GUP-slow follows that entry and KVM maps the text page
> > into the guest; a later zap_present_folio_ptes() does an
> > unbalanced folio_put() and frees the still-live text page.
> >
> > 4. That text page is reallocated as another page table while
> > text_poke()/the BPF JIT keep writing instruction bytes into it
> > through the ROX mapping. Instruction bytes are now PMD entries
> > with arbitrary pfns; some pass pmd_bad().
> >
> > 5. Reclaim: try_to_unmap_one() -> page_vma_mapped_walk() ->
> > pte_offset_map_lock() reads such a PMD, computes
> > __va(garbage pfn) as the PTE table, and set_pte_at() stores a
> > swap PTE there. If that pfn is the dentry_hashtable page, one
> > bucket head becomes 0x0fffffff0c930020-like.
> >
> > 6. Days later __d_lookup() hashes into that bucket and faults.
> >
> > Step 5 is the only writer of an ordinary swap type in the mm, and
> > the only step that can touch memory the allocator never owned.
> >
> > This is not speculation about the code. The same interleaving was
> > reported upstream with a KASAN reproducer, in the commit that first
> > tried to address it:
> >
> > commit 1aac65f3e651 ("x86/mm/pat: Take cpa_lock around large-page
> > collapse")
> >
> > BUG: KASAN: use-after-free in __change_page_attr+0x7cc/0x7e0
> > Write of size 8 at addr ffff888181139718 by task modprobe
> > ...
> > The buggy address belongs to the physical page:
> > pfn:0x181139 ... page_type: f2(table)
> >
> > Fixes: 41d88484c71c ("x86/mm/pat: restore large ROX pages after
> > fragmentation")
> > Signed-off-by: Denis V. Lunev <den@openvz.org>
> >
> >
> > Which stable releases carry the fixes
> > =====================================
> >
> > None of these are in 6.18.44.
> >
> > The fixes that matter are the init_mm mmap lock pair, a1c7570cedd0 and
> > d5d8b8662e6e: the collapse runs under the init_mm write lock and the
> > whole attribute change, including the lockless walk and the store
> > through the cached pointer, runs under the read lock. That excludes
> > both walker-vs-collapse and collapse-vs-collapse. 1587d3394e25 covers
> > the third walker: __text_poke() resolves the pages it patches with
> > vmalloc_to_page(), a lockless walk of the same execmem tables, and now
> > takes the init_mm read lock around it. Without that, a collapse under
> > a concurrent text_poke() returns NULL (the BUG_ON at
> > arch/x86/kernel/alternative.c:2576 that openSUSE hit) or, if the freed
> > table has already been reused, a wrong page that text_poke() then
> > writes instruction bytes into. 9e4a3ec3411b makes the split tables
> > real kernel page tables so their freeing is deferred. The earlier
> > cpa_lock patch, 1aac65f3e651, is superseded by the write lock and adds
> > nothing once those are applied.
> >
> > 6.18.52
> > 591b6fac9df3 x86/mm/pat: Take cpa_lock around large-page collapse
> > (upstream 1aac65f3e651)
> >
> > 6.18.53
> > 35820cf8dd52 x86/mm/pat: Acquire init_mm write lock on collapse to
> > avoid UAF (upstream a1c7570cedd0)
> > e21a9ea81426 x86/mm/pat: Acquire init_mm read lock on attribute
> > changes to avoid UAF (upstream d5d8b8662e6e)
> > e164f4a25e23 x86/mm/pat: Convert split_large_page() to use ptdescs
> > 5029589bb773 x86/mm/pat: Don't gate cpa_lock on
> > debug_pagealloc_enabled()
> > 84e0cd79d57f x86/mm/pat: Allocate split page tables as kernel page
> > tables (upstream 9e4a3ec3411b)
> > 281e6f536f2f x86/alternatives: Exclude text poking against
> > change_page_attr() (upstream 1587d3394e25)
> > 74a2626044de x86/mm: Fix and document DEBUG_PAGEALLOC
> > (upstream 7da514d819a0)
> >
> > 6.18.52 does not fix this: it only takes cpa_lock around the collapse,
> > and the walker never holds cpa_lock across its walk-then-store window.
> > The init_mm mmap lock pair that closes that window is in 6.18.53, which
> > is the first stable release with the complete set. Use 6.18.53 or
> > later.
> >
> > There is no clean interim mitigation on 6.18.44. CPA_COLLAPSE cannot be
> > disabled by a boot parameter.
> >
> >
> > How this reaches the symptoms you saw
> > =====================================
> >
> > Symptoms 1 and 2 follow directly. The stale store writes a PTE-shaped
> > qword into a page that has been reallocated. If that page is a page
> > cache folio for a mapped .so, eight bytes of its relocation table are
> > replaced and ld.so trips the R_X86_64_RELATIVE assertion while the file
> > on disk is intact. If it is an anonymous page that QEMU has just
> > populated as guest RAM on a migration destination, the guest sees eight
> > corrupt bytes. Both of these match "right after migration": the
> > destination host is populating gigabytes of guest RAM and allocating
> > page tables at maximum rate, which is exactly when a just-freed page
> > gets reused inside the race window.
> >
> > Symptom 3, this oops, needs one more step, because the dentry hash table
> > is memblock memory that is never freed and therefore cannot be the
> > directly reallocated page. The escalation is that the victim page is
> > itself a page table. Each of the following links is verified in the
> > 6.18.44 source, but I want to be clear that the end-to-end chain for
> > this particular oops is plausible rather than proven from a single
> > vmcore.
> >
> > Case A, the victim is a user PMD table. pmd_bad() is true for the
> > injected value, so the first user-mode touch faults and
> > mm/pgtable-generic.c:___pte_offset_map() heals it via pmd_clear_bad()
> > with a "bad pmd" pr_err. But a supervisor-side copy_to_user() is
> > permitted through a U=0 level. A read() or recvmsg() into guest RAM, or
> > vhost-net, or kvm_write_guest(), then has the hardware walker read a
> > qword of module text as a PTE. If that qword happens to have P and RW
> > set, the copied data is written to an arbitrary physical address. This
> > route leaves no taint.
> >
> > Case B, the victim is a user PTE table. x86 has no pte_bad(). GUP-fast
> > rejects a U=0 entry, but GUP-slow does not:
> > mm/gup.c:follow_page_pte() and mm/memory.c:__vm_normal_page() check
> > neither _PAGE_USER nor PageReserved, so a live execmem page is returned
> > and KVM maps kernel module text into a guest. A later
> > mm/memory.c:zap_present_folio_ptes() does folio_remove_rmap_ptes() and
> > folio_put() on a page that was never rmapped, dropping the refcount to
> > zero, printing "BUG: Bad page map" and releasing live ROX text into the
> > buddy allocator while text_poke() and the BPF JIT keep writing
> > instruction bytes into it. Recycled as a user PMD table, instruction
> > bytes are page table entries with arbitrary pfns that pass pmd_bad(),
> > pte_offset_map() computes __va() of an arbitrary physical page, and
> > try_to_unmap_one() writes a genuine host swap PTE into it.
> >
> > That last step is what the corrupt word looks like: a real host swap
> > PTE. The stray bit 5 is unexplained on this route too.
> >
> > One caveat that argues against case B on this specific host: the taint
> > is "G E" with no "B". "BUG: Bad page map" had not fired on that
> > machine before the crash. So either the taint-free case A route or a
> > direct stray write carries this particular chain, or the corrupting
> > event happened on a different boot. This limits, but does not refute,
> > the cascade. The corruption can sit in a rarely used dentry bucket for
> > a long time before something hashes into it, which fits the 22-day
> > uptime on this crash and your observation that the host crashes were
> > not correlated with migration.
> >
> >
> > Secondary findings
> > ==================
> >
> > These came up while looking and are worth knowing about, but none of
> > them explains this oops.
> >
> > 26505e1b5b54 ("KVM: SVM: make svm_flush_tlb_gva do a full asid flush if
> > NPT enabled") is mainline only and has not been backported to 6.18.y.
> > It is AMD-specific guest memory corruption via the Hyper-V PV TLB flush
> > path, so it only affects Windows guests running with hv-tlbflush=on. If
> > any of your crashing guests are Windows, this is worth backporting
> > separately. It is independent of the CPA race.
> >
> > 55ddbc2ca6d5 (upstream f7491d7c81db, "x86/mm: Fix user-space data loss
> > with MADV_FREE and THP") is a real bug in 6.18.44. pmd_modify() drops
> > _PAGE_DIRTY, which has been the case since v6.10, and it causes data
> > loss for MADV_FREE THP and writable file THP. It is fixed in 6.18.53.
> > It is data loss rather than a stray write, so it does not explain the
> > oops, but it is another reason to move off 6.18.44. Note that this one
> > needs THP, which you have now disabled.
> >
> > d053eb7e09e1 (upstream 1e75a8255f11, "iommu/amd: Wait for completion
> > instead of returning early in iommu_completion_wait()") is already in
> > 6.18.42 and therefore in the kernel that crashed. It is ruled out for
> > this oops, but it may be relevant to the earlier incidents you had on
> > 6.18.31 through 6.18.41.
> >
> >
> > Theories that were eliminated
> > =============================
> >
> > KVM NPT mapping at too large a level, or with the wrong base pfn. The
> > mapping level comes from the host page tables and the pfn from GUP;
> > KVM cannot reach memblock memory on its own.
> >
> > A host mm swap or migration PTE stored through a stale page table
> > pointer. All the store sites are bounded and the pointer provenance
> > checks out.
> >
> > A missed MMU notifier invalidation. Notifier ordering on the recovery
> > path is correct, and this cannot reach never-freed memory.
> >
> > NIC DMA to the wrong address. Only a teardown-time page_pool
> > use-after-free turned up, and the iommu/amd completion-wait fix is
> > already in 6.18.42.
> >
> > 105d04edbec8 (upstream 33192a26cddea, mm/huge_memory huge_zero_pfn
> > race). Real, but not present in 6.18.44, and it cannot reach memblock
> > memory.
> >
> > 0a25ee42e7d1 (upstream 3d679b7cb31f, KVM x86/mmu CMPXCHG when clearing
> > the Accessed bit in the TDP MMU). Not in 6.18.44, and benign for this.
> >
> > TDP MMU in-place huge page recovery. Structurally excluded: notifier
> > zaps take mmu_lock for write, recovery takes it for read.
> >
> > memblock/buddy physical aliasing. This would produce "Bad page state"
> > reports, which you have not seen. Worth confirming from the vmcore, see
> > below.
> >
> > Neither THP nor NUMA balancing is involved in the CPA race. Disabling
> > them was a reasonable precaution but it will not stop this. If you keep
> > seeing corruption with THP off, that is consistent with the diagnosis
> > rather than against it.
> >
> >
> > What would confirm this
> > =======================
> >
> > Log greps, across all affected hosts and all boots, not just the ones
> > that crashed:
> >
> > grep -i 'Bad page map' /var/log/messages*
> > grep -i 'bad pmd' /var/log/messages*
> > grep -i 'bad pud' /var/log/messages*
> > grep -i 'Bad page state' /var/log/messages*
> > grep -i 'CPA: called for zero pte' /var/log/messages*
> >
> > Any of these, particularly "bad pmd", is direct evidence that a freed
> > kernel PTE table was reused as a user page table. "CPA: called for zero
> > pte" would be the CPA walker itself tripping over a collapsed mapping.
> >
> > Questions:
> >
> > 1. swapon --show on the host, and inside the guests. Is there a swap
> > device with index 1 and a size of at least roughly 30.5GiB? That
> > tells us whether the corrupt word is a host swap PTE or a guest one,
> > which distinguishes case A from case B above.
> >
> > 2. From the vmcore, dump the 4K page containing 0xff2e6dbe0e51b440. If
> > the neighbouring words are also PTE-shaped, the page was being used
> > as a page table and the diagnosis above is confirmed. If only the
> > one word is corrupt, it was a single stray store.
> >
> > 3. What are CONFIG_DEBUG_VM, init_on_alloc, init_on_free and
> > page_poison set to in the production build versus the KASAN build?
> > free_page_is_bad() is gated on is_check_pages_enabled(), which needs
> > CONFIG_DEBUG_VM, so the production kernel would not report the bad
> > free even if it happened.
> >
> > 4. Are any of the crashing guests Windows, and is hv-tlbflush set on
> > them? That decides whether 26505e1b5b54 matters for you.
> >
> > 5. Has any corruption occurred since THP was disabled? If yes, that
> > supports the CPA race over your THP theory.
> >
> > 6. When you run the 6.18.54-rc1 test, please note that 6.18.52 is only
> > a partial fix, so if you have any results from a 6.18.52 kernel they
> > should not be treated as a clean run.
>
>
> --
> Ing. Nikola CIPRICH
> technický ředitel
>
> +420 591 166 214
> +420 777 093 799
> nikola.ciprich@linuxbox.cz
>
> www.linuxbox.cz
>
--
Ing. Nikola CIPRICH
technický ředitel
+420 591 166 214
+420 777 093 799
nikola.ciprich@linuxbox.cz
www.linuxbox.cz
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: hunting memory corruption bug in 6.18.x
2026-09-29 19:11 ` Nikola Ciprich
@ 2026-09-30 9:07 ` Lorenzo Stoakes (ARM)
2026-09-30 18:40 ` Nikola Ciprich
0 siblings, 1 reply; 16+ messages in thread
From: Lorenzo Stoakes (ARM) @ 2026-09-30 9:07 UTC (permalink / raw)
To: Nikola Ciprich
Cc: linux-mm, linux-kernel, akpm, david, Mike Rapoport, Dave Hansen,
Pedro Falcato, Kiryl Shutsemau, luizcap
On Tue, Sep 29, 2026 at 09:11:53PM +0200, Nikola Ciprich wrote:
> so he reverted back to 6.18.15
>
> then he got crash of another node (to which he was migrating VMs from the first updated node), that one
> was still running 6.18.15
>
> pretty unfortunate planned work outcome :-/
>
> but it confirms that both 6.18.15 and unfortunately also 6.18.54 are affected
Yeah 6.18.15 is expected, I'd not say reverting is really worthwhile honestly,
given what you've observed previously.
Maybe worth checking if commit 26505e1b5b54 ("KVM: SVM: make svm_flush_tlb_gva
do a full asid flush if NPT enabled") helps in that case?
--
Cheers, Lorenzo
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: hunting memory corruption bug in 6.18.x
2026-09-30 9:07 ` Lorenzo Stoakes (ARM)
@ 2026-09-30 18:40 ` Nikola Ciprich
2026-10-01 19:40 ` Nikola Ciprich
0 siblings, 1 reply; 16+ messages in thread
From: Nikola Ciprich @ 2026-09-30 18:40 UTC (permalink / raw)
To: Lorenzo Stoakes (ARM)
Cc: linux-mm, linux-kernel, akpm, david, Mike Rapoport, Dave Hansen,
Pedro Falcato, Kiryl Shutsemau, luizcap, pbonzini,
Nikola Ciprich
(CC Paolo Bonzini)
Hello Lorenzo,
>
> Yeah 6.18.15 is expected, I'd not say reverting is really worthwhile honestly,
> given what you've observed previously.
yes, I wasn't available at the time he was dealing with that..
>
> Maybe worth checking if commit 26505e1b5b54 ("KVM: SVM: make svm_flush_tlb_gva
> do a full asid flush if NPT enabled") helps in that case?
sure, I'll do that.. however, the patch doesn't apply cleanly on top of 6.18.54 at
all.. what do you guys recommend, is it OK to adjust the patch to this kernel
(I have to admit I'm able to do that, but without any deep knowledge of the subsystem)
or do you recommend to apply some of the previous patches?
tried going through them, but its ~134 commits affecting svm.c between
v6.18 and 26505e1b5b54
cheers
nik
>
> --
> Cheers, Lorenzo
>
--
Ing. Nikola CIPRICH
technický ředitel
+420 591 166 214
+420 777 093 799
nikola.ciprich@linuxbox.cz
www.linuxbox.cz
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: hunting memory corruption bug in 6.18.x
2026-09-30 18:40 ` Nikola Ciprich
@ 2026-10-01 19:40 ` Nikola Ciprich
2026-10-01 22:21 ` David Laight
2026-10-02 9:50 ` Lorenzo Stoakes (ARM)
0 siblings, 2 replies; 16+ messages in thread
From: Nikola Ciprich @ 2026-10-01 19:40 UTC (permalink / raw)
To: Lorenzo Stoakes (ARM)
Cc: linux-mm, linux-kernel, akpm, david, Mike Rapoport, Dave Hansen,
Pedro Falcato, Kiryl Shutsemau, luizcap, pbonzini,
Nikola Ciprich
Hello again,
the good (?) news is, in the meantime we got another crash on different machine
and I have a kdump including complete vmcore. This one was 6.18.53
here are some details:
Hardware: Supermicro AS-2024US-TRT / H12DSU-iN, BIOS 3.5 (2025-09-22). Dual-socket AMD EPYC 7343 16-Core, 32 CPUs, 1024 GB RAM.
analyzed with crash + matching vmlinux debuginfo:
Oops:
general protection fault, probably for non-canonical address 0xfffffff0c930038
RIP: __d_lookup+0x4a/0xc0
Comm: systemd PID: 1274600 CPU: 23
Call trace:
__d_lookup
lookup_fast
walk_component
link_path_walk
path_openat
do_filp_open
do_sys_openat2
__x64_sys_openat
do_syscall_64
Exception frame registers:
RAX: 0fffffff0c930020 RBX: 0fffffff0c930020 RCX: 000000000000000b
RDX: ffff985ffd223600 RSI: ffffb18962823d70 RDI: ffff989ddf09b5c0
RBP: 000000005560450b R8: 000000007fffffff R9: fefefefefefefeff
R10: 0000000000000000 R11: 93c3d2eb02a31dda R12: ffff989ddf09b5c0
R13: ffff989ddf09b5c0 R14: ffffb18962823d70 R15: 0000000000000000
Faulting instruction (cmp %ebp,0x18(%rbx)) dereferences RBX+0x18. RBX held
the non-canonical value 0x0fffffff0c930020, giving fault address
0x0fffffff0c930038. __d_lookup was walking the d_hash (hlist_bl) bucket;
RBX was the node pointer being dereferenced.
Observations from the vmcore:
The faulting value 0x0fffffff0c930020 is non-canonical and is not a mapped
kernel address:
crash> kmem 0x0fffffff0c930020
kmem: cannot determine page for fffffff0c930020
fffffff0c930020: physical address not found in mem map
The dentry being looked up (RDI/R12/R13 = 0xffff989ddf09b5c0) is intact and
well-formed:
name "app.slice", len 9, d_name.hash 0xCE973022 (consistent)
d_op = kernfs_dops; valid d_parent, d_inode, d_sb
d_hash.next = 0x0 (this node is the end of its bucket chain)
The target dentry and its hash chain in the dump show no corruption; the
chain terminates cleanly.
No page migration, compaction, or THP activity was in progress on any CPU
at panic. "bt -a" filtered for migrate*/compact*/khugepaged/kcompactd/
kswapd/split_huge*/folio*/d_move/rename returned nothing.
Automatic NUMA balancing was disabled at crash time (read from kernel memory):
crash> p sysctl_numa_balancing_mode
$ = 0
Top-level (PMD) transparent hugepage policy was "never" at crash time:
crash> p/x transparent_hugepage_flags
$ = 0x1c0
Bits set: 6 (DEFRAG_REQ_MADV), 7 (DEFRAG_KHUGEPAGED), 8 (USE_ZERO_PAGE).
Bits 0 (TRANSPARENT_HUGEPAGE_FLAG) and 1 (REQ_MADV_FLAG) are clear, i.e.
sysfs enabled = never. Per-order mTHP controls (huge_anon_orders_*) were not
inspected for this dump, so mTHP state is not asserted here.
No MCE/EDAC/hardware-error records are present in the kernel log for this host.
I can provide the full vmcore and the matching vmlinux/debuginfo on request, and
run further crash queries against it.
not sure if this is of any help?
with regards
nik
On Wed, Sep 30, 2026 at 08:40:23PM +0200, Nikola Ciprich wrote:
> (CC Paolo Bonzini)
>
> Hello Lorenzo,
>
> >
> > Yeah 6.18.15 is expected, I'd not say reverting is really worthwhile honestly,
> > given what you've observed previously.
> yes, I wasn't available at the time he was dealing with that..
>
>
> >
> > Maybe worth checking if commit 26505e1b5b54 ("KVM: SVM: make svm_flush_tlb_gva
> > do a full asid flush if NPT enabled") helps in that case?
> sure, I'll do that.. however, the patch doesn't apply cleanly on top of 6.18.54 at
> all.. what do you guys recommend, is it OK to adjust the patch to this kernel
> (I have to admit I'm able to do that, but without any deep knowledge of the subsystem)
> or do you recommend to apply some of the previous patches?
>
> tried going through them, but its ~134 commits affecting svm.c between
> v6.18 and 26505e1b5b54
>
> cheers
>
> nik
>
> >
> > --
> > Cheers, Lorenzo
> >
>
> --
> Ing. Nikola CIPRICH
> technický ředitel
>
> +420 591 166 214
> +420 777 093 799
> nikola.ciprich@linuxbox.cz
>
> www.linuxbox.cz
>
--
Ing. Nikola CIPRICH
technický ředitel
+420 591 166 214
+420 777 093 799
nikola.ciprich@linuxbox.cz
www.linuxbox.cz
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: hunting memory corruption bug in 6.18.x
2026-10-01 19:40 ` Nikola Ciprich
@ 2026-10-01 22:21 ` David Laight
2026-10-02 9:50 ` Lorenzo Stoakes (ARM)
1 sibling, 0 replies; 16+ messages in thread
From: David Laight @ 2026-10-01 22:21 UTC (permalink / raw)
To: Nikola Ciprich
Cc: Lorenzo Stoakes (ARM),
linux-mm, linux-kernel, akpm, david, Mike Rapoport, Dave Hansen,
Pedro Falcato, Kiryl Shutsemau, luizcap, pbonzini
On Thu, 1 Oct 2026 21:40:58 +0200
Nikola Ciprich <nikola.ciprich@linuxbox.cz> wrote:
> Hello again,
>
> the good (?) news is, in the meantime we got another crash on different machine
> and I have a kdump including complete vmcore. This one was 6.18.53
>
> here are some details:
>
> Hardware: Supermicro AS-2024US-TRT / H12DSU-iN, BIOS 3.5 (2025-09-22). Dual-socket AMD EPYC 7343 16-Core, 32 CPUs, 1024 GB RAM.
>
> analyzed with crash + matching vmlinux debuginfo:
>
> Oops:
> general protection fault, probably for non-canonical address 0xfffffff0c930038
> RIP: __d_lookup+0x4a/0xc0
> Comm: systemd PID: 1274600 CPU: 23
> Call trace:
> __d_lookup
> lookup_fast
> walk_component
> link_path_walk
> path_openat
> do_filp_open
> do_sys_openat2
> __x64_sys_openat
> do_syscall_64
>
> Exception frame registers:
> RAX: 0fffffff0c930020 RBX: 0fffffff0c930020 RCX: 000000000000000b
> RDX: ffff985ffd223600 RSI: ffffb18962823d70 RDI: ffff989ddf09b5c0
> RBP: 000000005560450b R8: 000000007fffffff R9: fefefefefefefeff
> R10: 0000000000000000 R11: 93c3d2eb02a31dda R12: ffff989ddf09b5c0
> R13: ffff989ddf09b5c0 R14: ffffb18962823d70 R15: 0000000000000000
>
> Faulting instruction (cmp %ebp,0x18(%rbx)) dereferences RBX+0x18. RBX held
> the non-canonical value 0x0fffffff0c930020, giving fault address
> 0x0fffffff0c930038. __d_lookup was walking the d_hash (hlist_bl) bucket;
> RBX was the node pointer being dereferenced.
Surprisingly it looks like your compile matches the one I built from head.
The crash seems to be from the 'if (dentry->d_name.hash != hash) read.
Annoyingly the list is followed with 'mov (%rbx),%rbx' so you don't get the
address of the previous item.
However the same bad address is in %rax.
That would rather imply that it is the first time around the loop and
the 'bad address' came from the hash table itself.
(Unless the exception code manages to corrupt %rax.)
The list being corrupt would have to be memory reuse (for something else)
and the rcu protection not working.
I've just noticed that the RAX and RBX values (and the code RPC offset)
exactly match those in your original report from 25-sep.
That can't be a coincidence.
Has to be some kind of 'smoking gun'.
Possibly scanning the entire dump for 0x0c930020 might show it being
used somewhere?
David
>
> Observations from the vmcore:
>
> The faulting value 0x0fffffff0c930020 is non-canonical and is not a mapped
> kernel address:
> crash> kmem 0x0fffffff0c930020
> kmem: cannot determine page for fffffff0c930020
> fffffff0c930020: physical address not found in mem map
>
> The dentry being looked up (RDI/R12/R13 = 0xffff989ddf09b5c0) is intact and
> well-formed:
> name "app.slice", len 9, d_name.hash 0xCE973022 (consistent)
> d_op = kernfs_dops; valid d_parent, d_inode, d_sb
> d_hash.next = 0x0 (this node is the end of its bucket chain)
>
> The target dentry and its hash chain in the dump show no corruption; the
> chain terminates cleanly.
>
> No page migration, compaction, or THP activity was in progress on any CPU
> at panic. "bt -a" filtered for migrate*/compact*/khugepaged/kcompactd/
> kswapd/split_huge*/folio*/d_move/rename returned nothing.
>
> Automatic NUMA balancing was disabled at crash time (read from kernel memory):
> crash> p sysctl_numa_balancing_mode
> $ = 0
>
> Top-level (PMD) transparent hugepage policy was "never" at crash time:
> crash> p/x transparent_hugepage_flags
> $ = 0x1c0
> Bits set: 6 (DEFRAG_REQ_MADV), 7 (DEFRAG_KHUGEPAGED), 8 (USE_ZERO_PAGE).
> Bits 0 (TRANSPARENT_HUGEPAGE_FLAG) and 1 (REQ_MADV_FLAG) are clear, i.e.
> sysfs enabled = never. Per-order mTHP controls (huge_anon_orders_*) were not
> inspected for this dump, so mTHP state is not asserted here.
>
> No MCE/EDAC/hardware-error records are present in the kernel log for this host.
>
> I can provide the full vmcore and the matching vmlinux/debuginfo on request, and
> run further crash queries against it.
>
> not sure if this is of any help?
>
> with regards
>
> nik
>
>
>
> On Wed, Sep 30, 2026 at 08:40:23PM +0200, Nikola Ciprich wrote:
> > (CC Paolo Bonzini)
> >
> > Hello Lorenzo,
> >
> > >
> > > Yeah 6.18.15 is expected, I'd not say reverting is really worthwhile honestly,
> > > given what you've observed previously.
> > yes, I wasn't available at the time he was dealing with that..
> >
> >
> > >
> > > Maybe worth checking if commit 26505e1b5b54 ("KVM: SVM: make svm_flush_tlb_gva
> > > do a full asid flush if NPT enabled") helps in that case?
> > sure, I'll do that.. however, the patch doesn't apply cleanly on top of 6.18.54 at
> > all.. what do you guys recommend, is it OK to adjust the patch to this kernel
> > (I have to admit I'm able to do that, but without any deep knowledge of the subsystem)
> > or do you recommend to apply some of the previous patches?
> >
> > tried going through them, but its ~134 commits affecting svm.c between
> > v6.18 and 26505e1b5b54
> >
> > cheers
> >
> > nik
> >
> > >
> > > --
> > > Cheers, Lorenzo
> > >
> >
> > --
> > Ing. Nikola CIPRICH
> > technický ředitel
> >
> > +420 591 166 214
> > +420 777 093 799
> > nikola.ciprich@linuxbox.cz
> >
> > www.linuxbox.cz
> >
>
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: hunting memory corruption bug in 6.18.x
2026-10-01 19:40 ` Nikola Ciprich
2026-10-01 22:21 ` David Laight
@ 2026-10-02 9:50 ` Lorenzo Stoakes (ARM)
2026-10-02 14:55 ` Nikola Ciprich
1 sibling, 1 reply; 16+ messages in thread
From: Lorenzo Stoakes (ARM) @ 2026-10-02 9:50 UTC (permalink / raw)
To: Nikola Ciprich
Cc: linux-mm, linux-kernel, akpm, david, Mike Rapoport, Dave Hansen,
Pedro Falcato, Kiryl Shutsemau, luizcap, pbonzini,
Borislav Petkov, Tal Zussman, Rik van Riel, Matt Fleming
[-- Attachment #1: Type: text/plain, Size: 6668 bytes --]
+cc Boris, Tal, Rik, Matt FYI - seems another instance of the INVLPGB bug.
On Thu, Oct 01, 2026 at 09:40:58PM +0200, Nikola Ciprich wrote:
> Hello again,
>
> the good (?) news is, in the meantime we got another crash on different machine
> and I have a kdump including complete vmcore. This one was 6.18.53
>
> here are some details:
>
> Hardware: Supermicro AS-2024US-TRT / H12DSU-iN, BIOS 3.5 (2025-09-22). Dual-socket AMD EPYC 7343 16-Core, 32 CPUs, 1024 GB RAM.
>
> analyzed with crash + matching vmlinux debuginfo:
Thanks that's useful.
We can rule out CPA at this point.
>
> Oops:
> general protection fault, probably for non-canonical address 0xfffffff0c930038
> RIP: __d_lookup+0x4a/0xc0
> Comm: systemd PID: 1274600 CPU: 23
> Call trace:
> __d_lookup
> lookup_fast
> walk_component
> link_path_walk
> path_openat
> do_filp_open
> do_sys_openat2
> __x64_sys_openat
> do_syscall_64
>
> Exception frame registers:
> RAX: 0fffffff0c930020 RBX: 0fffffff0c930020 RCX: 000000000000000b
> RDX: ffff985ffd223600 RSI: ffffb18962823d70 RDI: ffff989ddf09b5c0
> RBP: 000000005560450b R8: 000000007fffffff R9: fefefefefefefeff
> R10: 0000000000000000 R11: 93c3d2eb02a31dda R12: ffff989ddf09b5c0
> R13: ffff989ddf09b5c0 R14: ffffb18962823d70 R15: 0000000000000000
>
> Faulting instruction (cmp %ebp,0x18(%rbx)) dereferences RBX+0x18. RBX held
> the non-canonical value 0x0fffffff0c930020, giving fault address
> 0x0fffffff0c930038. __d_lookup was walking the d_hash (hlist_bl) bucket;
> RBX was the node pointer being dereferenced.
>
> Observations from the vmcore:
>
> The faulting value 0x0fffffff0c930020 is non-canonical and is not a mapped
> kernel address:
> crash> kmem 0x0fffffff0c930020
> kmem: cannot determine page for fffffff0c930020
> fffffff0c930020: physical address not found in mem map
>
> The dentry being looked up (RDI/R12/R13 = 0xffff989ddf09b5c0) is intact and
> well-formed:
> name "app.slice", len 9, d_name.hash 0xCE973022 (consistent)
> d_op = kernfs_dops; valid d_parent, d_inode, d_sb
> d_hash.next = 0x0 (this node is the end of its bucket chain)
>
> The target dentry and its hash chain in the dump show no corruption; the
> chain terminates cleanly.
>
> No page migration, compaction, or THP activity was in progress on any CPU
> at panic. "bt -a" filtered for migrate*/compact*/khugepaged/kcompactd/
> kswapd/split_huge*/folio*/d_move/rename returned nothing.
>
> Automatic NUMA balancing was disabled at crash time (read from kernel memory):
> crash> p sysctl_numa_balancing_mode
> $ = 0
>
> Top-level (PMD) transparent hugepage policy was "never" at crash time:
> crash> p/x transparent_hugepage_flags
> $ = 0x1c0
> Bits set: 6 (DEFRAG_REQ_MADV), 7 (DEFRAG_KHUGEPAGED), 8 (USE_ZERO_PAGE).
> Bits 0 (TRANSPARENT_HUGEPAGE_FLAG) and 1 (REQ_MADV_FLAG) are clear, i.e.
> sysfs enabled = never. Per-order mTHP controls (huge_anon_orders_*) were not
> inspected for this dump, so mTHP state is not asserted here.
>
> No MCE/EDAC/hardware-error records are present in the kernel log for this host.
>
> I can provide the full vmcore and the matching vmlinux/debuginfo on request, and
> run further crash queries against it.
>
> not sure if this is of any help?
Very useful.
It looks like AMD TLB invalidation issues are the leading likely cause here
then - one on the host side with INVLPGB, and a separate one in KVM with
INVLPGA.
And it looks like a hardware bug, unfortunately.
Support for this was merged in 6.15 which matches your kernel versions too.
It's actually not solved yet, but there are two workarounds that can be
applied here.
## Issue 1: INVLPGB
See [0] for a report of the same kind of thing (segfaults like yours),
and [1] for the proposed temporary workaround.
TL;DR: The mitigation for this is to update your kernel command line parameters
thusly:
6.18.45 and later (6.18.53, 6.18.54): tlbi=ipi
<6.18.45: clearcpuid=419
If this resolves it then it confirms that this is the issue.
There is also, usefully, a reproducer which should show corruption
in minutes rather than weeks, see:
https://gist.github.com/mfleming/ca26ad3f8d65a12d23d62fb176480fc1
https://lore.kernel.org/all/akzuETUv5XEFSRL6@matt-Precision-5490/
The fact you've seen this on a Milan machine (EPYC 7343) is new information
so that could be useful for the report, so if you can reproduce _without
the fix_ first that'd be very useful to know!
Also then try the fix and see if it reproduces afterwards.
If the reproducer doesn't work then I guess worth waiting to see if hosts
reproduce over a longer time period with the tlbi=ipi issue.
Also, which CPU is in the first crash host (the ASUS SP5 box)? That'd be
useful to know thanks!
## Issue 2: INVLPGA
This is a separate issue, as mentioned before - Red Hat have seen windows
guest memory corruption on AMD with hv-tlbflush which tlbi=ipi won't
touch.
I had the AI generate a patch for you that applies to 6.18.53 and even
6.18.44 to make your life easier :)
The patch is attached.
For a kernel tree with git:
git am 0001-KVM-SVM-flush_tlb_gva-full-asid-flush-6.18.y.patch
Or just plain raw code:
patch -p1 < 0001-KVM-SVM-flush_tlb_gva-full-asid-flush-6.18.y.patch
So please test issue 1 separately with the cmdline change, but if you're
still seeing issues especially with windows guests, then this KVM patch
should then ALSO be applied.
This is [2] which is merged upstream as commit 26505e1b5b54 ("KVM: SVM:
make svm_flush_tlb_gva do a full asid flush if NPT enabled").
## Crash + vmcore
Obviously do the above and especially try the reproducer in the lab! :)
But also with the vmcore, can you run these 4 commands and reply with the
output please?
crash> p dentry_hashtable
crash> rd -64 0xffff986002783000 512
crash> search -p 0x0fffffff0c930020
crash> vtop 0xffff986002783a50
crash> search -p -m 0xfff0000000000fff <the PHYSICAL value vtop prints>
## Mitigations for production
I suggest you apply both of the above for your production as it's _likely_
it will resolve the issue there.
The cmdline changes are perfectly safe but you should check the generated
patch, however!
I suspect it'll be fine but I can't guarantee it obviously LLMs
etc. (albeit it is a frontier model so at least as good as it can be ;)
## Attachments
I attach the backported fix as discussed above and also the AI's full debug
report FYI.
--
Cheers, Lorenzo
[0]:https://lore.kernel.org/all/CAAuFnRTyva1_3tsF3vrMBL+TLS1YL4EgUPT2c3O9k7A9hWUMnA@mail.gmail.com/
[1]:https://lore.kernel.org/all/20260914045823.GFaqd-7yDf6aK4JUQr@fat_crate.local/
[2]:https://lore.kernel.org/all/20260723094419.630204-1-pbonzini@redhat.com/
[-- Attachment #2: 0001-KVM-SVM-flush_tlb_gva-full-asid-flush-6.18.y.patch --]
[-- Type: text/plain, Size: 2587 bytes --]
From: Paolo Bonzini <pbonzini@redhat.com>
Subject: [PATCH 6.18.y] KVM: SVM: make svm_flush_tlb_gva do a full asid flush if NPT enabled
commit 26505e1b5b546e2fa9a0296b951ca158460c72d8 upstream.
Red Hat is seeing multiple reports of Windows memory corruptions
(and consequent BSODs) with hv-tlbflush=on, on AMD processors only
(Turin and Milan; none on Intel; none on Turin with a full ASID flush).
When NPT is enabled, replace the per-address INVLPGA in
svm_flush_tlb_gva() with a full flush of the current ASID. This covers
both callers of the flush_tlb_gva op: kvm_hv_vcpu_flush_tlb() (Hyper-V
PV TLB flush, i.e. hv-tlbflush) and kvm_mmu_invalidate_addr() (L1
INVLPGA for nested SVM, emulated #PF, emulated INVLPG). Shadow paging
keeps using INVLPGA.
[ Backport note for 6.18.y: reduced to the svm.c functional change.
Upstream also changes the kvm_x86_ops.flush_tlb_gva signature to
return a "full" flag so that kvm_hv_vcpu_flush_tlb() stops iterating
once a full flush has been requested, and touches vmx/main.c,
vmx/vmx.c, vmx/x86_ops.h, hyperv.c and mmu.c for that. That part is
a performance optimisation only and is omitted here: without it the
Hyper-V flush loop keeps calling svm_flush_tlb_asid() for each
remaining page, which only re-sets TLB_CONTROL_FLUSH_ASID and is
cheaper than the INVLPGA it replaces. Upstream's svm_flush_tlb_guest()
additionally marks VCPU_REG_ERAPS dirty; ERAPS virtualisation does not
exist in 6.18.y, whose .flush_tlb_guest is svm_flush_tlb_asid(), so
calling svm_flush_tlb_asid() directly is the exact equivalent. ]
Analyzed-by: Vitaly Kuznetsov <vkuznets@redhat.com>
Analyzed-by: Alexander Lougovski <alougovs@redhat.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
---
arch/x86/kvm/svm/svm.c | 13 ++++++++++++-
1 file changed, 12 insertions(+), 1 deletion(-)
diff --git a/arch/x86/kvm/svm/svm.c b/arch/x86/kvm/svm/svm.c
index a24a6871b693..5bfb72e550e2 100644
--- a/arch/x86/kvm/svm/svm.c
+++ b/arch/x86/kvm/svm/svm.c
@@ -4071,7 +4071,18 @@ static void svm_flush_tlb_gva(struct kvm_vcpu *vcpu, gva_t gva)
{
struct vcpu_svm *svm = to_svm(vcpu);
- invlpga(gva, svm->vmcb->control.asid);
+ /*
+ * INVLPGA has had errata on Genoa and Turin, and even on older
+ * generations there were reports of Windows BSODs if INVLPGA
+ * was used for Hyper-V tlbflush. Use it only for shadow paging
+ * where it seems to be okay.
+ */
+ if (!npt_enabled) {
+ invlpga(gva, svm->vmcb->control.asid);
+ return;
+ }
+
+ svm_flush_tlb_asid(vcpu);
}
static inline void sync_cr8_to_lapic(struct kvm_vcpu *vcpu)
[-- Attachment #3: debug-report.txt --]
[-- Type: text/plain, Size: 33526 bytes --]
Summary
=======
The 6.18.53 crash rules out the CPA collapse race I pointed at earlier.
6.18.53 carries the complete CPA series (a1c7570cedd0, d5d8b8662e6e,
1587d3394e25, 9e4a3ec3411b and the rest). The new oops is the same as
the 6.18.44 one: same RIP, __d_lookup+0x4a, and the same corrupt bucket
word, 0x0fffffff0c930020. It happened on a different host (Milan
instead of the SP5 box), with a different build and a different KASLR
layout. On top of that, 6.18.54 corrupted libglib in page cache and a
6.18.15 node crashed. The CPA bugs are real, but they are not this bug,
and my earlier diagnosis was wrong.
The leading candidate is now the AMD INVLPGB/TLBSYNC stale-TLB issue.
It was reported publicly in June and no kernel has a fix for it yet:
"PROBLEM: Probabilistic segfault on AMD hardware with INVLPGB"
Henrik Boving, 2026-06-19
Message-ID: <CAAuFnRTyva1_3tsF3vrMBL+TLS1YL4EgUPT2c3O9k7A9hWUMnA@mail.gmail.com>
https://lore.kernel.org/all/CAAuFnRTyva1_3tsF3vrMBL+TLS1YL4EgUPT2c3O9k7A9hWUMnA@mail.gmail.com/
- Henrik saw heap corruption on an EPYC 9455 (Turin) from 6.15 onwards.
He bisected it to CONFIG_BROADCAST_TLB_FLUSH.
- Matt Fleming posted a userspace reproducer on 2026-07-07. It reuses
VAs with munmap() + mmap(MAP_FIXED) under a rwlock. Reader threads
still see data from the previous mapping after the INVLPGB + TLBSYNC
flush has returned:
https://lore.kernel.org/all/akzuETUv5XEFSRL6@matt-Precision-5490/
- Rik reproduced it and posted a double-TLBSYNC diff, which is not
merged. He also noted that Meta saw elevated segfault rates on Turin
that the AMD-SB-3029 firmware fixed. Henrik's microcode (0x0B002162)
is newer than that fix, so firmware does not explain his case.
- Tal Zussman confirmed it on an EPYC 9965 (Turin Dense) on 2026-09-13.
It corrupts in the first round, and with clearcpuid=419 it passes 20
rounds:
https://lore.kernel.org/all/20260914022055.1639690-1-tz2294@columbia.edu/
- Borislav answered on the same day: "There will be an official thing
Soon(tm). In the meantime, tlbi=ipi":
https://lore.kernel.org/all/20260914045823.GFaqd-7yDf6aK4JUQr@fat_crate.local/
Every public confirmation so far is Zen5 (Turin and Turin Dense). Rik
asked whether Milan or Bergamo reproduce it, and nobody has answered.
Your second crash host is Milan (EPYC 7343, Zen3). A result from your
hosts, positive or negative, is therefore new information for the x86
maintainers.
It fits what you have seen:
- AMD only.
- 5.15 is fine and 6.18 is not. The broadcast flush code went in in
v6.15.
- QEMU does heavy mmap/munmap churn around migration.
- It is timing sensitive: you could not reproduce it with KASAN or
SLUB debugging enabled.
- Host .so files are corrupted in page cache while the files on disk
are intact.
I can't tie the dentry hash table damage to it directly. Further down
there is a hypothesis for that, clearly labelled, together with the
vmcore checks that would confirm or refute it.
What to do now
==============
Disable broadcast TLB flushing on every AMD host. No patch is needed.
6.18.45 and later (6.18.53, 6.18.54): tlbi=ipi
6.18.15 and 6.18.44: clearcpuid=419
tlbi= only exists from 6.18.45 onwards (upstream abe7c8b09bd7,
backported as 846b92e26c8a), so 6.18.15 and 6.18.44 have to use
clearcpuid=419. 419 is 13*32+3, i.e. X86_FEATURE_INVLPGB.
clearcpuid=419 works on 6.18.45+ as well, but tlbi=ipi does not taint
the kernel. These are the parameters Borislav and Tal used in the
thread above.
Details that matter:
- The parameter has to be exactly "tlbi=ipi". The handler only matches
"ipi" and returns 1 for anything else. "tlbi=off" is therefore
accepted silently and does nothing:
arch/x86/kernel/cpu/common.c (6.18.53):
static int __init tlbi_setup(char *str)
{
if (!strcmp(str, "ipi"))
setup_clear_cpu_cap(X86_FEATURE_INVLPGB);
return 1;
}
__setup("tlbi=", tlbi_setup);
- "clearcpuid=invlpgb" does not work on 6.18. X86_FEATURE_INVLPGB has no
name string in cpufeatures.h, so the boot log says "clearcpuid:
unknown CPU flag: invlpgb" and INVLPGB stays enabled. Use the number.
- With clearcpuid=419 the kernel prints "clearcpuid: force-disabling
CPU feature flag: 13:3", then the "setcpuid=/clearcpuid= in use ...
Tainting kernel" warning, and sets taint S. That is expected.
- Do not use "nopcid" as a workaround on 6.18.15. 44126343d58c
("x86/mm: Disable broadcast TLB flush when PCID is disabled") only
arrived in 6.18.35. Without it, nopcid leaves INVLPGB enabled, and
the first broadcast flush with a non-zero PCID takes a #GP in
broadcast_tlb_flush().
How to check that it is in effect:
- /proc/cpuinfo can't tell you. The flag has no name in 6.18, so
"invlpgb" never appears there, whether it is enabled or not.
- CONFIG_BROADCAST_TLB_FLUSH is "def_bool y" with "depends on
CPU_SUP_AMD && 64BIT" and no prompt. Every 6.18 x86-64 build with AMD
support has it:
grep BROADCAST_TLB_FLUSH /boot/config-$(uname -r)
- The definitive check is the capability word. You can read it with
crash or drgn and debuginfo, either on a live system or in a vmcore:
crash> p/x boot_cpu_data.x86_capability[13]
If bit 3 (0x8) is set, the kernel is using INVLPGB. If it is clear,
it is not.
- For tlbi=ipi, check /proc/cmdline. For clearcpuid=419, look for the
dmesg line above and for taint bit 2 (value & 4 in
/proc/sys/kernel/tainted).
- As a sanity check, the "TLB shootdowns" row in /proc/interrupts
should rise noticeably faster under the same load, because QEMU's
flushes go back to IPIs.
Testing whether your CPUs are affected, with Matt's reproducer:
https://gist.github.com/mfleming/ca26ad3f8d65a12d23d62fb176480fc1
https://lore.kernel.org/all/akzuETUv5XEFSRL6@matt-Precision-5490/
build: gcc -O2 -std=c11 -pthread -static -o repro-invlpgb <source>.c
run: ./repro-invlpgb --batch --rounds 20 --jobs 32 -d 5 -w 8 -m 2 -s 512 -q
Run it on a drained Milan host and a drained Genoa host, once with the
default boot and once with tlbi=ipi. If it fails on Zen3 or Zen4 with
the default boot and passes with tlbi=ipi, that is new information and
should go to the INVLPGB thread above. A pass with the default boot is
weaker evidence, because the reproducer was tuned on Zen5.
Even if the reproducer passes, running production with tlbi=ipi is
the real test. Given how intermittent this is, a clean run has to last
several times longer than your previous time to failure before it
means much.
Kernel versions and machines
============================
Crash 1: 6.18.44 (6.18.44lb9.01). ASUSTeK RS720A-E12-RS12 /
K14PP-D24, BIOS 2305 11/21/2025, AMD EPYC on SP5. Windows
guests only, no host swap, about 22 days of uptime.
pacemaker-controld in mkdir().
Crash 2: 6.18.53. Supermicro AS-2024US-TRT / H12DSU-iN, BIOS 3.5
(2025-09-22), 2x EPYC 7343 (Milan, Zen3), 1 TB RAM.
THP "never", NUMA balancing 0, no MCE/EDAC records. systemd
in openat(). The full vmcore is available.
6.18.54 production node (nrbphav4a): a libglib segfault storm within
minutes of receiving migrated VMs (see below).
6.18.15 node: crashed while VMs were being migrated to it. No
details were posted.
5.15.x: clean.
The CPU model of the first host is not in the report. I have been
assuming Genoa, but the board is SP5, which takes both Genoa (9004,
Zen4) and Turin (9005, Zen5). If that box is actually Turin, it falls
straight into the publicly confirmed set. Please send the "model name"
and "microcode" lines from /proc/cpuinfo for every affected host,
including nrbphav4a.
Stack traces
============
Crash 1 (6.18.44):
Oops: general protection fault, probably for non-canonical address
0xfffffff0c930038: 0000 [#1] SMP NOPTI
CPU: 23 UID: 189 PID: 7538 Comm: pacemaker-contr 6.18.44lb9.01 #1
RIP: 0010:__d_lookup+0x4a/0xc0
RAX: 0fffffff0c930020 RBX: 0fffffff0c930020 RCX: 0000000000000000
RDX: ff2e6dbe0d9b6000 RSI: ff7532a13699fe60 RDI: ff2e6d1e4e630d80
RBP: 000000000b654440
Call Trace:
<TASK>
d_lookup+0x27/0x50
lookup_dcache+0x1f/0x80
lookup_one_qstr_excl+0x1e/0xe0
filename_create+0xc4/0x160
do_mkdirat+0x5a/0x190
__x64_sys_mkdir+0x42/0x60
do_syscall_64+0x64/0xbf0
entry_SYSCALL_64_after_hwframe+0x76/0x7e
</TASK>
Crash 2 (6.18.53, from crash(8) on the vmcore):
general protection fault, probably for non-canonical address
0xfffffff0c930038
RIP: __d_lookup+0x4a/0xc0
Comm: systemd PID: 1274600 CPU: 23
RAX: 0fffffff0c930020 RBX: 0fffffff0c930020 RCX: 000000000000000b
RDX: ffff985ffd223600 RSI: ffffb18962823d70 RDI: ffff989ddf09b5c0
RBP: 000000005560450b
Call trace:
__d_lookup
lookup_fast
walk_component
link_path_walk
path_openat
do_filp_open
do_sys_openat2
__x64_sys_openat
do_syscall_64
In crash 2 the parent dentry (RDI, "app.slice", kernfs) is intact and
its own hash chain terminates cleanly.
Other relevant messages
=======================
On 6.18.54, nrbphav4a, right after receiving migrated VMs, every glib
user faulted at the same file offset:
qemu-system-x86[19014]: segfault at 0 ip 00007f9a3ee7e61e sp 00007ffd96294380 error 6 in libglib-2.0.so.0.6800.4[a461e,7f9a3edf7000+91000]
pacemaker-execd[17696]: segfault at 0 ip 00007f3c6ed6a61e sp 00007fffa95873d0 error 6 in libglib-2.0.so.0.6800.4[a461e,7f3c6ece3000+91000]
Code: 29 9c 24 80 00 00 00 0f 29 a4 24 90 00 00 00 0f 29 ac 24 a0 00 00 00 0f 29 b4 ff ff 00 00 00 0f 29 bc ff ff 00 00 00 64 00 00 <00> 00 28 00 00 00 48
The bytes around the fault are a run of movaps stores to the stack
followed by the stack-protector load. That pattern tells us what the
original bytes were:
file offset expected found
0xa460e 24 b0 ff ff movaps %xmm6,0xb0(%rsp)
0xa4616 24 c0 ff ff movaps %xmm7,0xc0(%rsp)
0xa461c 48 8b 04 25 00 00 00 00 mov %fs:0x28,%rax
The file on disk will confirm the expected column.
Every process maps the same page-cache page, so they all fault at the
same file offset. The damage consists of two 16-bit stores of 0xffff and one
32-bit store of 0, at +6, +6 and +4 of three consecutive 8-byte slots
(page offsets 0x608, 0x610 and 0x618).
These are ordinary CPU stores to small struct fields. They are not DMA
(NIC descriptor write-backs are 16 or 32 bytes) and they are not
page-walker A/D bit updates. This is what a CPU writing through a
translation that no longer belongs to the writer looks like. The
earlier ld.so relocation assertion and the libcrypto.so.3 corruption
are the same class of damage.
Register decode (short)
=======================
This is already on the list, so I'll keep it brief.
In both oopses RAX == RBX == 0x0fffffff0c930020. RAX is only written by
the initial load of the bucket head, so the fault is on the first
iteration of the loop:
fs/dcache.c:__d_lookup() {
struct hlist_bl_head *b = d_hash(hash);
...
hlist_bl_for_each_entry_rcu(dentry, node, b, d_hash) {
if (dentry->d_name.hash != hash) <-- faults on 0x18(%rbx)
continue;
The corrupt word is therefore the dentry_hashtable bucket head itself,
not a dentry. d_hash_shift is patched to 7, i.e. 2^25 buckets and a
256 MiB table:
crash 1: 0xff2e6dbe0d9b6000 + (0x0b654440 >> 7) * 8 = 0xff2e6dbe0e51b440
(page offset 0x440)
crash 2: 0xffff985ffd223600 + (0x5560450b >> 7) * 8 = 0xffff986002783a50
(page offset 0xa50)
The crash 2 line assumes that RDX holds the table base, as it does in
crash 1 (same code, same RIP). "p dentry_hashtable" in the vmcore will
confirm it.
The table comes from alloc_large_system_hash() with HASH_EARLY. That is
memblock memory, allocated at boot and never freed, and nothing
legitimately writes a non-pointer into it. So whatever wrote the word
used a wrong translation or a wrong physical address.
The same 64-bit value on two hosts, with different CPUs, builds and
KASLR layouts, means the data written is deterministic. As David said,
that can't be a coincidence.
What the word decodes to:
- As an x86-64 Linux swap PTE it fits exactly (type 1, offset
0x79b67f), with bit 5 set. But the first host has no swap and no
Linux guests, so a Linux swap PTE is an unlikely origin.
- It is not a KVM SPTE. SVM MMIO SPTEs have bit 0 set, and non-present
and frozen SPTEs have bit 63 set. It is also not an AVIC physical or
logical ID table entry, and not any VMCB field.
- As a Windows x64 software PTE (MMPTE_SOFTWARE) it reads Valid=0,
Protection=1 (MM_READONLY), Prototype=0, Transition=0,
PageFileLow=0 and PageFileHigh=0x0fffffff. That PageFileHigh is the
"lookup needed" marker 0xffffffff with bits 60-63 clear. However, the
UsedPageTableEntries (0x93), ShadowStack and Unused fields are also
non-zero. This fit is weak. It rests on public Windows internals
documentation, not on anything I can check in kernel sources.
The suspect code
================
commit 767ae437a32d644786c0779d0d54492ff9cbe574
Author: Rik van Riel <riel@surriel.com>
x86/mm: Add INVLPGB feature and Kconfig entry
Link: https://lore.kernel.org/r/20250226030129.530345-3-riel@surriel.com
This commit starts the v6.15 broadcast TLB flush series, and the tlbi=
switch (abe7c8b09bd7) names it as its Fixes: target. I am not claiming
a bug in this code. As far as I can tell, the hardware does not
provide the completion guarantee that the code relies on. The series
explains why 5.15 is good, 6.18 is bad and only AMD is affected.
> diff --git a/arch/x86/Kconfig.cpu b/arch/x86/Kconfig.cpu
> --- a/arch/x86/Kconfig.cpu
> +++ b/arch/x86/Kconfig.cpu
> @@ -334,6 +334,10 @@ menuconfig PROCESSOR_SELECT
[ ... ]
> +config BROADCAST_TLB_FLUSH
> + def_bool y
> + depends on CPU_SUP_AMD && 64BIT
There is no prompt, so a build can't opt out. The only switches are at
boot time.
> diff --git a/arch/x86/include/asm/cpufeatures.h b/arch/x86/include/asm/cpufeatures.h
> --- a/arch/x86/include/asm/cpufeatures.h
> +++ b/arch/x86/include/asm/cpufeatures.h
> @@ -338,6 +338,7 @@
[ ... ]
> #define X86_FEATURE_XSAVEERPTR (13*32+ 2) /* "xsaveerptr" Always save/restore FP error pointers */
> +#define X86_FEATURE_INVLPGB (13*32+ 3) /* INVLPGB and TLBSYNC instructions supported */
The comment has no quoted name. That is why the flag is invisible in
/proc/cpuinfo and why clearcpuid= only accepts it by number.
[ ... ]
commit 4afeb0ed1753ebcad93ee3b45427ce85e9c8ec40
Author: Rik van Riel <riel@surriel.com>
x86/mm: Enable broadcast TLB invalidation for multi-threaded processes
Link: https://lore.kernel.org/r/20250226030129.530345-11-riel@surriel.com
This commit moves ordinary user TLB flushes for large multi-threaded
processes onto INVLPGB. On a KVM host, that means QEMU.
> diff --git a/arch/x86/mm/tlb.c b/arch/x86/mm/tlb.c
> --- a/arch/x86/mm/tlb.c
> +++ b/arch/x86/mm/tlb.c
> @@ -430,6 +430,105 @@ static bool mm_needs_global_asid(struct mm_struct *mm, u16 asid)
[ ... ]
> +static void consider_global_asid(struct mm_struct *mm)
> +{
> + if (!cpu_feature_enabled(X86_FEATURE_INVLPGB))
> + return;
> +
> + /* Check every once in a while. */
> + if ((current->pid & 0x1f) != (jiffies & 0x1f))
> + return;
> +
> + /*
> + * Assign a global ASID if the process is active on
> + * 4 or more CPUs simultaneously.
> + */
> + if (mm_active_cpus_exceeds(mm, 3))
> + use_global_asid(mm);
> +}
A CPU running a vCPU thread in guest mode still has QEMU's mm loaded.
Any VM with four or more busy vCPUs therefore gets a global ASID
quickly.
[ ... ]
> +static void broadcast_tlb_flush(struct flush_tlb_info *info)
> +{
[ ... ]
> + } else do {
[ ... ]
> + invlpgb_flush_user_nr_nosync(kern_pcid(asid), addr, nr, pmd);
[ ... ]
> + } while (addr < info->end);
> +
> + finish_asid_transition(info);
> +
> + /* Wait for the INVLPGBs kicked off above to finish. */
> + __tlbsync();
> +}
The kernel's correctness argument rests on this TLBSYNC. Once it
returns, no CPU may still hold the old translation, and only after that
are the pages and page tables freed:
mm/mmu_gather.c:tlb_flush_mmu() {
tlb_flush_mmu_tlbonly(tlb); <-- flush_tlb_mm_range(): INVLPGB + TLBSYNC
tlb_flush_mmu_free(tlb); <-- pages and page tables freed
}
Matt's reproducer shows that on Zen5 another CPU can keep using the old
translation after TLBSYNC has returned.
> @@ -1260,9 +1359,12 @@ void flush_tlb_mm_range(struct mm_struct *mm, unsigned long start,
[ ... ]
> - if (cpumask_any_but(mm_cpumask(mm), cpu) < nr_cpu_ids) {
> + if (mm_global_asid(mm)) {
> + broadcast_tlb_flush(info);
> + } else if (cpumask_any_but(mm_cpumask(mm), cpu) < nr_cpu_ids) {
> info->trim_cpumask = should_trim_cpumask(mm);
> flush_tlb_multi(mm_cpumask(mm), info);
> + consider_global_asid(mm);
Once QEMU has a global ASID, every flush of its address space uses
INVLPGB + TLBSYNC and no IPI. That covers munmap(), MADV_DONTNEED
(balloon, free page reporting), mprotect() and page migration.
Global ASIDs aren't the only route. In 6.18 the batched unmap flush
used by reclaim and migration (arch_tlbbatch_flush()) issues an INVLPGB
for all non-global entries, for every process. Kernel range flushes use
INVLPGB as well.
Software audit. I went through the INVLPGB paths in 6.18.53 looking
for a kernel-side hole:
- munmap/zap with freed page tables
- the dynamic-to-global ASID transition
- global ASID reuse
- lazy CPUs and CPUs in the middle of a switch
- reclaim batching
- kernel range flushes
I did not find one:
- TLBSYNC runs on the issuing CPU, with preemption disabled, before
any page or page table is freed.
- INVLPGB reaches lazy CPUs.
- finish_asid_transition() sends IPIs to stragglers still on a dynamic
ASID.
- A global ASID is only reused after an INVLPGB of all non-global
entries plus a TLBSYNC.
The code is identical in 6.18.15, 6.18.44 and 6.18.53, and every
post-6.15 fix to it is in 6.18.53. KVM never issues INVLPGB and never
exposes it to guests.
How this could reach the dentry hash table (hypothesis)
=======================================================
Everything in this section is inference. None of it has been
demonstrated.
A stale leaf translation can only reach the page that was freed. That
covers the libglib and ld.so damage, where the freed page came back as
page cache, and guest RAM corruption. It cannot reach dentry_hashtable,
which the page allocator never owned.
A stale paging-structure entry can. Suppose the entry that survives on
another CPU is the cached PDE pointing at a PTE table that
free_pgtables() has just released. That CPU then walks whatever the
freed page now holds as if it were a PTE table. Its stores to any VA in
that 2 MiB region go to whatever PFNs the entries in the reused page
happen to name, and that includes boot-time memory.
The identical word follows if the data being stored is guest RAM. A
Windows page-table page commonly holds long runs of one identical
software PTE. Suppose QEMU copies such a guest page through the broken
walk, for example on migration receive or in a virtio/vhost copy into
guest memory. It then writes 512 copies of the same 8-byte value over
one host physical page.
If that page belongs to dentry_hashtable, every bucket in it reads
0x0fffffff0c930020. Any lookup that hashes into one of those 512
buckets then faults the way both oopses did, at whatever page offset it
lands on (0x440 in one crash, 0xa50 in the other). Two hosts running
Windows guests would end up with the same word.
CPU A (QEMU thread) CPU B (QEMU vCPU/IO/vhost
thread, same mm, global ASID)
----- -----
munmap() of a region
zap, free_pgtables()
tlb_finish_mmu()
flush_tlb_mm_range(freed_tables)
INVLPGB (PCID, VA range)
TLBSYNC returns
PTE table page P freed
still caches the PDE -> P
(the hardware issue)
P reallocated, e.g. as guest
RAM, now holding guest data
store to a VA in the same 2 MiB
region (VA reused by a new mmap)
walks P as a PTE table and
writes to the PFN it finds
a 4K copy of a Windows page
table page, 512 copies of
0x0fffffff0c930020, lands on
a dentry_hashtable page
days later:
__d_lookup() hashes into that
page and faults
There are weak points, and I want to be upfront about them:
- The public reproducer demonstrates stale leaf data only. This chain
needs a stale paging-structure entry.
- The Windows reading of the word is weak (see above).
- The guest mix on the Milan host isn't stated. If it runs no Windows
guests, the Windows part of this falls apart.
- The target PFN comes from whatever the reused page holds, so other
pages would be hit too. The dentry table is just large, read
constantly and never freed, which makes it a likely place to notice
the damage.
What the 6.18.53 vmcore can settle
==================================
B below is the crash 2 bucket address from above.
1. Bucket and physical page:
p dentry_hashtable
p d_hash_shift
eval 0xffff985ffd223600 + (0x5560450b >> 7) * 8 -> B
vtop B -> PHYS_B
kmem B
eval B & ~0xfff -> PAGE_B
2. The whole 4K page. This is the most important check:
rd -64 PAGE_B 512
Then do the same for the pages at PAGE_B - 0x1000 and PAGE_B +
0x1000 (compute them with eval first).
If the hypothesis holds, most of the 512 words in PAGE_B are
0x0fffffff0c930020, and the run starts and ends on the 4K boundary.
A few buckets may hold valid dentry pointers inserted after the
damage, and their chains should end in the same word.
If only the one word is bad and the rest are 0 or dentry pointers,
the hypothesis is refuted. In that case this was a single 8-byte
store, and the 4K copy story is wrong. The neighbouring pages show
whether there was more than one event.
3. Where else the word lives:
search -p 0x0fffffff0c930020
search -p -w 0x0c930020
Physical searches over 1 TB take a while. For each hit, use kmem -p
on the physical address to classify it: qemu anonymous memory (guest
RAM), page cache, slab, page table or reserved. If a hit is in guest
RAM, rd -64 the page and count the repeats.
The hypothesis predicts pages in Windows guests' RAM where the word
repeats in long runs (guest page tables), plus the damaged
hash-table page or pages. If the word exists only in the dentry
table, the guest-data origin loses its support.
4. Who points at the bucket's page:
search -p -m 0xfff0000000000fff PAGE_PHYS
PAGE_PHYS here is PHYS_B & ~0xfff. This matches any 8-byte value
whose bits 12-51 equal the page's physical address, i.e. any
PTE-shaped entry pointing at it. The direct map normally covers the
hash table with 2M or 1G leaves, so no 4K entry naming that page
should exist.
A PTE-shaped hit inside guest RAM that is laid out like a page table
would directly support the "guest data walked as a host page table"
route. Expect noise. Only hits with P and RW set and sensible low
bits are interesting. The page that served as the stale table may
have been reused again since, so finding no hit does not refute the
hypothesis.
5. Whether INVLPGB was in use at crash time:
p/x boot_cpu_data.x86_capability[13] # bit 3 (0x8) set: INVLPGB used
p global_asid_available # 2039 at boot (PTI built in),
# 4087 without; lower means
# global ASIDs were handed out
p last_global_asid
ps | grep -E 'qemu|vhost'
task -R mm <qemu pid>
p ((struct mm_struct *)<mm>)->context.global_asid # non-zero: broadcast
p cpu_tlbstate:a # loaded_mm_asid >= 6 is global
6. KVM and platform state:
mod -s kvm_amd
p npt_enabled
p nested
p avic
p x2avic_enabled
sys config | grep -E 'BROADCAST_TLB_FLUSH|DEBUG_VM|INIT_ON_(ALLOC|FREE)|PAGE_POISON'
log | grep -iE 'AVIC|AMD-Vi|microcode'
From a live host on the same build, please also send:
cat /sys/module/kvm_amd/parameters/{npt,nested,avic}
grep -E 'model name|microcode' /proc/cpuinfo | sort -u
virsh dumpxml <windows-vm> | grep -E 'hyperv|feature|<cpu'
Please also tell us whether the Windows guests run with VBS/HVCI
(msinfo32, "Virtualization-based security"), and which guest OSes run
on the Milan host.
KVM SVM INVLPGA (26505e1b5b54)
==============================
Posting: https://lore.kernel.org/all/20260723094419.630204-1-pbonzini@redhat.com/
This is a guest-memory fix. It is not the host corruption fix.
In 6.18.53, svm_flush_tlb_gva() (INVLPGA) only drops guest-virtual
translations. It is never the thing that protects a host page about to
be freed. Every path that takes a page away from a guest flushes
through kvm_flush_remote_tlbs() -> svm_flush_tlb_all() ->
TLB_CONTROL_FLUSH_ASID before the host page can be reused. That covers
mmu_notifier zaps (munmap, MADV_DONTNEED from balloon or free page
reporting), KSM, compaction and migration, memslot delete, dirty
logging, TDP page-table free and VM destroy.
GPA invalidations don't reach INVLPGA at all:
arch/x86/kvm/mmu/mmu.c:kvm_mmu_invalidate_addr() {
...
/* It's actually a GPA for vcpu->arch.guest_mmu. */
if (mmu != &vcpu->arch.guest_mmu) {
...
kvm_x86_call(flush_tlb_gva)(vcpu, addr);
}
A missed INVLPGA therefore leaves at worst a stale GVA->HPA entry
whose HPA still backs one of the guest's own GPAs. The result is
guest-internal corruption, i.e. Windows BSODs, which is what Red Hat
saw. It can't touch host page cache or the dentry table.
On 6.18.53 with NPT, INVLPGA is reached from:
- the Hyper-V PV TLB flush (hv-tlbflush). You said this never ran on
the crashed nodes.
- L1 INVLPGA under nested SVM (VBS/HVCI guests). This is redundant,
because every nested VMRUN and #VMEXIT already does a full ASID
flush.
- the rare emulated #PF and emulated INVLPG paths.
Applying it to 6.18.y: you don't need the ~134 svm.c commits. The
minimal backport changes only svm_flush_tlb_gva(), and
svm_flush_tlb_asid() is already defined above it in 6.18.
"git apply --check" passes on 6.18.53 and on 6.18.44. I have not
checked 6.18.54.
The rest of the upstream patch is a signature change that lets
kvm_hv_vcpu_flush_tlb() stop iterating early, which is a performance
optimisation only, and an ERAPS register mark that has no counterpart
in 6.18. The patch is attached as
0001-KVM-SVM-flush_tlb_gva-full-asid-flush-6.18.y.patch. The
functional change is:
@@ -4071,7 +4071,18 @@ static void svm_flush_tlb_gva(struct kvm_vcpu *vcpu, gva_t gva)
{
struct vcpu_svm *svm = to_svm(vcpu);
- invlpga(gva, svm->vmcb->control.asid);
+ /*
+ * INVLPGA has had errata on Genoa and Turin, and even on older
+ * generations there were reports of Windows BSODs if INVLPGA
+ * was used for Hyper-V tlbflush. Use it only for shadow paging
+ * where it seems to be okay.
+ */
+ if (!npt_enabled) {
+ invlpga(gva, svm->vmcb->control.asid);
+ return;
+ }
+
+ svm_flush_tlb_asid(vcpu);
}
It's worth carrying for the Windows guests, especially any with
hv-tlbflush or VBS. It won't stop the dentry or libglib corruption. The
host mitigation is tlbi=ipi (or clearcpuid=419), and that needs no
patch.
AVIC, the GA log and guest-side reports
=======================================
AVIC: ca2967de5a5b (v6.18-rc1) makes avic=auto turn AVIC and x2AVIC
on by default for Zen4+ with X2AVIC. 5.15 had avic=0.
arch/x86/kvm/svm/avic.c (6.18.53):
if (avic == AVIC_AUTO_MODE)
avic = boot_cpu_has(X86_FEATURE_X2AVIC) &&
(boot_cpu_data.x86 > 0x19 || cpu_feature_enabled(X86_FEATURE_ZEN4));
On a Zen4 or Zen5 host with X2AVIC, AVIC and IPI virtualisation (and,
with the AMD IOMMU in GA mode, device-posted interrupts) are therefore
on in 6.18 where they were off in 5.15. On the Milan 7343 host AVIC is
off unless avic=1 is set explicitly. That host crashed identically, so
AVIC is not needed for the dentry crash unless it was forced on there.
The /sys/module/kvm_amd/parameters/avic values from both hosts settle
this.
GA log UAF: 78684b65fcc0 ("KVM: SVM: Remove VM from the GA Log notifier
list before VM destruction", v7.3-rc1, not in 6.18.y) fixes a
use-after-free in avic_ga_log_notifier() at VM teardown. Sean's own
assessment in the commit is that it is "all but impossible to trigger".
I'd treat it as low priority and mention it only for completeness.
Guest crashes on the Milan host: Joris de Vries reported on kvm@
(2026-09-03, in the 26505e1b5b54 thread, Message-ID
<E6439A7D-77EA-475D-BEBC-6BB6EDA10A0C@gmail.com>) that on Zen3 (Ryzen
5900X, also on v7.3-rc1) guests take reserved-bit page faults from a
stale guest page-walk-cache entry. It happens after a 2 MiB-aligned
guest mapping is torn down and its table reused. No KVM-side
invalidation fixed it, but npt=0 did. That is guest-internal and
independent of host INVLPGB, so it can't explain host corruption. It
could be relevant to the guest crashes during migration on Milan. I
have not looked into it beyond reading the report.
Verified versus inferred
========================
Verified, from the code in the stable trees, the two oopses and the
cited threads:
- Both oopses read the same word as a dentry_hashtable bucket head on
the first loop iteration.
- dentry_hashtable is boot-time memblock memory that is never freed.
- 6.18.53 contains the complete CPA fix series, and it crashed.
- The INVLPGB code is identical in 6.18.15, 6.18.44 and 6.18.53. I
found no kernel-side ordering hole. TLBSYNC precedes every free.
- tlbi= exists from 6.18.45 and only "ipi" does anything.
clearcpuid=419 works on 6.18.15, 6.18.44 and 6.18.53, and
clearcpuid=invlpgb works on none of them. INVLPGB never shows in
/proc/cpuinfo on 6.18.
- 44126343d58c is missing from 6.18.15, so nopcid is unsafe there.
- KVM never issues INVLPGB. The INVLPGA path can only reach guest
memory.
- AVIC defaults to on for Zen4+ since v6.18-rc1.
- The INVLPGB stale-TLB issue is publicly reproduced and acknowledged
on Zen5, with no kernel fix yet. This is as reported in the lore
thread, not something I reproduced.
Inferred, not demonstrated:
- That the Zen5 issue exists on Zen3 and Zen4.
- That it is the cause here.
- The paging-structure route to the dentry table.
- That the word comes from Windows guest page tables.
- That the first host is Genoa. The board is SP5, so it could be Turin.
- The crash 2 bucket address. It assumes RDX is the table base, as in
crash 1.
Ruled out
=========
- CPA collapse race (41d88484c71c and its fixes): crash 2 is on
6.18.53, which has every fix.
- KVM SVM INVLPGA (26505e1b5b54): it only reaches guest memory. It is
still worth backporting for Windows guests (above).
- Host mm writing a swap or migration PTE through a stale pointer: the
first host has no swap, and an audit of every non-present PTE store
site found nothing.
- KVM NPT level or pfn errors, and TDP MMU huge page recovery: KVM can
only map PFNs that the host page tables hold, and the memblock table
is not one of them.
- THP and NUMA balancing: both were off on the Milan host (THP
"never", numa_balancing 0), and it still crashed.
- 55ddbc2ca6d5 (upstream f7491d7c81db, pmd_modify() dropping
_PAGE_DIRTY): this causes data loss, not a stray write. It is fixed
in 6.18.53.
- d053eb7e09e1 (upstream 1e75a8255f11, AMD IOMMU completion wait):
this has been in the tree since 6.18.42, and 6.18.15 crashed too.
- EFER.TCE (enabled since 6.15) versus pud_free_pmd_page() flushing
only one page: this is a structural hole, independent of INVLPGB.
However, it is only reachable through a >= 1 GiB ioremap, which is
unlikely on these hosts. clearcpuid=tce would rule it out if needed.
What would help most
====================
1. Boot production with tlbi=ipi (6.18.45+) or clearcpuid=419 (6.18.15
and 6.18.44). Check that it took with x86_capability[13], not with
/proc/cpuinfo.
2. Run Matt's reproducer on a drained Milan host and a drained Genoa
host, with and without tlbi=ipi, and post the result to the INVLPGB
thread.
3. Send the CPU model and microcode for every affected host, including
the first crash host and nrbphav4a.
4. From the 6.18.53 vmcore: the whole-page dump around the bucket, the
two searches for the word, the masked referrer search and the
global ASID state (above).
5. Send the kvm_amd parameters for both crash hosts, tell us which
guest OSes run on the Milan host, and whether the Windows guests use
VBS.
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: hunting memory corruption bug in 6.18.x
2026-10-02 9:50 ` Lorenzo Stoakes (ARM)
@ 2026-10-02 14:55 ` Nikola Ciprich
2026-10-02 19:27 ` Lorenzo Stoakes (ARM)
0 siblings, 1 reply; 16+ messages in thread
From: Nikola Ciprich @ 2026-10-02 14:55 UTC (permalink / raw)
To: Lorenzo Stoakes (ARM)
Cc: linux-mm, linux-kernel, akpm, david, Mike Rapoport, Dave Hansen,
Pedro Falcato, Kiryl Shutsemau, luizcap, pbonzini,
Borislav Petkov, Tal Zussman, Rik van Riel, Matt Fleming,
Nikola Ciprich
> Thanks that's useful.
>
> We can rule out CPA at this point.
>
> >
>
> Very useful.
>
> It looks like AMD TLB invalidation issues are the leading likely cause here
> then - one on the host side with INVLPGB, and a separate one in KVM with
> INVLPGA.
>
> And it looks like a hardware bug, unfortunately.
>
> Support for this was merged in 6.15 which matches your kernel versions too.
>
> It's actually not solved yet, but there are two workarounds that can be
> applied here.
>
> ## Issue 1: INVLPGB
>
> See [0] for a report of the same kind of thing (segfaults like yours),
> and [1] for the proposed temporary workaround.
>
> TL;DR: The mitigation for this is to update your kernel command line parameters
> thusly:
>
> 6.18.45 and later (6.18.53, 6.18.54): tlbi=ipi
> <6.18.45: clearcpuid=419
>
> If this resolves it then it confirms that this is the issue.
>
> There is also, usefully, a reproducer which should show corruption
> in minutes rather than weeks, see:
>
> https://gist.github.com/mfleming/ca26ad3f8d65a12d23d62fb176480fc1
> https://lore.kernel.org/all/akzuETUv5XEFSRL6@matt-Precision-5490/
>
> The fact you've seen this on a Milan machine (EPYC 7343) is new information
> so that could be useful for the report, so if you can reproduce _without
> the fix_ first that'd be very useful to know!
well, not so good news here.. I tried the reproducer on multiple
lab machines and also on one drained production box on which we've
experienced one crash and wasn't able to get a single hit so far.
Here's the list of machines:
hostname CPU RAM kernel microcode
pocstdv1b EPYC 9124 384G 6.18.31lb9.01 0x0a101158
labtest EPYC 9274F 64G 6.18.20lb9.01 0x0a101158
lbxovav6d EPYC 9124 1.5TB 6.18.31lb9.01 0x0a101158
prfjazv1g EPYC 7343 1TB 6.18.53lb9.02 0x0a0011de
nrbphav4a EPYC 7343 1TB 6.18.15lb9.01 0x0a0011de
I'll leave it running for few hours, and report back. If I'm able to reproduce,
I'll try recommended workarounds and report as well.
> ## Crash + vmcore
>
> Obviously do the above and especially try the reproducer in the lab! :)
>
> But also with the vmcore, can you run these 4 commands and reply with the
> output please?
>
> crash> p dentry_hashtable
> crash> rd -64 0xffff986002783000 512
> crash> search -p 0x0fffffff0c930020
> crash> vtop 0xffff986002783a50
> crash> search -p -m 0xfff0000000000fff <the PHYSICAL value vtop prints>
>
crash> p dentry_hashtable
dentry_hashtable = $1 = (struct hlist_bl_head *) 0xffffb18900a00000
crash> rd -64 0xffff986002783000 512
crash>
(empty output)
crash> search -p 0x0fffffff0c930020
103360440: fffffff0c930020
103360450: fffffff0c930020
114c32440: fffffff0c930020
114c32450: fffffff0c930020
1ae234400: fffffff0c930020
1ae234420: fffffff0c930020
1ae234430: fffffff0c930020
1ae234450: fffffff0c930020
24f0e2400: fffffff0c930020
24f0e2420: fffffff0c930020
24f0e2430: fffffff0c930020
24f0e2450: fffffff0c930020
342c40400: fffffff0c930020
342c40420: fffffff0c930020
342c40430: fffffff0c930020
342c40450: fffffff0c930020
77475b400: fffffff0c930020
77475b420: fffffff0c930020
77475b430: fffffff0c930020
77475b450: fffffff0c930020
b2fabe400: fffffff0c930020
b2fabe420: fffffff0c930020
b2fabe430: fffffff0c930020
b2fabe450: fffffff0c930020
ea2d7c400: fffffff0c930020
ea2d7c420: fffffff0c930020
ea2d7c430: fffffff0c930020
ea2d7c450: fffffff0c930020
ef7a2d400: fffffff0c930020
ef7a2d420: fffffff0c930020
ef7a2d430: fffffff0c930020
ef7a2d450: fffffff0c930020
ef7a90400: fffffff0c930020
ef7a90420: fffffff0c930020
ef7a90430: fffffff0c930020
ef7a90450: fffffff0c930020
10371fa400: fffffff0c930020
10371fa420: fffffff0c930020
10371fa430: fffffff0c930020
10371fa450: fffffff0c930020
...
(and lots of more addresses, truncated)
crash> vtop 0xffff986002783a50
VIRTUAL PHYSICAL
ffff986002783a50 582783a50
PGD DIRECTORY: ffffffffa0836000
PAGE DIRECTORY: 143a7801067
PUD: 143a7801c00 => 80000005800001e3
PAGE: 580000000 (1GB)
PTE PHYSICAL FLAGS
80000005800001e3 580000000 (PRESENT|RW|ACCESSED|DIRTY|PSE|GLOBAL|NX)
PAGE PHYSICAL MAPPING INDEX CNT FLAGS
fffff2715609e0c0 582783000 ffff996e5a8932d9 7f6a1589e 1 2ffff800020938 uptodate,dirty,lru,active,owner_2,swapbacked
crash> search -p -m 0xfff0000000000fff 582783a50
3c600964f0: 8000000582783067
11e238a54f0: 582783e67
should you need anything else, please let me know
> ## Mitigations for production
>
> I suggest you apply both of the above for your production as it's _likely_
> it will resolve the issue there.
>
> The cmdline changes are perfectly safe but you should check the generated
> patch, however!
>
> I suspect it'll be fine but I can't guarantee it obviously LLMs
> etc. (albeit it is a frontier model so at least as good as it can be ;)
ok, I'll try to reproduce first and then try those. thanks!
BR
nik
>
> ## Attachments
>
> I attach the backported fix as discussed above and also the AI's full debug
> report FYI.
>
> --
> Cheers, Lorenzo
>
> [0]:https://lore.kernel.org/all/CAAuFnRTyva1_3tsF3vrMBL+TLS1YL4EgUPT2c3O9k7A9hWUMnA@mail.gmail.com/
> [1]:https://lore.kernel.org/all/20260914045823.GFaqd-7yDf6aK4JUQr@fat_crate.local/
> [2]:https://lore.kernel.org/all/20260723094419.630204-1-pbonzini@redhat.com/
> From: Paolo Bonzini <pbonzini@redhat.com>
> Subject: [PATCH 6.18.y] KVM: SVM: make svm_flush_tlb_gva do a full asid flush if NPT enabled
>
> commit 26505e1b5b546e2fa9a0296b951ca158460c72d8 upstream.
>
> Red Hat is seeing multiple reports of Windows memory corruptions
> (and consequent BSODs) with hv-tlbflush=on, on AMD processors only
> (Turin and Milan; none on Intel; none on Turin with a full ASID flush).
>
> When NPT is enabled, replace the per-address INVLPGA in
> svm_flush_tlb_gva() with a full flush of the current ASID. This covers
> both callers of the flush_tlb_gva op: kvm_hv_vcpu_flush_tlb() (Hyper-V
> PV TLB flush, i.e. hv-tlbflush) and kvm_mmu_invalidate_addr() (L1
> INVLPGA for nested SVM, emulated #PF, emulated INVLPG). Shadow paging
> keeps using INVLPGA.
>
> [ Backport note for 6.18.y: reduced to the svm.c functional change.
> Upstream also changes the kvm_x86_ops.flush_tlb_gva signature to
> return a "full" flag so that kvm_hv_vcpu_flush_tlb() stops iterating
> once a full flush has been requested, and touches vmx/main.c,
> vmx/vmx.c, vmx/x86_ops.h, hyperv.c and mmu.c for that. That part is
> a performance optimisation only and is omitted here: without it the
> Hyper-V flush loop keeps calling svm_flush_tlb_asid() for each
> remaining page, which only re-sets TLB_CONTROL_FLUSH_ASID and is
> cheaper than the INVLPGA it replaces. Upstream's svm_flush_tlb_guest()
> additionally marks VCPU_REG_ERAPS dirty; ERAPS virtualisation does not
> exist in 6.18.y, whose .flush_tlb_guest is svm_flush_tlb_asid(), so
> calling svm_flush_tlb_asid() directly is the exact equivalent. ]
>
> Analyzed-by: Vitaly Kuznetsov <vkuznets@redhat.com>
> Analyzed-by: Alexander Lougovski <alougovs@redhat.com>
> Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
> ---
> arch/x86/kvm/svm/svm.c | 13 ++++++++++++-
> 1 file changed, 12 insertions(+), 1 deletion(-)
>
> diff --git a/arch/x86/kvm/svm/svm.c b/arch/x86/kvm/svm/svm.c
> index a24a6871b693..5bfb72e550e2 100644
> --- a/arch/x86/kvm/svm/svm.c
> +++ b/arch/x86/kvm/svm/svm.c
> @@ -4071,7 +4071,18 @@ static void svm_flush_tlb_gva(struct kvm_vcpu *vcpu, gva_t gva)
> {
> struct vcpu_svm *svm = to_svm(vcpu);
>
> - invlpga(gva, svm->vmcb->control.asid);
> + /*
> + * INVLPGA has had errata on Genoa and Turin, and even on older
> + * generations there were reports of Windows BSODs if INVLPGA
> + * was used for Hyper-V tlbflush. Use it only for shadow paging
> + * where it seems to be okay.
> + */
> + if (!npt_enabled) {
> + invlpga(gva, svm->vmcb->control.asid);
> + return;
> + }
> +
> + svm_flush_tlb_asid(vcpu);
> }
>
> static inline void sync_cr8_to_lapic(struct kvm_vcpu *vcpu)
> Summary
> =======
>
> The 6.18.53 crash rules out the CPA collapse race I pointed at earlier.
>
> 6.18.53 carries the complete CPA series (a1c7570cedd0, d5d8b8662e6e,
> 1587d3394e25, 9e4a3ec3411b and the rest). The new oops is the same as
> the 6.18.44 one: same RIP, __d_lookup+0x4a, and the same corrupt bucket
> word, 0x0fffffff0c930020. It happened on a different host (Milan
> instead of the SP5 box), with a different build and a different KASLR
> layout. On top of that, 6.18.54 corrupted libglib in page cache and a
> 6.18.15 node crashed. The CPA bugs are real, but they are not this bug,
> and my earlier diagnosis was wrong.
>
> The leading candidate is now the AMD INVLPGB/TLBSYNC stale-TLB issue.
> It was reported publicly in June and no kernel has a fix for it yet:
>
> "PROBLEM: Probabilistic segfault on AMD hardware with INVLPGB"
> Henrik Boving, 2026-06-19
> Message-ID: <CAAuFnRTyva1_3tsF3vrMBL+TLS1YL4EgUPT2c3O9k7A9hWUMnA@mail.gmail.com>
> https://lore.kernel.org/all/CAAuFnRTyva1_3tsF3vrMBL+TLS1YL4EgUPT2c3O9k7A9hWUMnA@mail.gmail.com/
>
> - Henrik saw heap corruption on an EPYC 9455 (Turin) from 6.15 onwards.
> He bisected it to CONFIG_BROADCAST_TLB_FLUSH.
>
> - Matt Fleming posted a userspace reproducer on 2026-07-07. It reuses
> VAs with munmap() + mmap(MAP_FIXED) under a rwlock. Reader threads
> still see data from the previous mapping after the INVLPGB + TLBSYNC
> flush has returned:
> https://lore.kernel.org/all/akzuETUv5XEFSRL6@matt-Precision-5490/
>
> - Rik reproduced it and posted a double-TLBSYNC diff, which is not
> merged. He also noted that Meta saw elevated segfault rates on Turin
> that the AMD-SB-3029 firmware fixed. Henrik's microcode (0x0B002162)
> is newer than that fix, so firmware does not explain his case.
>
> - Tal Zussman confirmed it on an EPYC 9965 (Turin Dense) on 2026-09-13.
> It corrupts in the first round, and with clearcpuid=419 it passes 20
> rounds:
> https://lore.kernel.org/all/20260914022055.1639690-1-tz2294@columbia.edu/
>
> - Borislav answered on the same day: "There will be an official thing
> Soon(tm). In the meantime, tlbi=ipi":
> https://lore.kernel.org/all/20260914045823.GFaqd-7yDf6aK4JUQr@fat_crate.local/
>
> Every public confirmation so far is Zen5 (Turin and Turin Dense). Rik
> asked whether Milan or Bergamo reproduce it, and nobody has answered.
> Your second crash host is Milan (EPYC 7343, Zen3). A result from your
> hosts, positive or negative, is therefore new information for the x86
> maintainers.
>
> It fits what you have seen:
>
> - AMD only.
> - 5.15 is fine and 6.18 is not. The broadcast flush code went in in
> v6.15.
> - QEMU does heavy mmap/munmap churn around migration.
> - It is timing sensitive: you could not reproduce it with KASAN or
> SLUB debugging enabled.
> - Host .so files are corrupted in page cache while the files on disk
> are intact.
>
> I can't tie the dentry hash table damage to it directly. Further down
> there is a hypothesis for that, clearly labelled, together with the
> vmcore checks that would confirm or refute it.
>
>
> What to do now
> ==============
>
> Disable broadcast TLB flushing on every AMD host. No patch is needed.
>
> 6.18.45 and later (6.18.53, 6.18.54): tlbi=ipi
> 6.18.15 and 6.18.44: clearcpuid=419
>
> tlbi= only exists from 6.18.45 onwards (upstream abe7c8b09bd7,
> backported as 846b92e26c8a), so 6.18.15 and 6.18.44 have to use
> clearcpuid=419. 419 is 13*32+3, i.e. X86_FEATURE_INVLPGB.
> clearcpuid=419 works on 6.18.45+ as well, but tlbi=ipi does not taint
> the kernel. These are the parameters Borislav and Tal used in the
> thread above.
>
> Details that matter:
>
> - The parameter has to be exactly "tlbi=ipi". The handler only matches
> "ipi" and returns 1 for anything else. "tlbi=off" is therefore
> accepted silently and does nothing:
>
> arch/x86/kernel/cpu/common.c (6.18.53):
>
> static int __init tlbi_setup(char *str)
> {
> if (!strcmp(str, "ipi"))
> setup_clear_cpu_cap(X86_FEATURE_INVLPGB);
>
> return 1;
> }
> __setup("tlbi=", tlbi_setup);
>
> - "clearcpuid=invlpgb" does not work on 6.18. X86_FEATURE_INVLPGB has no
> name string in cpufeatures.h, so the boot log says "clearcpuid:
> unknown CPU flag: invlpgb" and INVLPGB stays enabled. Use the number.
>
> - With clearcpuid=419 the kernel prints "clearcpuid: force-disabling
> CPU feature flag: 13:3", then the "setcpuid=/clearcpuid= in use ...
> Tainting kernel" warning, and sets taint S. That is expected.
>
> - Do not use "nopcid" as a workaround on 6.18.15. 44126343d58c
> ("x86/mm: Disable broadcast TLB flush when PCID is disabled") only
> arrived in 6.18.35. Without it, nopcid leaves INVLPGB enabled, and
> the first broadcast flush with a non-zero PCID takes a #GP in
> broadcast_tlb_flush().
>
> How to check that it is in effect:
>
> - /proc/cpuinfo can't tell you. The flag has no name in 6.18, so
> "invlpgb" never appears there, whether it is enabled or not.
>
> - CONFIG_BROADCAST_TLB_FLUSH is "def_bool y" with "depends on
> CPU_SUP_AMD && 64BIT" and no prompt. Every 6.18 x86-64 build with AMD
> support has it:
>
> grep BROADCAST_TLB_FLUSH /boot/config-$(uname -r)
>
> - The definitive check is the capability word. You can read it with
> crash or drgn and debuginfo, either on a live system or in a vmcore:
>
> crash> p/x boot_cpu_data.x86_capability[13]
>
> If bit 3 (0x8) is set, the kernel is using INVLPGB. If it is clear,
> it is not.
>
> - For tlbi=ipi, check /proc/cmdline. For clearcpuid=419, look for the
> dmesg line above and for taint bit 2 (value & 4 in
> /proc/sys/kernel/tainted).
>
> - As a sanity check, the "TLB shootdowns" row in /proc/interrupts
> should rise noticeably faster under the same load, because QEMU's
> flushes go back to IPIs.
>
> Testing whether your CPUs are affected, with Matt's reproducer:
>
> https://gist.github.com/mfleming/ca26ad3f8d65a12d23d62fb176480fc1
> https://lore.kernel.org/all/akzuETUv5XEFSRL6@matt-Precision-5490/
>
> build: gcc -O2 -std=c11 -pthread -static -o repro-invlpgb <source>.c
> run: ./repro-invlpgb --batch --rounds 20 --jobs 32 -d 5 -w 8 -m 2 -s 512 -q
>
> Run it on a drained Milan host and a drained Genoa host, once with the
> default boot and once with tlbi=ipi. If it fails on Zen3 or Zen4 with
> the default boot and passes with tlbi=ipi, that is new information and
> should go to the INVLPGB thread above. A pass with the default boot is
> weaker evidence, because the reproducer was tuned on Zen5.
>
> Even if the reproducer passes, running production with tlbi=ipi is
> the real test. Given how intermittent this is, a clean run has to last
> several times longer than your previous time to failure before it
> means much.
>
>
> Kernel versions and machines
> ============================
>
> Crash 1: 6.18.44 (6.18.44lb9.01). ASUSTeK RS720A-E12-RS12 /
> K14PP-D24, BIOS 2305 11/21/2025, AMD EPYC on SP5. Windows
> guests only, no host swap, about 22 days of uptime.
> pacemaker-controld in mkdir().
>
> Crash 2: 6.18.53. Supermicro AS-2024US-TRT / H12DSU-iN, BIOS 3.5
> (2025-09-22), 2x EPYC 7343 (Milan, Zen3), 1 TB RAM.
> THP "never", NUMA balancing 0, no MCE/EDAC records. systemd
> in openat(). The full vmcore is available.
>
> 6.18.54 production node (nrbphav4a): a libglib segfault storm within
> minutes of receiving migrated VMs (see below).
>
> 6.18.15 node: crashed while VMs were being migrated to it. No
> details were posted.
>
> 5.15.x: clean.
>
> The CPU model of the first host is not in the report. I have been
> assuming Genoa, but the board is SP5, which takes both Genoa (9004,
> Zen4) and Turin (9005, Zen5). If that box is actually Turin, it falls
> straight into the publicly confirmed set. Please send the "model name"
> and "microcode" lines from /proc/cpuinfo for every affected host,
> including nrbphav4a.
>
>
> Stack traces
> ============
>
> Crash 1 (6.18.44):
>
> Oops: general protection fault, probably for non-canonical address
> 0xfffffff0c930038: 0000 [#1] SMP NOPTI
> CPU: 23 UID: 189 PID: 7538 Comm: pacemaker-contr 6.18.44lb9.01 #1
> RIP: 0010:__d_lookup+0x4a/0xc0
> RAX: 0fffffff0c930020 RBX: 0fffffff0c930020 RCX: 0000000000000000
> RDX: ff2e6dbe0d9b6000 RSI: ff7532a13699fe60 RDI: ff2e6d1e4e630d80
> RBP: 000000000b654440
> Call Trace:
> <TASK>
> d_lookup+0x27/0x50
> lookup_dcache+0x1f/0x80
> lookup_one_qstr_excl+0x1e/0xe0
> filename_create+0xc4/0x160
> do_mkdirat+0x5a/0x190
> __x64_sys_mkdir+0x42/0x60
> do_syscall_64+0x64/0xbf0
> entry_SYSCALL_64_after_hwframe+0x76/0x7e
> </TASK>
>
> Crash 2 (6.18.53, from crash(8) on the vmcore):
>
> general protection fault, probably for non-canonical address
> 0xfffffff0c930038
> RIP: __d_lookup+0x4a/0xc0
> Comm: systemd PID: 1274600 CPU: 23
> RAX: 0fffffff0c930020 RBX: 0fffffff0c930020 RCX: 000000000000000b
> RDX: ffff985ffd223600 RSI: ffffb18962823d70 RDI: ffff989ddf09b5c0
> RBP: 000000005560450b
> Call trace:
> __d_lookup
> lookup_fast
> walk_component
> link_path_walk
> path_openat
> do_filp_open
> do_sys_openat2
> __x64_sys_openat
> do_syscall_64
>
> In crash 2 the parent dentry (RDI, "app.slice", kernfs) is intact and
> its own hash chain terminates cleanly.
>
>
> Other relevant messages
> =======================
>
> On 6.18.54, nrbphav4a, right after receiving migrated VMs, every glib
> user faulted at the same file offset:
>
> qemu-system-x86[19014]: segfault at 0 ip 00007f9a3ee7e61e sp 00007ffd96294380 error 6 in libglib-2.0.so.0.6800.4[a461e,7f9a3edf7000+91000]
> pacemaker-execd[17696]: segfault at 0 ip 00007f3c6ed6a61e sp 00007fffa95873d0 error 6 in libglib-2.0.so.0.6800.4[a461e,7f3c6ece3000+91000]
> Code: 29 9c 24 80 00 00 00 0f 29 a4 24 90 00 00 00 0f 29 ac 24 a0 00 00 00 0f 29 b4 ff ff 00 00 00 0f 29 bc ff ff 00 00 00 64 00 00 <00> 00 28 00 00 00 48
>
> The bytes around the fault are a run of movaps stores to the stack
> followed by the stack-protector load. That pattern tells us what the
> original bytes were:
>
> file offset expected found
> 0xa460e 24 b0 ff ff movaps %xmm6,0xb0(%rsp)
> 0xa4616 24 c0 ff ff movaps %xmm7,0xc0(%rsp)
> 0xa461c 48 8b 04 25 00 00 00 00 mov %fs:0x28,%rax
>
> The file on disk will confirm the expected column.
>
> Every process maps the same page-cache page, so they all fault at the
> same file offset. The damage consists of two 16-bit stores of 0xffff and one
> 32-bit store of 0, at +6, +6 and +4 of three consecutive 8-byte slots
> (page offsets 0x608, 0x610 and 0x618).
>
> These are ordinary CPU stores to small struct fields. They are not DMA
> (NIC descriptor write-backs are 16 or 32 bytes) and they are not
> page-walker A/D bit updates. This is what a CPU writing through a
> translation that no longer belongs to the writer looks like. The
> earlier ld.so relocation assertion and the libcrypto.so.3 corruption
> are the same class of damage.
>
>
> Register decode (short)
> =======================
>
> This is already on the list, so I'll keep it brief.
>
> In both oopses RAX == RBX == 0x0fffffff0c930020. RAX is only written by
> the initial load of the bucket head, so the fault is on the first
> iteration of the loop:
>
> fs/dcache.c:__d_lookup() {
> struct hlist_bl_head *b = d_hash(hash);
> ...
> hlist_bl_for_each_entry_rcu(dentry, node, b, d_hash) {
>
> if (dentry->d_name.hash != hash) <-- faults on 0x18(%rbx)
> continue;
>
> The corrupt word is therefore the dentry_hashtable bucket head itself,
> not a dentry. d_hash_shift is patched to 7, i.e. 2^25 buckets and a
> 256 MiB table:
>
> crash 1: 0xff2e6dbe0d9b6000 + (0x0b654440 >> 7) * 8 = 0xff2e6dbe0e51b440
> (page offset 0x440)
> crash 2: 0xffff985ffd223600 + (0x5560450b >> 7) * 8 = 0xffff986002783a50
> (page offset 0xa50)
>
> The crash 2 line assumes that RDX holds the table base, as it does in
> crash 1 (same code, same RIP). "p dentry_hashtable" in the vmcore will
> confirm it.
>
> The table comes from alloc_large_system_hash() with HASH_EARLY. That is
> memblock memory, allocated at boot and never freed, and nothing
> legitimately writes a non-pointer into it. So whatever wrote the word
> used a wrong translation or a wrong physical address.
>
> The same 64-bit value on two hosts, with different CPUs, builds and
> KASLR layouts, means the data written is deterministic. As David said,
> that can't be a coincidence.
>
> What the word decodes to:
>
> - As an x86-64 Linux swap PTE it fits exactly (type 1, offset
> 0x79b67f), with bit 5 set. But the first host has no swap and no
> Linux guests, so a Linux swap PTE is an unlikely origin.
>
> - It is not a KVM SPTE. SVM MMIO SPTEs have bit 0 set, and non-present
> and frozen SPTEs have bit 63 set. It is also not an AVIC physical or
> logical ID table entry, and not any VMCB field.
>
> - As a Windows x64 software PTE (MMPTE_SOFTWARE) it reads Valid=0,
> Protection=1 (MM_READONLY), Prototype=0, Transition=0,
> PageFileLow=0 and PageFileHigh=0x0fffffff. That PageFileHigh is the
> "lookup needed" marker 0xffffffff with bits 60-63 clear. However, the
> UsedPageTableEntries (0x93), ShadowStack and Unused fields are also
> non-zero. This fit is weak. It rests on public Windows internals
> documentation, not on anything I can check in kernel sources.
>
>
> The suspect code
> ================
>
> commit 767ae437a32d644786c0779d0d54492ff9cbe574
> Author: Rik van Riel <riel@surriel.com>
>
> x86/mm: Add INVLPGB feature and Kconfig entry
>
> Link: https://lore.kernel.org/r/20250226030129.530345-3-riel@surriel.com
>
> This commit starts the v6.15 broadcast TLB flush series, and the tlbi=
> switch (abe7c8b09bd7) names it as its Fixes: target. I am not claiming
> a bug in this code. As far as I can tell, the hardware does not
> provide the completion guarantee that the code relies on. The series
> explains why 5.15 is good, 6.18 is bad and only AMD is affected.
>
> > diff --git a/arch/x86/Kconfig.cpu b/arch/x86/Kconfig.cpu
> > --- a/arch/x86/Kconfig.cpu
> > +++ b/arch/x86/Kconfig.cpu
> > @@ -334,6 +334,10 @@ menuconfig PROCESSOR_SELECT
>
> [ ... ]
>
> > +config BROADCAST_TLB_FLUSH
> > + def_bool y
> > + depends on CPU_SUP_AMD && 64BIT
>
> There is no prompt, so a build can't opt out. The only switches are at
> boot time.
>
> > diff --git a/arch/x86/include/asm/cpufeatures.h b/arch/x86/include/asm/cpufeatures.h
> > --- a/arch/x86/include/asm/cpufeatures.h
> > +++ b/arch/x86/include/asm/cpufeatures.h
> > @@ -338,6 +338,7 @@
>
> [ ... ]
>
> > #define X86_FEATURE_XSAVEERPTR (13*32+ 2) /* "xsaveerptr" Always save/restore FP error pointers */
> > +#define X86_FEATURE_INVLPGB (13*32+ 3) /* INVLPGB and TLBSYNC instructions supported */
>
> The comment has no quoted name. That is why the flag is invisible in
> /proc/cpuinfo and why clearcpuid= only accepts it by number.
>
> [ ... ]
>
> commit 4afeb0ed1753ebcad93ee3b45427ce85e9c8ec40
> Author: Rik van Riel <riel@surriel.com>
>
> x86/mm: Enable broadcast TLB invalidation for multi-threaded processes
>
> Link: https://lore.kernel.org/r/20250226030129.530345-11-riel@surriel.com
>
> This commit moves ordinary user TLB flushes for large multi-threaded
> processes onto INVLPGB. On a KVM host, that means QEMU.
>
> > diff --git a/arch/x86/mm/tlb.c b/arch/x86/mm/tlb.c
> > --- a/arch/x86/mm/tlb.c
> > +++ b/arch/x86/mm/tlb.c
> > @@ -430,6 +430,105 @@ static bool mm_needs_global_asid(struct mm_struct *mm, u16 asid)
>
> [ ... ]
>
> > +static void consider_global_asid(struct mm_struct *mm)
> > +{
> > + if (!cpu_feature_enabled(X86_FEATURE_INVLPGB))
> > + return;
> > +
> > + /* Check every once in a while. */
> > + if ((current->pid & 0x1f) != (jiffies & 0x1f))
> > + return;
> > +
> > + /*
> > + * Assign a global ASID if the process is active on
> > + * 4 or more CPUs simultaneously.
> > + */
> > + if (mm_active_cpus_exceeds(mm, 3))
> > + use_global_asid(mm);
> > +}
>
> A CPU running a vCPU thread in guest mode still has QEMU's mm loaded.
> Any VM with four or more busy vCPUs therefore gets a global ASID
> quickly.
>
> [ ... ]
>
> > +static void broadcast_tlb_flush(struct flush_tlb_info *info)
> > +{
> [ ... ]
> > + } else do {
> [ ... ]
> > + invlpgb_flush_user_nr_nosync(kern_pcid(asid), addr, nr, pmd);
> [ ... ]
> > + } while (addr < info->end);
> > +
> > + finish_asid_transition(info);
> > +
> > + /* Wait for the INVLPGBs kicked off above to finish. */
> > + __tlbsync();
> > +}
>
> The kernel's correctness argument rests on this TLBSYNC. Once it
> returns, no CPU may still hold the old translation, and only after that
> are the pages and page tables freed:
>
> mm/mmu_gather.c:tlb_flush_mmu() {
> tlb_flush_mmu_tlbonly(tlb); <-- flush_tlb_mm_range(): INVLPGB + TLBSYNC
> tlb_flush_mmu_free(tlb); <-- pages and page tables freed
> }
>
> Matt's reproducer shows that on Zen5 another CPU can keep using the old
> translation after TLBSYNC has returned.
>
> > @@ -1260,9 +1359,12 @@ void flush_tlb_mm_range(struct mm_struct *mm, unsigned long start,
>
> [ ... ]
>
> > - if (cpumask_any_but(mm_cpumask(mm), cpu) < nr_cpu_ids) {
> > + if (mm_global_asid(mm)) {
> > + broadcast_tlb_flush(info);
> > + } else if (cpumask_any_but(mm_cpumask(mm), cpu) < nr_cpu_ids) {
> > info->trim_cpumask = should_trim_cpumask(mm);
> > flush_tlb_multi(mm_cpumask(mm), info);
> > + consider_global_asid(mm);
>
> Once QEMU has a global ASID, every flush of its address space uses
> INVLPGB + TLBSYNC and no IPI. That covers munmap(), MADV_DONTNEED
> (balloon, free page reporting), mprotect() and page migration.
>
> Global ASIDs aren't the only route. In 6.18 the batched unmap flush
> used by reclaim and migration (arch_tlbbatch_flush()) issues an INVLPGB
> for all non-global entries, for every process. Kernel range flushes use
> INVLPGB as well.
>
> Software audit. I went through the INVLPGB paths in 6.18.53 looking
> for a kernel-side hole:
>
> - munmap/zap with freed page tables
> - the dynamic-to-global ASID transition
> - global ASID reuse
> - lazy CPUs and CPUs in the middle of a switch
> - reclaim batching
> - kernel range flushes
>
> I did not find one:
>
> - TLBSYNC runs on the issuing CPU, with preemption disabled, before
> any page or page table is freed.
> - INVLPGB reaches lazy CPUs.
> - finish_asid_transition() sends IPIs to stragglers still on a dynamic
> ASID.
> - A global ASID is only reused after an INVLPGB of all non-global
> entries plus a TLBSYNC.
>
> The code is identical in 6.18.15, 6.18.44 and 6.18.53, and every
> post-6.15 fix to it is in 6.18.53. KVM never issues INVLPGB and never
> exposes it to guests.
>
>
> How this could reach the dentry hash table (hypothesis)
> =======================================================
>
> Everything in this section is inference. None of it has been
> demonstrated.
>
> A stale leaf translation can only reach the page that was freed. That
> covers the libglib and ld.so damage, where the freed page came back as
> page cache, and guest RAM corruption. It cannot reach dentry_hashtable,
> which the page allocator never owned.
>
> A stale paging-structure entry can. Suppose the entry that survives on
> another CPU is the cached PDE pointing at a PTE table that
> free_pgtables() has just released. That CPU then walks whatever the
> freed page now holds as if it were a PTE table. Its stores to any VA in
> that 2 MiB region go to whatever PFNs the entries in the reused page
> happen to name, and that includes boot-time memory.
>
> The identical word follows if the data being stored is guest RAM. A
> Windows page-table page commonly holds long runs of one identical
> software PTE. Suppose QEMU copies such a guest page through the broken
> walk, for example on migration receive or in a virtio/vhost copy into
> guest memory. It then writes 512 copies of the same 8-byte value over
> one host physical page.
>
> If that page belongs to dentry_hashtable, every bucket in it reads
> 0x0fffffff0c930020. Any lookup that hashes into one of those 512
> buckets then faults the way both oopses did, at whatever page offset it
> lands on (0x440 in one crash, 0xa50 in the other). Two hosts running
> Windows guests would end up with the same word.
>
> CPU A (QEMU thread) CPU B (QEMU vCPU/IO/vhost
> thread, same mm, global ASID)
> ----- -----
> munmap() of a region
> zap, free_pgtables()
> tlb_finish_mmu()
> flush_tlb_mm_range(freed_tables)
> INVLPGB (PCID, VA range)
> TLBSYNC returns
> PTE table page P freed
> still caches the PDE -> P
> (the hardware issue)
> P reallocated, e.g. as guest
> RAM, now holding guest data
> store to a VA in the same 2 MiB
> region (VA reused by a new mmap)
> walks P as a PTE table and
> writes to the PFN it finds
> a 4K copy of a Windows page
> table page, 512 copies of
> 0x0fffffff0c930020, lands on
> a dentry_hashtable page
> days later:
> __d_lookup() hashes into that
> page and faults
>
> There are weak points, and I want to be upfront about them:
>
> - The public reproducer demonstrates stale leaf data only. This chain
> needs a stale paging-structure entry.
> - The Windows reading of the word is weak (see above).
> - The guest mix on the Milan host isn't stated. If it runs no Windows
> guests, the Windows part of this falls apart.
> - The target PFN comes from whatever the reused page holds, so other
> pages would be hit too. The dentry table is just large, read
> constantly and never freed, which makes it a likely place to notice
> the damage.
>
>
> What the 6.18.53 vmcore can settle
> ==================================
>
> B below is the crash 2 bucket address from above.
>
> 1. Bucket and physical page:
>
> p dentry_hashtable
> p d_hash_shift
> eval 0xffff985ffd223600 + (0x5560450b >> 7) * 8 -> B
> vtop B -> PHYS_B
> kmem B
> eval B & ~0xfff -> PAGE_B
>
> 2. The whole 4K page. This is the most important check:
>
> rd -64 PAGE_B 512
>
> Then do the same for the pages at PAGE_B - 0x1000 and PAGE_B +
> 0x1000 (compute them with eval first).
>
>
> If the hypothesis holds, most of the 512 words in PAGE_B are
> 0x0fffffff0c930020, and the run starts and ends on the 4K boundary.
> A few buckets may hold valid dentry pointers inserted after the
> damage, and their chains should end in the same word.
>
> If only the one word is bad and the rest are 0 or dentry pointers,
> the hypothesis is refuted. In that case this was a single 8-byte
> store, and the 4K copy story is wrong. The neighbouring pages show
> whether there was more than one event.
>
> 3. Where else the word lives:
>
> search -p 0x0fffffff0c930020
> search -p -w 0x0c930020
>
> Physical searches over 1 TB take a while. For each hit, use kmem -p
> on the physical address to classify it: qemu anonymous memory (guest
> RAM), page cache, slab, page table or reserved. If a hit is in guest
> RAM, rd -64 the page and count the repeats.
>
> The hypothesis predicts pages in Windows guests' RAM where the word
> repeats in long runs (guest page tables), plus the damaged
> hash-table page or pages. If the word exists only in the dentry
> table, the guest-data origin loses its support.
>
> 4. Who points at the bucket's page:
>
> search -p -m 0xfff0000000000fff PAGE_PHYS
>
> PAGE_PHYS here is PHYS_B & ~0xfff. This matches any 8-byte value
> whose bits 12-51 equal the page's physical address, i.e. any
> PTE-shaped entry pointing at it. The direct map normally covers the
> hash table with 2M or 1G leaves, so no 4K entry naming that page
> should exist.
>
> A PTE-shaped hit inside guest RAM that is laid out like a page table
> would directly support the "guest data walked as a host page table"
> route. Expect noise. Only hits with P and RW set and sensible low
> bits are interesting. The page that served as the stale table may
> have been reused again since, so finding no hit does not refute the
> hypothesis.
>
> 5. Whether INVLPGB was in use at crash time:
>
> p/x boot_cpu_data.x86_capability[13] # bit 3 (0x8) set: INVLPGB used
> p global_asid_available # 2039 at boot (PTI built in),
> # 4087 without; lower means
> # global ASIDs were handed out
> p last_global_asid
> ps | grep -E 'qemu|vhost'
> task -R mm <qemu pid>
> p ((struct mm_struct *)<mm>)->context.global_asid # non-zero: broadcast
> p cpu_tlbstate:a # loaded_mm_asid >= 6 is global
>
> 6. KVM and platform state:
>
> mod -s kvm_amd
> p npt_enabled
> p nested
> p avic
> p x2avic_enabled
> sys config | grep -E 'BROADCAST_TLB_FLUSH|DEBUG_VM|INIT_ON_(ALLOC|FREE)|PAGE_POISON'
> log | grep -iE 'AVIC|AMD-Vi|microcode'
>
> From a live host on the same build, please also send:
>
> cat /sys/module/kvm_amd/parameters/{npt,nested,avic}
> grep -E 'model name|microcode' /proc/cpuinfo | sort -u
> virsh dumpxml <windows-vm> | grep -E 'hyperv|feature|<cpu'
>
> Please also tell us whether the Windows guests run with VBS/HVCI
> (msinfo32, "Virtualization-based security"), and which guest OSes run
> on the Milan host.
>
>
> KVM SVM INVLPGA (26505e1b5b54)
> ==============================
>
> Posting: https://lore.kernel.org/all/20260723094419.630204-1-pbonzini@redhat.com/
>
> This is a guest-memory fix. It is not the host corruption fix.
>
> In 6.18.53, svm_flush_tlb_gva() (INVLPGA) only drops guest-virtual
> translations. It is never the thing that protects a host page about to
> be freed. Every path that takes a page away from a guest flushes
> through kvm_flush_remote_tlbs() -> svm_flush_tlb_all() ->
> TLB_CONTROL_FLUSH_ASID before the host page can be reused. That covers
> mmu_notifier zaps (munmap, MADV_DONTNEED from balloon or free page
> reporting), KSM, compaction and migration, memslot delete, dirty
> logging, TDP page-table free and VM destroy.
>
> GPA invalidations don't reach INVLPGA at all:
>
> arch/x86/kvm/mmu/mmu.c:kvm_mmu_invalidate_addr() {
> ...
> /* It's actually a GPA for vcpu->arch.guest_mmu. */
> if (mmu != &vcpu->arch.guest_mmu) {
> ...
> kvm_x86_call(flush_tlb_gva)(vcpu, addr);
> }
>
> A missed INVLPGA therefore leaves at worst a stale GVA->HPA entry
> whose HPA still backs one of the guest's own GPAs. The result is
> guest-internal corruption, i.e. Windows BSODs, which is what Red Hat
> saw. It can't touch host page cache or the dentry table.
>
> On 6.18.53 with NPT, INVLPGA is reached from:
>
> - the Hyper-V PV TLB flush (hv-tlbflush). You said this never ran on
> the crashed nodes.
> - L1 INVLPGA under nested SVM (VBS/HVCI guests). This is redundant,
> because every nested VMRUN and #VMEXIT already does a full ASID
> flush.
> - the rare emulated #PF and emulated INVLPG paths.
>
> Applying it to 6.18.y: you don't need the ~134 svm.c commits. The
> minimal backport changes only svm_flush_tlb_gva(), and
> svm_flush_tlb_asid() is already defined above it in 6.18.
> "git apply --check" passes on 6.18.53 and on 6.18.44. I have not
> checked 6.18.54.
>
> The rest of the upstream patch is a signature change that lets
> kvm_hv_vcpu_flush_tlb() stop iterating early, which is a performance
> optimisation only, and an ERAPS register mark that has no counterpart
> in 6.18. The patch is attached as
> 0001-KVM-SVM-flush_tlb_gva-full-asid-flush-6.18.y.patch. The
> functional change is:
>
> @@ -4071,7 +4071,18 @@ static void svm_flush_tlb_gva(struct kvm_vcpu *vcpu, gva_t gva)
> {
> struct vcpu_svm *svm = to_svm(vcpu);
>
> - invlpga(gva, svm->vmcb->control.asid);
> + /*
> + * INVLPGA has had errata on Genoa and Turin, and even on older
> + * generations there were reports of Windows BSODs if INVLPGA
> + * was used for Hyper-V tlbflush. Use it only for shadow paging
> + * where it seems to be okay.
> + */
> + if (!npt_enabled) {
> + invlpga(gva, svm->vmcb->control.asid);
> + return;
> + }
> +
> + svm_flush_tlb_asid(vcpu);
> }
>
> It's worth carrying for the Windows guests, especially any with
> hv-tlbflush or VBS. It won't stop the dentry or libglib corruption. The
> host mitigation is tlbi=ipi (or clearcpuid=419), and that needs no
> patch.
>
>
> AVIC, the GA log and guest-side reports
> =======================================
>
> AVIC: ca2967de5a5b (v6.18-rc1) makes avic=auto turn AVIC and x2AVIC
> on by default for Zen4+ with X2AVIC. 5.15 had avic=0.
>
> arch/x86/kvm/svm/avic.c (6.18.53):
>
> if (avic == AVIC_AUTO_MODE)
> avic = boot_cpu_has(X86_FEATURE_X2AVIC) &&
> (boot_cpu_data.x86 > 0x19 || cpu_feature_enabled(X86_FEATURE_ZEN4));
>
> On a Zen4 or Zen5 host with X2AVIC, AVIC and IPI virtualisation (and,
> with the AMD IOMMU in GA mode, device-posted interrupts) are therefore
> on in 6.18 where they were off in 5.15. On the Milan 7343 host AVIC is
> off unless avic=1 is set explicitly. That host crashed identically, so
> AVIC is not needed for the dentry crash unless it was forced on there.
> The /sys/module/kvm_amd/parameters/avic values from both hosts settle
> this.
>
> GA log UAF: 78684b65fcc0 ("KVM: SVM: Remove VM from the GA Log notifier
> list before VM destruction", v7.3-rc1, not in 6.18.y) fixes a
> use-after-free in avic_ga_log_notifier() at VM teardown. Sean's own
> assessment in the commit is that it is "all but impossible to trigger".
> I'd treat it as low priority and mention it only for completeness.
>
> Guest crashes on the Milan host: Joris de Vries reported on kvm@
> (2026-09-03, in the 26505e1b5b54 thread, Message-ID
> <E6439A7D-77EA-475D-BEBC-6BB6EDA10A0C@gmail.com>) that on Zen3 (Ryzen
> 5900X, also on v7.3-rc1) guests take reserved-bit page faults from a
> stale guest page-walk-cache entry. It happens after a 2 MiB-aligned
> guest mapping is torn down and its table reused. No KVM-side
> invalidation fixed it, but npt=0 did. That is guest-internal and
> independent of host INVLPGB, so it can't explain host corruption. It
> could be relevant to the guest crashes during migration on Milan. I
> have not looked into it beyond reading the report.
>
>
> Verified versus inferred
> ========================
>
> Verified, from the code in the stable trees, the two oopses and the
> cited threads:
>
> - Both oopses read the same word as a dentry_hashtable bucket head on
> the first loop iteration.
> - dentry_hashtable is boot-time memblock memory that is never freed.
> - 6.18.53 contains the complete CPA fix series, and it crashed.
> - The INVLPGB code is identical in 6.18.15, 6.18.44 and 6.18.53. I
> found no kernel-side ordering hole. TLBSYNC precedes every free.
> - tlbi= exists from 6.18.45 and only "ipi" does anything.
> clearcpuid=419 works on 6.18.15, 6.18.44 and 6.18.53, and
> clearcpuid=invlpgb works on none of them. INVLPGB never shows in
> /proc/cpuinfo on 6.18.
> - 44126343d58c is missing from 6.18.15, so nopcid is unsafe there.
> - KVM never issues INVLPGB. The INVLPGA path can only reach guest
> memory.
> - AVIC defaults to on for Zen4+ since v6.18-rc1.
> - The INVLPGB stale-TLB issue is publicly reproduced and acknowledged
> on Zen5, with no kernel fix yet. This is as reported in the lore
> thread, not something I reproduced.
>
> Inferred, not demonstrated:
>
> - That the Zen5 issue exists on Zen3 and Zen4.
> - That it is the cause here.
> - The paging-structure route to the dentry table.
> - That the word comes from Windows guest page tables.
> - That the first host is Genoa. The board is SP5, so it could be Turin.
> - The crash 2 bucket address. It assumes RDX is the table base, as in
> crash 1.
>
>
> Ruled out
> =========
>
> - CPA collapse race (41d88484c71c and its fixes): crash 2 is on
> 6.18.53, which has every fix.
>
> - KVM SVM INVLPGA (26505e1b5b54): it only reaches guest memory. It is
> still worth backporting for Windows guests (above).
>
> - Host mm writing a swap or migration PTE through a stale pointer: the
> first host has no swap, and an audit of every non-present PTE store
> site found nothing.
>
> - KVM NPT level or pfn errors, and TDP MMU huge page recovery: KVM can
> only map PFNs that the host page tables hold, and the memblock table
> is not one of them.
>
> - THP and NUMA balancing: both were off on the Milan host (THP
> "never", numa_balancing 0), and it still crashed.
>
> - 55ddbc2ca6d5 (upstream f7491d7c81db, pmd_modify() dropping
> _PAGE_DIRTY): this causes data loss, not a stray write. It is fixed
> in 6.18.53.
>
> - d053eb7e09e1 (upstream 1e75a8255f11, AMD IOMMU completion wait):
> this has been in the tree since 6.18.42, and 6.18.15 crashed too.
>
> - EFER.TCE (enabled since 6.15) versus pud_free_pmd_page() flushing
> only one page: this is a structural hole, independent of INVLPGB.
> However, it is only reachable through a >= 1 GiB ioremap, which is
> unlikely on these hosts. clearcpuid=tce would rule it out if needed.
>
>
> What would help most
> ====================
>
> 1. Boot production with tlbi=ipi (6.18.45+) or clearcpuid=419 (6.18.15
> and 6.18.44). Check that it took with x86_capability[13], not with
> /proc/cpuinfo.
>
> 2. Run Matt's reproducer on a drained Milan host and a drained Genoa
> host, with and without tlbi=ipi, and post the result to the INVLPGB
> thread.
>
> 3. Send the CPU model and microcode for every affected host, including
> the first crash host and nrbphav4a.
>
> 4. From the 6.18.53 vmcore: the whole-page dump around the bucket, the
> two searches for the word, the masked referrer search and the
> global ASID state (above).
>
> 5. Send the kvm_amd parameters for both crash hosts, tell us which
> guest OSes run on the Milan host, and whether the Windows guests use
> VBS.
--
Ing. Nikola CIPRICH
technický ředitel
+420 591 166 214
+420 777 093 799
nikola.ciprich@linuxbox.cz
www.linuxbox.cz
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: hunting memory corruption bug in 6.18.x
2026-10-02 14:55 ` Nikola Ciprich
@ 2026-10-02 19:27 ` Lorenzo Stoakes (ARM)
0 siblings, 0 replies; 16+ messages in thread
From: Lorenzo Stoakes (ARM) @ 2026-10-02 19:27 UTC (permalink / raw)
To: Nikola Ciprich
Cc: linux-mm, linux-kernel, akpm, david, Mike Rapoport, Dave Hansen,
Pedro Falcato, Kiryl Shutsemau, luizcap, pbonzini,
Borislav Petkov, Tal Zussman, Rik van Riel, Matt Fleming
On Fri, Oct 02, 2026 at 04:55:39PM +0200, Nikola Ciprich wrote:
> > The fact you've seen this on a Milan machine (EPYC 7343) is new information
> > so that could be useful for the report, so if you can reproduce _without
> > the fix_ first that'd be very useful to know!
>
> well, not so good news here.. I tried the reproducer on multiple
> lab machines and also on one drained production box on which we've
> experienced one crash and wasn't able to get a single hit so far.
> Here's the list of machines:
>
> hostname CPU RAM kernel microcode
> pocstdv1b EPYC 9124 384G 6.18.31lb9.01 0x0a101158
> labtest EPYC 9274F 64G 6.18.20lb9.01 0x0a101158
> lbxovav6d EPYC 9124 1.5TB 6.18.31lb9.01 0x0a101158
> prfjazv1g EPYC 7343 1TB 6.18.53lb9.02 0x0a0011de
> nrbphav4a EPYC 7343 1TB 6.18.15lb9.01 0x0a0011de
>
> I'll leave it running for few hours, and report back. If I'm able to reproduce,
> I'll try recommended workarounds and report as well.
Thanks for trying that!
Yeah, it might have been tuned to the zen arch rather than milan so that could
be making it less effective unfortunately.
> (and lots of more addresses, truncated)
>
> crash> vtop 0xffff986002783a50
> VIRTUAL PHYSICAL
> ffff986002783a50 582783a50
>
> PGD DIRECTORY: ffffffffa0836000
> PAGE DIRECTORY: 143a7801067
> PUD: 143a7801c00 => 80000005800001e3
> PAGE: 580000000 (1GB)
>
> PTE PHYSICAL FLAGS
> 80000005800001e3 580000000 (PRESENT|RW|ACCESSED|DIRTY|PSE|GLOBAL|NX)
>
> PAGE PHYSICAL MAPPING INDEX CNT FLAGS
> fffff2715609e0c0 582783000 ffff996e5a8932d9 7f6a1589e 1 2ffff800020938 uptodate,dirty,lru,active,owner_2,swapbacked
>
> crash> search -p -m 0xfff0000000000fff 582783a50
> 3c600964f0: 8000000582783067
> 11e238a54f0: 582783e67
>
> should you need anything else, please let me know
Thanks! The LLM has informed me that it got that wrong but it was interesting
data (...!) which suggests guest memory page tables might have somehow ended up
there.
Could you try:
crash> p d_hash_shift
crash> eval (0x5560450b >> D) * 8 + 0xffffb18900a00000
D = the d_hash_shift value
crash> eval (B & 0xfffffffffffff000)
B = the hex result above
crash> vtop B
crash> rd -64 P 512
P = the hex result of the second eval
crash> search -p -m 0xfff0000000000fff PHYS
PHYS = the PHYSICAL column from vtop
And for the pages the physical search found:
crash> kmem 0x1ae234000
crash> kmem 0x24f0e2000
crash> kmem 0x103360000
crash> rd -p 0x1ae234000 512
Thanks!
--
Cheers, Lorenzo
^ permalink raw reply [flat|nested] 16+ messages in thread
end of thread, other threads:[~2026-10-02 19:27 UTC | newest]
Thread overview: 16+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-25 8:48 hunting memory corruption bug in 6.18.x Nikola Ciprich
2026-09-25 10:05 ` Lorenzo Stoakes (ARM)
2026-09-25 12:13 ` Lorenzo Stoakes (ARM)
2026-09-26 5:58 ` Nikola Ciprich
2026-09-26 9:32 ` Lorenzo Stoakes (ARM)
2026-09-28 9:05 ` Nikola Ciprich
2026-09-29 19:11 ` Nikola Ciprich
2026-09-30 9:07 ` Lorenzo Stoakes (ARM)
2026-09-30 18:40 ` Nikola Ciprich
2026-10-01 19:40 ` Nikola Ciprich
2026-10-01 22:21 ` David Laight
2026-10-02 9:50 ` Lorenzo Stoakes (ARM)
2026-10-02 14:55 ` Nikola Ciprich
2026-10-02 19:27 ` Lorenzo Stoakes (ARM)
2026-09-26 16:02 ` Luiz Capitulino
2026-09-28 8:47 ` Nikola Ciprich
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®