* Re: [PATCH 1/1] x86/mm: fix incomplete page-table invalidation with TCE
2026-10-05 5:23 [PATCH 1/1] x86/mm: fix incomplete page-table invalidation with TCE Lance Yang
@ 2026-10-05 5:47 ` Andrew Morton
2026-10-05 6:09 ` Pedro Falcato
` (2 subsequent siblings)
3 siblings, 0 replies; 20+ messages in thread
From: Andrew Morton @ 2026-10-05 5:47 UTC (permalink / raw)
To: Lance Yang
Cc: dave.hansen, luto, peterz, tglx, mingo, bp, x86, hpa, riel,
linux-kernel, qi.zheng, nadav.amit, thomas.lendacky, kernel-team,
linux-mm, brendan.jackman, jannh, mhklinux, andrew.cooper3,
Manali.Shukla, mingo, stable, toshi.kani, david,
mikhail.v.gavrilov, pfalcato
On Mon, 5 Oct 2026 13:23:02 +0800 Lance Yang <lance.yang@linux.dev> wrote:
> pud_free_pmd_page() uses a single-address invalidation to flush the
> paging-structure caches before freeing the page tables. With AMD TCE
> enabled, this only invalidates upper-level entries associated with the
> target address. Cached PMD entries for other addresses in the PUD range can
> still reference the PTE pages being freed.
>
> The AMD manual quoted in the commit enabling TCE says these instructions
> remove
>
> "only those upper-level entries that lead to the target PTE in the page
> table hierarchy, leaving unrelated upper-level entries intact."
>
> Even with all PTEs cleared, speculative page walks can cache present PMD
> entries after the earlier TLB purge.
>
> Use a full TLB flush before freeing the page tables on CPUs with TCE. Keep
> the single-address invalidation otherwise.
>
> Fixes: 440a65b7d25f ("x86/mm: Enable AMD translation cache extensions")
> Cc: stable@vger.kernel.org
Sorry, but my usual complaint applies.
Check the first 32 lines of
Documentation/process/stable-kernel-rules.rst. They're very simple,
I'm sure x86 people can immediately grasp the importance of this
change, but nobody else can. This includes -stable maintainers as well
as a large number of other downstream users of our work who are
wondering "why should I apply this to my kernel". Let's tell them!
iow, and not for the first time: when fixing a bug please fully
describe the userspace-visible runtime effects of that bug.
Thanks.
^ permalink raw reply [flat|nested] 20+ messages in thread* Re: [PATCH 1/1] x86/mm: fix incomplete page-table invalidation with TCE
2026-10-05 5:23 [PATCH 1/1] x86/mm: fix incomplete page-table invalidation with TCE Lance Yang
2026-10-05 5:47 ` Andrew Morton
@ 2026-10-05 6:09 ` Pedro Falcato
2026-10-05 7:29 ` Lance Yang
2026-10-05 6:38 ` Nadav Amit
2026-10-05 15:15 ` Rik van Riel
3 siblings, 1 reply; 20+ messages in thread
From: Pedro Falcato @ 2026-10-05 6:09 UTC (permalink / raw)
To: Lance Yang
Cc: dave.hansen, luto, peterz, tglx, mingo, bp, x86, hpa, riel,
linux-kernel, qi.zheng, nadav.amit, thomas.lendacky, kernel-team,
linux-mm, akpm, brendan.jackman, jannh, mhklinux, andrew.cooper3,
Manali.Shukla, mingo, stable, toshi.kani, david,
mikhail.v.gavrilov
On Mon, Oct 05, 2026 at 01:23:02PM +0800, Lance Yang wrote:
> pud_free_pmd_page() uses a single-address invalidation to flush the
> paging-structure caches before freeing the page tables. With AMD TCE
> enabled, this only invalidates upper-level entries associated with the
> target address. Cached PMD entries for other addresses in the PUD range can
> still reference the PTE pages being freed.
>
> The AMD manual quoted in the commit enabling TCE says these instructions
> remove
>
> "only those upper-level entries that lead to the target PTE in the page
> table hierarchy, leaving unrelated upper-level entries intact."
>
> Even with all PTEs cleared, speculative page walks can cache present PMD
> entries after the earlier TLB purge.
>
> Use a full TLB flush before freeing the page tables on CPUs with TCE. Keep
> the single-address invalidation otherwise.
>
> Fixes: 440a65b7d25f ("x86/mm: Enable AMD translation cache extensions")
> Cc: stable@vger.kernel.org
> Signed-off-by: Lance Yang <lance.yang@linux.dev>
I'm not sure this is correct. The PUD is clear. We invalidate the TLB, which
invalidates the translation caches for that walk. invlpg will notice the PUD
isn't present. I don't see a case where it can ever not clear the rest
of the translation caches for all leaves. And, in fact, by that point the CPU
can (does?) probably formally treat the PUD as the leaf.
Did you repro any bug related to this? The functionality is perhaps
underspecified in the AMD manual.
--
Pedro
^ permalink raw reply [flat|nested] 20+ messages in thread* Re: [PATCH 1/1] x86/mm: fix incomplete page-table invalidation with TCE
2026-10-05 6:09 ` Pedro Falcato
@ 2026-10-05 7:29 ` Lance Yang
2026-10-05 8:19 ` Andrew Cooper
2026-10-05 10:24 ` Pedro Falcato
0 siblings, 2 replies; 20+ messages in thread
From: Lance Yang @ 2026-10-05 7:29 UTC (permalink / raw)
To: Pedro Falcato
Cc: dave.hansen, luto, peterz, tglx, mingo, bp, x86, hpa, riel,
linux-kernel, qi.zheng, nadav.amit, thomas.lendacky, kernel-team,
linux-mm, akpm, brendan.jackman, jannh, mhklinux, andrew.cooper3,
Manali.Shukla, mingo, stable, toshi.kani, david,
mikhail.v.gavrilov
On 2026/10/5 14:09, Pedro Falcato wrote:
> On Mon, Oct 05, 2026 at 01:23:02PM +0800, Lance Yang wrote:
>> pud_free_pmd_page() uses a single-address invalidation to flush the
>> paging-structure caches before freeing the page tables. With AMD TCE
>> enabled, this only invalidates upper-level entries associated with the
>> target address. Cached PMD entries for other addresses in the PUD range can
>> still reference the PTE pages being freed.
>>
>> The AMD manual quoted in the commit enabling TCE says these instructions
>> remove
>>
>> "only those upper-level entries that lead to the target PTE in the page
>> table hierarchy, leaving unrelated upper-level entries intact."
>>
>> Even with all PTEs cleared, speculative page walks can cache present PMD
>> entries after the earlier TLB purge.
>>
>> Use a full TLB flush before freeing the page tables on CPUs with TCE. Keep
>> the single-address invalidation otherwise.
>>
>> Fixes: 440a65b7d25f ("x86/mm: Enable AMD translation cache extensions")
>> Cc: stable@vger.kernel.org
>> Signed-off-by: Lance Yang <lance.yang@linux.dev>
>
> I'm not sure this is correct. The PUD is clear. We invalidate the TLB, which
> invalidates the translation caches for that walk. invlpg will notice the PUD
> isn't present. I don't see a case where it can ever not clear the rest
> of the translation caches for all leaves. And, in fact, by that point the CPU
> can (does?) probably formally treat the PUD as the leaf.
IIUC, clearing the PUD in memory doesn't invalidate cached PMD entries by
itself. With TCE enabled, flushing one address only invalidates the entries
associated with that address ...
(That's how I read the manual, but AMD folks, please correct me if I'm
missing something.)
So couldn't other cached PMDs under the same PUD survive?
Later/speculative page walks could use one of those cached entries
without rereading the cleared PUD, and access a PTE page we've already
freed ... If that can happen, it sounds pretty serious ...
>
> Did you repro any bug related to this? The functionality is perhaps
> underspecified in the AMD manual.
TBH, I don't have a reproducer yet. Just LLM stumbled upon this while
I was investigating another memory corruption issue [1].
[1] https://lore.kernel.org/linux-mm/arY1Wq6R9OY20ans@pcnci.linuxbox.cz/#t
^ permalink raw reply [flat|nested] 20+ messages in thread* Re: [PATCH 1/1] x86/mm: fix incomplete page-table invalidation with TCE
2026-10-05 7:29 ` Lance Yang
@ 2026-10-05 8:19 ` Andrew Cooper
2026-10-05 8:32 ` Lance Yang
2026-10-05 10:24 ` Pedro Falcato
1 sibling, 1 reply; 20+ messages in thread
From: Andrew Cooper @ 2026-10-05 8:19 UTC (permalink / raw)
To: Lance Yang, Pedro Falcato
Cc: Andrew Cooper, dave.hansen, luto, peterz, tglx, mingo, bp, x86,
hpa, riel, linux-kernel, qi.zheng, nadav.amit, thomas.lendacky,
kernel-team, linux-mm, akpm, brendan.jackman, jannh, mhklinux,
Manali.Shukla, mingo, stable, toshi.kani, david,
mikhail.v.gavrilov
On 05/10/2026 8:29 am, Lance Yang wrote:
>>
>> Did you repro any bug related to this? The functionality is perhaps
>> underspecified in the AMD manual.
>
> TBH, I don't have a reproducer yet. Just LLM stumbled upon this while
> I was investigating another memory corruption issue [1].
>
> [1]
> https://lore.kernel.org/linux-mm/arY1Wq6R9OY20ans@pcnci.linuxbox.cz/#t
What hardware are you running on?
That looks like the Zen5 issue, for which you want either the latest
microcode out of linux-firmware and/or
https://lore.kernel.org/r/20261002211617.1001617-1-bp@kernel.org
TCE is a no-op in Zen1 and later, so unless you're on older hardware, it
won't be that.
~Andrew
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [PATCH 1/1] x86/mm: fix incomplete page-table invalidation with TCE
2026-10-05 8:19 ` Andrew Cooper
@ 2026-10-05 8:32 ` Lance Yang
2026-10-05 9:57 ` Andrew Cooper
0 siblings, 1 reply; 20+ messages in thread
From: Lance Yang @ 2026-10-05 8:32 UTC (permalink / raw)
To: Andrew Cooper, Pedro Falcato
Cc: dave.hansen, luto, peterz, tglx, mingo, bp, x86, hpa, riel,
linux-kernel, qi.zheng, nadav.amit, thomas.lendacky, kernel-team,
linux-mm, akpm, brendan.jackman, jannh, mhklinux, Manali.Shukla,
mingo, stable, toshi.kani, david, mikhail.v.gavrilov
On 2026/10/5 16:19, Andrew Cooper wrote:
> On 05/10/2026 8:29 am, Lance Yang wrote:
>>>
>>> Did you repro any bug related to this? The functionality is perhaps
>>> underspecified in the AMD manual.
>>
>> TBH, I don't have a reproducer yet. Just LLM stumbled upon this while
>> I was investigating another memory corruption issue [1].
>>
>> [1]
>> https://lore.kernel.org/linux-mm/arY1Wq6R9OY20ans@pcnci.linuxbox.cz/#t
>
> What hardware are you running on?
>
> That looks like the Zen5 issue, for which you want either the latest
> microcode out of linux-firmware and/or
> https://lore.kernel.org/r/20261002211617.1001617-1-bp@kernel.org
>
> TCE is a no-op in Zen1 and later, so unless you're on older hardware, it
> won't be that.
Just to clarify ... these are two separate issues.
I mentioned [1] only to explain how this came up while investigating
something else. I'm not claiming that TCE caused the corruption reported
there :)
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [PATCH 1/1] x86/mm: fix incomplete page-table invalidation with TCE
2026-10-05 8:32 ` Lance Yang
@ 2026-10-05 9:57 ` Andrew Cooper
2026-10-05 10:12 ` Lance Yang
0 siblings, 1 reply; 20+ messages in thread
From: Andrew Cooper @ 2026-10-05 9:57 UTC (permalink / raw)
To: Lance Yang, Pedro Falcato
Cc: Andrew Cooper, dave.hansen, luto, peterz, tglx, mingo, bp, x86,
hpa, riel, linux-kernel, qi.zheng, nadav.amit, thomas.lendacky,
kernel-team, linux-mm, akpm, brendan.jackman, jannh, mhklinux,
Manali.Shukla, mingo, stable, toshi.kani, david,
mikhail.v.gavrilov
On 05/10/2026 9:32 am, Lance Yang wrote:
>
>
> On 2026/10/5 16:19, Andrew Cooper wrote:
>> On 05/10/2026 8:29 am, Lance Yang wrote:
>>>>
>>>> Did you repro any bug related to this? The functionality is perhaps
>>>> underspecified in the AMD manual.
>>>
>>> TBH, I don't have a reproducer yet. Just LLM stumbled upon this while
>>> I was investigating another memory corruption issue [1].
>>>
>>> [1]
>>> https://lore.kernel.org/linux-mm/arY1Wq6R9OY20ans@pcnci.linuxbox.cz/#t
>>
>> What hardware are you running on?
>>
>> That looks like the Zen5 issue, for which you want either the latest
>> microcode out of linux-firmware and/or
>> https://lore.kernel.org/r/20261002211617.1001617-1-bp@kernel.org
>>
>> TCE is a no-op in Zen1 and later, so unless you're on older hardware, it
>> won't be that.
>
> Just to clarify ... these are two separate issues.
>
> I mentioned [1] only to explain how this came up while investigating
> something else. I'm not claiming that TCE caused the corruption reported
> there :)
Please can you answer the question. Which CPU are you seeing this on?
~Andrew
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [PATCH 1/1] x86/mm: fix incomplete page-table invalidation with TCE
2026-10-05 9:57 ` Andrew Cooper
@ 2026-10-05 10:12 ` Lance Yang
2026-10-05 14:22 ` Borislav Petkov
2026-10-05 15:22 ` Andrew Cooper
0 siblings, 2 replies; 20+ messages in thread
From: Lance Yang @ 2026-10-05 10:12 UTC (permalink / raw)
To: Andrew Cooper, Pedro Falcato
Cc: dave.hansen, luto, peterz, tglx, mingo, bp, x86, hpa, riel,
linux-kernel, qi.zheng, nadav.amit, thomas.lendacky, kernel-team,
linux-mm, akpm, brendan.jackman, jannh, mhklinux, Manali.Shukla,
mingo, stable, toshi.kani, david, mikhail.v.gavrilov
On 2026/10/5 17:57, Andrew Cooper wrote:
> On 05/10/2026 9:32 am, Lance Yang wrote:
>>
>>
>> On 2026/10/5 16:19, Andrew Cooper wrote:
>>> On 05/10/2026 8:29 am, Lance Yang wrote:
>>>>>
>>>>> Did you repro any bug related to this? The functionality is perhaps
>>>>> underspecified in the AMD manual.
>>>>
>>>> TBH, I don't have a reproducer yet. Just LLM stumbled upon this while
>>>> I was investigating another memory corruption issue [1].
>>>>
>>>> [1]
>>>> https://lore.kernel.org/linux-mm/arY1Wq6R9OY20ans@pcnci.linuxbox.cz/#t
>>>
>>> What hardware are you running on?
>>>
>>> That looks like the Zen5 issue, for which you want either the latest
>>> microcode out of linux-firmware and/or
>>> https://lore.kernel.org/r/20261002211617.1001617-1-bp@kernel.org
>>>
>>> TCE is a no-op in Zen1 and later, so unless you're on older hardware, it
>>> won't be that.
>>
>> Just to clarify ... these are two separate issues.
>>
>> I mentioned [1] only to explain how this came up while investigating
>> something else. I'm not claiming that TCE caused the corruption reported
>> there :)
>
> Please can you answer the question. Which CPU are you seeing this on?
>
Which CPU? None so far. As I said, I don't have a reproducer. I'm trying
to make sense of what the manual says and what the code does ...
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [PATCH 1/1] x86/mm: fix incomplete page-table invalidation with TCE
2026-10-05 10:12 ` Lance Yang
@ 2026-10-05 14:22 ` Borislav Petkov
2026-10-05 14:36 ` Lance Yang
2026-10-05 15:22 ` Andrew Cooper
1 sibling, 1 reply; 20+ messages in thread
From: Borislav Petkov @ 2026-10-05 14:22 UTC (permalink / raw)
To: Lance Yang
Cc: Andrew Cooper, Pedro Falcato, dave.hansen, luto, peterz, tglx,
mingo, x86, hpa, riel, linux-kernel, qi.zheng, nadav.amit,
thomas.lendacky, kernel-team, linux-mm, akpm, brendan.jackman,
jannh, mhklinux, Manali.Shukla, mingo, stable, toshi.kani, david,
mikhail.v.gavrilov
On Mon, Oct 05, 2026 at 06:12:45PM +0800, Lance Yang wrote:
> Which CPU? None so far. As I said, I don't have a reproducer. I'm trying
> to make sense of what the manual says and what the code does ...
So you're sending a patch for which you have absolutely no justification? And
CCing it to stable too?
And that patch hasn't been tested to actually address anything?
Just a wild hunch?
I wouldn't do that in the future if I were you.
HTH.
--
Regards/Gruss,
Boris.
https://people.kernel.org/tglx/notes-about-netiquette
^ permalink raw reply [flat|nested] 20+ messages in thread* Re: [PATCH 1/1] x86/mm: fix incomplete page-table invalidation with TCE
2026-10-05 14:22 ` Borislav Petkov
@ 2026-10-05 14:36 ` Lance Yang
2026-10-05 14:53 ` Borislav Petkov
0 siblings, 1 reply; 20+ messages in thread
From: Lance Yang @ 2026-10-05 14:36 UTC (permalink / raw)
To: Borislav Petkov
Cc: Andrew Cooper, Pedro Falcato, dave.hansen, luto, peterz, tglx,
mingo, x86, hpa, riel, linux-kernel, qi.zheng, nadav.amit,
thomas.lendacky, kernel-team, linux-mm, akpm, brendan.jackman,
jannh, mhklinux, Manali.Shukla, mingo, stable, toshi.kani, david,
mikhail.v.gavrilov
On 2026/10/5 22:22, Borislav Petkov wrote:
> On Mon, Oct 05, 2026 at 06:12:45PM +0800, Lance Yang wrote:
>> Which CPU? None so far. As I said, I don't have a reproducer. I'm trying
>> to make sense of what the manual says and what the code does ...
>
> So you're sending a patch for which you have absolutely no justification? And
> CCing it to stable too?
>
> And that patch hasn't been tested to actually address anything?
>
> Just a wild hunch?
>
> I wouldn't do that in the future if I were you.
>
> HTH.
I don't have an AMD machine to test this on. I spotted something
that looked wrong in the code and sent it out to see what others
think.
Maybe I'm reading the manual wrong, but is raising a concern
based on the code and the manual really just "a wild hunch"?
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [PATCH 1/1] x86/mm: fix incomplete page-table invalidation with TCE
2026-10-05 14:36 ` Lance Yang
@ 2026-10-05 14:53 ` Borislav Petkov
2026-10-05 14:58 ` Lance Yang
0 siblings, 1 reply; 20+ messages in thread
From: Borislav Petkov @ 2026-10-05 14:53 UTC (permalink / raw)
To: Lance Yang
Cc: Andrew Cooper, Pedro Falcato, dave.hansen, luto, peterz, tglx,
mingo, x86, hpa, riel, linux-kernel, qi.zheng, nadav.amit,
thomas.lendacky, kernel-team, linux-mm, akpm, brendan.jackman,
jannh, mhklinux, Manali.Shukla, mingo, stable, toshi.kani, david,
mikhail.v.gavrilov
On Mon, Oct 05, 2026 at 10:36:26PM +0800, Lance Yang wrote:
> I don't have an AMD machine to test this on. I spotted something
> that looked wrong in the code and sent it out to see what others
> think.
You send a normal mail - not a patch.
> Maybe I'm reading the manual wrong, but is raising a concern
> based on the code and the manual really just "a wild hunch"?
If it is formulated like a patch which is supposed to be even urgent, then
you're reading our process wrong.
https://docs.kernel.org/process/development-process.html
And since you have not tested it to actually fix anything, then this patch is
nothing but a wild hunch.
And to answer your question, no, it is not needed. Definitely not the TCE
angle.
Now, the speculative thing is another story but that whole story you're
referring to needs a lot more debugging.
But we do not apply patches on a wild hunch, without any definitive data that
they fix an issue and, if very much possible, with a clear explanation why
they fix it.
--
Regards/Gruss,
Boris.
https://people.kernel.org/tglx/notes-about-netiquette
^ permalink raw reply [flat|nested] 20+ messages in thread* Re: [PATCH 1/1] x86/mm: fix incomplete page-table invalidation with TCE
2026-10-05 14:53 ` Borislav Petkov
@ 2026-10-05 14:58 ` Lance Yang
0 siblings, 0 replies; 20+ messages in thread
From: Lance Yang @ 2026-10-05 14:58 UTC (permalink / raw)
To: Borislav Petkov
Cc: Andrew Cooper, Pedro Falcato, dave.hansen, luto, peterz, tglx,
mingo, x86, hpa, riel, linux-kernel, qi.zheng, nadav.amit,
thomas.lendacky, kernel-team, linux-mm, akpm, brendan.jackman,
jannh, mhklinux, Manali.Shukla, mingo, stable, toshi.kani, david,
mikhail.v.gavrilov
On 2026/10/5 22:53, Borislav Petkov wrote:
> On Mon, Oct 05, 2026 at 10:36:26PM +0800, Lance Yang wrote:
>> I don't have an AMD machine to test this on. I spotted something
>> that looked wrong in the code and sent it out to see what others
>> think.
>
> You send a normal mail - not a patch.
>
>> Maybe I'm reading the manual wrong, but is raising a concern
>> based on the code and the manual really just "a wild hunch"?
>
> If it is formulated like a patch which is supposed to be even urgent, then
> you're reading our process wrong.
>
> https://docs.kernel.org/process/development-process.html
>
> And since you have not tested it to actually fix anything, then this patch is
> nothing but a wild hunch.
>
> And to answer your question, no, it is not needed. Definitely not the TCE
> angle.
>
> Now, the speculative thing is another story but that whole story you're
> referring to needs a lot more debugging.
>
> But we do not apply patches on a wild hunch, without any definitive data that
> they fix an issue and, if very much possible, with a clear explanation why
> they fix it.
OK, let's drop the patch.
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [PATCH 1/1] x86/mm: fix incomplete page-table invalidation with TCE
2026-10-05 10:12 ` Lance Yang
2026-10-05 14:22 ` Borislav Petkov
@ 2026-10-05 15:22 ` Andrew Cooper
2026-10-05 15:36 ` Lance Yang
1 sibling, 1 reply; 20+ messages in thread
From: Andrew Cooper @ 2026-10-05 15:22 UTC (permalink / raw)
To: Lance Yang, Pedro Falcato
Cc: Andrew Cooper, dave.hansen, luto, peterz, tglx, mingo, bp, x86,
hpa, riel, linux-kernel, qi.zheng, nadav.amit, thomas.lendacky,
kernel-team, linux-mm, akpm, brendan.jackman, jannh, mhklinux,
Manali.Shukla, mingo, stable, toshi.kani, david,
mikhail.v.gavrilov
On 05/10/2026 11:12 am, Lance Yang wrote:
> On 2026/10/5 17:57, Andrew Cooper wrote:
>> On 05/10/2026 9:32 am, Lance Yang wrote:
>>> On 2026/10/5 16:19, Andrew Cooper wrote:
>>>> On 05/10/2026 8:29 am, Lance Yang wrote:
>>>>>>
>>>>>> Did you repro any bug related to this? The functionality is perhaps
>>>>>> underspecified in the AMD manual.
>>>>>
>>>>> TBH, I don't have a reproducer yet. Just LLM stumbled upon this while
>>>>> I was investigating another memory corruption issue [1].
>>>>>
>>>>> [1]
>>>>> https://lore.kernel.org/linux-mm/arY1Wq6R9OY20ans@pcnci.linuxbox.cz/#t
>>>>>
>>>>
>>>> What hardware are you running on?
>>>>
>>>> That looks like the Zen5 issue, for which you want either the latest
>>>> microcode out of linux-firmware and/or
>>>> https://lore.kernel.org/r/20261002211617.1001617-1-bp@kernel.org
>>>>
>>>> TCE is a no-op in Zen1 and later, so unless you're on older
>>>> hardware, it
>>>> won't be that.
>>>
>>> Just to clarify ... these are two separate issues.
>>>
>>> I mentioned [1] only to explain how this came up while investigating
>>> something else. I'm not claiming that TCE caused the corruption
>>> reported
>>> there :)
>>
>> Please can you answer the question. Which CPU are you seeing this on?
>>
>
> Which CPU? None so far. As I said, I don't have a reproducer. I'm trying
> to make sense of what the manual says and what the code does ...
I'm afraid that if you're trying to be helpful, you've had entirely the
opposite effect.
TLB handling is a complicated topic. What you've done is present what
is effectively a query about the AMD manual as if it were a bugfix for
an critical-sounding issue. You even sited a real bug-report for an
actually-critical issue, despite it turning out to have nothing to do
with your submission.
The patch is buggy. For starters, you should be checking is whether TCE
is enabled, not whether it's available on the system. This causes the
more expensive option to use used even when TCE is turned off.
But, AIUI INVLPG only flushes the whole structure cache because of a
windows bug which caused it to crash on a 486. AMD deliberately
introduced TCE to remove this overhead for every OS which didn't want
lumbering with a workaround for buggy windows.
Linux currently believes that it's TLB invalidation algorithm is
compatible with TCE, so at a bare minimum, you need to have some kind of
discussion on why you believe this not to be true before claiming that
it "might be unsafe because the manual says so".
It's fine to ask a question, and even ask "so shouldn't the code look
like this?" but such a patch needs a very clear RFC or QUESTION tag.
~Andrew
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [PATCH 1/1] x86/mm: fix incomplete page-table invalidation with TCE
2026-10-05 15:22 ` Andrew Cooper
@ 2026-10-05 15:36 ` Lance Yang
0 siblings, 0 replies; 20+ messages in thread
From: Lance Yang @ 2026-10-05 15:36 UTC (permalink / raw)
To: Andrew Cooper, Pedro Falcato
Cc: dave.hansen, luto, peterz, tglx, mingo, bp, x86, hpa, riel,
linux-kernel, qi.zheng, nadav.amit, thomas.lendacky, kernel-team,
linux-mm, akpm, brendan.jackman, jannh, mhklinux, Manali.Shukla,
mingo, stable, toshi.kani, david, mikhail.v.gavrilov
On 2026/10/5 23:22, Andrew Cooper wrote:
> On 05/10/2026 11:12 am, Lance Yang wrote:
>> On 2026/10/5 17:57, Andrew Cooper wrote:
>>> On 05/10/2026 9:32 am, Lance Yang wrote:
>>>> On 2026/10/5 16:19, Andrew Cooper wrote:
>>>>> On 05/10/2026 8:29 am, Lance Yang wrote:
>>>>>>>
>>>>>>> Did you repro any bug related to this? The functionality is perhaps
>>>>>>> underspecified in the AMD manual.
>>>>>>
>>>>>> TBH, I don't have a reproducer yet. Just LLM stumbled upon this while
>>>>>> I was investigating another memory corruption issue [1].
>>>>>>
>>>>>> [1]
>>>>>> https://lore.kernel.org/linux-mm/arY1Wq6R9OY20ans@pcnci.linuxbox.cz/#t
>>>>>>
>>>>>
>>>>> What hardware are you running on?
>>>>>
>>>>> That looks like the Zen5 issue, for which you want either the latest
>>>>> microcode out of linux-firmware and/or
>>>>> https://lore.kernel.org/r/20261002211617.1001617-1-bp@kernel.org
>>>>>
>>>>> TCE is a no-op in Zen1 and later, so unless you're on older
>>>>> hardware, it
>>>>> won't be that.
>>>>
>>>> Just to clarify ... these are two separate issues.
>>>>
>>>> I mentioned [1] only to explain how this came up while investigating
>>>> something else. I'm not claiming that TCE caused the corruption
>>>> reported
>>>> there :)
>>>
>>> Please can you answer the question. Which CPU are you seeing this on?
>>>
>>
>> Which CPU? None so far. As I said, I don't have a reproducer. I'm trying
>> to make sense of what the manual says and what the code does ...
>
> I'm afraid that if you're trying to be helpful, you've had entirely the
> opposite effect.
>
> TLB handling is a complicated topic. What you've done is present what
> is effectively a query about the AMD manual as if it were a bugfix for
> an critical-sounding issue. You even sited a real bug-report for an
> actually-critical issue, despite it turning out to have nothing to do
> with your submission.
>
> The patch is buggy. For starters, you should be checking is whether TCE
> is enabled, not whether it's available on the system. This causes the
> more expensive option to use used even when TCE is turned off.
>
> But, AIUI INVLPG only flushes the whole structure cache because of a
> windows bug which caused it to crash on a 486. AMD deliberately
> introduced TCE to remove this overhead for every OS which didn't want
> lumbering with a workaround for buggy windows.
>
> Linux currently believes that it's TLB invalidation algorithm is
> compatible with TCE, so at a bare minimum, you need to have some kind of
> discussion on why you believe this not to be true before claiming that
> it "might be unsafe because the manual says so".
>
>
> It's fine to ask a question, and even ask "so shouldn't the code look
> like this?" but such a patch needs a very clear RFC or QUESTION tag.
Thanks for explaining! Lesson learned. I should have made it clear
that this was a question about the manual ...
Let's drop the patch. I'll take another look.
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [PATCH 1/1] x86/mm: fix incomplete page-table invalidation with TCE
2026-10-05 7:29 ` Lance Yang
2026-10-05 8:19 ` Andrew Cooper
@ 2026-10-05 10:24 ` Pedro Falcato
2026-10-05 12:10 ` Lance Yang
1 sibling, 1 reply; 20+ messages in thread
From: Pedro Falcato @ 2026-10-05 10:24 UTC (permalink / raw)
To: Lance Yang
Cc: dave.hansen, luto, peterz, tglx, mingo, bp, x86, hpa, riel,
linux-kernel, qi.zheng, nadav.amit, thomas.lendacky, kernel-team,
linux-mm, akpm, brendan.jackman, jannh, mhklinux, andrew.cooper3,
Manali.Shukla, mingo, stable, toshi.kani, david,
mikhail.v.gavrilov
On Mon, Oct 05, 2026 at 03:29:22PM +0800, Lance Yang wrote:
>
>
> On 2026/10/5 14:09, Pedro Falcato wrote:
> > On Mon, Oct 05, 2026 at 01:23:02PM +0800, Lance Yang wrote:
> > > pud_free_pmd_page() uses a single-address invalidation to flush the
> > > paging-structure caches before freeing the page tables. With AMD TCE
> > > enabled, this only invalidates upper-level entries associated with the
> > > target address. Cached PMD entries for other addresses in the PUD range can
> > > still reference the PTE pages being freed.
> > >
> > > The AMD manual quoted in the commit enabling TCE says these instructions
> > > remove
> > >
> > > "only those upper-level entries that lead to the target PTE in the page
> > > table hierarchy, leaving unrelated upper-level entries intact."
> > >
> > > Even with all PTEs cleared, speculative page walks can cache present PMD
> > > entries after the earlier TLB purge.
> > >
> > > Use a full TLB flush before freeing the page tables on CPUs with TCE. Keep
> > > the single-address invalidation otherwise.
> > >
> > > Fixes: 440a65b7d25f ("x86/mm: Enable AMD translation cache extensions")
> > > Cc: stable@vger.kernel.org
> > > Signed-off-by: Lance Yang <lance.yang@linux.dev>
> >
> > I'm not sure this is correct. The PUD is clear. We invalidate the TLB, which
> > invalidates the translation caches for that walk. invlpg will notice the PUD
> > isn't present. I don't see a case where it can ever not clear the rest
> > of the translation caches for all leaves. And, in fact, by that point the CPU
> > can (does?) probably formally treat the PUD as the leaf.
>
> IIUC, clearing the PUD in memory doesn't invalidate cached PMD entries by
> itself. With TCE enabled, flushing one address only invalidates the entries
> associated with that address ...
>
> (That's how I read the manual, but AMD folks, please correct me if I'm
> missing something.)
>
> So couldn't other cached PMDs under the same PUD survive?
I don't read it as that. I read it as "flushing one address only invalidates
the entries on that path". So, if you flush one address, you'll flush the
whole translation cache for that range. And page table zapping agrees; if you
follow the code from zap_pte_range() -> pte_free_tlb(), it will do a single flush
for each PTE table (if the whole table is empty/non-present).
--
Pedro
^ permalink raw reply [flat|nested] 20+ messages in thread* Re: [PATCH 1/1] x86/mm: fix incomplete page-table invalidation with TCE
2026-10-05 10:24 ` Pedro Falcato
@ 2026-10-05 12:10 ` Lance Yang
0 siblings, 0 replies; 20+ messages in thread
From: Lance Yang @ 2026-10-05 12:10 UTC (permalink / raw)
To: Pedro Falcato
Cc: dave.hansen, luto, peterz, tglx, mingo, bp, x86, hpa, riel,
linux-kernel, qi.zheng, nadav.amit, thomas.lendacky, kernel-team,
linux-mm, akpm, brendan.jackman, jannh, mhklinux, andrew.cooper3,
Manali.Shukla, mingo, stable, toshi.kani, david,
mikhail.v.gavrilov
On 2026/10/5 18:24, Pedro Falcato wrote:
> On Mon, Oct 05, 2026 at 03:29:22PM +0800, Lance Yang wrote:
>>
>>
>> On 2026/10/5 14:09, Pedro Falcato wrote:
>>> On Mon, Oct 05, 2026 at 01:23:02PM +0800, Lance Yang wrote:
>>>> pud_free_pmd_page() uses a single-address invalidation to flush the
>>>> paging-structure caches before freeing the page tables. With AMD TCE
>>>> enabled, this only invalidates upper-level entries associated with the
>>>> target address. Cached PMD entries for other addresses in the PUD range can
>>>> still reference the PTE pages being freed.
>>>>
>>>> The AMD manual quoted in the commit enabling TCE says these instructions
>>>> remove
>>>>
>>>> "only those upper-level entries that lead to the target PTE in the page
>>>> table hierarchy, leaving unrelated upper-level entries intact."
>>>>
>>>> Even with all PTEs cleared, speculative page walks can cache present PMD
>>>> entries after the earlier TLB purge.
>>>>
>>>> Use a full TLB flush before freeing the page tables on CPUs with TCE. Keep
>>>> the single-address invalidation otherwise.
>>>>
>>>> Fixes: 440a65b7d25f ("x86/mm: Enable AMD translation cache extensions")
>>>> Cc: stable@vger.kernel.org
>>>> Signed-off-by: Lance Yang <lance.yang@linux.dev>
>>>
>>> I'm not sure this is correct. The PUD is clear. We invalidate the TLB, which
>>> invalidates the translation caches for that walk. invlpg will notice the PUD
>>> isn't present. I don't see a case where it can ever not clear the rest
>>> of the translation caches for all leaves. And, in fact, by that point the CPU
>>> can (does?) probably formally treat the PUD as the leaf.
>>
>> IIUC, clearing the PUD in memory doesn't invalidate cached PMD entries by
>> itself. With TCE enabled, flushing one address only invalidates the entries
>> associated with that address ...
>>
>> (That's how I read the manual, but AMD folks, please correct me if I'm
>> missing something.)
>>
>> So couldn't other cached PMDs under the same PUD survive?
>
> I don't read it as that. I read it as "flushing one address only invalidates
> the entries on that path". So, if you flush one address, you'll flush the
> whole translation cache for that range. And page table zapping agrees; if you
> follow the code from zap_pte_range() -> pte_free_tlb(), it will do a single flush
> for each PTE table (if the whole table is empty/non-present).
Let's wait for AMD folks to clarify whether a single-address flush is
sufficient :)
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [PATCH 1/1] x86/mm: fix incomplete page-table invalidation with TCE
2026-10-05 5:23 [PATCH 1/1] x86/mm: fix incomplete page-table invalidation with TCE Lance Yang
2026-10-05 5:47 ` Andrew Morton
2026-10-05 6:09 ` Pedro Falcato
@ 2026-10-05 6:38 ` Nadav Amit
2026-10-05 7:23 ` Lance Yang
2026-10-05 15:15 ` Rik van Riel
3 siblings, 1 reply; 20+ messages in thread
From: Nadav Amit @ 2026-10-05 6:38 UTC (permalink / raw)
To: Lance Yang
Cc: dave.hansen, luto, peterz, tglx, mingo, bp, x86, hpa, riel,
linux-kernel, qi.zheng, thomas.lendacky, kernel-team, linux-mm,
akpm, brendan.jackman, jannh, mhklinux, andrew.cooper3,
Manali.Shukla, mingo, stable, toshi.kani, david,
mikhail.v.gavrilov, pfalcato
>
>
> On 5 Oct 2026, at 8:23, Lance Yang <lance.yang@linux.dev> wrote:
>
> pud_free_pmd_page() uses a single-address invalidation to flush the
> paging-structure caches before freeing the page tables. With AMD TCE
> enabled, this only invalidates upper-level entries associated with the
> target address. Cached PMD entries for other addresses in the PUD range can
> still reference the PTE pages being freed.
>
> The AMD manual quoted in the commit enabling TCE says these instructions
> remove
>
> "only those upper-level entries that lead to the target PTE in the page
> table hierarchy, leaving unrelated upper-level entries intact."
>
> Even with all PTEs cleared, speculative page walks can cache present PMD
> entries after the earlier TLB purge.
>
> Use a full TLB flush before freeing the page tables on CPUs with TCE. Keep
> the single-address invalidation otherwise.
>
> Fixes: 440a65b7d25f ("x86/mm: Enable AMD translation cache extensions")
> Cc: stable@vger.kernel.org
> Signed-off-by: Lance Yang <lance.yang@linux.dev>
> ---
> arch/x86/mm/pgtable.c | 11 ++++++++++-
> 1 file changed, 10 insertions(+), 1 deletion(-)
>
> diff --git a/arch/x86/mm/pgtable.c b/arch/x86/mm/pgtable.c
> index 4a105f283cfb..6b7fa44f1bf6 100644
> --- a/arch/x86/mm/pgtable.c
> +++ b/arch/x86/mm/pgtable.c
> @@ -727,7 +727,16 @@ int pud_free_pmd_page(pud_t *pud, unsigned long addr)
> * via normal page walks. Make them unreachable
> * in cached mid-level walks too:
> */
> - flush_tlb_kernel_range(addr, addr + PAGE_SIZE-1);
> + if (boot_cpu_has(X86_FEATURE_TCE)) {
> + /*
> + * With TCE enabled, a single-address flush does not invalidate
> + * cached PMD entries for the rest of the PUD range.
> + */
> + flush_tlb_all();
> + } else {
> + /* INVLPG to clear all paging-structure caches */
> + flush_tlb_kernel_range(addr, addr + PAGE_SIZE-1);
> + }
>
It might be cleaner to replace flush_tlb_all() with:
flush_tlb_kernel_range(addr, addr + PUD_SIZE - 1);
While the flush-ceiling would usually end up doing a full flush, the
code would be easier to follow (the very least). Maybe adding stride
to kernel TLB range flushing would make sense in the future.
^ permalink raw reply [flat|nested] 20+ messages in thread* Re: [PATCH 1/1] x86/mm: fix incomplete page-table invalidation with TCE
2026-10-05 6:38 ` Nadav Amit
@ 2026-10-05 7:23 ` Lance Yang
0 siblings, 0 replies; 20+ messages in thread
From: Lance Yang @ 2026-10-05 7:23 UTC (permalink / raw)
To: Nadav Amit
Cc: dave.hansen, luto, peterz, tglx, mingo, bp, x86, hpa, riel,
linux-kernel, qi.zheng, thomas.lendacky, kernel-team, linux-mm,
akpm, brendan.jackman, jannh, mhklinux, andrew.cooper3,
Manali.Shukla, mingo, stable, toshi.kani, david,
mikhail.v.gavrilov, pfalcato
On 2026/10/5 14:38, Nadav Amit wrote:
>
>
>>
>>
>> On 5 Oct 2026, at 8:23, Lance Yang <lance.yang@linux.dev> wrote:
>>
>> pud_free_pmd_page() uses a single-address invalidation to flush the
>> paging-structure caches before freeing the page tables. With AMD TCE
>> enabled, this only invalidates upper-level entries associated with the
>> target address. Cached PMD entries for other addresses in the PUD range can
>> still reference the PTE pages being freed.
>>
>> The AMD manual quoted in the commit enabling TCE says these instructions
>> remove
>>
>> "only those upper-level entries that lead to the target PTE in the page
>> table hierarchy, leaving unrelated upper-level entries intact."
>>
>> Even with all PTEs cleared, speculative page walks can cache present PMD
>> entries after the earlier TLB purge.
>>
>> Use a full TLB flush before freeing the page tables on CPUs with TCE. Keep
>> the single-address invalidation otherwise.
>>
>> Fixes: 440a65b7d25f ("x86/mm: Enable AMD translation cache extensions")
>> Cc: stable@vger.kernel.org
>> Signed-off-by: Lance Yang <lance.yang@linux.dev>
>> ---
>> arch/x86/mm/pgtable.c | 11 ++++++++++-
>> 1 file changed, 10 insertions(+), 1 deletion(-)
>>
>> diff --git a/arch/x86/mm/pgtable.c b/arch/x86/mm/pgtable.c
>> index 4a105f283cfb..6b7fa44f1bf6 100644
>> --- a/arch/x86/mm/pgtable.c
>> +++ b/arch/x86/mm/pgtable.c
>> @@ -727,7 +727,16 @@ int pud_free_pmd_page(pud_t *pud, unsigned long addr)
>> * via normal page walks. Make them unreachable
>> * in cached mid-level walks too:
>> */
>> - flush_tlb_kernel_range(addr, addr + PAGE_SIZE-1);
>> + if (boot_cpu_has(X86_FEATURE_TCE)) {
>> + /*
>> + * With TCE enabled, a single-address flush does not invalidate
>> + * cached PMD entries for the rest of the PUD range.
>> + */
>> + flush_tlb_all();
>> + } else {
>> + /* INVLPG to clear all paging-structure caches */
>> + flush_tlb_kernel_range(addr, addr + PAGE_SIZE-1);
>> + }
>>
>
>
> It might be cleaner to replace flush_tlb_all() with:
>
> flush_tlb_kernel_range(addr, addr + PUD_SIZE - 1);
Looks much cleaner, Thanks!
> While the flush-ceiling would usually end up doing a full flush, the
> code would be easier to follow (the very least). Maybe adding stride
> to kernel TLB range flushing would make sense in the future.
Ack.
^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [PATCH 1/1] x86/mm: fix incomplete page-table invalidation with TCE
2026-10-05 5:23 [PATCH 1/1] x86/mm: fix incomplete page-table invalidation with TCE Lance Yang
` (2 preceding siblings ...)
2026-10-05 6:38 ` Nadav Amit
@ 2026-10-05 15:15 ` Rik van Riel
2026-10-05 15:30 ` Lance Yang
3 siblings, 1 reply; 20+ messages in thread
From: Rik van Riel @ 2026-10-05 15:15 UTC (permalink / raw)
To: Lance Yang, dave.hansen
Cc: luto, peterz, tglx, mingo, bp, x86, hpa, linux-kernel, qi.zheng,
nadav.amit, thomas.lendacky, kernel-team, linux-mm, akpm,
brendan.jackman, jannh, mhklinux, andrew.cooper3, Manali.Shukla,
mingo, stable, toshi.kani, david, mikhail.v.gavrilov, pfalcato
On Mon, 2026-10-05 at 13:23 +0800, Lance Yang wrote:
> pud_free_pmd_page() uses a single-address invalidation to flush the
> paging-structure caches before freeing the page tables. With AMD TCE
> enabled, this only invalidates upper-level entries associated with
> the
> target address. Cached PMD entries for other addresses in the PUD
> range can
> still reference the PTE pages being freed.
>
> The AMD manual quoted in the commit enabling TCE says these
> instructions
> remove
>
> "only those upper-level entries that lead to the target PTE in the
> page
> table hierarchy, leaving unrelated upper-level entries intact."
>
> Even with all PTEs cleared, speculative page walks can cache present
> PMD
> entries after the earlier TLB purge.
The comment above the function says it all. The TLB
range should already have been cleared by the time
pud_free_pmd_page() gets called:
/**
* pud_free_pmd_page - Clear PUD entry and free PMD page
* @pud: Pointer to a PUD
* @addr: Virtual address associated with PUD
*
* Context: The PUD range has been unmapped and TLB purged.
* Return: 1 if clearing the entry succeeded. 0 otherwise.
*
* NOTE: Callers must allow a single page allocation.
*/
int pud_free_pmd_page(pud_t *pud, unsigned long addr)
{
The PMD could have been (speculatively) loaded by the
CPU after the PTEs were freed, so that one PMD mapping
needs to be flushed here, but there should not be
anything else left to flush.
The code looks odd, but it's a good idea to always
ask your AI to draw up a full chain of events for
a bug to trigger, going all the way back to something
calling the mm from the outside.
--
All Rights Reversed.
^ permalink raw reply [flat|nested] 20+ messages in thread* Re: [PATCH 1/1] x86/mm: fix incomplete page-table invalidation with TCE
2026-10-05 15:15 ` Rik van Riel
@ 2026-10-05 15:30 ` Lance Yang
0 siblings, 0 replies; 20+ messages in thread
From: Lance Yang @ 2026-10-05 15:30 UTC (permalink / raw)
To: Rik van Riel, dave.hansen
Cc: luto, peterz, tglx, mingo, bp, x86, hpa, linux-kernel, qi.zheng,
nadav.amit, thomas.lendacky, kernel-team, linux-mm, akpm,
brendan.jackman, jannh, mhklinux, andrew.cooper3, Manali.Shukla,
mingo, stable, toshi.kani, david, mikhail.v.gavrilov, pfalcato
On 2026/10/5 23:15, Rik van Riel wrote:
> On Mon, 2026-10-05 at 13:23 +0800, Lance Yang wrote:
>> pud_free_pmd_page() uses a single-address invalidation to flush the
>> paging-structure caches before freeing the page tables. With AMD TCE
>> enabled, this only invalidates upper-level entries associated with
>> the
>> target address. Cached PMD entries for other addresses in the PUD
>> range can
>> still reference the PTE pages being freed.
>>
>> The AMD manual quoted in the commit enabling TCE says these
>> instructions
>> remove
>>
>> "only those upper-level entries that lead to the target PTE in the
>> page
>> table hierarchy, leaving unrelated upper-level entries intact."
>>
>> Even with all PTEs cleared, speculative page walks can cache present
>> PMD
>> entries after the earlier TLB purge.
>
> The comment above the function says it all. The TLB
> range should already have been cleared by the time
> pud_free_pmd_page() gets called:
>
> /**
> * pud_free_pmd_page - Clear PUD entry and free PMD page
> * @pud: Pointer to a PUD
> * @addr: Virtual address associated with PUD
> *
> * Context: The PUD range has been unmapped and TLB purged.
> * Return: 1 if clearing the entry succeeded. 0 otherwise.
> *
> * NOTE: Callers must allow a single page allocation.
> */
> int pud_free_pmd_page(pud_t *pud, unsigned long addr)
> {
>
> The PMD could have been (speculatively) loaded by the
> CPU after the PTEs were freed, so that one PMD mapping
> needs to be flushed here, but there should not be
> anything else left to flush.
>
> The code looks odd, but it's a good idea to always
> ask your AI to draw up a full chain of events for
> a bug to trigger, going all the way back to something
> calling the mm from the outside.
Thanks for looking into this! I'll take another look
and trace it through.
^ permalink raw reply [flat|nested] 20+ messages in thread