* [PATCH v2 1/5] x86/mm: Account for remote kernel TLB flush requests
2026-10-05 5:57 [PATCH v2 0/5] x86/mm: Allow preemption while waiting for kernel TLB flushes Chuyi Zhou
@ 2026-10-05 5:57 ` Chuyi Zhou
2026-10-05 5:57 ` [PATCH v2 2/5] x86/mm: Share the full TLB flush dispatch Chuyi Zhou
` (4 subsequent siblings)
5 siblings, 0 replies; 8+ messages in thread
From: Chuyi Zhou @ 2026-10-05 5:57 UTC (permalink / raw)
To: tglx, mingo, luto, peterz, paulmck, muchun.song, bp, dave.hansen,
pbonzini, bigeasy, clrkwllms, rostedt, nadav.amit
Cc: linux-kernel, Chuyi Zhou
NR_TLB_REMOTE_FLUSH counts requests to invalidate TLB entries on other
CPUs. flush_tlb_all() and native_flush_tlb_multi() account for these
requests, but flush_tlb_kernel_range() bypasses both functions and does
not update the counter. Kernel range flushes are therefore missing from
the sender-side statistics, including ranges promoted to a full flush.
Count each request in kernel_tlb_flush_all() and
kernel_tlb_flush_range(). Account for both IPI and INVLPGB backends, as
flush_tlb_all() already does, and count once per request regardless of
the number of target CPUs or invalidation instructions.
Signed-off-by: Chuyi Zhou <zhouchuyi@bytedance.com>
---
arch/x86/mm/tlb.c | 4 ++++
1 file changed, 4 insertions(+)
diff --git a/arch/x86/mm/tlb.c b/arch/x86/mm/tlb.c
index 0c4da320831c..da7c408f921c 100644
--- a/arch/x86/mm/tlb.c
+++ b/arch/x86/mm/tlb.c
@@ -1488,6 +1488,8 @@ static void do_kernel_range_flush(void *info)
static void kernel_tlb_flush_all(struct flush_tlb_info *info)
{
+ count_vm_tlb_event(NR_TLB_REMOTE_FLUSH);
+
if (cpu_feature_enabled(X86_FEATURE_INVLPGB))
invlpgb_flush_all();
else
@@ -1496,6 +1498,8 @@ static void kernel_tlb_flush_all(struct flush_tlb_info *info)
static void kernel_tlb_flush_range(struct flush_tlb_info *info)
{
+ count_vm_tlb_event(NR_TLB_REMOTE_FLUSH);
+
if (cpu_feature_enabled(X86_FEATURE_INVLPGB))
invlpgb_kernel_range_flush(info);
else
--
2.20.1
^ permalink raw reply [flat|nested] 8+ messages in thread* [PATCH v2 2/5] x86/mm: Share the full TLB flush dispatch
2026-10-05 5:57 [PATCH v2 0/5] x86/mm: Allow preemption while waiting for kernel TLB flushes Chuyi Zhou
2026-10-05 5:57 ` [PATCH v2 1/5] x86/mm: Account for remote kernel TLB flush requests Chuyi Zhou
@ 2026-10-05 5:57 ` Chuyi Zhou
2026-10-05 5:57 ` [PATCH v2 3/5] x86/mm: Extract the TLB range flush threshold check Chuyi Zhou
` (3 subsequent siblings)
5 siblings, 0 replies; 8+ messages in thread
From: Chuyi Zhou @ 2026-10-05 5:57 UTC (permalink / raw)
To: tglx, mingo, luto, peterz, paulmck, muchun.song, bp, dave.hansen,
pbonzini, bigeasy, clrkwllms, rostedt, nadav.amit
Cc: linux-kernel, Chuyi Zhou
flush_tlb_all() and kernel_tlb_flush_all() select the same INVLPGB or
IPI backend and account for the same sender-side event. Keeping two
implementations duplicates the backend selection and accounting.
Use flush_tlb_all() for kernel ranges promoted to a full flush and
remove kernel_tlb_flush_all(). Share the existing backend dispatch and
event accounting so each request is counted once.
Signed-off-by: Chuyi Zhou <zhouchuyi@bytedance.com>
---
arch/x86/mm/tlb.c | 12 +-----------
1 file changed, 1 insertion(+), 11 deletions(-)
diff --git a/arch/x86/mm/tlb.c b/arch/x86/mm/tlb.c
index da7c408f921c..5a1b96c61050 100644
--- a/arch/x86/mm/tlb.c
+++ b/arch/x86/mm/tlb.c
@@ -1486,16 +1486,6 @@ static void do_kernel_range_flush(void *info)
flush_tlb_one_kernel(addr);
}
-static void kernel_tlb_flush_all(struct flush_tlb_info *info)
-{
- count_vm_tlb_event(NR_TLB_REMOTE_FLUSH);
-
- if (cpu_feature_enabled(X86_FEATURE_INVLPGB))
- invlpgb_flush_all();
- else
- on_each_cpu(do_flush_tlb_all, NULL, 1);
-}
-
static void kernel_tlb_flush_range(struct flush_tlb_info *info)
{
count_vm_tlb_event(NR_TLB_REMOTE_FLUSH);
@@ -1515,7 +1505,7 @@ void flush_tlb_kernel_range(unsigned long start, unsigned long end)
TLB_GENERATION_INVALID);
if (info.end == TLB_FLUSH_ALL)
- kernel_tlb_flush_all(&info);
+ flush_tlb_all();
else
kernel_tlb_flush_range(&info);
}
--
2.20.1
^ permalink raw reply [flat|nested] 8+ messages in thread* [PATCH v2 3/5] x86/mm: Extract the TLB range flush threshold check
2026-10-05 5:57 [PATCH v2 0/5] x86/mm: Allow preemption while waiting for kernel TLB flushes Chuyi Zhou
2026-10-05 5:57 ` [PATCH v2 1/5] x86/mm: Account for remote kernel TLB flush requests Chuyi Zhou
2026-10-05 5:57 ` [PATCH v2 2/5] x86/mm: Share the full TLB flush dispatch Chuyi Zhou
@ 2026-10-05 5:57 ` Chuyi Zhou
2026-10-05 5:57 ` [PATCH v2 4/5] x86/mm: Decouple kernel TLB flushes from flush_tlb_info Chuyi Zhou
` (2 subsequent siblings)
5 siblings, 0 replies; 8+ messages in thread
From: Chuyi Zhou @ 2026-10-05 5:57 UTC (permalink / raw)
To: tglx, mingo, luto, peterz, paulmck, muchun.song, bp, dave.hansen,
pbonzini, bigeasy, clrkwllms, rostedt, nadav.amit
Cc: linux-kernel, Chuyi Zhou
The decision to replace a range flush with a full flush is embedded in
init_flush_tlb_info(). Both mm and kernel flushes use this policy, but
the kernel path only needs the range and the decision, without the
mm-specific descriptor initialization.
Extract the threshold predicate into tlb_range_exceeds_ceiling()
and use it in init_flush_tlb_info(). Preserve the unsigned range
arithmetic, right-shift rounding and strict greater-than comparison.
Descriptor initialization and both flush paths remain unchanged.
Signed-off-by: Chuyi Zhou <zhouchuyi@bytedance.com>
---
arch/x86/mm/tlb.c | 8 +++++++-
1 file changed, 7 insertions(+), 1 deletion(-)
diff --git a/arch/x86/mm/tlb.c b/arch/x86/mm/tlb.c
index 5a1b96c61050..0d03011c6825 100644
--- a/arch/x86/mm/tlb.c
+++ b/arch/x86/mm/tlb.c
@@ -1373,6 +1373,12 @@ void flush_tlb_multi(const struct cpumask *cpumask,
*/
unsigned long tlb_single_page_flush_ceiling __read_mostly = 33;
+static bool tlb_range_exceeds_ceiling(unsigned long start, unsigned long end,
+ unsigned int stride_shift)
+{
+ return ((end - start) >> stride_shift) > tlb_single_page_flush_ceiling;
+}
+
static void init_flush_tlb_info(struct flush_tlb_info *info,
struct mm_struct *mm,
unsigned long start, unsigned long end,
@@ -1383,7 +1389,7 @@ static void init_flush_tlb_info(struct flush_tlb_info *info,
* If the number of flushes is so large that a full flush
* would be faster, do a full flush.
*/
- if ((end - start) >> stride_shift > tlb_single_page_flush_ceiling) {
+ if (tlb_range_exceeds_ceiling(start, end, stride_shift)) {
start = 0;
end = TLB_FLUSH_ALL;
}
--
2.20.1
^ permalink raw reply [flat|nested] 8+ messages in thread* [PATCH v2 4/5] x86/mm: Decouple kernel TLB flushes from flush_tlb_info
2026-10-05 5:57 [PATCH v2 0/5] x86/mm: Allow preemption while waiting for kernel TLB flushes Chuyi Zhou
` (2 preceding siblings ...)
2026-10-05 5:57 ` [PATCH v2 3/5] x86/mm: Extract the TLB range flush threshold check Chuyi Zhou
@ 2026-10-05 5:57 ` Chuyi Zhou
2026-10-05 5:57 ` [PATCH v2 5/5] x86/mm: Re-enable preemption before waiting for kernel TLB flushes Chuyi Zhou
2026-10-05 6:50 ` [PATCH v2 0/5] x86/mm: Allow preemption while " Nadav Amit
5 siblings, 0 replies; 8+ messages in thread
From: Chuyi Zhou @ 2026-10-05 5:57 UTC (permalink / raw)
To: tglx, mingo, luto, peterz, paulmck, muchun.song, bp, dave.hansen,
pbonzini, bigeasy, clrkwllms, rostedt, nadav.amit
Cc: linux-kernel, Chuyi Zhou
Kernel range flushes only need the start and end addresses, but reuse
flush_tlb_info and its initialization of mm state, TLB generations and
the initiating CPU. None of those fields is consumed by the kernel
flush callbacks. In particular, initializing initiating_cpu imposes a
CPU-pinning requirement on a path that does not use it.
Select the full or range flush directly in flush_tlb_kernel_range()
using the shared threshold predicate. The interface accepts actual
address ranges; callers requesting an unconditional full flush can use
flush_tlb_all(). Keep init_flush_tlb_info() and its smp_processor_id()
check for the mm paths.
Pass start and end directly through the kernel range helpers. Package
them in a private kernel_tlb_range only for the IPI callback, retaining
the existing payload alignment. The synchronous on_each_cpu() call
keeps the stack descriptor alive until all callbacks have completed.
Retain the outer preemption guard in flush_tlb_kernel_range() so the
descriptor refactoring does not change the preemption behavior.
Signed-off-by: Chuyi Zhou <zhouchuyi@bytedance.com>
Link: https://lore.kernel.org/20260522104818.CbT5fyN8@linutronix.de/
---
arch/x86/mm/tlb.c | 40 ++++++++++++++++++++++++----------------
1 file changed, 24 insertions(+), 16 deletions(-)
diff --git a/arch/x86/mm/tlb.c b/arch/x86/mm/tlb.c
index 0d03011c6825..b55495765da8 100644
--- a/arch/x86/mm/tlb.c
+++ b/arch/x86/mm/tlb.c
@@ -1464,12 +1464,12 @@ void flush_tlb_all(void)
}
/* Flush an arbitrarily large range of memory with INVLPGB. */
-static void invlpgb_kernel_range_flush(struct flush_tlb_info *info)
+static void invlpgb_kernel_range_flush(unsigned long start, unsigned long end)
{
unsigned long addr, nr;
- for (addr = info->start; addr < info->end; addr += nr << PAGE_SHIFT) {
- nr = (info->end - addr) >> PAGE_SHIFT;
+ for (addr = start; addr < end; addr += nr << PAGE_SHIFT) {
+ nr = (end - addr) >> PAGE_SHIFT;
/*
* INVLPGB has a limit on the size of ranges it can
@@ -1482,38 +1482,46 @@ static void invlpgb_kernel_range_flush(struct flush_tlb_info *info)
__tlbsync();
}
+/* Preserve the alignment of the IPI payload shared with remote CPUs. */
+struct kernel_tlb_range {
+ unsigned long start;
+ unsigned long end;
+} __aligned(FLUSH_TLB_INFO_ALIGN);
+
static void do_kernel_range_flush(void *info)
{
- struct flush_tlb_info *f = info;
+ const struct kernel_tlb_range *range = info;
unsigned long addr;
/* flush range by one by one 'invlpg' */
- for (addr = f->start; addr < f->end; addr += PAGE_SIZE)
+ for (addr = range->start; addr < range->end; addr += PAGE_SIZE)
flush_tlb_one_kernel(addr);
}
-static void kernel_tlb_flush_range(struct flush_tlb_info *info)
+static void kernel_tlb_flush_range(unsigned long start, unsigned long end)
{
count_vm_tlb_event(NR_TLB_REMOTE_FLUSH);
- if (cpu_feature_enabled(X86_FEATURE_INVLPGB))
- invlpgb_kernel_range_flush(info);
- else
- on_each_cpu(do_kernel_range_flush, info, 1);
+ if (cpu_feature_enabled(X86_FEATURE_INVLPGB)) {
+ invlpgb_kernel_range_flush(start, end);
+ } else {
+ struct kernel_tlb_range range = {
+ .start = start,
+ .end = end,
+ };
+
+ on_each_cpu(do_kernel_range_flush, &range, 1);
+ }
}
void flush_tlb_kernel_range(unsigned long start, unsigned long end)
{
- struct flush_tlb_info info;
-
guard(preempt)();
- init_flush_tlb_info(&info, NULL, start, end, PAGE_SHIFT, false,
- TLB_GENERATION_INVALID);
- if (info.end == TLB_FLUSH_ALL)
+ if (tlb_range_exceeds_ceiling(start, end, PAGE_SHIFT))
flush_tlb_all();
else
- kernel_tlb_flush_range(&info);
+ kernel_tlb_flush_range(start, end);
}
/*
--
2.20.1
^ permalink raw reply [flat|nested] 8+ messages in thread* [PATCH v2 5/5] x86/mm: Re-enable preemption before waiting for kernel TLB flushes
2026-10-05 5:57 [PATCH v2 0/5] x86/mm: Allow preemption while waiting for kernel TLB flushes Chuyi Zhou
` (3 preceding siblings ...)
2026-10-05 5:57 ` [PATCH v2 4/5] x86/mm: Decouple kernel TLB flushes from flush_tlb_info Chuyi Zhou
@ 2026-10-05 5:57 ` Chuyi Zhou
2026-10-05 6:50 ` [PATCH v2 0/5] x86/mm: Allow preemption while " Nadav Amit
5 siblings, 0 replies; 8+ messages in thread
From: Chuyi Zhou @ 2026-10-05 5:57 UTC (permalink / raw)
To: tglx, mingo, luto, peterz, paulmck, muchun.song, bp, dave.hansen,
pbonzini, bigeasy, clrkwllms, rostedt, nadav.amit
Cc: linux-kernel, Chuyi Zhou
flush_tlb_kernel_range() uses on_each_cpu() to synchronously flush
kernel mappings on all online CPUs when using the IPI backend.
The synchronous wait can become longer as the number of online CPUs
grows, and a remote CPU with interrupts disabled can further delay
completion. Preemption remains disabled throughout this wait, delaying
higher-priority tasks on the initiating CPU. The outer guard prevents
the SMP layer from making its final completion wait preemptible.
The kernel range descriptor is private stack storage and remains valid
until on_each_cpu() returns. The SMP layer protects CPU selection, IPI
queueing and local callback execution, so the kernel flush caller does
not need to remain pinned while waiting for remote completion.
INVLPGB requires separate protection: TLBSYNC only waits for operations
issued on the same CPU. Disable preemption inside
invlpgb_kernel_range_flush() across the broadcast loop and TLBSYNC.
The full-flush helper invlpgb_flush_all() already provides this
protection.
Remove the outer preemption guard from flush_tlb_kernel_range() so the
IPI completion wait can be preempted when the caller's context allows
it. Keep both flush backends synchronous.
Signed-off-by: Chuyi Zhou <zhouchuyi@bytedance.com>
---
arch/x86/mm/tlb.c | 5 +++--
1 file changed, 3 insertions(+), 2 deletions(-)
diff --git a/arch/x86/mm/tlb.c b/arch/x86/mm/tlb.c
index b55495765da8..8bfcd66b639d 100644
--- a/arch/x86/mm/tlb.c
+++ b/arch/x86/mm/tlb.c
@@ -1468,6 +1468,9 @@ static void invlpgb_kernel_range_flush(unsigned long start, unsigned long end)
{
unsigned long addr, nr;
+ /* Keep the INVLPGB operations and TLBSYNC on the same CPU. */
+ guard(preempt)();
+
for (addr = start; addr < end; addr += nr << PAGE_SHIFT) {
nr = (end - addr) >> PAGE_SHIFT;
@@ -1516,8 +1519,6 @@ static void kernel_tlb_flush_range(unsigned long start, unsigned long end)
void flush_tlb_kernel_range(unsigned long start, unsigned long end)
{
- guard(preempt)();
-
if (tlb_range_exceeds_ceiling(start, end, PAGE_SHIFT))
flush_tlb_all();
else
--
2.20.1
^ permalink raw reply [flat|nested] 8+ messages in thread* Re: [PATCH v2 0/5] x86/mm: Allow preemption while waiting for kernel TLB flushes
2026-10-05 5:57 [PATCH v2 0/5] x86/mm: Allow preemption while waiting for kernel TLB flushes Chuyi Zhou
` (4 preceding siblings ...)
2026-10-05 5:57 ` [PATCH v2 5/5] x86/mm: Re-enable preemption before waiting for kernel TLB flushes Chuyi Zhou
@ 2026-10-05 6:50 ` Nadav Amit
2026-10-05 7:48 ` Chuyi Zhou
5 siblings, 1 reply; 8+ messages in thread
From: Nadav Amit @ 2026-10-05 6:50 UTC (permalink / raw)
To: Chuyi Zhou
Cc: tglx, mingo, luto, peterz, paulmck, muchun.song, bp, dave.hansen,
pbonzini, bigeasy, clrkwllms, rostedt, linux-kernel
> On 5 Oct 2026, at 8:57, Chuyi Zhou <zhouchuyi@bytedance.com> wrote:
>
> Changes in v2:
> - Patch 2: Reuse flush_tlb_all() for kernel ranges promoted to a full
> flush and remove kernel_tlb_flush_all() (Sebastian).
> - Patch 4: Drop the explicit TLB_FLUSH_ALL check on the end argument.
> Callers pass actual address ranges and can use flush_tlb_all() for
> unconditional full flushes (Sebastian).
Chuyi,
It all looks nice and clean, but I am not sure the end result is that
great. I think that instead of consolidating different TLB flush paths,
you break them further apart. Then you try to copy the logic from
userspace TLB flushed into the kernel code (TLB flush counting and such).
I think that perhaps a better path would be to further consolidate the
two instead of separating them. There is functionality that is missing
from kernel-space TLB flushes, and might be needed in the future.
For instance, you can see userspace TLB-flushing has a mechanism to
prevent TLB shootdown storm using TLB generations; and you see it
supports TLB-flushing stride. Now, the shootdown storm might be less
of an issue (for now?) but stride support is something you may want
eventually to support range flush with stride for stuff like [1].
Nadav
[1] https://lore.kernel.org/all/20261005052302.43042-1-lance.yang@linux.dev/
^ permalink raw reply [flat|nested] 8+ messages in thread* Re: [PATCH v2 0/5] x86/mm: Allow preemption while waiting for kernel TLB flushes
2026-10-05 6:50 ` [PATCH v2 0/5] x86/mm: Allow preemption while " Nadav Amit
@ 2026-10-05 7:48 ` Chuyi Zhou
0 siblings, 0 replies; 8+ messages in thread
From: Chuyi Zhou @ 2026-10-05 7:48 UTC (permalink / raw)
To: Nadav Amit
Cc: tglx, mingo, luto, peterz, paulmck, muchun.song, bp, dave.hansen,
pbonzini, bigeasy, clrkwllms, rostedt, linux-kernel
On 2026-10-05 2:50 p.m., Nadav Amit wrote:
>
>> On 5 Oct 2026, at 8:57, Chuyi Zhou <zhouchuyi@bytedance.com> wrote:
>>
>> Changes in v2:
>> - Patch 2: Reuse flush_tlb_all() for kernel ranges promoted to a full
>> flush and remove kernel_tlb_flush_all() (Sebastian).
>> - Patch 4: Drop the explicit TLB_FLUSH_ALL check on the end argument.
>> Callers pass actual address ranges and can use flush_tlb_all() for
>> unconditional full flushes (Sebastian).
>
> Chuyi,
>
> It all looks nice and clean, but I am not sure the end result is that
> great. I think that instead of consolidating different TLB flush paths,
> you break them further apart. Then you try to copy the logic from
> userspace TLB flushed into the kernel code (TLB flush counting and such).
>
> I think that perhaps a better path would be to further consolidate the
> two instead of separating them. There is functionality that is missing
> from kernel-space TLB flushes, and might be needed in the future.
>
> For instance, you can see userspace TLB-flushing has a mechanism to
> prevent TLB shootdown storm using TLB generations; and you see it
> supports TLB-flushing stride. Now, the shootdown storm might be less
> of an issue (for now?) but stride support is something you may want
> eventually to support range flush with stride for stuff like [1].
>
Hi Nadav,
The motivation for the descriptor split was that init_flush_tlb_info()
initializes initiating_cpu with smp_processor_id(). The kernel flush
callbacks do not use that field, but removing the outer preemption
guard would allow the initializer to run in a preemptible context.
We could retain flush_tlb_info and initialize the range fields directly
in the kernel path. On top of the first three patches, the following
change could replace patches 4 and 5:
--- a/arch/x86/mm/tlb.c
+++ b/arch/x86/mm/tlb.c
@@ -1466,10 +1466,13 @@
/* Flush an arbitrarily large range of memory with INVLPGB. */
static void invlpgb_kernel_range_flush(struct flush_tlb_info *info)
{
unsigned long addr, nr;
+ /* Keep the INVLPGB operations and TLBSYNC on the same CPU. */
+ guard(preempt)();
+
for (addr = info->start; addr < info->end; addr += nr << PAGE_SHIFT) {
nr = (info->end - addr) >> PAGE_SHIFT;
/*
* INVLPGB has a limit on the size of ranges it can
@@ -1502,17 +1505,16 @@
on_each_cpu(do_kernel_range_flush, info, 1);
}
void flush_tlb_kernel_range(unsigned long start, unsigned long end)
{
- struct flush_tlb_info info;
-
- guard(preempt)();
- init_flush_tlb_info(&info, NULL, start, end, PAGE_SHIFT, false,
- TLB_GENERATION_INVALID);
-
- if (info.end == TLB_FLUSH_ALL)
+ struct flush_tlb_info info = {
+ .start = start,
+ .end = end,
+ };
+
+ if (tlb_range_exceeds_ceiling(start, end, PAGE_SHIFT))
flush_tlb_all();
else
kernel_tlb_flush_range(&info);
}
The kernel helpers would continue to use flush_tlb_info, and the mm
paths would retain init_flush_tlb_info() and its smp_processor_id()
check. The synchronous on_each_cpu() call keeps the stack descriptor
valid until all callbacks complete.
This keeps the common descriptor and threshold policy, while allowing
preemption during the final IPI completion wait. It does not yet
provide the broader consolidation you suggested, but retains the
existing stride_shift field for future kernel stride support.
Would this smaller change be a better direction for the preemption
work, with further consolidation handled separately?
Thanks.
^ permalink raw reply [flat|nested] 8+ messages in thread