mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: "Chuyi Zhou" <zhouchuyi@bytedance.com>
To: "Nadav Amit" <nadav.amit@gmail.com>
Cc: <tglx@kernel.org>, <mingo@redhat.com>, <luto@kernel.org>,
	 <peterz@infradead.org>, <paulmck@kernel.org>,
	<muchun.song@linux.dev>,  <bp@alien8.de>,
	<dave.hansen@linux.intel.com>, <pbonzini@redhat.com>,
	 <bigeasy@linutronix.de>, <clrkwllms@kernel.org>,
	<rostedt@goodmis.org>,  <linux-kernel@vger.kernel.org>
Subject: Re: [PATCH v2 0/5] x86/mm: Allow preemption while waiting for kernel TLB flushes
Date: Mon, 5 Oct 2026 15:48:14 +0800	[thread overview]
Message-ID: <278e53d1-e0ca-4886-a762-ca98d1a21a69@bytedance.com> (raw)
In-Reply-To: <3520FAB5-83C3-4116-BF19-87D66048D993@gmail.com>

On 2026-10-05 2:50 p.m., Nadav Amit wrote:
> 
>> On 5 Oct 2026, at 8:57, Chuyi Zhou <zhouchuyi@bytedance.com> wrote:
>>
>> Changes in v2:
>>   - Patch 2: Reuse flush_tlb_all() for kernel ranges promoted to a full
>>     flush and remove kernel_tlb_flush_all() (Sebastian).
>>   - Patch 4: Drop the explicit TLB_FLUSH_ALL check on the end argument.
>>     Callers pass actual address ranges and can use flush_tlb_all() for
>>     unconditional full flushes (Sebastian).
> 
> Chuyi,
> 
> It all looks nice and clean, but I am not sure the end result is that
> great. I think that instead of consolidating different TLB flush paths,
> you break them further apart. Then you try to copy the logic from
> userspace TLB flushed into the kernel code (TLB flush counting and such).
> 
> I think that perhaps a better path would be to further consolidate the
> two instead of separating them. There is functionality that is missing
> from kernel-space TLB flushes, and might be needed in the future.
> 
> For instance, you can see userspace TLB-flushing has a mechanism to
> prevent TLB shootdown storm using TLB generations; and you see it
> supports TLB-flushing stride. Now, the shootdown storm might be less
> of an issue (for now?) but stride support is something you may want
> eventually to support range flush with stride for stuff like [1].
> 

Hi Nadav,

The motivation for the descriptor split was that init_flush_tlb_info()
initializes initiating_cpu with smp_processor_id(). The kernel flush
callbacks do not use that field, but removing the outer preemption
guard would allow the initializer to run in a preemptible context.

We could retain flush_tlb_info and initialize the range fields directly
in the kernel path. On top of the first three patches, the following
change could replace patches 4 and 5:

--- a/arch/x86/mm/tlb.c
+++ b/arch/x86/mm/tlb.c
@@ -1466,10 +1466,13 @@
  /* Flush an arbitrarily large range of memory with INVLPGB. */
  static void invlpgb_kernel_range_flush(struct flush_tlb_info *info)
  {
  	unsigned long addr, nr;

+	/* Keep the INVLPGB operations and TLBSYNC on the same CPU. */
+	guard(preempt)();
+
  	for (addr = info->start; addr < info->end; addr += nr << PAGE_SHIFT) {
  		nr = (info->end - addr) >> PAGE_SHIFT;

  		/*
  		 * INVLPGB has a limit on the size of ranges it can
@@ -1502,17 +1505,16 @@
  		on_each_cpu(do_kernel_range_flush, info, 1);
  }

  void flush_tlb_kernel_range(unsigned long start, unsigned long end)
  {
-	struct flush_tlb_info info;
-
-	guard(preempt)();
-	init_flush_tlb_info(&info, NULL, start, end, PAGE_SHIFT, false,
-			    TLB_GENERATION_INVALID);
-
-	if (info.end == TLB_FLUSH_ALL)
+	struct flush_tlb_info info = {
+		.start = start,
+		.end = end,
+	};
+
+	if (tlb_range_exceeds_ceiling(start, end, PAGE_SHIFT))
  		flush_tlb_all();
  	else
  		kernel_tlb_flush_range(&info);
  }


The kernel helpers would continue to use flush_tlb_info, and the mm
paths would retain init_flush_tlb_info() and its smp_processor_id()
check. The synchronous on_each_cpu() call keeps the stack descriptor
valid until all callbacks complete.

This keeps the common descriptor and threshold policy, while allowing
preemption during the final IPI completion wait. It does not yet
provide the broader consolidation you suggested, but retains the
existing stride_shift field for future kernel stride support.

Would this smaller change be a better direction for the preemption
work, with further consolidation handled separately?

Thanks.

  reply	other threads:[~2026-10-05  7:48 UTC|newest]

Thread overview: 9+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-05  5:57 Chuyi Zhou
2026-10-05  5:57 ` [PATCH v2 1/5] x86/mm: Account for remote kernel TLB flush requests Chuyi Zhou
2026-10-05  5:57 ` [PATCH v2 2/5] x86/mm: Share the full TLB flush dispatch Chuyi Zhou
2026-10-05  5:57 ` [PATCH v2 3/5] x86/mm: Extract the TLB range flush threshold check Chuyi Zhou
2026-10-05  5:57 ` [PATCH v2 4/5] x86/mm: Decouple kernel TLB flushes from flush_tlb_info Chuyi Zhou
2026-10-05  5:57 ` [PATCH v2 5/5] x86/mm: Re-enable preemption before waiting for kernel TLB flushes Chuyi Zhou
2026-10-05  6:50 ` [PATCH v2 0/5] x86/mm: Allow preemption while " Nadav Amit
2026-10-05  7:48   ` Chuyi Zhou [this message]
2026-10-05 16:36     ` Nadav Amit

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=278e53d1-e0ca-4886-a762-ca98d1a21a69@bytedance.com \
    --to=zhouchuyi@bytedance.com \
    --cc=bigeasy@linutronix.de \
    --cc=bp@alien8.de \
    --cc=clrkwllms@kernel.org \
    --cc=dave.hansen@linux.intel.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=luto@kernel.org \
    --cc=mingo@redhat.com \
    --cc=muchun.song@linux.dev \
    --cc=nadav.amit@gmail.com \
    --cc=paulmck@kernel.org \
    --cc=pbonzini@redhat.com \
    --cc=peterz@infradead.org \
    --cc=rostedt@goodmis.org \
    --cc=tglx@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®