mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: "Chuyi Zhou" <zhouchuyi@bytedance.com>
To: "Borislav Petkov" <bp@alien8.de>
Cc: <tglx@kernel.org>, <mingo@redhat.com>, <luto@kernel.org>,
	 <peterz@infradead.org>, <paulmck@kernel.org>,
	<muchun.song@linux.dev>,  <dave.hansen@linux.intel.com>,
	<pbonzini@redhat.com>,  <bigeasy@linutronix.de>,
	<clrkwllms@kernel.org>, <rostedt@goodmis.org>,
	 <nadav.amit@gmail.com>, <vkuznets@redhat.com>,
	 <linux-kernel@vger.kernel.org>
Subject: Re: [PATCH v10 00/14] Allow preemption during IPI completion waiting to improve real-time performance
Date: Fri, 24 Jul 2026 11:08:56 +0800	[thread overview]
Message-ID: <232b30b6-f555-4ae2-982d-b59172fcbb35@bytedance.com> (raw)
In-Reply-To: <20260723192657.GBamJrAcb_TCH5Ubog@fat_crate.local>

On 2026-07-24 3:26 a.m., Borislav Petkov wrote:
> On Thu, Jul 09, 2026 at 08:29:19PM +0800, Chuyi Zhou wrote:
>> In our production environments, latency-sensitive workloads (DPDK) are
>> configured with the highest priority to preempt lower-priority tasks at any
>> time. We discovered that DPDK's wake-up latency is primarily caused by the
>> current CPU having preemption disabled. Therefore, we collected the maximum
>> preemption disabled events within every 30-second interval and then
>> calculated the P50/P99 of these max preemption disabled events:
> 
> Out of curiosity, can you run your workloads on AMD Zen3 and newer which have
> TLBI support (this does away with the TLB flush IPIs). Do you see any
> improvement there?
> 
> Thx.
> 

Hi Boris,

Thanks for the suggestion.

If I get access to suitable AMD Zen 3 or newer production machines, I
would be happy to evaluate this series there.

On a system where X86_FEATURE_INVLPGB is enabled, I would expect the
direct benefit on the TLB flush paths to be smaller. For an mm which
already has a global ASID, flush_tlb_mm_range() uses
broadcast_tlb_flush(). arch_tlbbatch_flush() also uses
invlpgb_flush_all_nonglobals() directly for pending unmaps. Neither path
waits for remote TLB flush IPIs.

Those paths still have to remain on the CPU which issued INVLPGB until
TLBSYNC completes. Consequently, this series does not make the
INVLPGB/TLBSYNC section preemptible.

The benefit is not necessarily zero, though. A global ASID is assigned
opportunistically, currently for an mm active on at least four CPUs.
Before that assignment, flush_tlb_mm_range() can still fall back to the
IPI path. The same applies when INVLPGB is unavailable or disabled.

There is also an architecture-independent effect. The generic
smp_call_function*() changes keep the caller pinned only while selecting
the target CPUs, preparing the per-CPU state, queueing the callbacks and
sending the IPIs. The synchronous completion wait can then be 
preemptible when the caller's context otherwise permits it.

In particular, smp_call_function_many() and its wrappers no longer rely
on an outer preempt_disable() merely to protect the call-function
machinery. This benefits non-TLB synchronous IPI users as well. It also 
provides a basis for auditing other call-function IPI users and removing 
outer preemption-disabled regions which only existed because of the old
requirement. Callers which protect their own per-CPU state would, of
course, still need their local protection.

The data in the cover letter came from real production workloads. The
long-tail latency depends on the workload mix, interrupt-disabled
sections on remote CPUs and the number of target CPUs, so a short local
or synthetic test may not reproduce it accurately.

A more meaningful experiment would be a controlled, long-running
comparison between the baseline and patched kernels on comparable
production AMD Zen 3 or newer machines. We could then collect the same
30-second maximum preemption-disabled distributions and DPDK wake-up
latency over a sufficiently long period. That would also help separate
the remaining INVLPGB/TLBSYNC latency from latency caused by other
synchronous IPI users.

If I have the opportunity to perform such a production comparison, I
will report the results.

Thanks.

  reply	other threads:[~2026-07-24  3:09 UTC|newest]

Thread overview: 35+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-09 12:29 Chuyi Zhou
2026-07-09 12:29 ` [PATCH v10 01/14] smp: Disable preemption explicitly in __csd_lock_wait() Chuyi Zhou
2026-07-23 10:21   ` [tip: smp/core] " tip-bot2 for Chuyi Zhou
2026-07-09 12:29 ` [PATCH v10 02/14] smp: Enable preemption early in smp_call_function_single() Chuyi Zhou
2026-07-23 10:21   ` [tip: smp/core] " tip-bot2 for Chuyi Zhou
2026-07-09 12:29 ` [PATCH v10 03/14] smp: Refactor remote CPU selection in smp_call_function_any() Chuyi Zhou
2026-07-23 10:21   ` [tip: smp/core] " tip-bot2 for Chuyi Zhou
2026-07-09 12:29 ` [PATCH v10 04/14] smp: Use task-local IPI cpumask in smp_call_function_many_cond() Chuyi Zhou
2026-07-23 10:21   ` [tip: smp/core] " tip-bot2 for Chuyi Zhou
2026-07-09 12:29 ` [PATCH v10 05/14] smp: Alloc percpu csd data in smpcfd_prepare_cpu() only once Chuyi Zhou
2026-07-23 10:21   ` [tip: smp/core] " tip-bot2 for Chuyi Zhou
2026-07-09 12:29 ` [PATCH v10 06/14] smp: Enable preemption early in smp_call_function_many_cond() Chuyi Zhou
2026-07-23 10:21   ` [tip: smp/core] " tip-bot2 for Chuyi Zhou
2026-07-09 12:29 ` [PATCH v10 07/14] smp: Remove preempt_disable() from smp_call_function() Chuyi Zhou
2026-07-23 10:21   ` [tip: smp/core] " tip-bot2 for Chuyi Zhou
2026-07-09 12:29 ` [PATCH v10 08/14] smp: Remove preempt_disable() from on_each_cpu_cond_mask() Chuyi Zhou
2026-07-23 10:21   ` [tip: smp/core] " tip-bot2 for Chuyi Zhou
2026-07-09 12:29 ` [PATCH v10 09/14] scftorture: Remove preempt_disable() in scftorture_invoke_one() Chuyi Zhou
2026-07-23 10:21   ` [tip: smp/core] " tip-bot2 for Chuyi Zhou
2026-07-09 12:29 ` [PATCH v10 10/14] x86/mm: Factor out flush_tlb_info initialization Chuyi Zhou
2026-07-23 10:32   ` [tip: x86/mm] " tip-bot2 for Chuyi Zhou
2026-07-09 12:29 ` [PATCH v10 11/14] x86/mm: Cap flush_tlb_info alignment at 64 bytes Chuyi Zhou
2026-07-23 10:32   ` [tip: x86/mm] " tip-bot2 for Chuyi Zhou
2026-07-09 12:29 ` [PATCH v10 12/14] x86/mm: Move flush_tlb_info back to the stack Chuyi Zhou
2026-07-23 10:32   ` [tip: x86/mm] " tip-bot2 for Chuyi Zhou
2026-07-09 12:29 ` [PATCH v10 13/14] x86/kvm: Disable preemption in kvm_flush_tlb_multi() Chuyi Zhou
2026-07-23 10:32   ` [tip: x86/mm] " tip-bot2 for Chuyi Zhou
2026-07-09 12:29 ` [PATCH v10 14/14] x86/mm: Re-enable preemption before flush_tlb_multi() Chuyi Zhou
2026-07-23 10:32   ` [tip: x86/mm] " tip-bot2 for Chuyi Zhou
2026-07-16 14:14 ` [PATCH v10 00/14] Allow preemption during IPI completion waiting to improve real-time performance Chuyi Zhou
2026-07-16 21:18   ` Thomas Gleixner
2026-07-23 19:26 ` Borislav Petkov
2026-07-24  3:08   ` Chuyi Zhou [this message]
2026-08-11 21:22     ` Borislav Petkov
2026-08-12  3:56   ` Chuyi Zhou

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=232b30b6-f555-4ae2-982d-b59172fcbb35@bytedance.com \
    --to=zhouchuyi@bytedance.com \
    --cc=bigeasy@linutronix.de \
    --cc=bp@alien8.de \
    --cc=clrkwllms@kernel.org \
    --cc=dave.hansen@linux.intel.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=luto@kernel.org \
    --cc=mingo@redhat.com \
    --cc=muchun.song@linux.dev \
    --cc=nadav.amit@gmail.com \
    --cc=paulmck@kernel.org \
    --cc=pbonzini@redhat.com \
    --cc=peterz@infradead.org \
    --cc=rostedt@goodmis.org \
    --cc=tglx@kernel.org \
    --cc=vkuznets@redhat.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®