* [PATCH v3 0/2] alpha: complete the direct-mm shootdown fixes
@ 2026-10-09 21:19 Magnus Lindholm
2026-10-09 21:19 ` [PATCH v3 1/2] alpha: do not clear remote MMU contexts in migrate_flush_tlb_page() Magnus Lindholm
2026-10-09 21:19 ` [PATCH v3 2/2] alpha: do not skip TLB shootdown IPIs for kthread-borrowed mms Magnus Lindholm
0 siblings, 2 replies; 5+ messages in thread
From: Magnus Lindholm @ 2026-10-09 21:19 UTC (permalink / raw)
To: richard.henderson, mattst88, linux-kernel, linux-alpha; +Cc: linmag7
This is the second of two series. Apply all eight patches of "alpha: fix
stale TLB translations breaking copy-on-write and writeback" v4 first,
then these two patches. The base-commit and prerequisite-patch-id lines
below describe that complete ordering from for-next at 33465f6ab697.
The direct context-load fix formerly numbered v2 1/3 has moved to the
front of the stale-TLB series, where it is needed before the post-switch
hook can clear asn_lock. It is not duplicated here. This replaces the old
interleaved dependency with two series that can be applied in sequence.
Patch 1 removes the remote context-clearing loop after the migration
shootdown rendezvous. The callbacks have already updated their own slots;
clearing them remotely can hide a CPU still running the mm.
Patch 2 fixes the mm_users <= 1 shootdown shortcut. An explicit
kthread_use_mm() borrower takes mmgrab(), not mmget(), so mm_users does
not count it. Inspect remote context slots instead of clearing them, and
skip the IPIs only when no other CPU has a context. Together the two
patches remove runtime remote writes to the slots.
This revision addresses Matt Turner's review of v2:
- Add smp_mb() after context publication in __load_new_mm_context().
check_mmu_context() can publish after finish_lock_switch() has already
issued its barrier, so relying only on the scheduler and direct-switch
callers' barriers left that publisher unordered.
- Test current->mm in ipi_flush_tlb_mm() and ipi_flush_icache_page(),
matching the page-flush handler. A CPU holding only a lazy active_mm
then clears its slot rather than repeatedly publishing a new ASN.
An explicit borrower still has current->mm set and is flushed.
- Require the first series' need_new_asn change for newly allocated ASNs.
An IPI after finish_lock_switch() can zero such a slot while asn_lock
is set. Completing that handshake before returning to user space
prevents the shortcut from mistaking a live ASN for an inactive CPU.
- Move the direct-load patch to the first series and update the base,
prerequisite patch IDs, numbering and ordering explanations.
The scheduler publication remains ordered by finish_lock_switch(); the
new barrier in __load_new_mm_context() covers every caller of that helper.
The checking CPU orders its slot reads after its page-table changes with
smp_mb(). Stale slots may conservatively cause an extra round of IPIs,
after which a CPU not running the mm clears its own slot. The effect of
the new lazy-CPU handling on throughput has not been measured.
Testing
Cross-build validation for this revision covers arch/alpha/kernel/,
arch/alpha/mm/, kernel/sched/core.o and kernel/kthread.o with GCC 15.0.1:
ALPHA_GENERIC SMP and UP with COMPACTION=n, and SMP with COMPACTION=y,
RCU_STRICT_GRACE_PERIOD=y, PREEMPT_COUNT=y and NR_CPUS=4. These are compile
checks; all three passed. All ten patches pass checkpatch without warnings,
and sequential application reproduces the committed trees at both series
boundaries. No new boot or hardware runtime results are claimed.
The following are historical results reported for the previous versions,
not hardware results for this revision's barrier or handler changes.
On an ES40, EV68AL Tsunami, three CPUs and v7.2-rc6, the direct-load KUnit
test changed from 0/2 to 2/2 passing. Single-threaded fork throughput was
1360 forks/s with the old shortcut, 1050 with it removed, and 1360 with the
context-slot shortcut. Earlier regression runs passed glibc malloc-check
25/25 and four related tests 10/10 each, plus COW and writeback tests,
including continuous compaction with 359768 migrated folios.
Matt tested the direct-load change on an ES47 (EV7), v7.3-rc1, one CPU
online. His KUnit test passed 3/3, usercopy_kunit 4/4 and kunit_iov_iter
17/17, with a fork/COW check clean under compaction. Those runs did not
exercise remote-slot clearing or the shortcut's remote-CPU test. The
reviewed and tested direct-load patch now carries those tags in the first
series. Patch 1 here retains Matt's Reviewed-by; patch 2 needs review.
The callers are not limited to KUnit: vhost's kthread mode, USB gadget AIO
paths (including dummy_hcd setups) and vdpa_sim can borrow an mm too. Both
remaining patches retain Cc: stable, including the migration cleanup as
the shortcut's prerequisite.
The adjacent enter_lazy_tlb() page-table/ASN mismatch remains outside this
work. This series does not claim to fix every aspect of lazy TLB handling.
Thanks to Matt Turner for the review and EV7 testing.
v2: https://lore.kernel.org/linux-alpha/20261008195608.965266-1-linmag7@gmail.com/
v1: https://lore.kernel.org/linux-alpha/20260904162424.376504-1-linmag7@gmail.com/
Magnus Lindholm (2):
alpha: do not clear remote MMU contexts in migrate_flush_tlb_page()
alpha: do not skip TLB shootdown IPIs for kthread-borrowed mms
arch/alpha/include/asm/mmu_context.h | 2 +-
arch/alpha/kernel/smp.c | 52 ++++++++++++++--------------
arch/alpha/mm/fault.c | 4 ++-
arch/alpha/mm/tlbflush.c | 16 ---------
4 files changed, 30 insertions(+), 44 deletions(-)
base-commit: 33465f6ab697abceafd9a340067a544d551c4412
prerequisite-patch-id: d194ee1170a45793fa8533400dbc5599fc0e59ae
prerequisite-patch-id: 17df30a2da268e1bd07c7b44de4ef061cb744c50
prerequisite-patch-id: 94856bd5734d96521ef19e59cdb77af1d577fefa
prerequisite-patch-id: 7326de756f1a933cd3c44e4f7669c5a7c4519678
prerequisite-patch-id: 9e39c93f5b7267d1e8ffd176988803681ba03eeb
prerequisite-patch-id: c3f48cfa74e60b966610f88f73f7a393b5adf0b2
prerequisite-patch-id: 73a4faf81a8a0929ba8ba03a6d9a3286e0d9bd2b
prerequisite-patch-id: 55fb49cef84ea661b5340f14c56d73c04556f565
--
2.43.0
^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH v3 1/2] alpha: do not clear remote MMU contexts in migrate_flush_tlb_page()
2026-10-09 21:19 [PATCH v3 0/2] alpha: complete the direct-mm shootdown fixes Magnus Lindholm
@ 2026-10-09 21:19 ` Magnus Lindholm
2026-10-10 1:34 ` Matt Turner
2026-10-09 21:19 ` [PATCH v3 2/2] alpha: do not skip TLB shootdown IPIs for kthread-borrowed mms Magnus Lindholm
1 sibling, 1 reply; 5+ messages in thread
From: Magnus Lindholm @ 2026-10-09 21:19 UTC (permalink / raw)
To: richard.henderson, mattst88, linux-kernel, linux-alpha; +Cc: linmag7, stable
After the on_each_cpu() rendezvous, migrate_flush_tlb_page() walks the
other CPUs and zeroes their mm->context[cpu] when mm_users is at most one,
described as mimicking flush_tlb_mm()'s mm_users<=1 optimization.
It is not one. flush_tlb_mm() tests mm_users before deciding whether to
send the IPIs; here every CPU has already been visited and waited for, so
nothing is saved. The callback has also just set each CPU's own slot
correctly, so the loop only overwrites it, and does so while that CPU may
still be running in the address space.
Remove it. The shootdown shortcuts in flush_tlb_mm(), flush_tlb_page()
and flush_icache_user_page() clear remote slots too; once the next patch
removes those as well, every runtime update of mm->context[cpu] is made by
CPU cpu itself. That patch relies on the invariant when it reads those
slots to decide whether a shootdown can be skipped, so this one is tagged
for stable as its prerequisite.
Fixes: dd5712f3379c ("alpha: fix user-space corruption during memory compaction")
Cc: stable@vger.kernel.org
Reviewed-by: Matt Turner <mattst88@gmail.com>
Signed-off-by: Magnus Lindholm <linmag7@gmail.com>
Message-ID: <20261008195608.965266-3-linmag7@gmail.com>
---
arch/alpha/mm/tlbflush.c | 16 ----------------
1 file changed, 16 deletions(-)
diff --git a/arch/alpha/mm/tlbflush.c b/arch/alpha/mm/tlbflush.c
index ccbc317b9a34..239c72b8a741 100644
--- a/arch/alpha/mm/tlbflush.c
+++ b/arch/alpha/mm/tlbflush.c
@@ -90,22 +90,6 @@ void migrate_flush_tlb_page(struct vm_area_struct *vma, unsigned long addr)
*/
preempt_disable();
on_each_cpu(ipi_flush_mm_and_page, &d, 1);
-
- /*
- * mimic flush_tlb_mm()'s mm_users<=1 optimization.
- */
- if (atomic_read(&mm->mm_users) <= 1) {
-
- int cpu, this_cpu;
- this_cpu = smp_processor_id();
-
- for (cpu = 0; cpu < NR_CPUS; cpu++) {
- if (!cpu_online(cpu) || cpu == this_cpu)
- continue;
- if (READ_ONCE(mm->context[cpu]))
- WRITE_ONCE(mm->context[cpu], 0);
- }
- }
preempt_enable();
}
--
2.43.0
^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH v3 2/2] alpha: do not skip TLB shootdown IPIs for kthread-borrowed mms
2026-10-09 21:19 [PATCH v3 0/2] alpha: complete the direct-mm shootdown fixes Magnus Lindholm
2026-10-09 21:19 ` [PATCH v3 1/2] alpha: do not clear remote MMU contexts in migrate_flush_tlb_page() Magnus Lindholm
@ 2026-10-09 21:19 ` Magnus Lindholm
2026-10-10 1:35 ` Matt Turner
1 sibling, 1 reply; 5+ messages in thread
From: Magnus Lindholm @ 2026-10-09 21:19 UTC (permalink / raw)
To: richard.henderson, mattst88, linux-kernel, linux-alpha; +Cc: linmag7, stable
flush_tlb_mm(), flush_tlb_page() and flush_icache_user_page() skip the
shootdown IPI when mm_users <= 1, on the assumption that no other CPU can
be running the mm. A task that borrows an mm through kthread_use_mm()
takes mmgrab() rather than mmget(), so it never appears in mm_users, and
kthread_use_mm() may be handed an mm that is already the caller's
active_mm and loaded on that CPU. Such a CPU was not only left without
an IPI, its mm->context[cpu] was cleared from under it.
mm->context[cpu] is already the record of which CPUs hold an ASN for the
mm. Test it rather than clearing it: take the shortcut only when no
other CPU has one, and otherwise fall through to the IPI, which
invalidates those CPUs properly. A CPU that merely ran the mm in the
past also holds a context and now costs an IPI, which errs in the safe
direction.
Change ipi_flush_tlb_mm() and ipi_flush_icache_page() to test current->mm,
like ipi_flush_tlb_page(). A CPU that only retains the mm lazily then
clears its own slot through flush_tlb_other(), rather than publishing a
new ASN on every IPI. Once these historical slots are gone, the shortcut
is available again if mm_users remains at most one and no other CPU has
taken or kept a context. An explicit kthread_use_mm() borrower has
current->mm set and still receives the invalidate it needs.
Together with the preceding patch, dropping these clearing loops leaves
every runtime write to mm->context[] targeting the writing CPU's own slot,
so the array becomes a lockless publication of which CPUs may be using the
mm, with a single writer per slot. Mark the two stores that are now
observed from other CPUs; flush_tlb_other() already uses WRITE_ONCE().
The uniprocessor flush_icache_user_page() is left alone, as nothing reads
another CPU's slot there, and init_new_context() runs before the mm is
shared.
Order the slot reads after the page-table changes with smp_mb(). On the
publishing CPU, the store in ev5_switch_mm() precedes the scheduler's
post-switch barrier in finish_lock_switch(). Also issue smp_mb() after
publishing in __load_new_mm_context(), before loading and using the mm.
This covers every caller, including check_mmu_context(), which can
publish a replacement after finish_lock_switch() has already run.
This requires the preceding removal of remote slot clearing from
migrate_flush_tlb_page(). It also requires the stale-TLB series to set
need_new_asn on both the reused and newly allocated ASN paths. Otherwise
an IPI between finish_lock_switch() and check_mmu_context() can zero a
new slot while asn_lock is set, and the task can return to user space
with a live ASN that the shortcut cannot see.
Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2")
Cc: stable@vger.kernel.org
Signed-off-by: Magnus Lindholm <linmag7@gmail.com>
---
arch/alpha/include/asm/mmu_context.h | 2 +-
arch/alpha/kernel/smp.c | 52 ++++++++++++++--------------
arch/alpha/mm/fault.c | 4 ++-
3 files changed, 30 insertions(+), 28 deletions(-)
diff --git a/arch/alpha/include/asm/mmu_context.h b/arch/alpha/include/asm/mmu_context.h
index e5a5737506db..f8a42f42a5dc 100644
--- a/arch/alpha/include/asm/mmu_context.h
+++ b/arch/alpha/include/asm/mmu_context.h
@@ -158,7 +158,7 @@ ev5_switch_mm(struct mm_struct *prev_mm, struct mm_struct *next_mm,
mmc = next_mm->context[cpu];
if ((mmc ^ asn) & ~HARDWARE_ASN_MASK) {
mmc = __get_new_mm_context(next_mm, cpu);
- next_mm->context[cpu] = mmc;
+ WRITE_ONCE(next_mm->context[cpu], mmc);
}
#ifdef CONFIG_SMP
/* A deferred shootdown can also invalidate a newly allocated ASN. */
diff --git a/arch/alpha/kernel/smp.c b/arch/alpha/kernel/smp.c
index e21bc3920bec..11768aa292a4 100644
--- a/arch/alpha/kernel/smp.c
+++ b/arch/alpha/kernel/smp.c
@@ -625,12 +625,30 @@ static void
ipi_flush_tlb_mm(void *x)
{
struct mm_struct *mm = x;
- if (mm == current->active_mm && !asn_locked())
+ if (mm == current->mm && !asn_locked())
flush_tlb_current(mm);
else
flush_tlb_other(mm);
}
+/* True if a CPU other than this one holds an ASN for MM. */
+static bool
+mm_context_elsewhere(struct mm_struct *mm)
+{
+ int cpu, this_cpu = smp_processor_id();
+
+ /* Pairs with the barrier the publishing CPU issues before using MM. */
+ smp_mb();
+
+ for_each_online_cpu(cpu) {
+ if (cpu == this_cpu)
+ continue;
+ if (READ_ONCE(mm->context[cpu]))
+ return true;
+ }
+ return false;
+}
+
void
flush_tlb_mm(struct mm_struct *mm)
{
@@ -638,14 +656,8 @@ flush_tlb_mm(struct mm_struct *mm)
if (mm == current->active_mm) {
flush_tlb_current(mm);
- if (atomic_read(&mm->mm_users) <= 1) {
- int cpu, this_cpu = smp_processor_id();
- for (cpu = 0; cpu < NR_CPUS; cpu++) {
- if (!cpu_online(cpu) || cpu == this_cpu)
- continue;
- if (mm->context[cpu])
- mm->context[cpu] = 0;
- }
+ if (atomic_read(&mm->mm_users) <= 1 &&
+ !mm_context_elsewhere(mm)) {
preempt_enable();
return;
}
@@ -690,14 +702,8 @@ flush_tlb_page(struct vm_area_struct *vma, unsigned long addr)
/* As in ipi_flush_tlb_page(): a targeted tbi() needs MM current. */
if (mm == current->mm) {
flush_tlb_current_page(mm, vma, addr);
- if (atomic_read(&mm->mm_users) <= 1) {
- int cpu, this_cpu = smp_processor_id();
- for (cpu = 0; cpu < NR_CPUS; cpu++) {
- if (!cpu_online(cpu) || cpu == this_cpu)
- continue;
- if (mm->context[cpu])
- mm->context[cpu] = 0;
- }
+ if (atomic_read(&mm->mm_users) <= 1 &&
+ !mm_context_elsewhere(mm)) {
preempt_enable();
return;
}
@@ -728,7 +734,7 @@ static void
ipi_flush_icache_page(void *x)
{
struct mm_struct *mm = (struct mm_struct *) x;
- if (mm == current->active_mm && !asn_locked())
+ if (mm == current->mm && !asn_locked())
__load_new_mm_context(mm);
else
flush_tlb_other(mm);
@@ -747,14 +753,8 @@ flush_icache_user_page(struct vm_area_struct *vma, struct page *page,
if (mm == current->active_mm) {
__load_new_mm_context(mm);
- if (atomic_read(&mm->mm_users) <= 1) {
- int cpu, this_cpu = smp_processor_id();
- for (cpu = 0; cpu < NR_CPUS; cpu++) {
- if (!cpu_online(cpu) || cpu == this_cpu)
- continue;
- if (mm->context[cpu])
- mm->context[cpu] = 0;
- }
+ if (atomic_read(&mm->mm_users) <= 1 &&
+ !mm_context_elsewhere(mm)) {
preempt_enable();
return;
}
diff --git a/arch/alpha/mm/fault.c b/arch/alpha/mm/fault.c
index dfe427d93072..5b62bdeabe2a 100644
--- a/arch/alpha/mm/fault.c
+++ b/arch/alpha/mm/fault.c
@@ -69,7 +69,9 @@ __load_new_mm_context(struct mm_struct *next_mm)
struct pcb_struct *pcb;
mmc = __get_new_mm_context(next_mm, smp_processor_id());
- next_mm->context[smp_processor_id()] = mmc;
+ WRITE_ONCE(next_mm->context[smp_processor_id()], mmc);
+ /* Publish the context before this CPU can access the mm. */
+ smp_mb();
pcb = ¤t_thread_info()->pcb;
pcb->asn = mmc & HARDWARE_ASN_MASK;
--
2.43.0
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH v3 1/2] alpha: do not clear remote MMU contexts in migrate_flush_tlb_page()
2026-10-09 21:19 ` [PATCH v3 1/2] alpha: do not clear remote MMU contexts in migrate_flush_tlb_page() Magnus Lindholm
@ 2026-10-10 1:34 ` Matt Turner
0 siblings, 0 replies; 5+ messages in thread
From: Matt Turner @ 2026-10-10 1:34 UTC (permalink / raw)
To: Magnus Lindholm; +Cc: richard.henderson, linux-kernel, linux-alpha, stable
On Fri, Oct 9, 2026 at 11:19 PM Magnus Lindholm <linmag7@gmail.com> wrote:
> Fixes: dd5712f3379c ("alpha: fix user-space corruption during memory compaction")
> Cc: stable@vger.kernel.org
> Reviewed-by: Matt Turner <mattst88@gmail.com>
> Signed-off-by: Magnus Lindholm <linmag7@gmail.com>
> Message-ID: <20261008195608.965266-3-linmag7@gmail.com>
The Message-ID trailer is the v2 posting's and should be dropped.
The patch is unchanged from v2, so my Reviewed-by stands.
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH v3 2/2] alpha: do not skip TLB shootdown IPIs for kthread-borrowed mms
2026-10-09 21:19 ` [PATCH v3 2/2] alpha: do not skip TLB shootdown IPIs for kthread-borrowed mms Magnus Lindholm
@ 2026-10-10 1:35 ` Matt Turner
0 siblings, 0 replies; 5+ messages in thread
From: Matt Turner @ 2026-10-10 1:35 UTC (permalink / raw)
To: Magnus Lindholm; +Cc: richard.henderson, linux-kernel, linux-alpha, stable
On Fri, Oct 9, 2026 at 11:19 PM Magnus Lindholm <linmag7@gmail.com> wrote:
> Change ipi_flush_tlb_mm() and ipi_flush_icache_page() to test current->mm,
> like ipi_flush_tlb_page(). A CPU that only retains the mm lazily then
> clears its own slot through flush_tlb_other(), rather than publishing a
> new ASN on every IPI. Once these historical slots are gone, the shortcut
> is available again if mm_users remains at most one and no other CPU has
> taken or kept a context.
ipi_flush_mm_and_page() in arch/alpha/mm/tlbflush.c still tests
current->active_mm, so a lazy CPU allocates and publishes a new ASN on
every migration rendezvous. The next flush then finds that slot, sends
IPIs, and the lazy CPU clears it again. Nothing breaks, but with
compaction running the shortcut is lost after each migrated page, which
is the case this paragraph says is gone.
I think the handler can take the same test as the other three:
diff --git a/arch/alpha/mm/tlbflush.c b/arch/alpha/mm/tlbflush.c
--- a/arch/alpha/mm/tlbflush.c
+++ b/arch/alpha/mm/tlbflush.c
@@ -66,7 +66,7 @@ static void ipi_flush_mm_and_page(void *x)
struct tlb_mm_and_addr *d = x;
/* Part 1: mm context side (Alpha uses ASN/context as a key
mechanism). */
- if (d->mm == current->active_mm && !asn_locked())
+ if (d->mm == current->mm && !asn_locked())
__load_new_mm_context(d->mm);
else
flush_tlb_other(d->mm);
A lazy CPU then clears its slot and still does the tbi() below. Any way
back to the mm from there gets a new ASN: ev5_switch_mm() sees version
0, and kthread_use_mm() goes through the next == current path. I have
not tested that change.
The callers have the same asymmetry. flush_tlb_mm() and
flush_icache_user_page() test current->active_mm, so a kernel thread
that flushes an mm it only holds lazily publishes its own slot. That
one also takes the shortcut in that branch, so I would leave it alone,
but the changelog could say so.
> Order the slot reads after the page-table changes with smp_mb(). On the
> publishing CPU, the store in ev5_switch_mm() precedes the scheduler's
> post-switch barrier in finish_lock_switch().
The barrier there is the mb() in alpha's arch_spin_unlock(). The
scheduler does not promise a full barrier when it drops the rq lock, so
this is an alpha property, and nothing at the store says it is being
relied on. Could the changelog name it, and the store get a comment?
Something like:
mmc = __get_new_mm_context(next_mm, cpu);
+ /*
+ * Ordered before any use of the mm by the mb() in
+ * arch_spin_unlock(), when finish_lock_switch() drops
+ * the rq lock.
+ */
WRITE_ONCE(next_mm->context[cpu], mmc);
The rest looks right to me. I went through the window between
finish_lock_switch() and check_mmu_context(): an IPI there clears the
slot under asn_lock, a flush on another CPU can then take the shortcut,
and check_mmu_context() sees the zero and loads a new context behind
the new smp_mb(). The three points from my v2 review are addressed.
The fork numbers in the cover letter are from before the handler
change. It would be good to have them again for this version, along
with an SMP run under compaction.
With ipi_flush_mm_and_page() changed, or the changelog reworded to
match what it does:
Reviewed-by: Matt Turner <mattst88@gmail.com>
^ permalink raw reply [flat|nested] 5+ messages in thread
end of thread, other threads:[~2026-10-10 1:35 UTC | newest]
Thread overview: 5+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-09 21:19 [PATCH v3 0/2] alpha: complete the direct-mm shootdown fixes Magnus Lindholm
2026-10-09 21:19 ` [PATCH v3 1/2] alpha: do not clear remote MMU contexts in migrate_flush_tlb_page() Magnus Lindholm
2026-10-10 1:34 ` Matt Turner
2026-10-09 21:19 ` [PATCH v3 2/2] alpha: do not skip TLB shootdown IPIs for kthread-borrowed mms Magnus Lindholm
2026-10-10 1:35 ` Matt Turner
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®