* [PATCH v4 1/2] alpha: do not clear remote MMU contexts in migrate_flush_tlb_page()
2026-10-10 14:33 [PATCH v4 0/2] alpha: complete the direct-mm shootdown fixes Magnus Lindholm
@ 2026-10-10 14:33 ` Magnus Lindholm
2026-10-10 14:33 ` [PATCH v4 2/2] alpha: do not skip TLB shootdown IPIs for kthread-borrowed mms Magnus Lindholm
2026-10-10 23:33 ` [PATCH v4 0/2] alpha: complete the direct-mm shootdown fixes Matt Turner
2 siblings, 0 replies; 6+ messages in thread
From: Magnus Lindholm @ 2026-10-10 14:33 UTC (permalink / raw)
To: richard.henderson, mattst88, linux-kernel, linux-alpha; +Cc: linmag7, stable
After the on_each_cpu() rendezvous, migrate_flush_tlb_page() walks the
other CPUs and zeroes their mm->context[cpu] when mm_users is at most one,
described as mimicking flush_tlb_mm()'s mm_users<=1 optimization.
It is not one. flush_tlb_mm() tests mm_users before deciding whether to
send the IPIs; here every CPU has already been visited and waited for, so
nothing is saved. The callback has also just set each CPU's own slot
correctly, so the loop only overwrites it, and does so while that CPU may
still be running in the address space.
Remove it. The shootdown shortcuts in flush_tlb_mm(), flush_tlb_page()
and flush_icache_user_page() clear remote slots too; once the next patch
removes those as well, every runtime update of mm->context[cpu] is made by
CPU cpu itself. That patch relies on the invariant when it reads those
slots to decide whether a shootdown can be skipped, so this one is tagged
for stable as its prerequisite.
Fixes: dd5712f3379c ("alpha: fix user-space corruption during memory compaction")
Cc: stable@vger.kernel.org
Reviewed-by: Matt Turner <mattst88@gmail.com>
Signed-off-by: Magnus Lindholm <linmag7@gmail.com>
---
arch/alpha/mm/tlbflush.c | 16 ----------------
1 file changed, 16 deletions(-)
diff --git a/arch/alpha/mm/tlbflush.c b/arch/alpha/mm/tlbflush.c
index ccbc317b9a34..239c72b8a741 100644
--- a/arch/alpha/mm/tlbflush.c
+++ b/arch/alpha/mm/tlbflush.c
@@ -90,22 +90,6 @@ void migrate_flush_tlb_page(struct vm_area_struct *vma, unsigned long addr)
*/
preempt_disable();
on_each_cpu(ipi_flush_mm_and_page, &d, 1);
-
- /*
- * mimic flush_tlb_mm()'s mm_users<=1 optimization.
- */
- if (atomic_read(&mm->mm_users) <= 1) {
-
- int cpu, this_cpu;
- this_cpu = smp_processor_id();
-
- for (cpu = 0; cpu < NR_CPUS; cpu++) {
- if (!cpu_online(cpu) || cpu == this_cpu)
- continue;
- if (READ_ONCE(mm->context[cpu]))
- WRITE_ONCE(mm->context[cpu], 0);
- }
- }
preempt_enable();
}
--
2.43.0
^ permalink raw reply [flat|nested] 6+ messages in thread* [PATCH v4 2/2] alpha: do not skip TLB shootdown IPIs for kthread-borrowed mms
2026-10-10 14:33 [PATCH v4 0/2] alpha: complete the direct-mm shootdown fixes Magnus Lindholm
2026-10-10 14:33 ` [PATCH v4 1/2] alpha: do not clear remote MMU contexts in migrate_flush_tlb_page() Magnus Lindholm
@ 2026-10-10 14:33 ` Magnus Lindholm
2026-10-10 23:34 ` Matt Turner
2026-10-10 23:33 ` [PATCH v4 0/2] alpha: complete the direct-mm shootdown fixes Matt Turner
2 siblings, 1 reply; 6+ messages in thread
From: Magnus Lindholm @ 2026-10-10 14:33 UTC (permalink / raw)
To: richard.henderson, mattst88, linux-kernel, linux-alpha; +Cc: linmag7, stable
flush_tlb_mm(), flush_tlb_page() and flush_icache_user_page() skip the
shootdown IPI when mm_users <= 1, on the assumption that no other CPU can
be running the mm. A task that borrows an mm through kthread_use_mm()
takes mmgrab() rather than mmget(), so it never appears in mm_users, and
kthread_use_mm() may be handed an mm that is already the caller's
active_mm and loaded on that CPU. Such a CPU was not only left without
an IPI, its mm->context[cpu] was cleared from under it.
mm->context[cpu] is already the record of which CPUs hold an ASN for the
mm. Test it rather than clearing it: take the shortcut only when no
other CPU has one, and otherwise fall through to the IPI, which
invalidates those CPUs properly. A CPU that merely ran the mm in the
past also holds a context and now costs an IPI, which errs in the safe
direction.
Change ipi_flush_tlb_mm(), ipi_flush_icache_page() and the migration
callback ipi_flush_mm_and_page() to test current->mm, like
ipi_flush_tlb_page(). A CPU that only retains the mm lazily then clears
its own slot through flush_tlb_other(), rather than publishing a new ASN
on every IPI or migration rendezvous. The migration callback still does
its per-VA tbi() afterwards. Returning to the mm allocates a fresh ASN
through ev5_switch_mm(), or through the direct-load path when
kthread_use_mm() switches current's mm.
Once these historical slots are gone, the shortcut is available again
if mm_users remains at most one and no other CPU has taken or kept a
context. An explicit kthread_use_mm() borrower has current->mm set and
still receives the invalidate it needs.
The caller-side tests in flush_tlb_mm() and flush_icache_user_page()
remain current->active_mm. A kernel thread that flushes its lazily held
mm can therefore publish its own context slot. Those branches also
perform the shortcut check, which excludes the calling CPU's slot, so
leave them unchanged. This is distinct from a remote lazy CPU keeping
its slot populated through repeated callbacks.
Together with the preceding patch, dropping these clearing loops leaves
every runtime write to mm->context[] targeting the writing CPU's own slot,
so the array becomes a lockless publication of which CPUs may be using the
mm, with a single writer per slot. Mark the two stores that are now
observed from other CPUs; flush_tlb_other() already uses WRITE_ONCE().
The uniprocessor flush_icache_user_page() is left alone, as nothing reads
another CPU's slot there, and init_new_context() runs before the mm is
shared.
Order the slot reads after the page-table changes with smp_mb(). On the
publishing CPU, the store in ev5_switch_mm() precedes the mb() in Alpha's
arch_spin_unlock(), when finish_lock_switch() drops the rq lock. This
relies on Alpha's unlock implementation, not a generic scheduler promise
of a full barrier on unlock; document that reliance at the store.
Also issue smp_mb() after
publishing in __load_new_mm_context(), before loading and using the mm.
This covers every caller, including check_mmu_context(), which can
publish a replacement after finish_lock_switch() has already run.
This requires the preceding removal of remote slot clearing from
migrate_flush_tlb_page(). It also requires the stale-TLB series to set
need_new_asn on both the reused and newly allocated ASN paths. Otherwise
an IPI between finish_lock_switch() and check_mmu_context() can zero a
new slot while asn_lock is set, and the task can return to user space
with a live ASN that the shortcut cannot see.
Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2")
Cc: stable@vger.kernel.org
Reviewed-by: Matt Turner <mattst88@gmail.com>
Signed-off-by: Magnus Lindholm <linmag7@gmail.com>
---
arch/alpha/include/asm/mmu_context.h | 7 +++-
arch/alpha/kernel/smp.c | 52 ++++++++++++++--------------
arch/alpha/mm/fault.c | 4 ++-
arch/alpha/mm/tlbflush.c | 2 +-
4 files changed, 36 insertions(+), 29 deletions(-)
diff --git a/arch/alpha/include/asm/mmu_context.h b/arch/alpha/include/asm/mmu_context.h
index e5a5737506db..46b9b82f3895 100644
--- a/arch/alpha/include/asm/mmu_context.h
+++ b/arch/alpha/include/asm/mmu_context.h
@@ -158,7 +158,12 @@ ev5_switch_mm(struct mm_struct *prev_mm, struct mm_struct *next_mm,
mmc = next_mm->context[cpu];
if ((mmc ^ asn) & ~HARDWARE_ASN_MASK) {
mmc = __get_new_mm_context(next_mm, cpu);
- next_mm->context[cpu] = mmc;
+ /*
+ * Ordered before any use of the mm by the mb() in
+ * arch_spin_unlock(), when finish_lock_switch() drops
+ * the rq lock.
+ */
+ WRITE_ONCE(next_mm->context[cpu], mmc);
}
#ifdef CONFIG_SMP
/* A deferred shootdown can also invalidate a newly allocated ASN. */
diff --git a/arch/alpha/kernel/smp.c b/arch/alpha/kernel/smp.c
index e21bc3920bec..11768aa292a4 100644
--- a/arch/alpha/kernel/smp.c
+++ b/arch/alpha/kernel/smp.c
@@ -625,12 +625,30 @@ static void
ipi_flush_tlb_mm(void *x)
{
struct mm_struct *mm = x;
- if (mm == current->active_mm && !asn_locked())
+ if (mm == current->mm && !asn_locked())
flush_tlb_current(mm);
else
flush_tlb_other(mm);
}
+/* True if a CPU other than this one holds an ASN for MM. */
+static bool
+mm_context_elsewhere(struct mm_struct *mm)
+{
+ int cpu, this_cpu = smp_processor_id();
+
+ /* Pairs with the barrier the publishing CPU issues before using MM. */
+ smp_mb();
+
+ for_each_online_cpu(cpu) {
+ if (cpu == this_cpu)
+ continue;
+ if (READ_ONCE(mm->context[cpu]))
+ return true;
+ }
+ return false;
+}
+
void
flush_tlb_mm(struct mm_struct *mm)
{
@@ -638,14 +656,8 @@ flush_tlb_mm(struct mm_struct *mm)
if (mm == current->active_mm) {
flush_tlb_current(mm);
- if (atomic_read(&mm->mm_users) <= 1) {
- int cpu, this_cpu = smp_processor_id();
- for (cpu = 0; cpu < NR_CPUS; cpu++) {
- if (!cpu_online(cpu) || cpu == this_cpu)
- continue;
- if (mm->context[cpu])
- mm->context[cpu] = 0;
- }
+ if (atomic_read(&mm->mm_users) <= 1 &&
+ !mm_context_elsewhere(mm)) {
preempt_enable();
return;
}
@@ -690,14 +702,8 @@ flush_tlb_page(struct vm_area_struct *vma, unsigned long addr)
/* As in ipi_flush_tlb_page(): a targeted tbi() needs MM current. */
if (mm == current->mm) {
flush_tlb_current_page(mm, vma, addr);
- if (atomic_read(&mm->mm_users) <= 1) {
- int cpu, this_cpu = smp_processor_id();
- for (cpu = 0; cpu < NR_CPUS; cpu++) {
- if (!cpu_online(cpu) || cpu == this_cpu)
- continue;
- if (mm->context[cpu])
- mm->context[cpu] = 0;
- }
+ if (atomic_read(&mm->mm_users) <= 1 &&
+ !mm_context_elsewhere(mm)) {
preempt_enable();
return;
}
@@ -728,7 +734,7 @@ static void
ipi_flush_icache_page(void *x)
{
struct mm_struct *mm = (struct mm_struct *) x;
- if (mm == current->active_mm && !asn_locked())
+ if (mm == current->mm && !asn_locked())
__load_new_mm_context(mm);
else
flush_tlb_other(mm);
@@ -747,14 +753,8 @@ flush_icache_user_page(struct vm_area_struct *vma, struct page *page,
if (mm == current->active_mm) {
__load_new_mm_context(mm);
- if (atomic_read(&mm->mm_users) <= 1) {
- int cpu, this_cpu = smp_processor_id();
- for (cpu = 0; cpu < NR_CPUS; cpu++) {
- if (!cpu_online(cpu) || cpu == this_cpu)
- continue;
- if (mm->context[cpu])
- mm->context[cpu] = 0;
- }
+ if (atomic_read(&mm->mm_users) <= 1 &&
+ !mm_context_elsewhere(mm)) {
preempt_enable();
return;
}
diff --git a/arch/alpha/mm/fault.c b/arch/alpha/mm/fault.c
index dfe427d93072..5b62bdeabe2a 100644
--- a/arch/alpha/mm/fault.c
+++ b/arch/alpha/mm/fault.c
@@ -69,7 +69,9 @@ __load_new_mm_context(struct mm_struct *next_mm)
struct pcb_struct *pcb;
mmc = __get_new_mm_context(next_mm, smp_processor_id());
- next_mm->context[smp_processor_id()] = mmc;
+ WRITE_ONCE(next_mm->context[smp_processor_id()], mmc);
+ /* Publish the context before this CPU can access the mm. */
+ smp_mb();
pcb = ¤t_thread_info()->pcb;
pcb->asn = mmc & HARDWARE_ASN_MASK;
diff --git a/arch/alpha/mm/tlbflush.c b/arch/alpha/mm/tlbflush.c
index 239c72b8a741..a2aef90e6234 100644
--- a/arch/alpha/mm/tlbflush.c
+++ b/arch/alpha/mm/tlbflush.c
@@ -66,7 +66,7 @@ static void ipi_flush_mm_and_page(void *x)
struct tlb_mm_and_addr *d = x;
/* Part 1: mm context side (Alpha uses ASN/context as a key mechanism). */
- if (d->mm == current->active_mm && !asn_locked())
+ if (d->mm == current->mm && !asn_locked())
__load_new_mm_context(d->mm);
else
flush_tlb_other(d->mm);
--
2.43.0
^ permalink raw reply [flat|nested] 6+ messages in thread* Re: [PATCH v4 0/2] alpha: complete the direct-mm shootdown fixes
2026-10-10 14:33 [PATCH v4 0/2] alpha: complete the direct-mm shootdown fixes Magnus Lindholm
2026-10-10 14:33 ` [PATCH v4 1/2] alpha: do not clear remote MMU contexts in migrate_flush_tlb_page() Magnus Lindholm
2026-10-10 14:33 ` [PATCH v4 2/2] alpha: do not skip TLB shootdown IPIs for kthread-borrowed mms Magnus Lindholm
@ 2026-10-10 23:33 ` Matt Turner
2 siblings, 0 replies; 6+ messages in thread
From: Matt Turner @ 2026-10-10 23:33 UTC (permalink / raw)
To: Magnus Lindholm; +Cc: richard.henderson, linux-kernel, linux-alpha
On Sat, Oct 10, 2026 at 16:33 +0200, Magnus Lindholm wrote:
> This is the second of two series. Apply all eight patches of "alpha: fix
> stale TLB translations breaking copy-on-write and writeback" v4 first,
> then these two patches. The base-commit and prerequisite-patch-id lines
> below describe that ordering from for-next at 33465f6ab697.
The first two prerequisite-patch-id lines (d194ee1170a4, 17df30a2da26) do
not match what is in for-next. for-next carries the v5 versions of "load
the MMU context when switch_mm() switches the current task" and "run
check_mmu_context() from finish_arch_post_lock_switch()". The other six
match, and both patches here apply unchanged on top, so this only matters
to someone following the cover letter literally.
> Fresh fork-throughput measurements and an SMP run under compaction are
> still pending for this revision. Targeted Alpha SMP, UP and SMP with
> compaction builds passed, as did a full ES40 CONFIG_COMPACTION=y build.
Thanks for being explicit about this. The current->mm test in
ipi_flush_mm_and_page() is the one change in v4 that no run has
exercised, and the series is in for-next already. I would like to see an
SMP soak with CONFIG_COMPACTION=y and compaction actually triggered
before this goes further.
I went through v4 again and both tags stand. In particular I checked
that the only caller reaching the ev5_switch_mm() store is
context_switch(); kthread_use_mm(), sched_force_init_mm() and
idle_task_exit() all pass current and take the __load_new_mm_context()
branch, so the arch_spin_unlock() argument covers every publication from
that store. Two small comments on patch 2.
^ permalink raw reply [flat|nested] 6+ messages in thread