* [PATCH v2 0/3] alpha: load the MMU context on a direct mm switch
@ 2026-10-08 19:55 Magnus Lindholm
2026-10-08 19:55 ` [PATCH v2 1/3] alpha: load the MMU context when switch_mm() switches the current task Magnus Lindholm
` (2 more replies)
0 siblings, 3 replies; 7+ messages in thread
From: Magnus Lindholm @ 2026-10-08 19:55 UTC (permalink / raw)
To: richard.henderson, mattst88, linux-kernel, linux-alpha; +Cc: linmag7
Three fixes for how alpha handles an mm that is switched in directly
rather than by the scheduler.
Patch 1 makes those direct mm switches actually load the MMU context, which
they currently do not: kthread_use_mm() and sched_force_init_mm() call
switch_mm_irqs_off() directly rather than through the scheduler, so the
hardware context is never installed and the task keeps running under
whatever was loaded before. Patch 2 removes a redundant clearing loop in
migrate_flush_tlb_page() that patch 3 depends on being gone. Patch 3 stops
the TLB shootdown shortcut from skipping a CPU that is borrowing an mm
through kthread_use_mm(), using mm->context[cpu] as a lockless publication
of which CPUs may be using the mm; between them patches 2 and 3 remove the
remote context clearing that would otherwise invalidate that test.
Patch 1 comes first here because it also has to come first relative to the
stale-TLB series; see Order below. It was patch 3 of v1, which is how the
stale-TLB v3 cover letter refers to it.
Reproducing it. A KUnit test was written for patch 1; source can be made
available on request. It uses kunit_attach_mm(), which calls
kthread_use_mm(), and then compares the loaded context against current->mm:
# alpha_use_mm_loads_context: EXPECTATION FAILED
Expected pcb->ptbr == mm_to_ptbr(current->mm), but
pcb->ptbr == 384 (0x180)
mm_to_ptbr(current->mm) == 12555 (0x310b)
The kernel thread is running on the page tables it had before the switch.
Here that was 0x180, swapper_pg_dir, but it need not be: Matt Turner ran
an equivalent test on an ES47 and saw a stale ptbr that was the page table
of a user process that had run on that CPU earlier. The kthread can then
read and write that process's memory where the stale mappings allow it,
and translations taken that way can end up tagged with the borrowed mm's
ASN.
The pcb.asn check in the same test passes, which is the signature -
ev5_switch_mm() does write the ASN field, so it is the load that is
missing rather than the bookkeeping. That also means a passing pcb.asn
comparison says nothing about what the hardware has installed: writing the
PCB field is not the PAL context reload.
Do not try to detect this by dereferencing the borrowed mm's user
addresses. Where the stale tables have nothing mapped the access faults,
and do_page_fault() resolves faults against current->mm without reloading
the context, so it can fault again on return instead of reporting
anything.
Testing. ES40, EV68AL (21264C) Tsunami, 3 CPUs, v7.2-rc6.
before after
KUnit alpha_mmu_context 0 of 2 pass 2 of 2 pass
Patch 3 was measured rather than argued. Dropping the shortcut outright
is the obvious fix and costs far too much, so it tests mm->context[cpu]
instead:
fork/s, single-threaded shortcut as before 1360
shortcut removed 1050 -23%
patch 3 1360
Medians of seven, seven and twelve runs, spread 1352-1367, 1046-1052 and
1340-1366. The unchanged throughput shows the shortcut remains effective
for this workload, and neither the barrier nor the context marking shows
above the noise on this machine.
No regressions in the wider suite: glibc malloc-check 25/25 and four
related tests 10/10 each, the copy-on-write and writeback reproducers
from the previous series clean, and the same again under continuous
compaction with 359768 folios migrated during the run.
Matt Turner tested the series on an ES47 (EV7), running v7.3-rc1 with the
stale-TLB series applied - v2 at the time, which is code-identical to the
v3 now posted - with and without these patches: his own KUnit test 3 of 3
including user accesses either side of a sleep, usercopy_kunit 4 of 4,
kunit_iov_iter 17 of 17, and a fork/COW check under continuous
compact_memory clean with fork throughput unchanged. One CPU was online
for that run, so patch 3's shortcut had no other CPU to look at and
patches 2 and 3 are not covered by it.
One adjacent problem is deliberately not addressed. enter_lazy_tlb() sets
pcb.ptbr for the borrowed mm without setting pcb.asn, so a kernel thread
switched in through PAL_swpctx loads one address space's page tables
against another's ASN. I have not found an observable failure from that
mismatch outside kernel-thread user accesses, which go through
kthread_use_mm() and are a matched pair after patch 1. Changing it would
mean touching the VPTB self-map behaviour on every lazy switch, with no
reproducer to justify the risk.
Reachability. v1 argued that KUnit was the only thing on alpha that
reaches kthread_use_mm(), and that was wrong. vhost has a kthread worker
mode, vhost_run_work_kthread_list(), that calls it;
CONFIG_VHOST_ENABLE_FORK_OWNER_CONTROL defaults to y and userspace selects
the mode with VHOST_SET_FORK_FROM_OWNER or the fork_from_owner_default
module parameter. The USB gadget f_fs and gadgetfs AIO completion paths
call it, and dummy_hcd supplies a controller on any architecture. vdpa_sim
calls it. KUnit's kunit_attach_mm() reaches it as well. Since the failure
is a cross-process memory access, all three patches now carry Cc: stable.
Thanks to Matt Turner for catching this.
Order. Patch 1 stands alone: against v7.3-rc1 it applies to a plain tree,
it does not depend on patches 2 and 3, and it does not depend on the
stale-TLB series. It should go in before that series' first patch. That
patch adds a finish_arch_post_lock_switch() hook which, on an ordinary
alpha build where preemptible() is a compile-time 0, clears asn_lock as
kthread_use_mm() returns - while the borrowed mm's context is still not
installed. A shootdown IPI landing in that window then issues a targeted
tbi() against the loaded context instead of taking the conservative
flush_tlb_other() path, and leaves mm->context[cpu] valid. Patch 1 here
closes that window by installing the context at the direct switch, which is
what its cover letter says should happen:
https://lore.kernel.org/linux-alpha/20260923074903.862898-1-linmag7@gmail.com/
Patches 2 and 3 do need that series. Patch 3 rewrites the same
flush_tlb_page() and flush_tlb_mm() shortcuts it touches, and needs patch 2
as well: while migrate_flush_tlb_page() still zeroes remote context slots, a
CPU running a borrowed mm can be cleared out of the array and skipped by
exactly the test patch 3 adds. The prerequisite-patch-id lines below record
that series; they are unchanged from v1, because v3 of it changed no code.
Changes since v1:
- reordered: the direct context load is patch 1 rather than patch 3, so
that it comes first where both series are applied. See Order above.
- Cc: stable on all three patches. The v1 claim that only KUnit reached
kthread_use_mm() on alpha does not hold; see Reachability above.
- Describe the failure as a possible cross-process memory access rather
than a fault loop, and stop implying the stale ptbr is always
swapper_pg_dir.
- Patch 3 says what becomes of a stale context slot: the shootdown makes
a remote CPU that no longer has the mm active clear its own slot in
flush_tlb_other(), so once those are gone the shortcut is available
again, as long as mm_users stays at most one and no other CPU has taken
or kept a context.
- Patches 2 and 3 no longer credit patch 2 alone with making every
mm->context[] write local to the writing CPU; the three shortcuts patch
3 rewrites clear remote slots too, so it takes both patches.
- Patch 1's comment names kthread_use_mm() and sched_force_init_mm(), so
the next reader knows why the scheduler never takes that branch.
- Patch 1 carries Matt Turner's Tested-by.
- No functional change: the only code difference from v1 is that comment.
v1: https://lore.kernel.org/linux-alpha/20260904162424.376504-1-linmag7@gmail.com/
Magnus Lindholm (3):
alpha: load the MMU context when switch_mm() switches the current task
alpha: do not clear remote MMU contexts in migrate_flush_tlb_page()
alpha: do not skip TLB shootdown IPIs for kthread-borrowed mms
arch/alpha/include/asm/mmu_context.h | 14 ++++++--
arch/alpha/kernel/smp.c | 48 ++++++++++++++--------------
arch/alpha/mm/fault.c | 2 +-
arch/alpha/mm/tlbflush.c | 16 ----------
4 files changed, 37 insertions(+), 43 deletions(-)
base-commit: cee9395acd8043be0644b25c34bfa86623f2b935
prerequisite-patch-id: 6e9d8cf8190164a8cbe34a879d5a18f23bc9bd61
prerequisite-patch-id: 94856bd5734d96521ef19e59cdb77af1d577fefa
prerequisite-patch-id: 9e39c93f5b7267d1e8ffd176988803681ba03eeb
prerequisite-patch-id: 7326de756f1a933cd3c44e4f7669c5a7c4519678
prerequisite-patch-id: c3f48cfa74e60b966610f88f73f7a393b5adf0b2
prerequisite-patch-id: 73a4faf81a8a0929ba8ba03a6d9a3286e0d9bd2b
prerequisite-patch-id: 55fb49cef84ea661b5340f14c56d73c04556f565
--
2.43.0
^ permalink raw reply [flat|nested] 7+ messages in thread* [PATCH v2 1/3] alpha: load the MMU context when switch_mm() switches the current task
2026-10-08 19:55 [PATCH v2 0/3] alpha: load the MMU context on a direct mm switch Magnus Lindholm
@ 2026-10-08 19:55 ` Magnus Lindholm
2026-10-08 21:25 ` Matt Turner
2026-10-08 19:55 ` [PATCH v2 2/3] alpha: do not clear remote MMU contexts in migrate_flush_tlb_page() Magnus Lindholm
2026-10-08 19:55 ` [PATCH v2 3/3] alpha: do not skip TLB shootdown IPIs for kthread-borrowed mms Magnus Lindholm
2 siblings, 1 reply; 7+ messages in thread
From: Magnus Lindholm @ 2026-10-08 19:55 UTC (permalink / raw)
To: richard.henderson, mattst88, linux-kernel, linux-alpha; +Cc: linmag7, stable
ev5_switch_mm() only prepares the incoming PCB. The context is installed
by PAL_swpctx, which alpha_switch_to() issues against that PCB on the way
out of the scheduler.
Two callers reach switch_mm_irqs_off() without going through
alpha_switch_to(): kthread_use_mm(), which borrows an mm for the current
kernel thread, and sched_force_init_mm() on the CPU-hotplug teardown path.
Neither explicitly loads the context, so the task can carry on running
under whatever was loaded before while current->mm says otherwise.
The stale page-table root need not be swapper_pg_dir; it may belong to a
user process that ran on the CPU earlier, and the kthread's user accesses
can then read and write that process's memory wherever the stale mappings
allow. Translations taken that way can also end up tagged with the
borrowed mm's ASN: ev5_switch_mm() writes that ASN into the PCB, which the
next PAL_swpctx installs against the stale ptbr. An address the stale
tables do not map faults instead, and since do_page_fault() resolves
faults against current->mm without reloading the context, the same access
can fault again on return.
sched_force_init_mm() needs CONFIG_HOTPLUG_CPU, which alpha does not
support, so kthread_use_mm() is the only one of the two reachable in
practice; the fix below tests the caller's identity rather than
special-casing either one.
The scheduler passes the incoming task, which is not current until
alpha_switch_to() runs; both direct callers pass current. Test for that
and load the context the way activate_mm() does. Both hold interrupts
disabled across switch_mm_irqs_off(), so this completes before any
shootdown can be taken and needs no asn_lock handshake.
Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2")
Cc: <stable@vger.kernel.org>
Signed-off-by: Magnus Lindholm <linmag7@gmail.com>
Tested-by: Matt Turner <mattst88@gmail.com>
---
arch/alpha/include/asm/mmu_context.h | 12 +++++++++++-
1 file changed, 11 insertions(+), 1 deletion(-)
diff --git a/arch/alpha/include/asm/mmu_context.h b/arch/alpha/include/asm/mmu_context.h
index 825d3b9605c9..6ecc8a0676cb 100644
--- a/arch/alpha/include/asm/mmu_context.h
+++ b/arch/alpha/include/asm/mmu_context.h
@@ -130,6 +130,8 @@ __get_new_mm_context(struct mm_struct *mm, long cpu)
return next;
}
+extern void __load_new_mm_context(struct mm_struct *);
+
__EXTERN_INLINE void
ev5_switch_mm(struct mm_struct *prev_mm, struct mm_struct *next_mm,
struct task_struct *next)
@@ -139,6 +141,15 @@ ev5_switch_mm(struct mm_struct *prev_mm, struct mm_struct *next_mm,
unsigned long mmc;
long cpu = smp_processor_id();
+ /*
+ * kthread_use_mm() and sched_force_init_mm() switch current's mm
+ * without alpha_switch_to(), which is what loads the context.
+ */
+ if (next == current) {
+ __load_new_mm_context(next_mm);
+ return;
+ }
+
#ifdef CONFIG_SMP
cpu_data[cpu].asn_lock = 1;
barrier();
@@ -160,7 +171,6 @@ ev5_switch_mm(struct mm_struct *prev_mm, struct mm_struct *next_mm,
task_thread_info(next)->pcb.asn = mmc & HARDWARE_ASN_MASK;
}
-extern void __load_new_mm_context(struct mm_struct *);
asmlinkage void do_page_fault(unsigned long address, unsigned long mmcsr,
long cause, struct pt_regs *regs);
--
2.43.0
^ permalink raw reply [flat|nested] 7+ messages in thread* Re: [PATCH v2 1/3] alpha: load the MMU context when switch_mm() switches the current task
2026-10-08 19:55 ` [PATCH v2 1/3] alpha: load the MMU context when switch_mm() switches the current task Magnus Lindholm
@ 2026-10-08 21:25 ` Matt Turner
0 siblings, 0 replies; 7+ messages in thread
From: Matt Turner @ 2026-10-08 21:25 UTC (permalink / raw)
To: Magnus Lindholm; +Cc: richard.henderson, linux-kernel, linux-alpha, stable
On Thu, Oct 8, 2026 at 3:56 PM Magnus Lindholm <linmag7@gmail.com> wrote:
>
> ev5_switch_mm() only prepares the incoming PCB. The context is installed
> by PAL_swpctx, which alpha_switch_to() issues against that PCB on the way
> out of the scheduler.
>
> Two callers reach switch_mm_irqs_off() without going through
> alpha_switch_to(): kthread_use_mm(), which borrows an mm for the current
> kernel thread, and sched_force_init_mm() on the CPU-hotplug teardown path.
> Neither explicitly loads the context, so the task can carry on running
> under whatever was loaded before while current->mm says otherwise.
>
> The stale page-table root need not be swapper_pg_dir; it may belong to a
> user process that ran on the CPU earlier, and the kthread's user accesses
> can then read and write that process's memory wherever the stale mappings
> allow. Translations taken that way can also end up tagged with the
> borrowed mm's ASN: ev5_switch_mm() writes that ASN into the PCB, which the
> next PAL_swpctx installs against the stale ptbr. An address the stale
> tables do not map faults instead, and since do_page_fault() resolves
> faults against current->mm without reloading the context, the same access
> can fault again on return.
>
> sched_force_init_mm() needs CONFIG_HOTPLUG_CPU, which alpha does not
> support, so kthread_use_mm() is the only one of the two reachable in
> practice; the fix below tests the caller's identity rather than
> special-casing either one.
>
> The scheduler passes the incoming task, which is not current until
> alpha_switch_to() runs; both direct callers pass current. Test for that
> and load the context the way activate_mm() does. Both hold interrupts
> disabled across switch_mm_irqs_off(), so this completes before any
> shootdown can be taken and needs no asn_lock handshake.
>
> Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2")
> Cc: <stable@vger.kernel.org>
> Signed-off-by: Magnus Lindholm <linmag7@gmail.com>
> Tested-by: Matt Turner <mattst88@gmail.com>
Reviewed-by: Matt Turner <mattst88@gmail.com>
^ permalink raw reply [flat|nested] 7+ messages in thread
* [PATCH v2 2/3] alpha: do not clear remote MMU contexts in migrate_flush_tlb_page()
2026-10-08 19:55 [PATCH v2 0/3] alpha: load the MMU context on a direct mm switch Magnus Lindholm
2026-10-08 19:55 ` [PATCH v2 1/3] alpha: load the MMU context when switch_mm() switches the current task Magnus Lindholm
@ 2026-10-08 19:55 ` Magnus Lindholm
2026-10-08 21:19 ` Matt Turner
2026-10-08 19:55 ` [PATCH v2 3/3] alpha: do not skip TLB shootdown IPIs for kthread-borrowed mms Magnus Lindholm
2 siblings, 1 reply; 7+ messages in thread
From: Magnus Lindholm @ 2026-10-08 19:55 UTC (permalink / raw)
To: richard.henderson, mattst88, linux-kernel, linux-alpha; +Cc: linmag7, stable
After the on_each_cpu() rendezvous, migrate_flush_tlb_page() walks the
other CPUs and zeroes their mm->context[cpu] when mm_users is at most one,
described as mimicking flush_tlb_mm()'s mm_users<=1 optimization.
It is not one. flush_tlb_mm() tests mm_users before deciding whether to
send the IPIs; here every CPU has already been visited and waited for, so
nothing is saved. The callback has also just set each CPU's own slot
correctly, so the loop only overwrites it, and does so while that CPU may
still be running in the address space.
Remove it. The shootdown shortcuts in flush_tlb_mm(), flush_tlb_page()
and flush_icache_user_page() clear remote slots too; once the next patch
removes those as well, every runtime update of mm->context[cpu] is made by
CPU cpu itself. That patch relies on the invariant when it reads those
slots to decide whether a shootdown can be skipped, so this one is tagged
for stable as its prerequisite.
Fixes: dd5712f3379c ("alpha: fix user-space corruption during memory compaction")
Cc: <stable@vger.kernel.org>
Signed-off-by: Magnus Lindholm <linmag7@gmail.com>
---
arch/alpha/mm/tlbflush.c | 16 ----------------
1 file changed, 16 deletions(-)
diff --git a/arch/alpha/mm/tlbflush.c b/arch/alpha/mm/tlbflush.c
index ccbc317b9a34..239c72b8a741 100644
--- a/arch/alpha/mm/tlbflush.c
+++ b/arch/alpha/mm/tlbflush.c
@@ -90,22 +90,6 @@ void migrate_flush_tlb_page(struct vm_area_struct *vma, unsigned long addr)
*/
preempt_disable();
on_each_cpu(ipi_flush_mm_and_page, &d, 1);
-
- /*
- * mimic flush_tlb_mm()'s mm_users<=1 optimization.
- */
- if (atomic_read(&mm->mm_users) <= 1) {
-
- int cpu, this_cpu;
- this_cpu = smp_processor_id();
-
- for (cpu = 0; cpu < NR_CPUS; cpu++) {
- if (!cpu_online(cpu) || cpu == this_cpu)
- continue;
- if (READ_ONCE(mm->context[cpu]))
- WRITE_ONCE(mm->context[cpu], 0);
- }
- }
preempt_enable();
}
--
2.43.0
^ permalink raw reply [flat|nested] 7+ messages in thread* Re: [PATCH v2 2/3] alpha: do not clear remote MMU contexts in migrate_flush_tlb_page()
2026-10-08 19:55 ` [PATCH v2 2/3] alpha: do not clear remote MMU contexts in migrate_flush_tlb_page() Magnus Lindholm
@ 2026-10-08 21:19 ` Matt Turner
0 siblings, 0 replies; 7+ messages in thread
From: Matt Turner @ 2026-10-08 21:19 UTC (permalink / raw)
To: Magnus Lindholm; +Cc: richard.henderson, linux-kernel, linux-alpha, stable
On Thu, Oct 8, 2026 at 3:56 PM Magnus Lindholm <linmag7@gmail.com> wrote:
>
> After the on_each_cpu() rendezvous, migrate_flush_tlb_page() walks the
> other CPUs and zeroes their mm->context[cpu] when mm_users is at most one,
> described as mimicking flush_tlb_mm()'s mm_users<=1 optimization.
>
> It is not one. flush_tlb_mm() tests mm_users before deciding whether to
> send the IPIs; here every CPU has already been visited and waited for, so
> nothing is saved. The callback has also just set each CPU's own slot
> correctly, so the loop only overwrites it, and does so while that CPU may
> still be running in the address space.
>
> Remove it. The shootdown shortcuts in flush_tlb_mm(), flush_tlb_page()
> and flush_icache_user_page() clear remote slots too; once the next patch
> removes those as well, every runtime update of mm->context[cpu] is made by
> CPU cpu itself. That patch relies on the invariant when it reads those
> slots to decide whether a shootdown can be skipped, so this one is tagged
> for stable as its prerequisite.
>
> Fixes: dd5712f3379c ("alpha: fix user-space corruption during memory compaction")
> Cc: <stable@vger.kernel.org>
> Signed-off-by: Magnus Lindholm <linmag7@gmail.com>
Reviewed-by: Matt Turner <mattst88@gmail.com>
^ permalink raw reply [flat|nested] 7+ messages in thread
* [PATCH v2 3/3] alpha: do not skip TLB shootdown IPIs for kthread-borrowed mms
2026-10-08 19:55 [PATCH v2 0/3] alpha: load the MMU context on a direct mm switch Magnus Lindholm
2026-10-08 19:55 ` [PATCH v2 1/3] alpha: load the MMU context when switch_mm() switches the current task Magnus Lindholm
2026-10-08 19:55 ` [PATCH v2 2/3] alpha: do not clear remote MMU contexts in migrate_flush_tlb_page() Magnus Lindholm
@ 2026-10-08 19:55 ` Magnus Lindholm
2026-10-08 21:23 ` Matt Turner
2 siblings, 1 reply; 7+ messages in thread
From: Magnus Lindholm @ 2026-10-08 19:55 UTC (permalink / raw)
To: richard.henderson, mattst88, linux-kernel, linux-alpha; +Cc: linmag7, stable
flush_tlb_mm(), flush_tlb_page() and flush_icache_user_page() skip the
shootdown IPI when mm_users <= 1, on the assumption that no other CPU can
be running the mm. A task that borrows an mm through kthread_use_mm()
takes mmgrab() rather than mmget(), so it never appears in mm_users, and
kthread_use_mm() may be handed an mm that is already the caller's
active_mm and loaded on that CPU. Such a CPU was not only left without
an IPI, its mm->context[cpu] was cleared from under it.
mm->context[cpu] is already the record of which CPUs hold an ASN for the
mm. Test it rather than clearing it: take the shortcut only when no
other CPU has one, and otherwise fall through to the IPI, which
invalidates those CPUs properly. A CPU that merely ran the mm in the
past also holds a context and now costs an IPI, which errs in the safe
direction.
Those historical contexts do not cost an IPI forever. When the shootdown
reaches a CPU that no longer has the mm active, ipi_flush_tlb_mm() takes
the flush_tlb_other() path, which clears that CPU's own slot. Once the
stale slots are gone the shortcut is available again, as long as mm_users
stays at most one and no other CPU has since taken or kept a context. A
CPU that does have the mm active keeps its slot, and that is the CPU the
IPI is there for.
Together with the preceding patch, dropping these clearing loops leaves
every runtime write to mm->context[] targeting the writing CPU's own slot,
so the array becomes a lockless publication of which CPUs may be using the
mm, with a single writer per slot. Mark the two stores that are now
observed from other CPUs; flush_tlb_other() already uses WRITE_ONCE().
The uniprocessor flush_icache_user_page() is left alone, as nothing reads
another CPU's slot there, and init_new_context() runs before the mm is
shared.
The reads also need ordering against the changes that led to the flush.
A CPU that publishes a context is in turn ordered before it goes on to
access the mm, by two different barriers: kthread_use_mm() issues one
through mmdrop_lazy_tlb() for the initial direct switch, and if the task
later migrates, the context published from ev5_switch_mm() is followed by
the scheduler's post-switch barrier, on alpha the mb() in
arch_spin_unlock() from finish_lock_switch(), as described in
Documentation/scheduler/membarrier.rst. So a CPU either already holds a
context and is seen here, or it allocates a fresh one before going on to
use the mm, and a fresh ASN carries nothing over from the previous
context.
This depends on the preceding patch: while migrate_flush_tlb_page() still
zeroes remote slots, a CPU running a borrowed mm can be cleared out of the
array and skipped by exactly the test added here.
Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2")
Cc: <stable@vger.kernel.org>
Signed-off-by: Magnus Lindholm <linmag7@gmail.com>
---
arch/alpha/include/asm/mmu_context.h | 2 +-
arch/alpha/kernel/smp.c | 48 ++++++++++++++--------------
arch/alpha/mm/fault.c | 2 +-
3 files changed, 26 insertions(+), 26 deletions(-)
diff --git a/arch/alpha/include/asm/mmu_context.h b/arch/alpha/include/asm/mmu_context.h
index 6ecc8a0676cb..9cbcc4aa62e3 100644
--- a/arch/alpha/include/asm/mmu_context.h
+++ b/arch/alpha/include/asm/mmu_context.h
@@ -158,7 +158,7 @@ ev5_switch_mm(struct mm_struct *prev_mm, struct mm_struct *next_mm,
mmc = next_mm->context[cpu];
if ((mmc ^ asn) & ~HARDWARE_ASN_MASK) {
mmc = __get_new_mm_context(next_mm, cpu);
- next_mm->context[cpu] = mmc;
+ WRITE_ONCE(next_mm->context[cpu], mmc);
}
#ifdef CONFIG_SMP
else
diff --git a/arch/alpha/kernel/smp.c b/arch/alpha/kernel/smp.c
index e21bc3920bec..c4d12d8312a7 100644
--- a/arch/alpha/kernel/smp.c
+++ b/arch/alpha/kernel/smp.c
@@ -631,6 +631,24 @@ ipi_flush_tlb_mm(void *x)
flush_tlb_other(mm);
}
+/* True if a CPU other than this one holds an ASN for MM. */
+static bool
+mm_context_elsewhere(struct mm_struct *mm)
+{
+ int cpu, this_cpu = smp_processor_id();
+
+ /* Pairs with the barrier the publishing CPU issues before using MM. */
+ smp_mb();
+
+ for_each_online_cpu(cpu) {
+ if (cpu == this_cpu)
+ continue;
+ if (READ_ONCE(mm->context[cpu]))
+ return true;
+ }
+ return false;
+}
+
void
flush_tlb_mm(struct mm_struct *mm)
{
@@ -638,14 +656,8 @@ flush_tlb_mm(struct mm_struct *mm)
if (mm == current->active_mm) {
flush_tlb_current(mm);
- if (atomic_read(&mm->mm_users) <= 1) {
- int cpu, this_cpu = smp_processor_id();
- for (cpu = 0; cpu < NR_CPUS; cpu++) {
- if (!cpu_online(cpu) || cpu == this_cpu)
- continue;
- if (mm->context[cpu])
- mm->context[cpu] = 0;
- }
+ if (atomic_read(&mm->mm_users) <= 1 &&
+ !mm_context_elsewhere(mm)) {
preempt_enable();
return;
}
@@ -690,14 +702,8 @@ flush_tlb_page(struct vm_area_struct *vma, unsigned long addr)
/* As in ipi_flush_tlb_page(): a targeted tbi() needs MM current. */
if (mm == current->mm) {
flush_tlb_current_page(mm, vma, addr);
- if (atomic_read(&mm->mm_users) <= 1) {
- int cpu, this_cpu = smp_processor_id();
- for (cpu = 0; cpu < NR_CPUS; cpu++) {
- if (!cpu_online(cpu) || cpu == this_cpu)
- continue;
- if (mm->context[cpu])
- mm->context[cpu] = 0;
- }
+ if (atomic_read(&mm->mm_users) <= 1 &&
+ !mm_context_elsewhere(mm)) {
preempt_enable();
return;
}
@@ -747,14 +753,8 @@ flush_icache_user_page(struct vm_area_struct *vma, struct page *page,
if (mm == current->active_mm) {
__load_new_mm_context(mm);
- if (atomic_read(&mm->mm_users) <= 1) {
- int cpu, this_cpu = smp_processor_id();
- for (cpu = 0; cpu < NR_CPUS; cpu++) {
- if (!cpu_online(cpu) || cpu == this_cpu)
- continue;
- if (mm->context[cpu])
- mm->context[cpu] = 0;
- }
+ if (atomic_read(&mm->mm_users) <= 1 &&
+ !mm_context_elsewhere(mm)) {
preempt_enable();
return;
}
diff --git a/arch/alpha/mm/fault.c b/arch/alpha/mm/fault.c
index a9816bbc9f34..7bbba9010dac 100644
--- a/arch/alpha/mm/fault.c
+++ b/arch/alpha/mm/fault.c
@@ -45,7 +45,7 @@ __load_new_mm_context(struct mm_struct *next_mm)
struct pcb_struct *pcb;
mmc = __get_new_mm_context(next_mm, smp_processor_id());
- next_mm->context[smp_processor_id()] = mmc;
+ WRITE_ONCE(next_mm->context[smp_processor_id()], mmc);
pcb = ¤t_thread_info()->pcb;
pcb->asn = mmc & HARDWARE_ASN_MASK;
--
2.43.0
^ permalink raw reply [flat|nested] 7+ messages in thread* Re: [PATCH v2 3/3] alpha: do not skip TLB shootdown IPIs for kthread-borrowed mms
2026-10-08 19:55 ` [PATCH v2 3/3] alpha: do not skip TLB shootdown IPIs for kthread-borrowed mms Magnus Lindholm
@ 2026-10-08 21:23 ` Matt Turner
0 siblings, 0 replies; 7+ messages in thread
From: Matt Turner @ 2026-10-08 21:23 UTC (permalink / raw)
To: Magnus Lindholm; +Cc: richard.henderson, linux-kernel, linux-alpha, stable
On Thu, Oct 08, 2026 at 09:55:18PM +0200, Magnus Lindholm wrote:
> A
> CPU that does have the mm active keeps its slot, and that is the CPU the
> IPI is there for.
That includes a CPU that only holds the mm lazily, idle since the task
migrated away. ipi_flush_tlb_mm() and ipi_flush_icache_page() test
active_mm, so it takes flush_tlb_current() and publishes a new ASN
every time. It stays visible, and each flush_tlb_mm() from the
single-threaded owner sends IPIs again until that CPU runs some other
user task. munmap(), mprotect(), fork() and exit all go through there.
Testing current->mm in those two handlers, as ipi_flush_tlb_page() now
does, would send the lazy CPU to flush_tlb_other() instead. I have not
measured this.
> A CPU that publishes a context is in turn ordered before it goes on to
> access the mm, by two different barriers:
check_mmu_context() is a third publisher. Its __load_new_mm_context()
takes a slot from zero to nonzero after finish_lock_switch() has
dropped the rq lock, and nothing orders that store before the return
to user space. An smp_mb() after the WRITE_ONCE() in
__load_new_mm_context() would cover every caller.
> So a CPU either already holds a
> context and is seen here, or it allocates a fresh one before going on to
> use the mm, and a fresh ASN carries nothing over from the previous
> context.
There is one way for a CPU to run the mm with its slot at zero. When
ev5_switch_mm() allocates a new context it does not set need_new_asn.
With the stale-TLB series, finish_task_switch() enables interrupts
before check_mmu_context() runs. An IPI for the mm taken there sees
asn_locked() and calls flush_tlb_other(), which zeroes the slot.
check_mmu_context() then finds need_new_asn clear and does not reload,
and the task returns to user space on a live ASN that
mm_context_elsewhere() cannot see.
Setting need_new_asn in both branches of ev5_switch_mm() closes it,
since check_mmu_context() only reloads when the slot is zero.
^ permalink raw reply [flat|nested] 7+ messages in thread
end of thread, other threads:[~2026-10-08 21:25 UTC | newest]
Thread overview: 7+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-08 19:55 [PATCH v2 0/3] alpha: load the MMU context on a direct mm switch Magnus Lindholm
2026-10-08 19:55 ` [PATCH v2 1/3] alpha: load the MMU context when switch_mm() switches the current task Magnus Lindholm
2026-10-08 21:25 ` Matt Turner
2026-10-08 19:55 ` [PATCH v2 2/3] alpha: do not clear remote MMU contexts in migrate_flush_tlb_page() Magnus Lindholm
2026-10-08 21:19 ` Matt Turner
2026-10-08 19:55 ` [PATCH v2 3/3] alpha: do not skip TLB shootdown IPIs for kthread-borrowed mms Magnus Lindholm
2026-10-08 21:23 ` Matt Turner
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®