mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [RFC PATCH v2 0/3] locking, mm: Add atomic allocator trylocks on RT
@ 2026-10-05  7:06 Karl Mehltretter
  2026-10-05  7:06 ` [RFC PATCH v2 1/3] locking/rtmutex: Support atomic PREEMPT_RT spin trylocks Karl Mehltretter
                   ` (2 more replies)
  0 siblings, 3 replies; 5+ messages in thread
From: Karl Mehltretter @ 2026-10-05  7:06 UTC (permalink / raw)
  To: Peter Zijlstra, Thomas Gleixner, Sebastian Andrzej Siewior,
	Andrew Morton, Vlastimil Babka, Harry Yoo, Alexei Starovoitov
  Cc: Karl Mehltretter, Ingo Molnar, Will Deacon, Boqun Feng,
	Waiman Long, Jonathan Corbet, David Hildenbrand, Johannes Weiner,
	Shakeel Butt, David Stevens, Daniel Borkmann, Andrii Nakryiko,
	Martin KaFai Lau, Shuah Khan, Amery Hung, Swaraj Gaikwad,
	Clark Williams, Steven Rostedt, linux-kernel, linux-doc,
	linux-mm, linux-rt-devel, bpf, linux-kselftest, cgroups

On PREEMPT_RT, a BPF task-storage program attached to sched_waking
can deadlock when kmalloc_nolock() obtains an rtmutex-backed allocator
spinlock while try_to_wake_up() holds p->pi_lock. Releasing the allocator
lock can enter priority-inheritance or wakeup code and re-enter scheduler
locking.

The first RFC [1] rejected every non-preemptible caller. That prevents the
deadlock, but it also rejects BPF arena allocation and faults under the
arena's ordinary raw lock.

This RFC instead adds an atomic owner state for bounded PREEMPT_RT spinlock
trylocks. Atomic acquisition succeeds only from the completely free state.
A regular waiter sets HAS_WAITERS before waiting, which prevents a later
atomic owner from barging. Atomic release preserves HAS_WAITERS and does
not enter priority inheritance or wake a task. Preemption and local
interrupts remain disabled for the atomic-owner section.

The design tradeoff is that a regular waiter cannot boost an atomic owner
and spins with interrupts disabled until the bounded allocator section
finishes. I would value locking review of whether that owner state and
handoff are acceptable, or whether the no-lock allocator should instead
fail in these contexts.

Patch 2 uses the new operation for global SLUB and page-allocator locks. It
avoids regular per-CPU RT local-lock slow paths and reuses centralized
objcg credit when the per-CPU stock is unavailable. Patch 3 adds a BPF
selftest for task-storage allocation from hrtimer_start while the hrtimer
base raw lock is held.

The series has one prerequisite, recorded by prerequisite-patch-id in this
cover letter:

  mm/page_alloc: skip shuffling and reporting for no-lock frees

That independent fix has been posted as a normal patch [2]. It keeps a
successful no-lock page free out of allocator shuffling and page-reporting
notification. It is separate because the issue begins with the v6.15
free_pages_nolock() API rather than the v7.0 slab regression addressed by
patch 2.

Patch 2 should also be evaluated with David Stevens's pending memory.high
deferral fix [3]. There is no build dependency, but bypassing the per-CPU
stock can make a no-lock charge reach that pre-existing schedule_work()
hazard more often.

The pre-rebase version of these atomic-owner changes passed four-vCPU
x86-64 PREEMPT_RT QEMU in release and lockdep/debug-rtmutex builds. Tests
completed 5,000 forced waiter handoffs without barging, kept asynchronous
IPIs out of atomic-owner sections and passed the BPF hrtimer workload.

After rebasing onto current mainline and the prerequisite, the affected
locking and MM objects build with PREEMPT_RT and lockdep. I have not
repeated the runtime campaigns for this RFC rebase.

If this direction is accepted, patches 1 and 2 would need joint stable
backports for v7.0 and later.

Changes since the RFC v1:

  - replace the blanket context rejection with an atomic rtmutex owner
  - preserve local IRQ state across a successful atomic trylock
  - cover the global slab, page allocator and memcg-cache paths
  - bound shared objcg credit when the per-CPU stock is skipped
  - add forced-handoff, caller-attribution and BPF hrtimer tests
  - keep the independent no-lock page-free fix as a prerequisite

[1] https://lore.kernel.org/r/20260919171443.90512-1-kmehltretter@gmail.com
[2] https://lore.kernel.org/r/20261005063515.6312-1-kmehltretter@gmail.com
[3] https://lore.kernel.org/r/20260904173145.2028377-1-stevensd@google.com

Karl Mehltretter (3):
  locking/rtmutex: Support atomic PREEMPT_RT spin trylocks
  mm: use atomic RT trylocks for no-lock allocation
  selftests/bpf: exercise task storage from hrtimer_start

 Documentation/locking/rt-mutex.rst            |  40 ++--
 include/linux/rtmutex.h                       |   4 +-
 include/linux/spinlock.h                      |   3 +
 include/linux/spinlock_rt.h                   |  28 ++-
 kernel/locking/rtmutex.c                      | 211 ++++++++++--------
 kernel/locking/spinlock_rt.c                  |  66 +++++-
 mm/internal.h                                 |  27 ++-
 mm/memcontrol.c                               |  90 ++++++--
 mm/page_alloc.c                               |  33 ++-
 mm/slub.c                                     |  36 +--
 .../bpf/prog_tests/task_storage_hrtimer.c     |  50 +++++
 .../bpf/progs/task_storage_hrtimer.c          |  48 ++++
 12 files changed, 479 insertions(+), 157 deletions(-)
 create mode 100644 tools/testing/selftests/bpf/prog_tests/task_storage_hrtimer.c
 create mode 100644 tools/testing/selftests/bpf/progs/task_storage_hrtimer.c


base-commit: e767a4ea70a3992c37ed604157d32f0dfbf9b1e3
prerequisite-patch-id: 33838040c410e5de0aef855a2719a092b561a5c4
-- 
2.53.0


^ permalink raw reply	[flat|nested] 5+ messages in thread

* [RFC PATCH v2 1/3] locking/rtmutex: Support atomic PREEMPT_RT spin trylocks
  2026-10-05  7:06 [RFC PATCH v2 0/3] locking, mm: Add atomic allocator trylocks on RT Karl Mehltretter
@ 2026-10-05  7:06 ` Karl Mehltretter
  2026-10-05  7:06 ` [RFC PATCH v2 2/3] mm: use atomic RT trylocks for no-lock allocation Karl Mehltretter
  2026-10-05  7:06 ` [RFC PATCH v2 3/3] selftests/bpf: exercise task storage from hrtimer_start Karl Mehltretter
  2 siblings, 0 replies; 5+ messages in thread
From: Karl Mehltretter @ 2026-10-05  7:06 UTC (permalink / raw)
  To: Peter Zijlstra, Thomas Gleixner, Sebastian Andrzej Siewior,
	Andrew Morton, Vlastimil Babka, Harry Yoo, Alexei Starovoitov
  Cc: Karl Mehltretter, Ingo Molnar, Will Deacon, Boqun Feng,
	Waiman Long, Jonathan Corbet, David Hildenbrand, Johannes Weiner,
	Shakeel Butt, David Stevens, Daniel Borkmann, Andrii Nakryiko,
	Martin KaFai Lau, Shuah Khan, Amery Hung, Swaraj Gaikwad,
	Clark Williams, Steven Rostedt, linux-kernel, linux-doc,
	linux-mm, linux-rt-devel, bpf, linux-kselftest, cgroups

No-lock allocation may run with preemption disabled or while a raw lock is
held.  On PREEMPT_RT, allocator spinlocks are backed by rtmutexes.  A
successful spin_trylock() can therefore enter priority inheritance or
wakeup code when the lock is released, which is not safe for these callers.

Add spin_trylock_nolock_irqsave() for bounded sections which cannot enter
the rtmutex slow path.  Preemptible callers retain the regular RT trylock.
For a non-preemptible caller, acquire only an uncontended lock, mark the
owner as atomic, and keep preemption and local interrupts disabled until
release.

A regular waiter sets HAS_WAITERS before waiting for an atomic owner.  This
prevents another atomic owner from barging ahead.  The atomic owner
releases to NULL | HAS_WAITERS without entering priority inheritance or
waking a task.  Disabling interrupts while the atomic owner holds the lock
also prevents a waiter from spinning through hardirq work on the owner's
CPU.

The protected section must not block and must remain bounded because a
regular waiter spins with interrupts disabled and cannot boost an atomic
owner.  Document that a successful acquisition must be released with
spin_unlock_irqrestore(), and keep the MM-only entry point unexported.

Pass the real unlock call site to lockdep for both regular and irqrestore
unlocks.  Also make the debug slow unlock publish the free state with
cmpxchg after dropping wait_lock, so an atomic owner cannot reuse an
embedded lock before the old unlock has finished with it.  Use the same
slow-unlock helper in debug and non-debug builds.

Assisted-by: LLM
Signed-off-by: Karl Mehltretter <kmehltretter@gmail.com>
---
 Documentation/locking/rt-mutex.rst |  40 ++++--
 include/linux/rtmutex.h            |   4 +-
 include/linux/spinlock.h           |   3 +
 include/linux/spinlock_rt.h        |  28 +++-
 kernel/locking/rtmutex.c           | 211 +++++++++++++++++------------
 kernel/locking/spinlock_rt.c       |  66 ++++++++-
 6 files changed, 239 insertions(+), 113 deletions(-)

diff --git a/Documentation/locking/rt-mutex.rst b/Documentation/locking/rt-mutex.rst
index 3b5097a380e6..53011c917710 100644
--- a/Documentation/locking/rt-mutex.rst
+++ b/Documentation/locking/rt-mutex.rst
@@ -46,21 +46,26 @@ is used]
 The state of the rt-mutex is tracked via the owner field of the rt-mutex
 structure:
 
-lock->owner holds the task_struct pointer of the owner. Bit 0 is used to
-keep track of the "lock has waiters" state:
-
- ============ ======= ================================================
- owner        bit0    Notes
- ============ ======= ================================================
- NULL         0       lock is free (fast acquire possible)
- NULL         1       lock is free and has waiters and the top waiter
-		      is going to take the lock [1]_
- taskpointer  0       lock is held (fast release possible)
- taskpointer  1       lock is held and has waiters [2]_
- ============ ======= ================================================
-
-The fast atomic compare exchange based acquire and release is only
-possible when bit 0 of lock->owner is 0.
+lock->owner holds the task_struct pointer of the owner. Bit 0 tracks the
+"lock has waiters" state. Bit 1 marks an atomic owner used by PREEMPT_RT
+spinlock trylocks whose callers cannot enter the priority inheritance or
+wakeup paths:
+
+ ============ ======= ======= ================================================
+ owner        bit1    bit0    Notes
+ ============ ======= ======= ================================================
+ NULL         0       0       lock is free (fast acquire possible)
+ NULL         0       1       lock is free and has waiters and the top waiter
+			      is going to take the lock [1]_
+ taskpointer  0       0       lock is held (fast release possible)
+ taskpointer  0       1       lock is held and has waiters [2]_
+ taskpointer  1       0       lock is held by an atomic owner
+ taskpointer  1       1       lock is held by an atomic owner while a slow
+			      path excludes new atomic owners [3]_
+ ============ ======= ======= ================================================
+
+The regular fast atomic compare exchange based acquire and release is only
+possible when both flag bits of lock->owner are clear.
 
 .. [1] It also can be a transitional state when grabbing the lock
        with ->wait_lock is held. To prevent any fast path cmpxchg to the lock,
@@ -72,6 +77,11 @@ possible when bit 0 of lock->owner is 0.
        To prevent a cmpxchg of the owner releasing the lock, we need to
        set this bit before looking at the lock.
 
+.. [3] The slow path sets bit 0 before waiting for the atomic owner. The
+       atomic owner then releases to ``NULL | HAS_WAITERS`` without entering
+       priority inheritance or wakeup handling. This lets the slow path take
+       the lock and prevents another atomic owner from barging ahead of it.
+
 BTW, there is still technically a "Pending Owner", it's just not called
 that anymore. The pending owner happens to be the top_waiter of a lock
 that has no owner and has been woken up to grab the lock.
diff --git a/include/linux/rtmutex.h b/include/linux/rtmutex.h
index 9e1f012f89db..e889a6d1d0be 100644
--- a/include/linux/rtmutex.h
+++ b/include/linux/rtmutex.h
@@ -46,12 +46,14 @@ static inline bool rt_mutex_base_is_locked(struct rt_mutex_base *lock)
 
 #ifdef CONFIG_RT_MUTEXES
 #define RT_MUTEX_HAS_WAITERS	1UL
+#define RT_MUTEX_OWNER_ATOMIC	2UL
+#define RT_MUTEX_OWNER_MASK	(RT_MUTEX_HAS_WAITERS | RT_MUTEX_OWNER_ATOMIC)
 
 static inline struct task_struct *rt_mutex_owner(struct rt_mutex_base *lock)
 {
 	unsigned long owner = (unsigned long) data_race(READ_ONCE(lock->owner));
 
-	return (struct task_struct *) (owner & ~RT_MUTEX_HAS_WAITERS);
+	return (struct task_struct *)(owner & ~RT_MUTEX_OWNER_MASK);
 }
 #endif
 extern void rt_mutex_base_init(struct rt_mutex_base *rtb);
diff --git a/include/linux/spinlock.h b/include/linux/spinlock.h
index 3d405cc4c121..e1f4804efce5 100644
--- a/include/linux/spinlock.h
+++ b/include/linux/spinlock.h
@@ -444,6 +444,9 @@ static __always_inline bool _spin_trylock_irqsave(spinlock_t *lock, unsigned lon
 }
 #define spin_trylock_irqsave(lock, flags) _spin_trylock_irqsave(lock, &(flags))
 
+#define spin_trylock_nolock_irqsave(lock, flags) \
+	spin_trylock_irqsave(lock, flags)
+
 static __always_inline int spin_trylock_irq_disable(spinlock_t *lock)
 	__cond_acquires(true, lock) __no_context_analysis
 {
diff --git a/include/linux/spinlock_rt.h b/include/linux/spinlock_rt.h
index 560d06384e0c..b913b8b57b04 100644
--- a/include/linux/spinlock_rt.h
+++ b/include/linux/spinlock_rt.h
@@ -35,9 +35,14 @@ extern void rt_spin_lock(spinlock_t *lock) __acquires(lock);
 extern void rt_spin_lock_nested(spinlock_t *lock, int subclass)	__acquires(lock);
 extern void rt_spin_lock_nest_lock(spinlock_t *lock, struct lockdep_map *nest_lock) __acquires(lock);
 extern void rt_spin_unlock(spinlock_t *lock)	__releases(lock);
+extern void rt_spin_unlock_irqrestore(spinlock_t *lock, unsigned long flags)
+	__releases(lock);
 extern void rt_spin_lock_unlock(spinlock_t *lock);
 extern int rt_spin_trylock_bh(spinlock_t *lock) __cond_acquires(true, lock);
 extern int rt_spin_trylock(spinlock_t *lock) __cond_acquires(true, lock);
+extern int rt_spin_trylock_nolock_irqsave(spinlock_t *lock,
+					  unsigned long *flags)
+	__cond_acquires(true, lock);
 
 static __always_inline void spin_lock(spinlock_t *lock)
 	__acquires(lock)
@@ -138,7 +143,7 @@ static __always_inline void spin_unlock_irqrestore(spinlock_t *lock,
 						   unsigned long flags)
 	__releases(lock)
 {
-	rt_spin_unlock(lock);
+	rt_spin_unlock_irqrestore(lock, flags);
 }
 
 #define spin_trylock(lock)	rt_spin_trylock(lock)
@@ -161,6 +166,27 @@ static __always_inline bool _spin_trylock_irqsave(spinlock_t *lock, unsigned lon
 }
 #define spin_trylock_irqsave(lock, flags) _spin_trylock_irqsave(lock, &(flags))
 
+/*
+ * For bounded no-lock allocator sections. A non-preemptible caller can
+ * acquire only an uncontended lock. Success creates an atomic owner, leaves
+ * local interrupts disabled, and adds a preemption-disable level. NMI and
+ * hard interrupt callers fail.
+ *
+ * The section must not block. A regular waiter spins with interrupts disabled
+ * and cannot boost the atomic owner. Pair every successful call with
+ * spin_unlock_irqrestore(lock, flags); spin_unlock() does not restore the
+ * interrupt state of an atomic owner.
+ */
+static __always_inline bool
+_spin_trylock_nolock_irqsave(spinlock_t *lock, unsigned long *flags)
+	__cond_acquires(true, lock)
+{
+	return rt_spin_trylock_nolock_irqsave(lock, flags);
+}
+
+#define spin_trylock_nolock_irqsave(lock, flags) \
+	_spin_trylock_nolock_irqsave(lock, &(flags))
+
 #define spin_is_contended(lock)		(((void)(lock), 0))
 
 static inline int spin_is_locked(spinlock_t *lock)
diff --git a/kernel/locking/rtmutex.c b/kernel/locking/rtmutex.c
index 4728631ae719..342028cd682f 100644
--- a/kernel/locking/rtmutex.c
+++ b/kernel/locking/rtmutex.c
@@ -68,18 +68,23 @@ static inline int __ww_mutex_check_kill(struct rt_mutex *lock,
 /*
  * lock->owner state tracking:
  *
- * lock->owner holds the task_struct pointer of the owner. Bit 0
- * is used to keep track of the "lock has waiters" state.
+ * lock->owner holds the task_struct pointer of the owner. Bit 0 is used to
+ * keep track of the "lock has waiters" state. Bit 1 identifies an atomic
+ * owner which cannot participate in priority inheritance. Atomic ownership
+ * is used only by PREEMPT_RT spinlock trylocks whose caller cannot block.
  *
- * owner	bit0
- * NULL		0	lock is free (fast acquire possible)
- * NULL		1	lock is free and has waiters and the top waiter
- *				is going to take the lock*
- * taskpointer	0	lock is held (fast release possible)
- * taskpointer	1	lock is held and has waiters**
+ * owner       bit1 bit0
+ * NULL        0    0    lock is free (fast acquire possible)
+ * NULL        0    1    lock is free and has waiters; the top waiter
+ *                       is going to take the lock*
+ * taskpointer 0    0    lock is held (fast release possible)
+ * taskpointer 0    1    lock is held and has waiters**
+ * taskpointer 1    0    lock is held by an atomic owner
+ * taskpointer 1    1    lock is held by an atomic owner while a slow path
+ *                       excludes further atomic acquisitions***
  *
- * The fast atomic compare exchange based acquire and release is only
- * possible when bit 0 of lock->owner is 0.
+ * The regular fast atomic compare exchange based acquire and release is only
+ * possible when both flag bits in lock->owner are 0.
  *
  * (*) It also can be a transitional state when grabbing the lock
  * with ->wait_lock is held. To prevent any fast path cmpxchg to the lock,
@@ -90,8 +95,50 @@ static inline int __ww_mutex_check_kill(struct rt_mutex *lock,
  * waiters. This can happen when grabbing the lock in the slow path.
  * To prevent a cmpxchg of the owner releasing the lock, we need to
  * set this bit before looking at the lock.
+ *
+ * (***) The slow path sets the waiters bit while holding ->wait_lock, then
+ * waits for the atomic owner to release the lock to NULL|HAS_WAITERS. The
+ * atomic owner never enters the PI machinery and can release without taking
+ * ->wait_lock or waking a task. It cannot be preempted while holding the lock.
  */
 
+static __always_inline bool
+rt_mutex_atomic_owner(struct rt_mutex_base *lock)
+{
+	return (unsigned long)READ_ONCE(lock->owner) & RT_MUTEX_OWNER_ATOMIC;
+}
+
+static __always_inline bool
+rt_mutex_atomic_try_acquire(struct rt_mutex_base *lock)
+{
+	struct task_struct *old = NULL;
+	struct task_struct *new = (struct task_struct *)
+		((unsigned long)current | RT_MUTEX_OWNER_ATOMIC);
+
+	return try_cmpxchg_acquire(&lock->owner, &old, new);
+}
+
+static __always_inline bool
+rt_mutex_atomic_release(struct rt_mutex_base *lock)
+{
+	struct task_struct *old = READ_ONCE(lock->owner);
+	struct task_struct *new;
+	unsigned long owner;
+
+	for (;;) {
+		owner = (unsigned long)old;
+		if (WARN_ON_ONCE((owner & ~RT_MUTEX_OWNER_MASK) !=
+				 (unsigned long)current) ||
+		    WARN_ON_ONCE(!(owner & RT_MUTEX_OWNER_ATOMIC)))
+			return false;
+
+		new = (struct task_struct *)(owner & RT_MUTEX_HAS_WAITERS);
+		if (try_cmpxchg_release(&lock->owner, &old, new))
+			return true;
+		cpu_relax();
+	}
+}
+
 static __always_inline struct task_struct *
 rt_mutex_owner_encode(struct rt_mutex_base *lock, struct task_struct *owner)
 	__must_hold(&lock->wait_lock)
@@ -214,6 +261,37 @@ fixup_rt_mutex_waiters(struct rt_mutex_base *lock, bool acquire_lock)
 	}
 }
 
+/*
+ * Callers hold ->wait_lock, which serializes slow-path updates. This cmpxchg
+ * also arbitrates with regular lockless fast-path owner transitions when
+ * enabled and with atomic owner transitions, which do not take ->wait_lock.
+ * Once HAS_WAITERS is set, no new lockless acquisition can succeed. If an
+ * atomic owner was already present, wait until it drops its task pointer while
+ * preserving HAS_WAITERS.
+ */
+static __always_inline void mark_rt_mutex_waiters(struct rt_mutex_base *lock)
+	__must_hold(&lock->wait_lock)
+{
+	unsigned long *p = (unsigned long *)&lock->owner;
+	unsigned long owner, new;
+
+	owner = READ_ONCE(*p);
+	for (;;) {
+		new = owner | RT_MUTEX_HAS_WAITERS;
+		if (try_cmpxchg_relaxed(p, &owner, new))
+			break;
+		cpu_relax();
+	}
+
+	/*
+	 * The cmpxchg above is relaxed to avoid back-to-back ACQUIRE operations
+	 * in the event of contention. Ensure the successful cmpxchg is visible.
+	 */
+	smp_mb__after_atomic();
+
+	smp_cond_load_relaxed(p, !(VAL & RT_MUTEX_OWNER_ATOMIC));
+}
+
 /*
  * We can speed up the acquire/release, if there's no debugging state to be
  * set up.
@@ -238,34 +316,45 @@ static __always_inline bool rt_mutex_cmpxchg_release(struct rt_mutex_base *lock,
 	return try_cmpxchg_release(&lock->owner, &old, new);
 }
 
-/*
- * Callers must hold the ->wait_lock -- which is the whole purpose as we force
- * all future threads that attempt to [Rmw] the lock to the slowpath. As such
- * relaxed semantics suffice.
- */
-static __always_inline void mark_rt_mutex_waiters(struct rt_mutex_base *lock)
+#else
+static __always_inline bool rt_mutex_cmpxchg_acquire(struct rt_mutex_base *lock,
+						     struct task_struct *old,
+						     struct task_struct *new)
 {
-	unsigned long *p = (unsigned long *) &lock->owner;
-	unsigned long owner, new;
+	return false;
+}
 
-	owner = READ_ONCE(*p);
-	do {
-		new = owner | RT_MUTEX_HAS_WAITERS;
-	} while (!try_cmpxchg_relaxed(p, &owner, new));
+static int __sched rt_mutex_slowtrylock(struct rt_mutex_base *lock);
 
+static __always_inline bool rt_mutex_try_acquire(struct rt_mutex_base *lock)
+{
 	/*
-	 * The cmpxchg loop above is relaxed to avoid back-to-back ACQUIRE
-	 * operations in the event of contention. Ensure the successful
-	 * cmpxchg is visible.
+	 * With debug enabled rt_mutex_cmpxchg trylock() will always fail.
+	 *
+	 * Avoid unconditionally taking the slow path by using
+	 * rt_mutex_slow_trylock() which is covered by the debug code and can
+	 * acquire a non-contended rtmutex.
 	 */
-	smp_mb__after_atomic();
+	return rt_mutex_slowtrylock(lock);
 }
 
+static __always_inline bool rt_mutex_cmpxchg_release(struct rt_mutex_base *lock,
+						     struct task_struct *old,
+						     struct task_struct *new)
+{
+	return false;
+}
+
+#endif
+
 /*
  * Safe fastpath aware unlock:
  * 1) Clear the waiters bit
  * 2) Drop lock->wait_lock
  * 3) Try to unlock the lock with cmpxchg
+ *
+ * Atomic trylock also bypasses wait_lock, so debug builds need the same
+ * cmpxchg arbitration even though the regular fast path is disabled.
  */
 static __always_inline bool unlock_rt_mutex_safe(struct rt_mutex_base *lock,
 						 unsigned long flags)
@@ -299,59 +388,9 @@ static __always_inline bool unlock_rt_mutex_safe(struct rt_mutex_base *lock,
 	 *					lock(wait_lock);
 	 *					acquire(lock);
 	 */
-	return rt_mutex_cmpxchg_release(lock, owner, NULL);
-}
-
-#else
-static __always_inline bool rt_mutex_cmpxchg_acquire(struct rt_mutex_base *lock,
-						     struct task_struct *old,
-						     struct task_struct *new)
-{
-	return false;
-
-}
-
-static int __sched rt_mutex_slowtrylock(struct rt_mutex_base *lock);
-
-static __always_inline bool rt_mutex_try_acquire(struct rt_mutex_base *lock)
-{
-	/*
-	 * With debug enabled rt_mutex_cmpxchg trylock() will always fail.
-	 *
-	 * Avoid unconditionally taking the slow path by using
-	 * rt_mutex_slow_trylock() which is covered by the debug code and can
-	 * acquire a non-contended rtmutex.
-	 */
-	return rt_mutex_slowtrylock(lock);
-}
-
-static __always_inline bool rt_mutex_cmpxchg_release(struct rt_mutex_base *lock,
-						     struct task_struct *old,
-						     struct task_struct *new)
-{
-	return false;
-}
-
-static __always_inline void mark_rt_mutex_waiters(struct rt_mutex_base *lock)
-	__must_hold(&lock->wait_lock)
-{
-	lock->owner = (struct task_struct *)
-			((unsigned long)lock->owner | RT_MUTEX_HAS_WAITERS);
+	return try_cmpxchg_release(&lock->owner, &owner, NULL);
 }
 
-/*
- * Simple slow path only version: lock->owner is protected by lock->wait_lock.
- */
-static __always_inline bool unlock_rt_mutex_safe(struct rt_mutex_base *lock,
-						 unsigned long flags)
-	__releases(lock->wait_lock)
-{
-	lock->owner = NULL;
-	raw_spin_unlock_irqrestore(&lock->wait_lock, flags);
-	return true;
-}
-#endif
-
 static __always_inline int __waiter_prio(struct task_struct *task)
 {
 	int prio = task->prio;
@@ -1432,9 +1471,8 @@ static void __sched rt_mutex_slowunlock(struct rt_mutex_base *lock)
 	debug_rt_mutex_unlock(lock);
 
 	/*
-	 * We must be careful here if the fast path is enabled. If we
-	 * have no waiters queued we cannot set owner to NULL here
-	 * because of:
+	 * If there are no waiters queued, owner still cannot be set to NULL
+	 * while wait_lock is held because a lockless acquisition can race:
 	 *
 	 * foo->lock->owner = NULL;
 	 *			rtmutex_lock(foo->lock);   <- fast path
@@ -1444,10 +1482,9 @@ static void __sched rt_mutex_slowunlock(struct rt_mutex_base *lock)
 	 *				kfree(foo);
 	 * raw_spin_unlock(foo->lock->wait_lock);
 	 *
-	 * So for the fastpath enabled kernel:
-	 *
-	 * Nothing can set the waiters bit as long as we hold
-	 * lock->wait_lock. So we do the following sequence:
+	 * The regular fast path can do this when enabled. Atomic trylock can do
+	 * this in debug builds too. Nothing can set the waiters bit as long as
+	 * wait_lock is held, so use the following sequence in either case:
 	 *
 	 *	owner = rt_mutex_owner(lock);
 	 *	clear_rt_mutex_waiters(lock);
@@ -1455,12 +1492,6 @@ static void __sched rt_mutex_slowunlock(struct rt_mutex_base *lock)
 	 *	if (cmpxchg(&lock->owner, owner, 0) == owner)
 	 *		return;
 	 *	goto retry;
-	 *
-	 * The fastpath disabled variant is simple as all access to
-	 * lock->owner is serialized by lock->wait_lock:
-	 *
-	 *	lock->owner = NULL;
-	 *	raw_spin_unlock(&lock->wait_lock);
 	 */
 	while (!rt_mutex_has_waiters(lock)) {
 		/* Drops lock->wait_lock ! */
diff --git a/kernel/locking/spinlock_rt.c b/kernel/locking/spinlock_rt.c
index 1d5e1b3c60bf..27f263296ac2 100644
--- a/kernel/locking/spinlock_rt.c
+++ b/kernel/locking/spinlock_rt.c
@@ -75,9 +75,22 @@ void __sched rt_spin_lock_nest_lock(spinlock_t *lock,
 EXPORT_SYMBOL(rt_spin_lock_nest_lock);
 #endif
 
-void __sched rt_spin_unlock(spinlock_t *lock) __releases(RCU)
+static __always_inline bool
+__rt_spin_unlock(spinlock_t *lock, unsigned long ip)
 {
-	spin_release(&lock->dep_map, _RET_IP_);
+	spin_release(&lock->dep_map, ip);
+
+	if (unlikely(rt_mutex_atomic_owner(&lock->lock))) {
+		/*
+		 * Atomic trylock holders have preemption disabled and never
+		 * entered the rtmutex PI machinery. Leave a marked slow path
+		 * the transitional HAS_WAITERS state and let it acquire the lock.
+		 */
+		WARN_ON_ONCE(!rt_mutex_atomic_release(&lock->lock));
+		rcu_read_unlock();
+		return true;
+	}
+
 	migrate_enable();
 
 	if (unlikely(!rt_mutex_cmpxchg_release(&lock->lock, current, NULL)))
@@ -100,9 +113,26 @@ void __sched rt_spin_unlock(spinlock_t *lock) __releases(RCU)
 	 *			    UAF ->	  rt_mutex_cmpxchg_release(&p->lock.lock...)
 	 */
 	rcu_read_unlock();
+	return false;
+}
+
+void __sched rt_spin_unlock(spinlock_t *lock) __releases(RCU)
+{
+	if (__rt_spin_unlock(lock, _RET_IP_))
+		preempt_enable();
 }
 EXPORT_SYMBOL(rt_spin_unlock);
 
+void __sched rt_spin_unlock_irqrestore(spinlock_t *lock, unsigned long flags)
+	__releases(RCU)
+{
+	if (__rt_spin_unlock(lock, _RET_IP_)) {
+		local_irq_restore(flags);
+		preempt_enable();
+	}
+}
+EXPORT_SYMBOL(rt_spin_unlock_irqrestore);
+
 /*
  * Wait for the lock to get unlocked: instead of polling for an unlock
  * (like raw spinlocks do), lock and unlock, to force the kernel to
@@ -115,7 +145,8 @@ void __sched rt_spin_lock_unlock(spinlock_t *lock)
 }
 EXPORT_SYMBOL(rt_spin_lock_unlock);
 
-static __always_inline int __rt_spin_trylock(spinlock_t *lock)
+static __always_inline int
+__rt_spin_trylock(spinlock_t *lock, unsigned long ip)
 {
 	int ret = 1;
 
@@ -123,7 +154,7 @@ static __always_inline int __rt_spin_trylock(spinlock_t *lock)
 		ret = rt_mutex_slowtrylock(&lock->lock);
 
 	if (ret) {
-		spin_acquire(&lock->dep_map, 0, 1, _RET_IP_);
+		spin_acquire(&lock->dep_map, 0, 1, ip);
 		rcu_read_lock();
 		migrate_disable();
 	}
@@ -132,16 +163,39 @@ static __always_inline int __rt_spin_trylock(spinlock_t *lock)
 
 int __sched rt_spin_trylock(spinlock_t *lock)
 {
-	return __rt_spin_trylock(lock);
+	return __rt_spin_trylock(lock, _RET_IP_);
 }
 EXPORT_SYMBOL(rt_spin_trylock);
 
+int __sched rt_spin_trylock_nolock_irqsave(spinlock_t *lock,
+					   unsigned long *flags)
+{
+	if (unlikely(in_nmi() || in_hardirq()))
+		return 0;
+	if (preemptible()) {
+		*flags = 0;
+		return __rt_spin_trylock(lock, _RET_IP_);
+	}
+
+	local_irq_save(*flags);
+	preempt_disable();
+	if (!rt_mutex_atomic_try_acquire(&lock->lock)) {
+		preempt_enable();
+		local_irq_restore(*flags);
+		return 0;
+	}
+
+	spin_acquire(&lock->dep_map, 0, 1, _RET_IP_);
+	rcu_read_lock();
+	return 1;
+}
+
 int __sched rt_spin_trylock_bh(spinlock_t *lock)
 {
 	int ret;
 
 	local_bh_disable();
-	ret = __rt_spin_trylock(lock);
+	ret = __rt_spin_trylock(lock, _RET_IP_);
 	if (!ret)
 		local_bh_enable();
 	return ret;
-- 
2.53.0


^ permalink raw reply	[flat|nested] 5+ messages in thread

* [RFC PATCH v2 2/3] mm: use atomic RT trylocks for no-lock allocation
  2026-10-05  7:06 [RFC PATCH v2 0/3] locking, mm: Add atomic allocator trylocks on RT Karl Mehltretter
  2026-10-05  7:06 ` [RFC PATCH v2 1/3] locking/rtmutex: Support atomic PREEMPT_RT spin trylocks Karl Mehltretter
@ 2026-10-05  7:06 ` Karl Mehltretter
  2026-10-05  7:06 ` [RFC PATCH v2 3/3] selftests/bpf: exercise task storage from hrtimer_start Karl Mehltretter
  2 siblings, 0 replies; 5+ messages in thread
From: Karl Mehltretter @ 2026-10-05  7:06 UTC (permalink / raw)
  To: Peter Zijlstra, Thomas Gleixner, Sebastian Andrzej Siewior,
	Andrew Morton, Vlastimil Babka, Harry Yoo, Alexei Starovoitov
  Cc: Karl Mehltretter, Ingo Molnar, Will Deacon, Boqun Feng,
	Waiman Long, Jonathan Corbet, David Hildenbrand, Johannes Weiner,
	Shakeel Butt, David Stevens, Daniel Borkmann, Andrii Nakryiko,
	Martin KaFai Lau, Shuah Khan, Amery Hung, Swaraj Gaikwad,
	Clark Williams, Steven Rostedt, linux-kernel, linux-doc,
	linux-mm, linux-rt-devel, bpf, linux-kselftest, cgroups

Commit 073d5f156292 ("slab: simplify kmalloc_nolock()") allowed no-lock
slab allocation from preempt-disabled sections again on PREEMPT_RT. The
remaining global allocator locks are backed by rtmutexes. Releasing one
after a successful trylock can enter the rtmutex priority inheritance and
wakeup paths.

BPF task storage from sched_waking can reach this while try_to_wake_up()
holds p->pi_lock:

  try_to_wake_up()
    raw_spin_lock_irqsave(&p->pi_lock)
      bpf_task_storage_get()
        kmalloc_nolock()
          spin_trylock_irqsave(&n->list_lock)
          spin_unlock_irqrestore(&n->list_lock)
            rtmutex wakeup paths

The wakeup from the allocator unlock can re-enter scheduler locking and
deadlock.

Use spin_trylock_nolock_irqsave() for the global SLUB and page allocator
locks. Its atomic owner state avoids the rtmutex slow path for
non-preemptible RT callers. Preemptible callers keep regular RT spinlock
semantics.

Skip per-CPU allocator and memcg caches for non-preemptible RT callers.
Their regular RT local-lock slow paths can take current->pi_lock.

Skipping the per-CPU objcg stock on every accounted no-lock allocation
would otherwise charge a full page each time. Matching frees add prepaid
bytes to objcg->nr_charged_bytes, but the next allocation did not consume
them. Reuse and refill those centralized bytes with cmpxchg when the local
stock cannot be used, and return complete pages once the retained credit
crosses the existing thresholds.

The page allocator has accepted the same contexts since its no-lock
interface was introduced, so apply the atomic trylock rule to both slab
and page no-lock paths. The prerequisite page-free fix keeps successful
no-lock frees out of allocator shuffling and page-reporting notification.

Bypassing the local objcg stock can make a no-lock charge reach
memory.high handling more often. Apply this after David Stevens's pending
change which defers memory.high work from contexts where spinning is not
allowed. Otherwise the existing schedule_work() hazard remains. There is
no build dependency between the changes.

Fixes: 073d5f156292 ("slab: simplify kmalloc_nolock()")
Link: https://lore.kernel.org/r/20260919171443.90512-1-kmehltretter@gmail.com
Link: https://lore.kernel.org/r/20260904173145.2028377-1-stevensd@google.com
Assisted-by: LLM
Signed-off-by: Karl Mehltretter <kmehltretter@gmail.com>
---
 mm/internal.h   | 27 ++++++++++-----
 mm/memcontrol.c | 90 +++++++++++++++++++++++++++++++++++++++----------
 mm/page_alloc.c | 33 +++++++++++++++---
 mm/slub.c       | 36 +++++++++++++-------
 4 files changed, 142 insertions(+), 44 deletions(-)

diff --git a/mm/internal.h b/mm/internal.h
index 38b1165212c9..5b8cabaf86df 100644
--- a/mm/internal.h
+++ b/mm/internal.h
@@ -1631,18 +1631,27 @@ static inline void mm_prepare_for_swap_entries(struct mm_struct *mm)
 	}
 }
 
+/*
+ * On PREEMPT_RT, local_trylock() uses a regular rtmutex-backed spinlock.
+ * Its slow path can take current->pi_lock. Skip these caches when a no-lock
+ * caller is already non-preemptible.
+ */
+#define mm_local_trylock_nolock(lock)					\
+({									\
+	bool __locked = false;						\
+									\
+	if (!IS_ENABLED(CONFIG_PREEMPT_RT) || preemptible())		\
+		__locked = local_trylock(lock);				\
+	__locked;							\
+})
+
 static inline bool can_spin_trylock(void)
 {
 	/*
-	 * In PREEMPT_RT spin_trylock() will call raw_spin_lock() which is
-	 * unsafe in NMI. If spin_trylock() is called from hard IRQ the current
-	 * task may be waiting for one rt_spin_lock, but rt_spin_trylock() will
-	 * mark the task as the owner of another rt_spin_lock which will
-	 * confuse PI logic, so return immediately if called from hard IRQ or
-	 * NMI.
-	 *
-	 * Note, irqs_disabled() case is ok. spin_trylock() can be called
-	 * from raw_spin_lock_irqsave region.
+	 * PREEMPT_RT no-lock allocation has not been validated in NMI or hard
+	 * IRQ context. Other atomic contexts use
+	 * spin_trylock_nolock_irqsave(), which does not enter the rtmutex PI
+	 * machinery.
 	 */
 	if (IS_ENABLED(CONFIG_PREEMPT_RT) && (in_nmi() || in_hardirq()))
 		return false;
diff --git a/mm/memcontrol.c b/mm/memcontrol.c
index 856a7d07586c..03a83664f99b 100644
--- a/mm/memcontrol.c
+++ b/mm/memcontrol.c
@@ -2129,7 +2129,7 @@ static bool consume_stock(struct mem_cgroup *memcg, unsigned int nr_pages)
 	int i;
 
 	if (nr_pages > MEMCG_CHARGE_BATCH ||
-	    !local_trylock(&memcg_stock.lock))
+	    !mm_local_trylock_nolock(&memcg_stock.lock))
 		return ret;
 
 	stock = this_cpu_ptr(&memcg_stock);
@@ -2243,7 +2243,7 @@ static void refill_stock(struct mem_cgroup *memcg, unsigned int nr_pages)
 	VM_WARN_ON_ONCE(mem_cgroup_is_root(memcg));
 
 	if (nr_pages > MEMCG_CHARGE_BATCH ||
-	    !local_trylock(&memcg_stock.lock)) {
+	    !mm_local_trylock_nolock(&memcg_stock.lock)) {
 		/*
 		 * In case of larger than batch refill or unlikely failure to
 		 * lock the percpu memcg_stock.lock, uncharge memcg directly.
@@ -3232,7 +3232,7 @@ void __memcg_kmem_uncharge_page(struct page *page, int order)
 
 static struct obj_stock_pcp *trylock_stock(void)
 {
-	if (local_trylock(&obj_stock.lock))
+	if (mm_local_trylock_nolock(&obj_stock.lock))
 		return this_cpu_ptr(&obj_stock);
 
 	return NULL;
@@ -3336,14 +3336,64 @@ static bool __consume_obj_stock(struct obj_cgroup *objcg,
 	return false;
 }
 
-static bool consume_obj_stock(struct obj_cgroup *objcg, unsigned int nr_bytes)
+/*
+ * A failed per-CPU stock trylock normally falls back to charging another
+ * page. No-lock callers can fail that trylock on every allocation, so reuse
+ * the centralized prepaid bytes before charging again. Paired frees return
+ * the bytes here and keep the retained charge bounded instead of adding one
+ * page per allocation.
+ */
+static bool consume_obj_stock_shared(struct obj_cgroup *objcg,
+				     unsigned int nr_bytes)
+{
+	int old = atomic_read(&objcg->nr_charged_bytes);
+
+	do {
+		if (old < 0 || (unsigned int)old < nr_bytes)
+			return false;
+	} while (!atomic_try_cmpxchg(&objcg->nr_charged_bytes, &old,
+				     old - nr_bytes));
+
+	return true;
+}
+
+static unsigned int refill_obj_stock_shared(struct obj_cgroup *objcg,
+					    unsigned int nr_bytes,
+					    bool allow_uncharge)
+{
+	int old = atomic_read(&objcg->nr_charged_bytes);
+	unsigned int new, nr_pages;
+	u64 total;
+
+	do {
+		if (WARN_ON_ONCE(old < 0))
+			return 0;
+
+		total = (unsigned int)old + (u64)nr_bytes;
+		nr_pages = 0;
+		if ((allow_uncharge && total > PAGE_SIZE) || total > U16_MAX) {
+			nr_pages = total >> PAGE_SHIFT;
+			new = total & (PAGE_SIZE - 1);
+		} else {
+			new = total;
+		}
+	} while (!atomic_try_cmpxchg(&objcg->nr_charged_bytes, &old, new));
+
+	return nr_pages;
+}
+
+static bool consume_obj_stock(struct obj_cgroup *objcg, unsigned int nr_bytes,
+			      gfp_t gfp)
 {
 	struct obj_stock_pcp *stock;
 	bool ret = false;
 
 	stock = trylock_stock();
-	if (!stock)
+	if (!stock) {
+		if (!gfpflags_allow_spinning(gfp))
+			ret = consume_obj_stock_shared(objcg, nr_bytes);
 		return ret;
+	}
 
 	ret = __consume_obj_stock(objcg, stock, nr_bytes);
 	unlock_stock(stock);
@@ -3462,9 +3512,8 @@ static void __refill_obj_stock(struct obj_cgroup *objcg,
 	int i, slot = -1, empty_slot = -1;
 
 	if (!stock) {
-		nr_pages = nr_bytes >> PAGE_SHIFT;
-		nr_bytes = nr_bytes & (PAGE_SIZE - 1);
-		atomic_add(nr_bytes, &objcg->nr_charged_bytes);
+		nr_pages = refill_obj_stock_shared(objcg, nr_bytes,
+						   allow_uncharge);
 		goto out;
 	}
 
@@ -3548,18 +3597,18 @@ int obj_cgroup_charge(struct obj_cgroup *objcg, gfp_t gfp, size_t size)
 	size_t remainder;
 	int ret;
 
-	if (likely(consume_obj_stock(objcg, size)))
+	if (likely(consume_obj_stock(objcg, size, gfp)))
 		return 0;
 
 	/*
-	 * In theory, objcg->nr_charged_bytes can have enough
-	 * pre-charged bytes to satisfy the allocation. However,
+	 * For callers which allow spinning, objcg->nr_charged_bytes can have
+	 * enough pre-charged bytes to satisfy the allocation. However,
 	 * flushing objcg->nr_charged_bytes requires two atomic
 	 * operations, and objcg->nr_charged_bytes can't be big.
 	 * The shared objcg->nr_charged_bytes can also become a
 	 * performance bottleneck if all tasks of the same memcg are
-	 * trying to update it. So it's better to ignore it and try
-	 * grab some new pages. The stock's nr_bytes will be flushed to
+	 * trying to update it. So it's better for those callers to ignore it
+	 * and try to grab some new pages. The stock's nr_bytes will be flushed to
 	 * objcg->nr_charged_bytes later on when objcg changes.
 	 *
 	 * The stock's nr_bytes may contain enough pre-charged bytes
@@ -3570,9 +3619,9 @@ int obj_cgroup_charge(struct obj_cgroup *objcg, gfp_t gfp, size_t size)
 	 * page uncharge right after a page charge, we set the
 	 * allow_uncharge flag to false when calling refill_obj_stock()
 	 * to temporarily allow the pre-charged bytes to exceed the page
-	 * size limit. The maximum reachable value of the pre-charged
-	 * bytes is (sizeof(object) + PAGE_SIZE - 2) if there is no data
-	 * race.
+	 * size limit. No-lock callers already tried the centralized bytes
+	 * above. The maximum reachable value of the pre-charged bytes is
+	 * (sizeof(object) + PAGE_SIZE - 2) if there is no data race.
 	 */
 	ret = __obj_cgroup_charge(objcg, gfp, size, &remainder);
 	if (!ret && remainder)
@@ -3640,6 +3689,7 @@ bool __memcg_slab_post_alloc_hook(struct kmem_cache *s, struct list_lru *lru,
 		unsigned long obj_exts;
 		struct slabobj_ext *obj_ext;
 		struct obj_stock_pcp *stock;
+		bool consumed;
 
 		slab = virt_to_slab(p[i]);
 
@@ -3662,7 +3712,13 @@ bool __memcg_slab_post_alloc_hook(struct kmem_cache *s, struct list_lru *lru,
 		 * between iterations, with a more complicated undo
 		 */
 		stock = trylock_stock();
-		if (!stock || !__consume_obj_stock(objcg, stock, obj_size)) {
+		if (stock)
+			consumed = __consume_obj_stock(objcg, stock, obj_size);
+		else
+			consumed = !gfpflags_allow_spinning(flags) &&
+				   consume_obj_stock_shared(objcg, obj_size);
+
+		if (!consumed) {
 			size_t remainder;
 
 			unlock_stock(stock);
diff --git a/mm/page_alloc.c b/mm/page_alloc.c
index 3e21dc90b858..617997ed40ba 100644
--- a/mm/page_alloc.c
+++ b/mm/page_alloc.c
@@ -142,6 +142,15 @@ static DEFINE_MUTEX(pcp_batch_high_lock);
 	pcpu_task_unpin();						\
 })
 
+#define pcp_spin_trylock_nolock(ptr)					\
+({									\
+	struct per_cpu_pages *_ret = NULL;				\
+									\
+	if (!IS_ENABLED(CONFIG_PREEMPT_RT) || preemptible())		\
+		_ret = pcp_spin_trylock(ptr);				\
+	_ret;								\
+})
+
 /*
  * On CONFIG_SMP=n the UP implementation of spin_trylock() never fails and thus
  * is not compatible with our locking scheme. However we do not need pcp for
@@ -154,6 +163,13 @@ static DEFINE_MUTEX(pcp_batch_high_lock);
 
 #define pcp_spin_unlock(ptr)		\
 		BUG_ON(1)
+
+#define pcp_spin_trylock_nolock(ptr)		\
+({						\
+	(void)(ptr);				\
+	NULL;					\
+})
+
 #endif
 
 /*
@@ -1562,7 +1578,8 @@ static void free_one_page(struct zone *zone, struct page *page,
 	unsigned long flags;
 
 	if (unlikely(fpi_flags & FPI_NOLOCK)) {
-		if (!can_spin_trylock() || !spin_trylock_irqsave(&zone->lock, flags)) {
+		if (!can_spin_trylock() ||
+		    !spin_trylock_nolock_irqsave(&zone->lock, flags)) {
 			add_page_to_zone_llist(zone, page, order);
 			return;
 		}
@@ -2544,7 +2561,7 @@ static int rmqueue_bulk(struct zone *zone, unsigned int order,
 	int i;
 
 	if (unlikely(alloc_flags & ALLOC_NOLOCK)) {
-		if (!spin_trylock_irqsave(&zone->lock, flags))
+		if (!spin_trylock_nolock_irqsave(&zone->lock, flags))
 			return 0;
 	} else {
 		spin_lock_irqsave(&zone->lock, flags);
@@ -2986,7 +3003,10 @@ static void __free_frozen_pages(struct page *page, unsigned int order,
 		add_page_to_zone_llist(zone, page, order);
 		return;
 	}
-	pcp = pcp_spin_trylock(zone->per_cpu_pageset);
+	if (unlikely(fpi_flags & FPI_NOLOCK))
+		pcp = pcp_spin_trylock_nolock(zone->per_cpu_pageset);
+	else
+		pcp = pcp_spin_trylock(zone->per_cpu_pageset);
 	if (pcp) {
 		if (!free_frozen_page_commit(zone, pcp, page, migratetype,
 						order, fpi_flags))
@@ -3231,7 +3251,7 @@ struct page *rmqueue_buddy(struct zone *preferred_zone, struct zone *zone,
 	do {
 		page = NULL;
 		if (unlikely(alloc_flags & ALLOC_NOLOCK)) {
-			if (!spin_trylock_irqsave(&zone->lock, flags))
+			if (!spin_trylock_nolock_irqsave(&zone->lock, flags))
 				return NULL;
 		} else {
 			spin_lock_irqsave(&zone->lock, flags);
@@ -3380,7 +3400,10 @@ static struct page *rmqueue_pcplist(struct zone *preferred_zone,
 	struct page *page;
 
 	/* spin_trylock may fail due to a parallel drain or IRQ reentrancy. */
-	pcp = pcp_spin_trylock(zone->per_cpu_pageset);
+	if (unlikely(alloc_flags & ALLOC_NOLOCK))
+		pcp = pcp_spin_trylock_nolock(zone->per_cpu_pageset);
+	else
+		pcp = pcp_spin_trylock(zone->per_cpu_pageset);
 	if (!pcp)
 		return NULL;
 
diff --git a/mm/slub.c b/mm/slub.c
index 544cff39762c..5281f13f210b 100644
--- a/mm/slub.c
+++ b/mm/slub.c
@@ -3149,7 +3149,7 @@ static struct slab_sheaf *barn_get_empty_sheaf(struct node_barn *barn,
 
 	if (likely(allow_spin))
 		spin_lock_irqsave(&barn->lock, flags);
-	else if (!spin_trylock_irqsave(&barn->lock, flags))
+	else if (!spin_trylock_nolock_irqsave(&barn->lock, flags))
 		return NULL;
 
 	if (likely(barn->nr_empty)) {
@@ -3238,7 +3238,7 @@ barn_replace_empty_sheaf(struct node_barn *barn, struct slab_sheaf *empty,
 
 	if (likely(allow_spin))
 		spin_lock_irqsave(&barn->lock, flags);
-	else if (!spin_trylock_irqsave(&barn->lock, flags))
+	else if (!spin_trylock_nolock_irqsave(&barn->lock, flags))
 		return NULL;
 
 	if (likely(barn->nr_full)) {
@@ -3274,7 +3274,7 @@ barn_replace_full_sheaf(struct node_barn *barn, struct slab_sheaf *full,
 
 	if (likely(allow_spin))
 		spin_lock_irqsave(&barn->lock, flags);
-	else if (!spin_trylock_irqsave(&barn->lock, flags))
+	else if (!spin_trylock_nolock_irqsave(&barn->lock, flags))
 		return ERR_PTR(-EBUSY);
 
 	if (likely(barn->nr_empty)) {
@@ -3784,7 +3784,7 @@ static void *alloc_single_from_new_slab(struct kmem_cache *s, struct slab *slab,
 	n = get_node(s, slab_nid(slab));
 	if (allow_spin) {
 		spin_lock_irqsave(&n->list_lock, flags);
-	} else if (!spin_trylock_irqsave(&n->list_lock, flags)) {
+	} else if (!spin_trylock_nolock_irqsave(&n->list_lock, flags)) {
 		/*
 		 * Unlucky, discard newly allocated slab.
 		 * The slab is not fully free, but it's fine as
@@ -3829,7 +3829,7 @@ static bool get_partial_node_bulk(struct kmem_cache *s,
 
 	if (allow_spin)
 		spin_lock_irqsave(&n->list_lock, flags);
-	else if (!spin_trylock_irqsave(&n->list_lock, flags))
+	else if (!spin_trylock_nolock_irqsave(&n->list_lock, flags))
 		return false;
 
 	list_for_each_entry_safe(slab, slab2, &n->partial, slab_list) {
@@ -3902,7 +3902,7 @@ static void *get_from_partial_node(struct kmem_cache *s,
 
 	if (alloc_flags_allow_spinning(ac->alloc_flags))
 		spin_lock_irqsave(&n->list_lock, flags);
-	else if (!spin_trylock_irqsave(&n->list_lock, flags))
+	else if (!spin_trylock_nolock_irqsave(&n->list_lock, flags))
 		return NULL;
 	list_for_each_entry_safe(slab, slab2, &n->partial, slab_list) {
 
@@ -4506,7 +4506,7 @@ static unsigned int alloc_from_new_slab(struct kmem_cache *s, struct slab *slab,
 
 		if (allow_spin) {
 			spin_lock_irqsave(&n->list_lock, flags);
-		} else if (!spin_trylock_irqsave(&n->list_lock, flags)) {
+		} else if (!spin_trylock_nolock_irqsave(&n->list_lock, flags)) {
 			/*
 			 * Unlucky, discard newly allocated slab.
 			 * The slab is not fully free, but it's fine as
@@ -4702,6 +4702,15 @@ bool slab_post_alloc_hook(struct kmem_cache *s, gfp_t flags, size_t size,
 	return memcg_slab_post_alloc_hook(s, flags, size, p, ac);
 }
 
+static __always_inline bool
+cpu_sheaves_trylock(struct kmem_cache *s, bool allow_spin)
+{
+	if (allow_spin)
+		return local_trylock(&s->cpu_sheaves->lock);
+
+	return mm_local_trylock_nolock(&s->cpu_sheaves->lock);
+}
+
 /*
  * Replace the empty main sheaf with a (at least partially) full sheaf.
  *
@@ -4827,6 +4836,7 @@ static __fastpath_inline
 void *alloc_from_pcs(struct kmem_cache *s, gfp_t gfp, unsigned int alloc_flags, int node)
 {
 	struct slub_percpu_sheaves *pcs;
+	bool allow_spin = alloc_flags_allow_spinning(alloc_flags);
 	bool node_requested;
 	void *object;
 
@@ -4841,7 +4851,7 @@ void *alloc_from_pcs(struct kmem_cache *s, gfp_t gfp, unsigned int alloc_flags,
 		return NULL;
 	}
 
-	if (!local_trylock(&s->cpu_sheaves->lock))
+	if (!cpu_sheaves_trylock(s, allow_spin))
 		return NULL;
 
 	pcs = this_cpu_ptr(s->cpu_sheaves);
@@ -5477,7 +5487,7 @@ static void *__kmalloc_nolock_noprof(DECL_TOKEN_PARAMS(size, token), gfp_t gfp_f
 		 * But debug caches don't use that and only rely on
 		 * kmem_cache_node->list_lock, so kmalloc_nolock() can attempt
 		 * to allocate from debug caches by
-		 * spin_trylock_irqsave(&n->list_lock, ...)
+		 * spin_trylock_nolock_irqsave(&n->list_lock, ...)
 		 */
 		return NULL;
 
@@ -5980,7 +5990,7 @@ __pcs_replace_full_main(struct kmem_cache *s, struct slub_percpu_sheaves *pcs,
 	if (!sheaf_try_flush_main(s))
 		return NULL;
 
-	if (!local_trylock(&s->cpu_sheaves->lock))
+	if (!cpu_sheaves_trylock(s, allow_spin))
 		return NULL;
 
 	pcs = this_cpu_ptr(s->cpu_sheaves);
@@ -6016,7 +6026,7 @@ bool free_to_pcs(struct kmem_cache *s, void *object, bool allow_spin)
 {
 	struct slub_percpu_sheaves *pcs;
 
-	if (!local_trylock(&s->cpu_sheaves->lock))
+	if (!cpu_sheaves_trylock(s, allow_spin))
 		return false;
 
 	pcs = this_cpu_ptr(s->cpu_sheaves);
@@ -6120,7 +6130,7 @@ bool __kfree_rcu_sheaf(struct kmem_cache *s, void *obj, unsigned int free_flags)
 	if (!IS_ENABLED(CONFIG_PREEMPT_RT))
 		lock_map_acquire_try(&kfree_rcu_sheaf_map);
 
-	if (!local_trylock(&s->cpu_sheaves->lock))
+	if (!cpu_sheaves_trylock(s, allow_spin))
 		goto fail;
 
 	pcs = this_cpu_ptr(s->cpu_sheaves);
@@ -6163,7 +6173,7 @@ bool __kfree_rcu_sheaf(struct kmem_cache *s, void *obj, unsigned int free_flags)
 		if (!empty)
 			goto fail;
 
-		if (!local_trylock(&s->cpu_sheaves->lock)) {
+		if (!cpu_sheaves_trylock(s, allow_spin)) {
 			__free_empty_sheaf(s, empty, free_flags);
 			goto fail;
 		}
-- 
2.53.0


^ permalink raw reply	[flat|nested] 5+ messages in thread

* [RFC PATCH v2 3/3] selftests/bpf: exercise task storage from hrtimer_start
  2026-10-05  7:06 [RFC PATCH v2 0/3] locking, mm: Add atomic allocator trylocks on RT Karl Mehltretter
  2026-10-05  7:06 ` [RFC PATCH v2 1/3] locking/rtmutex: Support atomic PREEMPT_RT spin trylocks Karl Mehltretter
  2026-10-05  7:06 ` [RFC PATCH v2 2/3] mm: use atomic RT trylocks for no-lock allocation Karl Mehltretter
@ 2026-10-05  7:06 ` Karl Mehltretter
  2026-10-05  7:53   ` bot+bpf-ci
  2 siblings, 1 reply; 5+ messages in thread
From: Karl Mehltretter @ 2026-10-05  7:06 UTC (permalink / raw)
  To: Peter Zijlstra, Thomas Gleixner, Sebastian Andrzej Siewior,
	Andrew Morton, Vlastimil Babka, Harry Yoo, Alexei Starovoitov
  Cc: Karl Mehltretter, Ingo Molnar, Will Deacon, Boqun Feng,
	Waiman Long, Jonathan Corbet, David Hildenbrand, Johannes Weiner,
	Shakeel Butt, David Stevens, Daniel Borkmann, Andrii Nakryiko,
	Martin KaFai Lau, Shuah Khan, Amery Hung, Swaraj Gaikwad,
	Clark Williams, Steven Rostedt, linux-kernel, linux-doc,
	linux-mm, linux-rt-devel, bpf, linux-kselftest, cgroups

No BPF selftest creates task-local storage from a tracepoint which runs
while an hrtimer base raw lock is held.  On PREEMPT_RT, this can exercise
allocator rtmutexes inside a raw-lock section.

Attach a program to tp_btf/hrtimer_start for the test process.  Create and
delete task storage on every event and trigger 1,000 events with
timerfd_settime().  Detach the program before reading its counters, then
require at least one successful allocation and balanced attempt and
outcome counters.  Allocation failure remains valid for the best-effort
no-lock API.

Assisted-by: LLM
Signed-off-by: Karl Mehltretter <kmehltretter@gmail.com>
---
 .../bpf/prog_tests/task_storage_hrtimer.c     | 50 +++++++++++++++++++
 .../bpf/progs/task_storage_hrtimer.c          | 48 ++++++++++++++++++
 2 files changed, 98 insertions(+)
 create mode 100644 tools/testing/selftests/bpf/prog_tests/task_storage_hrtimer.c
 create mode 100644 tools/testing/selftests/bpf/progs/task_storage_hrtimer.c

diff --git a/tools/testing/selftests/bpf/prog_tests/task_storage_hrtimer.c b/tools/testing/selftests/bpf/prog_tests/task_storage_hrtimer.c
new file mode 100644
index 000000000000..216e66263bf6
--- /dev/null
+++ b/tools/testing/selftests/bpf/prog_tests/task_storage_hrtimer.c
@@ -0,0 +1,50 @@
+// SPDX-License-Identifier: GPL-2.0
+#include <sys/timerfd.h>
+#include <unistd.h>
+
+#include <test_progs.h>
+#include "task_storage_hrtimer.skel.h"
+
+#define TRIGGER_COUNT 1000
+
+void test_task_storage_hrtimer(void)
+{
+	struct itimerspec timer = {
+		.it_value.tv_nsec = 1000000,
+	};
+	struct task_storage_hrtimer *skel;
+	int err, fd = -1, i;
+
+	skel = task_storage_hrtimer__open_and_load();
+	if (!ASSERT_OK_PTR(skel, "skel_open_and_load"))
+		return;
+
+	skel->bss->target_tgid = getpid();
+	err = task_storage_hrtimer__attach(skel);
+	if (!ASSERT_OK(err, "skel_attach"))
+		goto cleanup;
+
+	fd = timerfd_create(CLOCK_MONOTONIC, TFD_CLOEXEC);
+	if (!ASSERT_GE(fd, 0, "timerfd_create"))
+		goto cleanup;
+
+	for (i = 0; i < TRIGGER_COUNT; i++) {
+		err = timerfd_settime(fd, 0, &timer, NULL);
+		if (!ASSERT_OK(err, "timerfd_settime"))
+			goto cleanup;
+	}
+
+	task_storage_hrtimer__detach(skel);
+
+	ASSERT_GE(skel->bss->seen, TRIGGER_COUNT, "seen");
+	ASSERT_EQ(skel->bss->attempted, skel->bss->seen, "attempted");
+	ASSERT_GT(skel->bss->succeeded, 0, "succeeded");
+	ASSERT_EQ(skel->bss->delete_errors, 0, "delete_errors");
+	ASSERT_EQ(skel->bss->attempted,
+		  skel->bss->succeeded + skel->bss->failed, "attempts");
+
+cleanup:
+	if (fd >= 0)
+		close(fd);
+	task_storage_hrtimer__destroy(skel);
+}
diff --git a/tools/testing/selftests/bpf/progs/task_storage_hrtimer.c b/tools/testing/selftests/bpf/progs/task_storage_hrtimer.c
new file mode 100644
index 000000000000..a53c966f15b5
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/task_storage_hrtimer.c
@@ -0,0 +1,48 @@
+// SPDX-License-Identifier: GPL-2.0
+#include "vmlinux.h"
+#include <bpf/bpf_helpers.h>
+#include <bpf/bpf_tracing.h>
+
+char _license[] SEC("license") = "GPL";
+
+struct {
+	__uint(type, BPF_MAP_TYPE_TASK_STORAGE);
+	__uint(map_flags, BPF_F_NO_PREALLOC);
+	__type(key, int);
+	__type(value, long);
+} task_storage SEC(".maps");
+
+u32 target_tgid;
+u64 seen;
+u64 attempted;
+u64 succeeded;
+u64 failed;
+u64 delete_errors;
+
+SEC("tp_btf/hrtimer_start")
+int BPF_PROG(on_hrtimer_start, struct hrtimer *timer,
+	     enum hrtimer_mode mode, bool was_armed)
+{
+	struct task_struct *task;
+	long *storage;
+
+	if (bpf_get_current_pid_tgid() >> 32 != target_tgid)
+		return 0;
+
+	seen++;
+	attempted++;
+	task = bpf_get_current_task_btf();
+	storage = bpf_task_storage_get(&task_storage, task, NULL,
+				       BPF_LOCAL_STORAGE_GET_F_CREATE);
+	if (!storage) {
+		failed++;
+		return 0;
+	}
+
+	if (bpf_task_storage_delete(&task_storage, task))
+		delete_errors++;
+	else
+		succeeded++;
+
+	return 0;
+}
-- 
2.53.0


^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [RFC PATCH v2 3/3] selftests/bpf: exercise task storage from hrtimer_start
  2026-10-05  7:06 ` [RFC PATCH v2 3/3] selftests/bpf: exercise task storage from hrtimer_start Karl Mehltretter
@ 2026-10-05  7:53   ` bot+bpf-ci
  0 siblings, 0 replies; 5+ messages in thread
From: bot+bpf-ci @ 2026-10-05  7:53 UTC (permalink / raw)
  To: kmehltretter, peterz, tglx, bigeasy, akpm, vbabka, harry, ast
  Cc: kmehltretter, mingo, will, boqun, longman, corbet, david, hannes,
	shakeel.butt, stevensd, daniel, andrii, martin.lau, shuah,
	ameryhung, swarajgaikwad1925, clrkwllms, rostedt, linux-kernel,
	linux-doc, linux-mm, linux-rt-devel, bpf, linux-kselftest,
	cgroups, ast, andrii, daniel, martin.lau, eddyz87, yonghong.song,
	mason, ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 2685 bytes --]

> selftests/bpf: exercise task storage from hrtimer_start

TL;DR: Test-organization question only: should the new hrtimer task-storage
selftest be a subtest of task_local_storage.c instead of a separate file? The
review finds the test logic correct and the standalone file defensible.

> diff --git a/tools/testing/selftests/bpf/prog_tests/task_storage_hrtimer.c b/tools/testing/selftests/bpf/prog_tests/task_storage_hrtimer.c
> new file mode 100644
> index 000000000000..216e66263bf6
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/prog_tests/task_storage_hrtimer.c
> @@ -0,0 +1,50 @@
> +// SPDX-License-Identifier: GPL-2.0
> +#include <sys/timerfd.h>
> +#include <unistd.h>
> +
> +#include <test_progs.h>
> +#include "task_storage_hrtimer.skel.h"
> +
> +#define TRIGGER_COUNT 1000
> +
> +void test_task_storage_hrtimer(void)
> +{

A subsystem pattern flags this as potentially concerning: this adds a new
top-level test file for a narrow variant of behavior that an existing test
in the same directory already covers more generally.

prog_tests/task_local_storage.c already groups the general task-local
storage tests as subtests, including test_nodeadlock(), backed by
progs/task_storage_nodeadlock.c.  That subtest performs the same
bpf_task_storage_get(..., BPF_LOCAL_STORAGE_GET_F_CREATE) followed by
bpf_task_storage_delete() sequence and counts failures, just from a
different hook (lsm.s/socket_post_create).

The only difference here is the attach point, tp_btf/hrtimer_start, which
reaches the task storage allocator under the hrtimer base raw lock on
PREEMPT_RT.

Should this be a new subtest of test_task_local_storage() instead of a
separate prog_tests/task_storage_hrtimer.c?

For balance, the standalone file may well be defensible.  It uses its own
skeleton, as the existing task_local_storage.c subtests also do, so folding
it in would mostly move the test function rather than share setup.  It also
exercises a path that no existing selftest reaches: none of the current
programs that use task storage attach to tracepoints emitted from inside the
hrtimer code, and the programs that do use hrtimer tracepoints
(timer_start_deadlock.c, test_vmlinux.c) do not use task storage.

The test logic itself looks correct: the tp_btf prototype matches
TP_PROTO(hrtimer, mode, was_armed), and timerfd_settime() with a non-zero
it_value emits one trace_hrtimer_start under cpu_base->lock, so the counter
assertions are consistent.


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/37277547264

^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-10-05  7:53 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-05  7:06 [RFC PATCH v2 0/3] locking, mm: Add atomic allocator trylocks on RT Karl Mehltretter
2026-10-05  7:06 ` [RFC PATCH v2 1/3] locking/rtmutex: Support atomic PREEMPT_RT spin trylocks Karl Mehltretter
2026-10-05  7:06 ` [RFC PATCH v2 2/3] mm: use atomic RT trylocks for no-lock allocation Karl Mehltretter
2026-10-05  7:06 ` [RFC PATCH v2 3/3] selftests/bpf: exercise task storage from hrtimer_start Karl Mehltretter
2026-10-05  7:53   ` bot+bpf-ci

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®