mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH 0/2] mm: preserve nolock context through memcg cleanup
@ 2026-10-01  4:40 Karl Mehltretter
  2026-10-01  4:40 ` [PATCH 1/2] mm/slub: preserve no-lock freeing after memcg charge failure Karl Mehltretter
  2026-10-01  4:40 ` [PATCH 2/2] mm/memcontrol: defer final objcg release from no-lock frees Karl Mehltretter
  0 siblings, 2 replies; 4+ messages in thread
From: Karl Mehltretter @ 2026-10-01  4:40 UTC (permalink / raw)
  To: Vlastimil Babka, Harry Yoo, Johannes Weiner, Michal Hocko,
	Roman Gushchin, Shakeel Butt, Andrew Morton
  Cc: Karl Mehltretter, Hao Li, Christoph Lameter, David Rientjes,
	Muchun Song, Sebastian Andrzej Siewior, Clark Williams,
	Steven Rostedt, Alexei Starovoitov, cgroups, linux-mm,
	linux-kernel, linux-rt-devel

The no-lock allocation and free APIs can still enter regular locking
through two memcg cleanup paths.

Patch 1 handles rollback after a post-allocation charge failure. If
kmalloc_nolock() obtains an object but its memcg charge fails,
memcg_slab_post_alloc_hook() frees the object through the regular SLUB
path. Preserve the selected allocation mode by using kfree_nolock()
when the allocation flags disallow spinning.

Patch 2 handles the final reference to a killed object cgroup.
kfree_nolock() and free_pages_nolock() can drop that reference in the
caller's context, after which obj_cgroup_release() takes regular locks.
Queue the objcg on a lockless list and finish its teardown from normal
irq_work. On PREEMPT_RT that work runs in irq_workd task context.

These paths were found while discussing the following RFC and during
the subsequent investigation of no-lock allocation and free paths:

https://lore.kernel.org/r/20260919171443.90512-1-kmehltretter@gmail.com

Based on 40288c9206c17. Each patch's forced cleanup test passed on
x86-64 release and lockdep PREEMPT_RT QEMU, arm64 and ARM32 SMP QEMU,
and Pi 400 hardware. The objcg test accounted for all 256 queued and
deferred releases. The series also passed non-RT builds and arena tests.

Karl Mehltretter (2):
  mm/slub: preserve no-lock freeing after memcg charge failure
  mm/memcontrol: defer final objcg release from no-lock frees

 include/linux/memcontrol.h |  1 +
 mm/memcontrol.c            | 29 ++++++++++++++++++++++++++---
 mm/slub.c                  |  5 ++++-
 3 files changed, 31 insertions(+), 4 deletions(-)


base-commit: 40288c9206c17eb66a603262e06a58d300d0f279
-- 
2.53.0

^ permalink raw reply	[flat|nested] 4+ messages in thread

* [PATCH 1/2] mm/slub: preserve no-lock freeing after memcg charge failure
  2026-10-01  4:40 [PATCH 0/2] mm: preserve nolock context through memcg cleanup Karl Mehltretter
@ 2026-10-01  4:40 ` Karl Mehltretter
  2026-10-07 17:40   ` Harry Yoo
  2026-10-01  4:40 ` [PATCH 2/2] mm/memcontrol: defer final objcg release from no-lock frees Karl Mehltretter
  1 sibling, 1 reply; 4+ messages in thread
From: Karl Mehltretter @ 2026-10-01  4:40 UTC (permalink / raw)
  To: Vlastimil Babka, Harry Yoo, Johannes Weiner, Michal Hocko,
	Roman Gushchin, Shakeel Butt, Andrew Morton
  Cc: Karl Mehltretter, Hao Li, Christoph Lameter, David Rientjes,
	Muchun Song, Sebastian Andrzej Siewior, Clark Williams,
	Steven Rostedt, Alexei Starovoitov, cgroups, linux-mm,
	linux-kernel, linux-rt-devel

kmalloc_nolock() can obtain an object before its memcg
post-allocation charge fails.  For a single object,
memcg_slab_post_alloc_hook() rolls the allocation back through
memcg_alloc_abort_single().  That enters the regular SLUB free path,
which can take a sleeping list_lock on PREEMPT_RT even though the caller
selected a no-lock allocation.

Use kfree_nolock() when the allocation flags disallow spinning.  This
keeps the SLUB object rollback on the no-lock free path.  The failed
allocation continues to return NULL.

Fixes: af92793e52c3 ("slab: Introduce kmalloc_nolock() and kfree_nolock().")
Assisted-by: LLM
Signed-off-by: Karl Mehltretter <kmehltretter@gmail.com>
---
 mm/slub.c | 5 ++++-
 1 file changed, 4 insertions(+), 1 deletion(-)

diff --git a/mm/slub.c b/mm/slub.c
index 54ec125033571..a1f08338102e2 100644
--- a/mm/slub.c
+++ b/mm/slub.c
@@ -2519,7 +2519,10 @@ bool memcg_slab_post_alloc_hook(struct kmem_cache *s, gfp_t flags,
 		return true;
 
 	if (likely(size == 1)) {
-		memcg_alloc_abort_single(s, *p);
+		if (alloc_flags_allow_spinning(ac->alloc_flags))
+			memcg_alloc_abort_single(s, *p);
+		else
+			kfree_nolock(*p);
 		*p = NULL;
 	} else {
 		kmem_cache_free_bulk(s, size, p);
-- 
2.53.0

^ permalink raw reply	[flat|nested] 4+ messages in thread

* [PATCH 2/2] mm/memcontrol: defer final objcg release from no-lock frees
  2026-10-01  4:40 [PATCH 0/2] mm: preserve nolock context through memcg cleanup Karl Mehltretter
  2026-10-01  4:40 ` [PATCH 1/2] mm/slub: preserve no-lock freeing after memcg charge failure Karl Mehltretter
@ 2026-10-01  4:40 ` Karl Mehltretter
  1 sibling, 0 replies; 4+ messages in thread
From: Karl Mehltretter @ 2026-10-01  4:40 UTC (permalink / raw)
  To: Vlastimil Babka, Harry Yoo, Johannes Weiner, Michal Hocko,
	Roman Gushchin, Shakeel Butt, Andrew Morton
  Cc: Karl Mehltretter, Hao Li, Christoph Lameter, David Rientjes,
	Muchun Song, Sebastian Andrzej Siewior, Clark Williams,
	Steven Rostedt, Alexei Starovoitov, cgroups, linux-mm,
	linux-kernel, linux-rt-devel

kfree_nolock() and free_pages_nolock() can drop the last reference to a
killed object cgroup.  percpu_ref then invokes obj_cgroup_release() in
the caller's context.  The callback can uncharge pages.  It then takes
objcg_lock, exits the percpu reference, and schedules an RCU free.  A
no-lock free can therefore enter regular locking.  On PREEMPT_RT this
can take a sleeping lock while the caller holds a raw scheduler lock.

Make the release callback add the object cgroup to an NMI-safe lockless
list and queue normal irq_work.  Drain the list outside the no-lock
caller's context.  On PREEMPT_RT normal irq work runs in irq_workd task
context.  On non-RT the work remains safe to run from hard interrupt
context, as required by the existing release path.

Fixes: af92793e52c3 ("slab: Introduce kmalloc_nolock() and kfree_nolock().")
Assisted-by: LLM
Signed-off-by: Karl Mehltretter <kmehltretter@gmail.com>
---
 include/linux/memcontrol.h |  1 +
 mm/memcontrol.c            | 29 ++++++++++++++++++++++++++---
 2 files changed, 27 insertions(+), 3 deletions(-)

diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h
index 7d1c0ce189a88..c755d946430b6 100644
--- a/include/linux/memcontrol.h
+++ b/include/linux/memcontrol.h
@@ -186,6 +186,7 @@ struct obj_cgroup {
 	struct percpu_ref refcnt;
 	struct mem_cgroup *memcg;
 	atomic_t nr_charged_bytes;
+	struct llist_node release_node;
 	union {
 		struct list_head list; /* protected by objcg_lock */
 		struct rcu_head rcu;
diff --git a/mm/memcontrol.c b/mm/memcontrol.c
index 856a7d07586cc..63b18c7f1965f 100644
--- a/mm/memcontrol.c
+++ b/mm/memcontrol.c
@@ -33,6 +33,8 @@
 #include <linux/sched/mm.h>
 #include <linux/shmem_fs.h>
 #include <linux/hugetlb.h>
+#include <linux/irq_work.h>
+#include <linux/llist.h>
 #include <linux/pagemap.h>
 #include <linux/folio_batch.h>
 #include <linux/vm_event_item.h>
@@ -146,9 +148,12 @@ static void memcg_uncharge_kmem(struct mem_cgroup *memcg, unsigned int nr_pages)
 		memcg_uncharge(memcg, nr_pages);
 }
 
-static void obj_cgroup_release(struct percpu_ref *ref)
+static LLIST_HEAD(objcg_release_list);
+static void obj_cgroup_release_workfn(struct irq_work *work);
+static DEFINE_IRQ_WORK(objcg_release_work, obj_cgroup_release_workfn);
+
+static void obj_cgroup_release_one(struct obj_cgroup *objcg)
 {
-	struct obj_cgroup *objcg = container_of(ref, struct obj_cgroup, refcnt);
 	unsigned int nr_bytes;
 	unsigned int nr_pages;
 	unsigned long flags;
@@ -189,10 +194,28 @@ static void obj_cgroup_release(struct percpu_ref *ref)
 	list_del(&objcg->list);
 	spin_unlock_irqrestore(&objcg_lock, flags);
 
-	percpu_ref_exit(ref);
+	percpu_ref_exit(&objcg->refcnt);
 	kfree_rcu(objcg, rcu);
 }
 
+static void obj_cgroup_release_workfn(struct irq_work *work)
+{
+	struct llist_node *node;
+	struct obj_cgroup *objcg, *next;
+
+	node = llist_del_all(&objcg_release_list);
+	llist_for_each_entry_safe(objcg, next, node, release_node)
+		obj_cgroup_release_one(objcg);
+}
+
+static void obj_cgroup_release(struct percpu_ref *ref)
+{
+	struct obj_cgroup *objcg = container_of(ref, struct obj_cgroup, refcnt);
+
+	llist_add(&objcg->release_node, &objcg_release_list);
+	irq_work_queue(&objcg_release_work);
+}
+
 static struct obj_cgroup *obj_cgroup_alloc(void)
 {
 	struct obj_cgroup *objcg;
-- 
2.53.0

^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: [PATCH 1/2] mm/slub: preserve no-lock freeing after memcg charge failure
  2026-10-01  4:40 ` [PATCH 1/2] mm/slub: preserve no-lock freeing after memcg charge failure Karl Mehltretter
@ 2026-10-07 17:40   ` Harry Yoo
  0 siblings, 0 replies; 4+ messages in thread
From: Harry Yoo @ 2026-10-07 17:40 UTC (permalink / raw)
  To: Karl Mehltretter
  Cc: Vlastimil Babka, Johannes Weiner, Michal Hocko, Roman Gushchin,
	Shakeel Butt, Andrew Morton, Hao Li, Christoph Lameter,
	David Rientjes, Muchun Song, Sebastian Andrzej Siewior,
	Clark Williams, Steven Rostedt, Alexei Starovoitov, cgroups,
	linux-mm, linux-kernel, linux-rt-devel

On Thu, Oct 01, 2026 at 06:40:55AM +0200, Karl Mehltretter wrote:
> kmalloc_nolock() can obtain an object before its memcg
> post-allocation charge fails.  For a single object,
> memcg_slab_post_alloc_hook() rolls the allocation back through
> memcg_alloc_abort_single().  That enters the regular SLUB free path,
> which can take a sleeping list_lock on PREEMPT_RT even though the caller
> selected a no-lock allocation.
> 
> Use kfree_nolock() when the allocation flags disallow spinning.  This
> keeps the SLUB object rollback on the no-lock free path.  The failed
> allocation continues to return NULL.

When this happens, the object is not charged by memcg.

The kernel should not invoke memcg_slab_free_hook() (called by
kfree_nolock()) for a slab object that is not charged by memcg.

> Fixes: af92793e52c3 ("slab: Introduce kmalloc_nolock() and kfree_nolock().")
> Assisted-by: LLM
> Signed-off-by: Karl Mehltretter <kmehltretter@gmail.com>
> ---
>  mm/slub.c | 5 ++++-
>  1 file changed, 4 insertions(+), 1 deletion(-)
> 
> diff --git a/mm/slub.c b/mm/slub.c
> index 54ec125033571..a1f08338102e2 100644
> --- a/mm/slub.c
> +++ b/mm/slub.c
> @@ -2519,7 +2519,10 @@ bool memcg_slab_post_alloc_hook(struct kmem_cache *s, gfp_t flags,
>  		return true;
>  
>  	if (likely(size == 1)) {
> -		memcg_alloc_abort_single(s, *p);
> +		if (alloc_flags_allow_spinning(ac->alloc_flags))
> +			memcg_alloc_abort_single(s, *p);
> +		else
> +			kfree_nolock(*p);
>  		*p = NULL;
>  	} else {
>  		kmem_cache_free_bulk(s, size, p);
> -- 
> 2.53.0

-- 
Cheers,
Harry / Hyeonggon

^ permalink raw reply	[flat|nested] 4+ messages in thread

end of thread, other threads:[~2026-10-07 17:40 UTC | newest]

Thread overview: 4+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-01  4:40 [PATCH 0/2] mm: preserve nolock context through memcg cleanup Karl Mehltretter
2026-10-01  4:40 ` [PATCH 1/2] mm/slub: preserve no-lock freeing after memcg charge failure Karl Mehltretter
2026-10-07 17:40   ` Harry Yoo
2026-10-01  4:40 ` [PATCH 2/2] mm/memcontrol: defer final objcg release from no-lock frees Karl Mehltretter

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®