* [PATCH] mm/slab: sample slab_state once in kmem_cache_destroy()
@ 2026-09-30 19:51 Imre Kaloz
2026-09-30 21:26 ` Jose A. Perez de Azpillaga
2026-10-01 14:22 ` Harry Yoo
0 siblings, 2 replies; 4+ messages in thread
From: Imre Kaloz @ 2026-09-30 19:51 UTC (permalink / raw)
To: Vlastimil Babka, Harry Yoo, Andrew Morton
Cc: Hao Li, Christoph Lameter, David Rientjes, Roman Gushchin,
Jann Horn, linux-mm, linux-kernel, stable
kmem_cache_destroy() tests slab_state >= FULL for sysfs_slab_unlink()
and again in kmem_cache_release() for sysfs_slab_release(), with the
cache already off slab_caches in between. If slab_late_init() runs in
that gap it sets slab_state to FULL but never calls sysfs_slab_add() for
the unlinked cache, so kmem_cache_release() ends up in kobject_put() on
a kobject that was never initialized:
WARNING: lib/kobject.c:734 at kobject_put+0x64/0x2c0, CPU#1: kworker/u8:3/55
kobject: '(null)' ((____ptrval____)): is not initialized, yet kobject_put() is being called.
refcount_t: underflow; use-after-free.
kobject_put+0x64/0x2c0
sysfs_slab_release+0xc/0x20
kmem_cache_destroy+0x104/0x1e0
bioset_exit+0x13c/0x1e0
disk_release+0x54/0x140
put_disk+0x18/0x40
floppy_async_init+0xbec/0xd10
Seen on sparc64 at boot, where the asynchronous floppy init tears down
its bio slab while the late initcalls are running.
Read slab_state once, under slab_mutex which slab_late_init() holds when
it sets FULL, and use the result for both the unlink and the release.
The kobject flags (state_initialized, state_in_sysfs) were considered as
the key instead, but slab_state is what cache creation and
slab_late_init() decide on.
debugfs_slab_release() is not part of the race: it only looks the cache
up by name in the debugfs root and does nothing before that root exists.
Fixes: 4ec10268ed98 ("mm, slab: unlink slabinfo, sysfs and debugfs immediately")
Cc: stable@vger.kernel.org
Signed-off-by: Imre Kaloz <kaloz@kernel.org>
---
mm/slab_common.c | 17 +++++++++++++----
1 file changed, 13 insertions(+), 4 deletions(-)
diff --git a/mm/slab_common.c b/mm/slab_common.c
index 7223a7596dab..de11edfd1a0b 100644
--- a/mm/slab_common.c
+++ b/mm/slab_common.c
@@ -515,10 +515,10 @@ EXPORT_SYMBOL(kmem_buckets_create);
* and release of the kobject does not need slab_mutex or cpu_hotplug_lock
* protection. So they are now done without holding those locks.
*/
-static void kmem_cache_release(struct kmem_cache *s)
+static void kmem_cache_release(struct kmem_cache *s, bool sysfs_ready)
{
kfence_shutdown_cache(s);
- if (__is_defined(SLAB_SUPPORTS_SYSFS) && slab_state >= FULL)
+ if (__is_defined(SLAB_SUPPORTS_SYSFS) && sysfs_ready)
sysfs_slab_release(s);
else
slab_kmem_cache_release(s);
@@ -533,6 +533,7 @@ void slab_kmem_cache_release(struct kmem_cache *s)
void kmem_cache_destroy(struct kmem_cache *s)
{
+ bool sysfs_ready;
int err;
if (unlikely(!s) || !kasan_check_byte(s))
@@ -580,10 +581,18 @@ void kmem_cache_destroy(struct kmem_cache *s)
list_del(&s->list);
+ /*
+ * slab_late_init() sets slab_state to FULL under slab_mutex and adds
+ * sysfs entries only for caches still on the list. Sample the state
+ * here, so that a cache unlinked before that point is not handed to
+ * sysfs_slab_release() with an uninitialized kobject.
+ */
+ sysfs_ready = slab_state >= FULL;
+
mutex_unlock(&slab_mutex);
cpus_read_unlock();
- if (slab_state >= FULL)
+ if (sysfs_ready)
sysfs_slab_unlink(s);
debugfs_slab_release(s);
@@ -593,7 +602,7 @@ void kmem_cache_destroy(struct kmem_cache *s)
if (s->flags & SLAB_TYPESAFE_BY_RCU)
rcu_barrier();
- kmem_cache_release(s);
+ kmem_cache_release(s, sysfs_ready);
}
EXPORT_SYMBOL(kmem_cache_destroy);
base-commit: 72d3fcf802c45d00b300f25b848a93c3a2bd7c7e
--
2.47.3
^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [PATCH] mm/slab: sample slab_state once in kmem_cache_destroy()
2026-09-30 19:51 [PATCH] mm/slab: sample slab_state once in kmem_cache_destroy() Imre Kaloz
@ 2026-09-30 21:26 ` Jose A. Perez de Azpillaga
2026-10-01 14:22 ` Harry Yoo
1 sibling, 0 replies; 4+ messages in thread
From: Jose A. Perez de Azpillaga @ 2026-09-30 21:26 UTC (permalink / raw)
To: Imre Kaloz
Cc: Vlastimil Babka, Harry Yoo, Andrew Morton, Hao Li,
Christoph Lameter, David Rientjes, Roman Gushchin, Jann Horn,
linux-mm, linux-kernel, stable
On Wed, Sep 30, 2026 at 09:51:09PM +0200, Imre Kaloz wrote:
> kmem_cache_destroy() tests slab_state >= FULL for sysfs_slab_unlink()
> and again in kmem_cache_release() for sysfs_slab_release(), with the
> cache already off slab_caches in between. If slab_late_init() runs in
> that gap it sets slab_state to FULL but never calls sysfs_slab_add() for
> the unlinked cache, so kmem_cache_release() ends up in kobject_put() on
> a kobject that was never initialized:
>
> WARNING: lib/kobject.c:734 at kobject_put+0x64/0x2c0, CPU#1: kworker/u8:3/55
> kobject: '(null)' ((____ptrval____)): is not initialized, yet kobject_put() is being called.
> refcount_t: underflow; use-after-free.
> kobject_put+0x64/0x2c0
> sysfs_slab_release+0xc/0x20
> kmem_cache_destroy+0x104/0x1e0
> bioset_exit+0x13c/0x1e0
> disk_release+0x54/0x140
> put_disk+0x18/0x40
> floppy_async_init+0xbec/0xd10
>
> Seen on sparc64 at boot, where the asynchronous floppy init tears down
> its bio slab while the late initcalls are running.
>
> Read slab_state once, under slab_mutex which slab_late_init() holds when
> it sets FULL, and use the result for both the unlink and the release.
> The kobject flags (state_initialized, state_in_sysfs) were considered as
> the key instead, but slab_state is what cache creation and
> slab_late_init() decide on.
the enum has nothing between UP and FULL, and the create path skips
sysfs_slab_add() while slab_state <= UP, so both tests are the one that
decides whether the kobject exists, and a false value cannot skip a
kobject_put() that was owed.
it would read better to me if the changelog said that. otherwise LGTM.
Reviewed-by: Jose A. Perez de Azpillaga <azpijr@gmail.com>
> debugfs_slab_release() is not part of the race: it only looks the cache
> up by name in the debugfs root and does nothing before that root exists.
>
> Fixes: 4ec10268ed98 ("mm, slab: unlink slabinfo, sysfs and debugfs immediately")
> Cc: stable@vger.kernel.org
> Signed-off-by: Imre Kaloz <kaloz@kernel.org>
--
cheers,
jose a. p-a
^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [PATCH] mm/slab: sample slab_state once in kmem_cache_destroy()
2026-09-30 19:51 [PATCH] mm/slab: sample slab_state once in kmem_cache_destroy() Imre Kaloz
2026-09-30 21:26 ` Jose A. Perez de Azpillaga
@ 2026-10-01 14:22 ` Harry Yoo
2026-10-01 15:42 ` Imre Kaloz
1 sibling, 1 reply; 4+ messages in thread
From: Harry Yoo @ 2026-10-01 14:22 UTC (permalink / raw)
To: Imre Kaloz
Cc: Vlastimil Babka, Andrew Morton, Hao Li, Christoph Lameter,
David Rientjes, Roman Gushchin, Jann Horn, linux-mm,
linux-kernel, stable
Hi Imre,
Thanks for catching and fixing this!
On Wed, Sep 30, 2026 at 09:51:09PM +0200, Imre Kaloz wrote:
> kmem_cache_destroy() tests slab_state >= FULL for sysfs_slab_unlink()
> and again in kmem_cache_release() for sysfs_slab_release(), with the
> cache already off slab_caches in between. If slab_late_init() runs in
> that gap it sets slab_state to FULL but never calls sysfs_slab_add() for
> the unlinked cache, so kmem_cache_release() ends up in kobject_put() on
> a kobject that was never initialized:
>
> WARNING: lib/kobject.c:734 at kobject_put+0x64/0x2c0, CPU#1: kworker/u8:3/55
> kobject: '(null)' ((____ptrval____)): is not initialized, yet kobject_put() is being called.
> refcount_t: underflow; use-after-free.
> kobject_put+0x64/0x2c0
> sysfs_slab_release+0xc/0x20
> kmem_cache_destroy+0x104/0x1e0
> bioset_exit+0x13c/0x1e0
> disk_release+0x54/0x140
> put_disk+0x18/0x40
> floppy_async_init+0xbec/0xd10
>
> Seen on sparc64 at boot, where the asynchronous floppy init tears down
> its bio slab while the late initcalls are running.
>
> Read slab_state once, under slab_mutex which slab_late_init() holds when
> it sets FULL, and use the result for both the unlink and the release.
> The kobject flags (state_initialized, state_in_sysfs) were considered as
> the key instead, but slab_state is what cache creation and
> slab_late_init() decide on.
>
> debugfs_slab_release() is not part of the race: it only looks the cache
> up by name in the debugfs root and does nothing before that root exists.
>
> Fixes: 4ec10268ed98 ("mm, slab: unlink slabinfo, sysfs and debugfs immediately")
Overall looks good to me, but could you please explain why it's
not relevant before this commit? pre-4ec10268 still reads slab_state
outside slab_mutex.
> Cc: stable@vger.kernel.org
> Signed-off-by: Imre Kaloz <kaloz@kernel.org>
--
Cheers,
Harry / Hyeonggon
^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [PATCH] mm/slab: sample slab_state once in kmem_cache_destroy()
2026-10-01 14:22 ` Harry Yoo
@ 2026-10-01 15:42 ` Imre Kaloz
0 siblings, 0 replies; 4+ messages in thread
From: Imre Kaloz @ 2026-10-01 15:42 UTC (permalink / raw)
To: Harry Yoo
Cc: Vlastimil Babka, Andrew Morton, Hao Li, Christoph Lameter,
David Rientjes, Roman Gushchin, Jann Horn, linux-mm,
linux-kernel, stable
Hi Harry,
On Thu, 1 Oct 2026, Harry Yoo wrote:
> Hi Imre,
> Thanks for catching and fixing this!
>
> On Wed, Sep 30, 2026 at 09:51:09PM +0200, Imre Kaloz wrote:
>> kmem_cache_destroy() tests slab_state >= FULL for sysfs_slab_unlink()
>> and again in kmem_cache_release() for sysfs_slab_release(), with the
>> cache already off slab_caches in between. If slab_late_init() runs in
>> that gap it sets slab_state to FULL but never calls sysfs_slab_add() for
>> the unlinked cache, so kmem_cache_release() ends up in kobject_put() on
>> a kobject that was never initialized:
>>
>> WARNING: lib/kobject.c:734 at kobject_put+0x64/0x2c0, CPU#1: kworker/u8:3/55
>> kobject: '(null)' ((____ptrval____)): is not initialized, yet kobject_put() is being called.
>> refcount_t: underflow; use-after-free.
>> kobject_put+0x64/0x2c0
>> sysfs_slab_release+0xc/0x20
>> kmem_cache_destroy+0x104/0x1e0
>> bioset_exit+0x13c/0x1e0
>> disk_release+0x54/0x140
>> put_disk+0x18/0x40
>> floppy_async_init+0xbec/0xd10
>>
>> Seen on sparc64 at boot, where the asynchronous floppy init tears down
>> its bio slab while the late initcalls are running.
>>
>> Read slab_state once, under slab_mutex which slab_late_init() holds when
>> it sets FULL, and use the result for both the unlink and the release.
>> The kobject flags (state_initialized, state_in_sysfs) were considered as
>> the key instead, but slab_state is what cache creation and
>> slab_late_init() decide on.
>>
>> debugfs_slab_release() is not part of the race: it only looks the cache
>> up by name in the debugfs root and does nothing before that root exists.
>>
>> Fixes: 4ec10268ed98 ("mm, slab: unlink slabinfo, sysfs and debugfs immediately")
>
> Overall looks good to me, but could you please explain why it's
> not relevant before this commit? pre-4ec10268 still reads slab_state
> outside slab_mutex.
>
The outside-mutex read is older, yes. What 4ec10268ed98 changed is that
unlink and release no longer share it.
Before that commit, kmem_cache_release() did one test and used it for
both:
if (slab_state >= FULL) {
sysfs_slab_unlink(s);
sysfs_slab_release(s);
} else {
slab_kmem_cache_release(s);
}
So the two could not disagree. 4ec10268 moved sysfs_slab_unlink() into
kmem_cache_destroy(), after list_del() and after dropping slab_mutex,
and left a second slab_state test in kmem_cache_release() for
sysfs_slab_release(). The cache is already off slab_caches in between.
That is the window in the warning: the first test sees < FULL, so the
kobject is never linked; slab_late_init() then sets FULL under
slab_mutex and calls sysfs_slab_add() only for caches still on the
list; the second test sees FULL and kobject_put()s a kobject that
kobject_init() never ran on.
A single read cannot produce that split decision, which is why Fixes:
points at 4ec10268ed98.
There is a related older window: after list_del() and before that
single read, slab_sysfs_init() could set FULL, skip this cache, and
the single read would then call both unlink and release on an
uninitialized kobject. I have not hit that, and this patch does not
close it.
Best,
Imre
^ permalink raw reply [flat|nested] 4+ messages in thread
end of thread, other threads:[~2026-10-01 15:43 UTC | newest]
Thread overview: 4+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-30 19:51 [PATCH] mm/slab: sample slab_state once in kmem_cache_destroy() Imre Kaloz
2026-09-30 21:26 ` Jose A. Perez de Azpillaga
2026-10-01 14:22 ` Harry Yoo
2026-10-01 15:42 ` Imre Kaloz
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®