mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH net-next v6 0/8] net: skb: isolate skb data area allocations into a separate bucket
@ 2026-10-06  9:20 Kees Cook
  2026-10-06  9:20 ` [PATCH net-next v6 1/8] mm/slab: Mark the kmem_buckets_create() context as a Context: section Kees Cook
                   ` (7 more replies)
  0 siblings, 8 replies; 9+ messages in thread
From: Kees Cook @ 2026-10-06  9:20 UTC (permalink / raw)
  To: Vlastimil Babka
  Cc: Kees Cook, Harry Yoo, David S. Miller, Andrew Morton, Hao Li,
	Christoph Lameter, David Rientjes, Roman Gushchin, Pedro Falcato,
	Kuniyuki Iwashima, Christian Brauner, Jan Kara, Johannes Weiner,
	Michal Hocko, Shakeel Butt, Muchun Song, Eric Dumazet,
	Jakub Kicinski, Paolo Abeni, Simon Horman, Willem de Bruijn,
	Jason Xing, cgroups, netdev, linux-mm, linux-kernel,
	linux-hardening

Hi!

This gets the buckets able to handle memcg (GFP_KERNEL_ACCOUNT) with
isolation (since it's common due to AF_UNIX), and GFP_DMA with fall back
(since it's rare). It gave me an excuse to build out bucket kunit tests
too, and that (and LLM review) found a couple other issues that needed
fixing too, including msg_msg allocations going uncharged to their memcg
when CONFIG_SLAB_BUCKETS=n (fixed in 2/8).

Harry, on your v4 question[1] about bucket users giving their own
alignment: I tried that in v5, but a set's allocations don't always
come from its own caches. With CONFIG_SLAB_BUCKETS=n, after a failed
kmem_buckets_create(), and for the DMA and reclaimable fallbacks, they
come from the general kmalloc caches, which can only give kmalloc()'s
alignment. So v6 goes back to mirroring the kmalloc cache's alignment,
and drops the ctor and flags arguments for the same reason. And the whole
exploration made me realize I had a completely wrong understanding of
how memcg worked. :P

The bulk of this is mm/slab, but the final patch is netdev, which Paolo
acked in v4, so I'm hoping this whole series can go via slab?

Thanks!

-Kees

 v6:
 - drop v5's 6/7 ("Let a bucket set handle __GFP_ACCOUNT") and its
   kmem_buckets_create_types(): memcg charges each object in whatever
   cache serves it, so accounted allocations can stay in a set's single
   row of caches, and the fallback now covers only DMA, reclaimable, and
   no-obj-ext allocations (Sashiko)
 - 2/8: new: account msg_msg with GFP_KERNEL_ACCOUNT again; with
   CONFIG_SLAB_BUCKETS=n it went uncharged, since its accounting lived in
   SLAB_ACCOUNT on bucket caches that are not created (Sashiko)
 - 3/8: new: drop the ctor and flags arguments from kmem_buckets_create();
   neither reaches allocations that fall back to the general kmalloc
   caches, and no caller needs them any more (Sashiko)
 - 4/8: go back to v4's form: no alignment argument, and each bucket cache
   takes the alignment of the kmalloc cache it mirrors, since the fallbacks
   to kmalloc can give no other (Sashiko, Harry)
 - 5/8: say in the teardown comment that cache sharing comes from kmalloc
   rounding sizes up to a larger class (Sashiko)
 - 6/8: drop the explicit alignment tests; check that each size lands in
   the cache of the size kmalloc() rounds it up to, not just a big enough
   one; and in the destroy test, assert on the allocation, skip when KFENCE
   serves it, and tear the set down through its KUnit cleanup action
   (Sashiko)
 - 7/8: keep __GFP_ACCOUNT allocations in the set, document what an
   allocation that falls back loses, and test the reclaimable fallback
   (Sashiko)
 - 8/8: create the skb_data set with kmem_buckets_create(), and make the
   comment above kmalloc_reserve() name no allocator (Sashiko)
 - v5..v6 diff: https://git.kernel.org/pub/scm/linux/kernel/git/kees/linux.git/diff/?id=dev/v7.3-rc2/skb-buckets/v6&id2=dev/v7.3-rc2/skb-buckets/v5
 v5: https://lore.kernel.org/all/20261002231120.late.500-kees@kernel.org/
 v4: https://lore.kernel.org/all/20260921075811.too.775-kees@kernel.org/
 v3: https://lore.kernel.org/all/20260702170728.168755-1-pfalcato@suse.de/

[1] https://lore.kernel.org/all/arViR2Miz61-3fV4@thinkstation/

Kees Cook (7):
  mm/slab: Mark the kmem_buckets_create() context as a Context: section
  ipc, msg: Account msg_msg allocations with GFP_KERNEL_ACCOUNT again
  mm/slab: Drop the ctor and flags arguments from kmem_buckets_create()
  mm/slab: Give bucket caches the alignment of the caches they mirror
  mm/slab: Add kmem_buckets_destroy()
  mm/slab: Add tests for the existing kmem_buckets behaviour
  mm/slab: Provide kmalloc type fallback for bucket allocations

Pedro Falcato (1):
  net: skb: isolate skb data area allocations into a separate bucket

 include/linux/slab.h   |   6 +-
 mm/slab.h              |  19 ++-
 ipc/msgutil.c          |   8 +-
 lib/tests/slub_kunit.c | 291 +++++++++++++++++++++++++++++++++++++++++
 mm/slab_common.c       |  74 ++++++++---
 mm/util.c              |   2 +-
 net/core/skbuff.c      |  10 +-
 7 files changed, 378 insertions(+), 32 deletions(-)

-- 
2.55.0


^ permalink raw reply	[flat|nested] 9+ messages in thread

* [PATCH net-next v6 1/8] mm/slab: Mark the kmem_buckets_create() context as a Context: section
  2026-10-06  9:20 [PATCH net-next v6 0/8] net: skb: isolate skb data area allocations into a separate bucket Kees Cook
@ 2026-10-06  9:20 ` Kees Cook
  2026-10-06  9:20 ` [PATCH net-next v6 2/8] ipc, msg: Account msg_msg allocations with GFP_KERNEL_ACCOUNT again Kees Cook
                   ` (6 subsequent siblings)
  7 siblings, 0 replies; 9+ messages in thread
From: Kees Cook @ 2026-10-06  9:20 UTC (permalink / raw)
  To: Vlastimil Babka
  Cc: Kees Cook, Harry Yoo, Andrew Morton, Hao Li, Christoph Lameter,
	David Rientjes, Roman Gushchin, linux-mm, Pedro Falcato,
	Kuniyuki Iwashima, linux-hardening, linux-kernel

kmem_buckets_create() had the same sentence about calling context as
kmem_cache_create(), but lacked the "Context:" prefix, so kernel-doc
rendered in the body instead of as a "context" section.

Give it the missing prefix. Additionally fix the "a interrupt" typo
__kmem_cache_create_args() had.

Fixes: b32801d1255be ("mm/slab: Introduce kmem_buckets_create() and family")
Assisted-by: LLM
Reviewed-by: Harry Yoo (Meta) <harry@kernel.org>
Acked-by: Pedro Falcato <pfalcato@suse.de>
Reviewed-by: Hao Li <hao.li@linux.dev>
Signed-off-by: Kees Cook <kees@kernel.org>
---
 mm/slab_common.c | 4 ++--
 1 file changed, 2 insertions(+), 2 deletions(-)

diff --git a/mm/slab_common.c b/mm/slab_common.c
index b19ba1b31484..270408ce5a9d 100644
--- a/mm/slab_common.c
+++ b/mm/slab_common.c
@@ -311,7 +311,7 @@ __kmem_cache_alias(const char *name, unsigned int size, slab_flags_t flags,
  * &SLAB_TYPESAFE_BY_RCU - Slab page (not individual objects) freeing delayed
  * by a grace period - see the full description before using.
  *
- * Context: Cannot be called within a interrupt, but can be interrupted.
+ * Context: Cannot be called within an interrupt, but can be interrupted.
  *
  * Return: a pointer to the cache on success, NULL on failure.
  */
@@ -422,7 +422,7 @@ static struct kmem_cache *kmem_buckets_cache __ro_after_init;
  *		to/from userspace.
  * @ctor: A constructor for the objects, run when new allocations are made.
  *
- * Cannot be called within an interrupt, but can be interrupted.
+ * Context: Cannot be called within an interrupt, but can be interrupted.
  *
  * Return: a pointer to the cache on success, NULL on failure. When
  * CONFIG_SLAB_BUCKETS is not enabled, ZERO_SIZE_PTR is returned, and
-- 
2.55.0


^ permalink raw reply	[flat|nested] 9+ messages in thread

* [PATCH net-next v6 2/8] ipc, msg: Account msg_msg allocations with GFP_KERNEL_ACCOUNT again
  2026-10-06  9:20 [PATCH net-next v6 0/8] net: skb: isolate skb data area allocations into a separate bucket Kees Cook
  2026-10-06  9:20 ` [PATCH net-next v6 1/8] mm/slab: Mark the kmem_buckets_create() context as a Context: section Kees Cook
@ 2026-10-06  9:20 ` Kees Cook
  2026-10-06  9:20 ` [PATCH net-next v6 3/8] mm/slab: Drop the ctor and flags arguments from kmem_buckets_create() Kees Cook
                   ` (5 subsequent siblings)
  7 siblings, 0 replies; 9+ messages in thread
From: Kees Cook @ 2026-10-06  9:20 UTC (permalink / raw)
  To: Vlastimil Babka
  Cc: Kees Cook, Christian Brauner, Jan Kara, Andrew Morton,
	Roman Gushchin, Johannes Weiner, Michal Hocko, Shakeel Butt,
	Muchun Song, cgroups, linux-mm, Pedro Falcato, Kuniyuki Iwashima,
	linux-hardening, linux-kernel

Since commit 734bbc1c97ea7 ("ipc, msg: Use dedicated slab buckets for
alloc_msg()"), alloc_msg() allocates with GFP_KERNEL, and a msg_msg is
accounted only through SLAB_ACCOUNT on its bucket caches. With
CONFIG_SLAB_BUCKETS=n, kmem_buckets_create() creates no caches and
kmem_buckets_alloc() is a GFP_KERNEL kmalloc() from the general caches,
which do not account it; the same happens when kmem_buckets_create()
fails. Either way, the allocation is not charged to the sender's memory
cgroup.

Allocate with GFP_KERNEL_ACCOUNT again, as before that commit. Memcg
charges such an allocation in whichever cache serves it, so drop the
SLAB_ACCOUNT, which no longer adds anything.

Build tested ARCH=x86_64 defconfig with GCC 16.2.0, with
CONFIG_SLAB_BUCKETS as y and n.

Fixes: 734bbc1c97ea7 ("ipc, msg: Use dedicated slab buckets for alloc_msg()")
Assisted-by: LLM
Signed-off-by: Kees Cook <kees@kernel.org>
---
 ipc/msgutil.c | 5 +++--
 1 file changed, 3 insertions(+), 2 deletions(-)

diff --git a/ipc/msgutil.c b/ipc/msgutil.c
index e28f0cecb2ec..1ba8e59cb255 100644
--- a/ipc/msgutil.c
+++ b/ipc/msgutil.c
@@ -43,7 +43,7 @@ static kmem_buckets *msg_buckets __ro_after_init;
 
 static int __init init_msg_buckets(void)
 {
-	msg_buckets = kmem_buckets_create("msg_msg", SLAB_ACCOUNT,
+	msg_buckets = kmem_buckets_create("msg_msg", 0,
 					  sizeof(struct msg_msg),
 					  DATALEN_MSG, NULL);
 
@@ -58,7 +58,8 @@ static struct msg_msg *alloc_msg(size_t len)
 	size_t alen;
 
 	alen = min(len, DATALEN_MSG);
-	msg = kmem_buckets_alloc(msg_buckets, sizeof(*msg) + alen, GFP_KERNEL);
+	msg = kmem_buckets_alloc(msg_buckets, sizeof(*msg) + alen,
+				 GFP_KERNEL_ACCOUNT);
 	if (msg == NULL)
 		return NULL;
 
-- 
2.55.0


^ permalink raw reply	[flat|nested] 9+ messages in thread

* [PATCH net-next v6 3/8] mm/slab: Drop the ctor and flags arguments from kmem_buckets_create()
  2026-10-06  9:20 [PATCH net-next v6 0/8] net: skb: isolate skb data area allocations into a separate bucket Kees Cook
  2026-10-06  9:20 ` [PATCH net-next v6 1/8] mm/slab: Mark the kmem_buckets_create() context as a Context: section Kees Cook
  2026-10-06  9:20 ` [PATCH net-next v6 2/8] ipc, msg: Account msg_msg allocations with GFP_KERNEL_ACCOUNT again Kees Cook
@ 2026-10-06  9:20 ` Kees Cook
  2026-10-06  9:20 ` [PATCH net-next v6 4/8] mm/slab: Give bucket caches the alignment of the caches they mirror Kees Cook
                   ` (4 subsequent siblings)
  7 siblings, 0 replies; 9+ messages in thread
From: Kees Cook @ 2026-10-06  9:20 UTC (permalink / raw)
  To: Vlastimil Babka
  Cc: Kees Cook, Harry Yoo, Andrew Morton, Hao Li, Christoph Lameter,
	David Rientjes, Roman Gushchin, linux-mm, Pedro Falcato,
	Kuniyuki Iwashima, linux-hardening, linux-kernel

A bucket set passes its constructor and slab flags to its own caches,
but its allocations do not always come from them. With
CONFIG_SLAB_BUCKETS=n, kmem_buckets_create() returns ZERO_SIZE_PTR, and
when creating the set fails it returns NULL. Either way
kmem_buckets_alloc() is served by the general kmalloc caches, which have
neither, so a caller cannot depend on them. msg_msg depended on
SLAB_ACCOUNT, which it no longer passes since accounting through
GFP_KERNEL_ACCOUNT instead.

No caller passes a constructor or flags; drop both arguments. The set's
caches keep SLAB_NO_MERGE, which kmem_buckets_create() always added.

Assisted-by: LLM
Signed-off-by: Kees Cook <kees@kernel.org>
---
 include/linux/slab.h |  5 ++---
 ipc/msgutil.c        |  5 ++---
 mm/slab_common.c     | 14 ++++----------
 mm/util.c            |  2 +-
 4 files changed, 9 insertions(+), 17 deletions(-)

diff --git a/include/linux/slab.h b/include/linux/slab.h
index cda126def67a..31f97e2579a7 100644
--- a/include/linux/slab.h
+++ b/include/linux/slab.h
@@ -890,9 +890,8 @@ void *kmem_cache_alloc_lru_noprof(struct kmem_cache *s, struct list_lru *lru,
 bool kmem_cache_charge(void *objp, gfp_t gfpflags);
 void kmem_cache_free(struct kmem_cache *s, void *objp);
 
-kmem_buckets *kmem_buckets_create(const char *name, slab_flags_t flags,
-				  unsigned int useroffset, unsigned int usersize,
-				  void (*ctor)(void *));
+kmem_buckets *kmem_buckets_create(const char *name, unsigned int useroffset,
+				  unsigned int usersize);
 
 /*
  * Bulk allocation and freeing operations. These are accelerated in an
diff --git a/ipc/msgutil.c b/ipc/msgutil.c
index 1ba8e59cb255..10ce3087b089 100644
--- a/ipc/msgutil.c
+++ b/ipc/msgutil.c
@@ -43,9 +43,8 @@ static kmem_buckets *msg_buckets __ro_after_init;
 
 static int __init init_msg_buckets(void)
 {
-	msg_buckets = kmem_buckets_create("msg_msg", 0,
-					  sizeof(struct msg_msg),
-					  DATALEN_MSG, NULL);
+	msg_buckets = kmem_buckets_create("msg_msg", sizeof(struct msg_msg),
+					  DATALEN_MSG);
 
 	return 0;
 }
diff --git a/mm/slab_common.c b/mm/slab_common.c
index 270408ce5a9d..f8bb70d76eb4 100644
--- a/mm/slab_common.c
+++ b/mm/slab_common.c
@@ -415,12 +415,10 @@ static struct kmem_cache *kmem_buckets_cache __ro_after_init;
  *			 allocations via kmem_buckets_alloc()
  * @name: A prefix string which is used in /proc/slabinfo to identify this
  *	  cache. The individual caches with have their sizes as the suffix.
- * @flags: SLAB flags (see kmem_cache_create() for details).
  * @useroffset: Starting offset within an allocation that may be copied
  *		to/from userspace.
  * @usersize: How many bytes, starting at @useroffset, may be copied
  *		to/from userspace.
- * @ctor: A constructor for the objects, run when new allocations are made.
  *
  * Context: Cannot be called within an interrupt, but can be interrupted.
  *
@@ -429,10 +427,8 @@ static struct kmem_cache *kmem_buckets_cache __ro_after_init;
  * subsequent calls to kmem_buckets_alloc() will fall back to kmalloc().
  * (i.e. callers only need to check for NULL on failure.)
  */
-kmem_buckets *kmem_buckets_create(const char *name, slab_flags_t flags,
-				  unsigned int useroffset,
-				  unsigned int usersize,
-				  void (*ctor)(void *))
+kmem_buckets *kmem_buckets_create(const char *name, unsigned int useroffset,
+				  unsigned int usersize)
 {
 	unsigned long mask = 0;
 	unsigned int idx;
@@ -455,8 +451,6 @@ kmem_buckets *kmem_buckets_create(const char *name, slab_flags_t flags,
 	if (WARN_ON(!b))
 		return NULL;
 
-	flags |= SLAB_NO_MERGE;
-
 	for (idx = 0; idx < ARRAY_SIZE(kmalloc_caches[KMALLOC_NORMAL]); idx++) {
 		char *short_size, *cache_name;
 		unsigned int cache_useroffset, cache_usersize;
@@ -487,8 +481,8 @@ kmem_buckets *kmem_buckets_create(const char *name, slab_flags_t flags,
 			if (WARN_ON(!cache_name))
 				goto fail;
 			(*b)[aligned_idx] = kmem_cache_create_usercopy(cache_name, size,
-					0, flags, cache_useroffset,
-					cache_usersize, ctor);
+					0, SLAB_NO_MERGE, cache_useroffset,
+					cache_usersize, NULL);
 			kfree(cache_name);
 			if (WARN_ON(!(*b)[aligned_idx]))
 				goto fail;
diff --git a/mm/util.c b/mm/util.c
index bf0513d1d3d0..0cd125f1ea99 100644
--- a/mm/util.c
+++ b/mm/util.c
@@ -199,7 +199,7 @@ static kmem_buckets *user_buckets __ro_after_init;
 
 static int __init init_user_buckets(void)
 {
-	user_buckets = kmem_buckets_create("memdup_user", 0, 0, INT_MAX, NULL);
+	user_buckets = kmem_buckets_create("memdup_user", 0, INT_MAX);
 
 	return 0;
 }
-- 
2.55.0


^ permalink raw reply	[flat|nested] 9+ messages in thread

* [PATCH net-next v6 4/8] mm/slab: Give bucket caches the alignment of the caches they mirror
  2026-10-06  9:20 [PATCH net-next v6 0/8] net: skb: isolate skb data area allocations into a separate bucket Kees Cook
                   ` (2 preceding siblings ...)
  2026-10-06  9:20 ` [PATCH net-next v6 3/8] mm/slab: Drop the ctor and flags arguments from kmem_buckets_create() Kees Cook
@ 2026-10-06  9:20 ` Kees Cook
  2026-10-06  9:20 ` [PATCH net-next v6 5/8] mm/slab: Add kmem_buckets_destroy() Kees Cook
                   ` (3 subsequent siblings)
  7 siblings, 0 replies; 9+ messages in thread
From: Kees Cook @ 2026-10-06  9:20 UTC (permalink / raw)
  To: Vlastimil Babka
  Cc: Kees Cook, Harry Yoo, Andrew Morton, Hao Li, Christoph Lameter,
	David Rientjes, Roman Gushchin, linux-mm, Pedro Falcato,
	Kuniyuki Iwashima, linux-hardening, linux-kernel

A bucket set is created with kmem_cache_create_usercopy(..., align = 0),
so calculate_alignment() falls back to arch_slab_minalign(), typically 8
bytes. The general kmalloc caches it stands in for are created through
create_boot_cache(), which starts from ARCH_KMALLOC_MINALIGN and raises
it to the largest power-of-two divisor of the size:

	if (flags & SLAB_KMALLOC)
		align = max(align, 1U << (ffs(size) - 1));

This is only a problem when slab metadata is enabled with
CONFIG_KASAN=y, CONFIG_SLUB_DEBUG_ON=y, or "slab_debug=...", because
metadata changes the stride size off a power of two, for example:

	size 128:  bucket align=8 size=224  | kmalloc align=128  size=384
	size 512:  bucket align=8 size=608  | kmalloc align=512  size=1536
	size 2048: bucket align=8 size=2144 | kmalloc align=2048 size=6144

So bucket allocations will fail the IS_ALIGNED(p, ARCH_DMA_MINALIGN)
check, potentially creating problems for non-coherent DMA situation.

Take the alignment from the kmalloc cache being mirrored, which is where
the size and the name suffix already come from, so a caller moving from
kmalloc() keeps the alignment it had.

Don't take an alignment from the caller instead. With
CONFIG_SLAB_BUCKETS=n, or when kmem_buckets_create() fails,
kmem_buckets_alloc() is served by the general kmalloc caches, which give
kmalloc()'s alignment and no other.

Fixes: b32801d1255be ("mm/slab: Introduce kmem_buckets_create() and family")
Assisted-by: LLM
Signed-off-by: Kees Cook <kees@kernel.org>
---
 mm/slab_common.c | 3 ++-
 1 file changed, 2 insertions(+), 1 deletion(-)

diff --git a/mm/slab_common.c b/mm/slab_common.c
index f8bb70d76eb4..885aafa23da7 100644
--- a/mm/slab_common.c
+++ b/mm/slab_common.c
@@ -481,7 +481,8 @@ kmem_buckets *kmem_buckets_create(const char *name, unsigned int useroffset,
 			if (WARN_ON(!cache_name))
 				goto fail;
 			(*b)[aligned_idx] = kmem_cache_create_usercopy(cache_name, size,
-					0, SLAB_NO_MERGE, cache_useroffset,
+					kmalloc_caches[KMALLOC_NORMAL][idx]->align,
+					SLAB_NO_MERGE, cache_useroffset,
 					cache_usersize, NULL);
 			kfree(cache_name);
 			if (WARN_ON(!(*b)[aligned_idx]))
-- 
2.55.0


^ permalink raw reply	[flat|nested] 9+ messages in thread

* [PATCH net-next v6 5/8] mm/slab: Add kmem_buckets_destroy()
  2026-10-06  9:20 [PATCH net-next v6 0/8] net: skb: isolate skb data area allocations into a separate bucket Kees Cook
                   ` (3 preceding siblings ...)
  2026-10-06  9:20 ` [PATCH net-next v6 4/8] mm/slab: Give bucket caches the alignment of the caches they mirror Kees Cook
@ 2026-10-06  9:20 ` Kees Cook
  2026-10-06  9:20 ` [PATCH net-next v6 6/8] mm/slab: Add tests for the existing kmem_buckets behaviour Kees Cook
                   ` (2 subsequent siblings)
  7 siblings, 0 replies; 9+ messages in thread
From: Kees Cook @ 2026-10-06  9:20 UTC (permalink / raw)
  To: Vlastimil Babka
  Cc: Kees Cook, Harry Yoo, Andrew Morton, Hao Li, Christoph Lameter,
	David Rientjes, Roman Gushchin, linux-mm, Pedro Falcato,
	Kuniyuki Iwashima, linux-hardening, linux-kernel

kmem_buckets_create() intentionally had no "destroy" counterpart. Every
caller has lived in core kernel code and creates its set once at boot,
so nothing has needed to take one down. However, KUnit tests may be
built module, so we need it now to support the coming tests.

Some caches have size aliases, so the same pointer is stored at more
then one index, so we have to save it, clear all matching instances, and
then free the saved cache pointer. (This is what the bitmap was tracking
before in the "allocation failed" error path.)

When CONFIG_SLAB_BUCKETS=n the whole body compiles away, matching the
ZERO_SIZE_PTR that kmem_buckets_create() hands back in that configuration.

Link: https://lore.kernel.org/all/20240809073309.2134488-1-kees@kernel.org/
Assisted-by: LLM
Signed-off-by: Kees Cook <kees@kernel.org>
---
 include/linux/slab.h |  1 +
 mm/slab_common.c     | 50 +++++++++++++++++++++++++++++++++++++-------
 2 files changed, 44 insertions(+), 7 deletions(-)

diff --git a/include/linux/slab.h b/include/linux/slab.h
index 31f97e2579a7..034d4d0ece00 100644
--- a/include/linux/slab.h
+++ b/include/linux/slab.h
@@ -892,6 +892,7 @@ void kmem_cache_free(struct kmem_cache *s, void *objp);
 
 kmem_buckets *kmem_buckets_create(const char *name, unsigned int useroffset,
 				  unsigned int usersize);
+void kmem_buckets_destroy(kmem_buckets *bucket);
 
 /*
  * Bulk allocation and freeing operations. These are accelerated in an
diff --git a/mm/slab_common.c b/mm/slab_common.c
index 885aafa23da7..fb1dd15953a7 100644
--- a/mm/slab_common.c
+++ b/mm/slab_common.c
@@ -430,12 +430,9 @@ static struct kmem_cache *kmem_buckets_cache __ro_after_init;
 kmem_buckets *kmem_buckets_create(const char *name, unsigned int useroffset,
 				  unsigned int usersize)
 {
-	unsigned long mask = 0;
 	unsigned int idx;
 	kmem_buckets *b;
 
-	BUILD_BUG_ON(ARRAY_SIZE(kmalloc_caches[KMALLOC_NORMAL]) > BITS_PER_LONG);
-
 	/*
 	 * When the separate buckets API is not built in, just return
 	 * a non-NULL value for the kmem_buckets pointer, which will be
@@ -487,7 +484,6 @@ kmem_buckets *kmem_buckets_create(const char *name, unsigned int useroffset,
 			kfree(cache_name);
 			if (WARN_ON(!(*b)[aligned_idx]))
 				goto fail;
-			set_bit(aligned_idx, &mask);
 		}
 		if (idx != aligned_idx)
 			(*b)[idx] = (*b)[aligned_idx];
@@ -496,14 +492,54 @@ kmem_buckets *kmem_buckets_create(const char *name, unsigned int useroffset,
 	return b;
 
 fail:
-	for_each_set_bit(idx, &mask, ARRAY_SIZE(kmalloc_caches[KMALLOC_NORMAL]))
-		kmem_cache_destroy((*b)[idx]);
-	kmem_cache_free(kmem_buckets_cache, b);
+	kmem_buckets_destroy(b);
 
 	return NULL;
 }
 EXPORT_SYMBOL(kmem_buckets_create);
 
+/**
+ * kmem_buckets_destroy - Destroy a set of caches made by kmem_buckets_create()
+ * @bucket: The set to destroy, which may be NULL.
+ *
+ * Destroys each cache in @bucket and then frees @bucket itself. As for
+ * kmem_cache_destroy(), every object allocated from @bucket must have been
+ * freed beforehand, and @bucket must not be used afterwards.
+ *
+ * Context: Process context. May sleep, as kmem_cache_destroy() takes the
+ *	    slab mutex and can wait on RCU callbacks for each cache.
+ */
+void kmem_buckets_destroy(kmem_buckets *bucket)
+{
+	unsigned int idx, i;
+
+	if (!IS_ENABLED(CONFIG_SLAB_BUCKETS) || ZERO_OR_NULL_PTR(bucket))
+		return;
+
+	for (idx = 0; idx < ARRAY_SIZE(kmalloc_caches[KMALLOC_NORMAL]); idx++) {
+		struct kmem_cache *cache = (*bucket)[idx];
+
+		if (!cache)
+			continue;
+
+		/*
+		 * Sizes that kmalloc rounds up to a larger size class share
+		 * that class's cache, which kmem_buckets_create() then stores
+		 * at each of their indices.
+		 * Drop every reference to it before destroying it, so that no
+		 * later pass reads a pointer to a cache that is already gone.
+		 */
+		for (i = idx; i < ARRAY_SIZE(kmalloc_caches[KMALLOC_NORMAL]); i++)
+			if ((*bucket)[i] == cache)
+				(*bucket)[i] = NULL;
+
+		kmem_cache_destroy(cache);
+	}
+
+	kmem_cache_free(kmem_buckets_cache, bucket);
+}
+EXPORT_SYMBOL(kmem_buckets_destroy);
+
 /*
  * For a given kmem_cache, kmem_cache_destroy() should only be called
  * once or there will be a use-after-free problem. The actual deletion
-- 
2.55.0


^ permalink raw reply	[flat|nested] 9+ messages in thread

* [PATCH net-next v6 6/8] mm/slab: Add tests for the existing kmem_buckets behaviour
  2026-10-06  9:20 [PATCH net-next v6 0/8] net: skb: isolate skb data area allocations into a separate bucket Kees Cook
                   ` (4 preceding siblings ...)
  2026-10-06  9:20 ` [PATCH net-next v6 5/8] mm/slab: Add kmem_buckets_destroy() Kees Cook
@ 2026-10-06  9:20 ` Kees Cook
  2026-10-06  9:20 ` [PATCH net-next v6 7/8] mm/slab: Provide kmalloc type fallback for bucket allocations Kees Cook
  2026-10-06  9:20 ` [PATCH net-next v6 8/8] net: skb: isolate skb data area allocations into a separate bucket Kees Cook
  7 siblings, 0 replies; 9+ messages in thread
From: Kees Cook @ 2026-10-06  9:20 UTC (permalink / raw)
  To: Vlastimil Babka
  Cc: Kees Cook, Harry Yoo, Andrew Morton, Hao Li, Christoph Lameter,
	David Rientjes, Roman Gushchin, linux-mm, Pedro Falcato,
	Kuniyuki Iwashima, linux-hardening, linux-kernel

kmem_buckets has had no test coverage since it was added. Add tests,
including stuff unique to the bucket design:

 - A bucket allocation comes from a cache of the set's own, and that
   cache carries SLAB_NO_MERGE.

 - Each size is served by a cache of the set, of the size kmalloc()
   rounds it up to, including 96 and 192, which are not powers of two and
   which kmalloc shares with a larger class on some configurations. Sizes
   above KMALLOC_MAX_CACHE_SIZE go to the page allocator instead, bucket
   set or not.

 - Each cache is aligned like the kmalloc cache it mirrors.

 - With CONFIG_SLAB_BUCKETS=n, kmem_buckets_create() still returns a
   non-NULL (zero size alloc pointer), so that callers only have to check
   for failure, and allocations through it come from the general caches.

 - Destroying a set takes its caches down rather than only freeing the
   set, which is what a module creating one on each load depends on.
   A set freed without its caches would leave the names taken, and the
   next load would warn about every one of them. This test skips when
   KFENCE serves its allocation, since a KFENCE object is in none of the
   cache's slabs for the teardown to find.

The tests skip rather than compile out, which is useful for testing
the CONFIG_SLAB_BUCKETS=n behaviors.

Built and tests passing (with expected skips) on ARCH=x86_64 defconfig
with GCC 16.2.0, with CONFIG_SLAB_BUCKETS as y and n.

Assisted-by: LLM
Signed-off-by: Kees Cook <kees@kernel.org>
---
 lib/tests/slub_kunit.c | 228 +++++++++++++++++++++++++++++++++++++++++
 1 file changed, 228 insertions(+)

diff --git a/lib/tests/slub_kunit.c b/lib/tests/slub_kunit.c
index e3b63f0338d5..a2a15a49c5d7 100644
--- a/lib/tests/slub_kunit.c
+++ b/lib/tests/slub_kunit.c
@@ -1,6 +1,7 @@
 // SPDX-License-Identifier: GPL-2.0
 #include <kunit/test.h>
 #include <kunit/test-bug.h>
+#include <kunit/resource.h>
 #include <linux/mm.h>
 #include <linux/slab.h>
 #include <linux/module.h>
@@ -9,6 +10,7 @@
 #include <linux/delay.h>
 #include <linux/perf_event.h>
 #include <linux/kprobes.h>
+#include <linux/kfence.h>
 #include "../mm/slab.h"
 
 static struct kunit_resource resource;
@@ -474,6 +476,227 @@ static int test_init(struct kunit *test)
 	return 0;
 }
 
+/* Destroy buckets on test exit so a failed KUNIT_ASSERT_*() doesn't leak. */
+KUNIT_DEFINE_ACTION_WRAPPER(destroy_buckets, kmem_buckets_destroy, kmem_buckets *);
+
+#define KUNIT_ASSERT_BUCKETS_CREATED(test, b)					\
+	do {									\
+		KUNIT_ASSERT_NOT_NULL(test, b);					\
+		KUNIT_ASSERT_EQ(test, 0,					\
+				kunit_add_action_or_reset(test,			\
+							  destroy_buckets, b)); \
+	} while (0)
+
+/*
+ * The cache an allocation came from, or NULL if it came from no cache at
+ * all, e.g. a size too big for any of them is served by the page allocator.
+ */
+static struct kmem_cache *cache_of(void *p)
+{
+	struct slab *slab = virt_to_slab(p);
+
+	return slab ? slab->slab_cache : NULL;
+}
+
+/*
+ * A bucket set exists to keep its allocations out of the caches everything
+ * else uses, so check the two things that make that true: they come from a
+ * cache of the set's own, and that cache is never merged into another.
+ */
+static void test_kmem_buckets_isolation(struct kunit *test)
+{
+	struct kmem_cache *bucket_cache, *general_cache;
+	kmem_buckets *b;
+	void *p, *q;
+
+	if (!IS_ENABLED(CONFIG_SLAB_BUCKETS))
+		kunit_skip(test, "needs CONFIG_SLAB_BUCKETS");
+
+	b = kmem_buckets_create("isolated_buckets", 0, INT_MAX);
+	KUNIT_ASSERT_BUCKETS_CREATED(test, b);
+
+	/*
+	 * Free each allocation before asserting on the next one: the cache
+	 * outlives its objects, so nothing below needs them, and an assertion
+	 * that leaves one behind would make the deferred teardown report a
+	 * cache that is still in use.
+	 */
+	p = kmem_buckets_alloc(b, 128, GFP_KERNEL);
+	KUNIT_ASSERT_NOT_NULL(test, p);
+	bucket_cache = cache_of(p);
+	kfree(p);
+	KUNIT_ASSERT_NOT_NULL(test, bucket_cache);
+
+	KUNIT_EXPECT_TRUE_MSG(test, strstarts(bucket_cache->name, "isolated_buckets-"),
+			      "expected a bucket cache, got %s", bucket_cache->name);
+
+	/*
+	 * Cache merging is on by default, and a bucket cache merged into a
+	 * same-sized general one would quietly undo the whole separation.
+	 */
+	KUNIT_EXPECT_TRUE(test, bucket_cache->flags & SLAB_NO_MERGE);
+
+	q = kmalloc(128, GFP_KERNEL);
+	KUNIT_ASSERT_NOT_NULL(test, q);
+	general_cache = cache_of(q);
+	kfree(q);
+	KUNIT_ASSERT_NOT_NULL(test, general_cache);
+
+	KUNIT_EXPECT_PTR_NE(test, bucket_cache, general_cache);
+}
+
+/*
+ * Each size is served by a cache of the set, of the size kmalloc() rounds it
+ * up to, including the size classes that are not powers of two, which kmalloc
+ * shares with a larger class on some configurations. Sizes past the largest
+ * cache are served by the page allocator, bucket set or not.
+ */
+static void test_kmem_buckets_sizes(struct kunit *test)
+{
+	static const size_t sizes[] = { 8, 96, 192, 1024, 4096 };
+	struct kmem_cache *c;
+	kmem_buckets *b;
+	void *p;
+	int i;
+
+	if (!IS_ENABLED(CONFIG_SLAB_BUCKETS))
+		kunit_skip(test, "needs CONFIG_SLAB_BUCKETS");
+
+	b = kmem_buckets_create("sized_buckets", 0, INT_MAX);
+	KUNIT_ASSERT_BUCKETS_CREATED(test, b);
+
+	for (i = 0; i < ARRAY_SIZE(sizes); i++) {
+		p = kmem_buckets_alloc(b, sizes[i], GFP_KERNEL);
+		KUNIT_ASSERT_NOT_NULL(test, p);
+		c = cache_of(p);
+		kfree(p);
+		KUNIT_ASSERT_NOT_NULL(test, c);
+
+		KUNIT_EXPECT_TRUE_MSG(test, strstarts(c->name, "sized_buckets-"),
+				      "size %zu: expected a bucket cache, got %s",
+				      sizes[i], c->name);
+		KUNIT_EXPECT_EQ_MSG(test, c->object_size,
+				    kmalloc_size_roundup(sizes[i]),
+				    "size %zu: served by %s", sizes[i], c->name);
+	}
+
+	/* Too big for any cache: a folio from the page allocator, not a slab. */
+	p = kmem_buckets_alloc(b, KMALLOC_MAX_CACHE_SIZE + 1, GFP_KERNEL);
+	KUNIT_ASSERT_NOT_NULL(test, p);
+	c = cache_of(p);
+	kfree(p);
+
+	KUNIT_EXPECT_NULL(test, c);
+}
+
+/*
+ * A bucket cache stands in for a kmalloc cache, so it has to be aligned like
+ * one. The DMA layer decides whether a buffer needs bouncing from its size,
+ * on the grounds that a kmalloc cache of that size is already aligned for
+ * the device, so a weaker alignment here is not something a caller can see
+ * coming. Without slab debugging the size implies the alignment and this
+ * holds either way; with it, only the cache's own alignment does.
+ */
+static void test_kmem_buckets_alignment(struct kunit *test)
+{
+	static const size_t sizes[] = { 128, 512, 2048 };
+	struct kmem_cache *bucket_cache, *general_cache;
+	kmem_buckets *b;
+	void *p;
+	int i;
+
+	if (!IS_ENABLED(CONFIG_SLAB_BUCKETS))
+		kunit_skip(test, "needs CONFIG_SLAB_BUCKETS");
+
+	b = kmem_buckets_create("aligned_buckets", 0, INT_MAX);
+	KUNIT_ASSERT_BUCKETS_CREATED(test, b);
+
+	for (i = 0; i < ARRAY_SIZE(sizes); i++) {
+		p = kmem_buckets_alloc(b, sizes[i], GFP_KERNEL);
+		KUNIT_ASSERT_NOT_NULL(test, p);
+		bucket_cache = cache_of(p);
+		KUNIT_EXPECT_TRUE_MSG(test,
+				      IS_ALIGNED((unsigned long)p, ARCH_DMA_MINALIGN),
+				      "size %zu: object %p is not %d byte aligned",
+				      sizes[i], p, (int)ARCH_DMA_MINALIGN);
+		kfree(p);
+
+		p = kmalloc(sizes[i], GFP_KERNEL);
+		KUNIT_ASSERT_NOT_NULL(test, p);
+		general_cache = cache_of(p);
+		kfree(p);
+
+		KUNIT_ASSERT_NOT_NULL(test, bucket_cache);
+		KUNIT_ASSERT_NOT_NULL(test, general_cache);
+		KUNIT_EXPECT_EQ_MSG(test, bucket_cache->align, general_cache->align,
+				    "size %zu: bucket cache aligned to %u, %s to %u",
+				    sizes[i], bucket_cache->align,
+				    general_cache->name, general_cache->align);
+	}
+}
+
+/*
+ * With the feature compiled out, kmem_buckets_create() still returns
+ * something non-NULL so that callers only have to check for failure, and
+ * allocations through it work (i.e. come from the general caches).
+ */
+static void test_kmem_buckets_disabled(struct kunit *test)
+{
+	kmem_buckets *b;
+	struct kmem_cache *c;
+	void *p;
+
+	if (IS_ENABLED(CONFIG_SLAB_BUCKETS))
+		kunit_skip(test, "only meaningful without CONFIG_SLAB_BUCKETS");
+
+	b = kmem_buckets_create("disabled_buckets", 0, INT_MAX);
+	KUNIT_ASSERT_BUCKETS_CREATED(test, b);
+
+	p = kmem_buckets_alloc(b, 128, GFP_KERNEL);
+	KUNIT_ASSERT_NOT_NULL(test, p);
+	c = cache_of(p);
+	kfree(p);
+	KUNIT_ASSERT_NOT_NULL(test, c);
+
+	KUNIT_EXPECT_TRUE_MSG(test, !strstarts(c->name, "disabled_buckets-"),
+			      "expected a general cache, got %s", c->name);
+}
+
+/* Destroying a set has to take its caches down, not just free the set. */
+static void test_kmem_buckets_destroy(struct kunit *test)
+{
+	kmem_buckets *b;
+	void *p;
+
+	if (!IS_ENABLED(CONFIG_SLAB_BUCKETS))
+		kunit_skip(test, "needs CONFIG_SLAB_BUCKETS");
+
+	b = kmem_buckets_create("destroyed_buckets", 0, INT_MAX);
+	KUNIT_ASSERT_BUCKETS_CREATED(test, b);
+
+	/*
+	 * Deliberately leaked, as test_leak_destroy() leaks its own: the
+	 * teardown below has to find it. kmem_cache_destroy() unlists the
+	 * cache either way, so the name is still released.
+	 */
+	p = kmem_buckets_alloc(b, 128, GFP_KERNEL);
+	KUNIT_ASSERT_NOT_NULL(test, p);
+
+	/*
+	 * A KFENCE object is in none of the cache's slabs, so the teardown
+	 * would not find it to report.
+	 */
+	if (is_kfence_address(p)) {
+		kfree(p);
+		kunit_skip(test, "the allocation came from KFENCE");
+	}
+
+	/* Tear the set down now, rather than at exit, to check the report. */
+	kunit_release_action(test, destroy_buckets, b);
+
+	KUNIT_EXPECT_EQ(test, 2, slab_errors);
+}
+
 static struct kunit_case test_cases[] = {
 	KUNIT_CASE(test_clobber_zone),
 
@@ -495,6 +718,11 @@ static struct kunit_case test_cases[] = {
 #if defined(CONFIG_KPROBES) && defined(CONFIG_SMP)
 	KUNIT_CASE_SLOW(test_kmalloc_nolock_and_friends_kprobe),
 #endif
+	KUNIT_CASE(test_kmem_buckets_isolation),
+	KUNIT_CASE(test_kmem_buckets_sizes),
+	KUNIT_CASE(test_kmem_buckets_alignment),
+	KUNIT_CASE(test_kmem_buckets_disabled),
+	KUNIT_CASE(test_kmem_buckets_destroy),
 	{}
 };
 
-- 
2.55.0


^ permalink raw reply	[flat|nested] 9+ messages in thread

* [PATCH net-next v6 7/8] mm/slab: Provide kmalloc type fallback for bucket allocations
  2026-10-06  9:20 [PATCH net-next v6 0/8] net: skb: isolate skb data area allocations into a separate bucket Kees Cook
                   ` (5 preceding siblings ...)
  2026-10-06  9:20 ` [PATCH net-next v6 6/8] mm/slab: Add tests for the existing kmem_buckets behaviour Kees Cook
@ 2026-10-06  9:20 ` Kees Cook
  2026-10-06  9:20 ` [PATCH net-next v6 8/8] net: skb: isolate skb data area allocations into a separate bucket Kees Cook
  7 siblings, 0 replies; 9+ messages in thread
From: Kees Cook @ 2026-10-06  9:20 UTC (permalink / raw)
  To: Vlastimil Babka
  Cc: Kees Cook, Harry Yoo, Andrew Morton, Hao Li, Christoph Lameter,
	David Rientjes, Roman Gushchin, linux-mm, Pedro Falcato,
	Kuniyuki Iwashima, linux-hardening, Johannes Weiner,
	Michal Hocko, Shakeel Butt, Muchun Song, cgroups, linux-kernel

kmem_buckets_create() clones kmalloc_caches[KMALLOC_NORMAL].
kmalloc_slab() figures out the kmalloc type the caller asks for, but
then ignored it whenever a bucket set was in use, returning a normal
cache regardless. That breaks an allocation that needs other pages: a
GFP_DMA allocation would not get memory from ZONE_DMA, and a
__GFP_RECLAIMABLE one would miss the reclaimable caches. None of the
current users do this, so there is no problem, but it makes adding new
users fragile. For example, skb data[1] needs to handle GFP_DMA (rarely).

Send those allocations to the general caches instead, so nothing breaks
and regular allocations remain isolated in the set. Accounted
allocations stay in the set: memcg charges each object on its own, in
any cache, so a bucket cache serves them as well as kmalloc-cg-* does.

Built and tests pass with ARCH=x86_64 defconfig with GCC 16.2.0, with
CONFIG_SLAB_BUCKETS as y and n, and with CONFIG_MEMCG as y, n, and y
with "cgroup.memory=nokmem".

Assisted-by: LLM
Link: https://lore.kernel.org/all/04debe19-bbe8-4b5f-9668-753d1f97832d@redhat.com/ [1]
Signed-off-by: Kees Cook <kees@kernel.org>
---
 mm/slab.h              | 19 +++++++++++--
 lib/tests/slub_kunit.c | 63 ++++++++++++++++++++++++++++++++++++++++++
 mm/slab_common.c       |  5 ++++
 3 files changed, 85 insertions(+), 2 deletions(-)

diff --git a/mm/slab.h b/mm/slab.h
index 8fd6835e4235..af39a4e47c9e 100644
--- a/mm/slab.h
+++ b/mm/slab.h
@@ -421,6 +421,22 @@ static inline unsigned int size_index_elem(unsigned int bytes)
 	return (bytes - 1) / 8;
 }
 
+/*
+ * Which set of buckets to use for the given kmalloc_cache_type. A bucket set
+ * mirrors the KMALLOC_NORMAL caches, and also serves accounted allocations:
+ * memcg charges each object on its own, in any cache. Types that need
+ * different pages (DMA, reclaimable) or no obj_exts fall back to the
+ * general caches.
+ */
+static inline kmem_buckets *
+kmalloc_choose_bucket(kmem_buckets *bucket, enum kmalloc_cache_type type)
+{
+	if (bucket && (type <= KMALLOC_PARTITION_END || type == KMALLOC_CGROUP))
+		return bucket;
+
+	return &kmalloc_caches[type];
+}
+
 /*
  * Find the kmem_cache structure that serves a given size of
  * allocation
@@ -438,8 +454,7 @@ kmalloc_slab(size_t size, kmem_buckets *b, gfp_t flags, kmalloc_token_t token,
 	if (alloc_flags & SLAB_ALLOC_NO_OBJ_EXT)
 		type = KMALLOC_NO_OBJ_EXT;
 
-	if (!b)
-		b = &kmalloc_caches[type];
+	b = kmalloc_choose_bucket(b, type);
 	if (size <= 192)
 		index = kmalloc_size_index[size_index_elem(size)];
 	else
diff --git a/lib/tests/slub_kunit.c b/lib/tests/slub_kunit.c
index a2a15a49c5d7..1e6fcbbf8409 100644
--- a/lib/tests/slub_kunit.c
+++ b/lib/tests/slub_kunit.c
@@ -697,6 +697,68 @@ static void test_kmem_buckets_destroy(struct kunit *test)
 	KUNIT_EXPECT_EQ(test, 2, slab_errors);
 }
 
+/*
+ * A bucket set mirrors the normal kmalloc caches, which can serve accounted
+ * allocations too, so those stay in the set. An allocation that needs other
+ * pages (DMA or reclaimable) has to come from the general caches. Check that
+ * it does, rather than being served a cache that does not satisfy what the
+ * flags asked for.
+ */
+static void test_kmem_buckets_type_fallback(struct kunit *test)
+{
+	struct kmem_cache *c;
+	kmem_buckets *b;
+	void *p;
+
+	if (!IS_ENABLED(CONFIG_SLAB_BUCKETS))
+		kunit_skip(test, "needs CONFIG_SLAB_BUCKETS");
+
+	b = kmem_buckets_create("test_buckets", 0, INT_MAX);
+	KUNIT_ASSERT_BUCKETS_CREATED(test, b);
+
+	/* A plain allocation stays isolated in the bucket set. */
+	p = kmem_buckets_alloc(b, 128, GFP_KERNEL);
+	KUNIT_ASSERT_NOT_NULL(test, p);
+	c = cache_of(p);
+	kfree(p);
+	KUNIT_ASSERT_NOT_NULL(test, c);
+
+	KUNIT_EXPECT_TRUE_MSG(test, strstarts(c->name, "test_buckets-"),
+			      "expected a bucket cache, got %s", c->name);
+
+	/* One that needs ZONE_DMA cannot, so it falls back. */
+	if (IS_ENABLED(CONFIG_ZONE_DMA)) {
+		p = kmem_buckets_alloc(b, 128, GFP_KERNEL | GFP_DMA);
+		KUNIT_ASSERT_NOT_NULL(test, p);
+		c = cache_of(p);
+		kfree(p);
+		KUNIT_ASSERT_NOT_NULL(test, c);
+
+		KUNIT_EXPECT_TRUE_MSG(test, strstarts(c->name, "dma-kmalloc-"),
+				      "expected a DMA cache, got %s", c->name);
+	}
+
+	/* Nor can one that is reclaimable. */
+	p = kmem_buckets_alloc(b, 128, GFP_KERNEL | __GFP_RECLAIMABLE);
+	KUNIT_ASSERT_NOT_NULL(test, p);
+	c = cache_of(p);
+	kfree(p);
+	KUNIT_ASSERT_NOT_NULL(test, c);
+
+	KUNIT_EXPECT_TRUE_MSG(test, strstarts(c->name, "kmalloc-rcl-"),
+			      "expected a reclaimable cache, got %s", c->name);
+
+	/* An accounted allocation stays in the set; memcg charges it there. */
+	p = kmem_buckets_alloc(b, 128, GFP_KERNEL | __GFP_ACCOUNT);
+	KUNIT_ASSERT_NOT_NULL(test, p);
+	c = cache_of(p);
+	kfree(p);
+	KUNIT_ASSERT_NOT_NULL(test, c);
+
+	KUNIT_EXPECT_TRUE_MSG(test, strstarts(c->name, "test_buckets-"),
+			      "expected a bucket cache, got %s", c->name);
+}
+
 static struct kunit_case test_cases[] = {
 	KUNIT_CASE(test_clobber_zone),
 
@@ -723,6 +785,7 @@ static struct kunit_case test_cases[] = {
 	KUNIT_CASE(test_kmem_buckets_alignment),
 	KUNIT_CASE(test_kmem_buckets_disabled),
 	KUNIT_CASE(test_kmem_buckets_destroy),
+	KUNIT_CASE(test_kmem_buckets_type_fallback),
 	{}
 };
 
diff --git a/mm/slab_common.c b/mm/slab_common.c
index fb1dd15953a7..6fa02b4ff775 100644
--- a/mm/slab_common.c
+++ b/mm/slab_common.c
@@ -420,6 +420,11 @@ static struct kmem_cache *kmem_buckets_cache __ro_after_init;
  * @usersize: How many bytes, starting at @useroffset, may be copied
  *		to/from userspace.
  *
+ * Accounted (__GFP_ACCOUNT) allocations are served by the set like any
+ * other. Allocations that need DMA or reclaimable memory are served by the
+ * general kmalloc caches instead, without the set's isolation or usercopy
+ * region.
+ *
  * Context: Cannot be called within an interrupt, but can be interrupted.
  *
  * Return: a pointer to the cache on success, NULL on failure. When
-- 
2.55.0


^ permalink raw reply	[flat|nested] 9+ messages in thread

* [PATCH net-next v6 8/8] net: skb: isolate skb data area allocations into a separate bucket
  2026-10-06  9:20 [PATCH net-next v6 0/8] net: skb: isolate skb data area allocations into a separate bucket Kees Cook
                   ` (6 preceding siblings ...)
  2026-10-06  9:20 ` [PATCH net-next v6 7/8] mm/slab: Provide kmalloc type fallback for bucket allocations Kees Cook
@ 2026-10-06  9:20 ` Kees Cook
  7 siblings, 0 replies; 9+ messages in thread
From: Kees Cook @ 2026-10-06  9:20 UTC (permalink / raw)
  To: Vlastimil Babka
  Cc: Kees Cook, Pedro Falcato, David S. Miller, Eric Dumazet,
	Jakub Kicinski, Paolo Abeni, Simon Horman, Willem de Bruijn,
	Jason Xing, netdev, Kuniyuki Iwashima, linux-hardening, linux-mm,
	linux-kernel

From: Pedro Falcato <pfalcato@suse.de>

SKB data area allocations (as done from alloc_skb()) use kmalloc().
These allocations can be variably sized and their contents can be more
or less controlled from userspace, which makes them useful for attackers
that want to overwrite a use-after-free'd object from the same kmalloc slab
(which often just requires the sizes to roughly match into the same kmalloc
bucket). [0] is an easy example of an exploit that uses netlink skb
allocation to target another similarly-sized accidentally freed object.

While other mitigations like CONFIG_RANDOM_KMALLOC_CACHES exist, these are
probabilistic. Use the existing kmem buckets API to further isolate these
allocations in a guaranteed fashion, when CONFIG_SLAB_BUCKETS=y.

AF_UNIX sets sk_allocation to GFP_KERNEL_ACCOUNT, and those skb data
areas, the ones most worth isolating, stay in the set, where memcg
charges them as it would in the general caches. GFP_DMA falls back to
the general caches, being passed to an skb allocator only by rare
devices.

Link: https://github.com/google/security-research/blob/master/pocs/linux/kernelctf/CVE-2023-4207_lts_cos_mitigation_2/docs/exploit.md [0]
Reviewed-by: Kees Cook <kees@kernel.org>
Signed-off-by: Pedro Falcato <pfalcato@suse.de>
Acked-by: Paolo Abeni <pabeni@redhat.com>
Signed-off-by: Kees Cook <kees@kernel.org>
---
 net/core/skbuff.c | 10 +++++++---
 1 file changed, 7 insertions(+), 3 deletions(-)

diff --git a/net/core/skbuff.c b/net/core/skbuff.c
index 966af3beed94..6f6b5f4cb39f 100644
--- a/net/core/skbuff.c
+++ b/net/core/skbuff.c
@@ -586,6 +586,8 @@ struct sk_buff *napi_build_skb(void *data, unsigned int frag_size)
 }
 EXPORT_SYMBOL(napi_build_skb);
 
+static kmem_buckets *skb_data_buckets __ro_after_init;
+
 static void *kmalloc_pfmemalloc(size_t obj_size, gfp_t flags, int node)
 {
 	if (!gfp_pfmemalloc_allowed(flags))
@@ -593,11 +595,12 @@ static void *kmalloc_pfmemalloc(size_t obj_size, gfp_t flags, int node)
 	if (!obj_size)
 		return kmem_cache_alloc_node(net_hotdata.skb_small_head_cache,
 					     flags, node);
-	return kmalloc_node_track_caller(obj_size, flags, node);
+	return kmem_buckets_alloc_node_track_caller(skb_data_buckets, obj_size,
+						    flags, node);
 }
 
 /*
- * kmalloc_reserve is a wrapper around kmalloc_node_track_caller that tells
+ * kmalloc_reserve is a wrapper around a caller-tracked kmalloc that tells
  * the caller if emergency pfmemalloc reserves are being used. If it is and
  * the socket is later found to be SOCK_MEMALLOC then PFMEMALLOC reserves
  * may be used. Otherwise, the packet data may be discarded until enough
@@ -634,7 +637,7 @@ static void *kmalloc_reserve(unsigned int *size, gfp_t flags, int node,
 	 * Try a regular allocation, when that fails and we're not entitled
 	 * to the reserves, fail.
 	 */
-	obj = kmalloc_node_track_caller(obj_size,
+	obj = kmem_buckets_alloc_node_track_caller(skb_data_buckets, obj_size,
 					flags | __GFP_NOMEMALLOC | __GFP_NOWARN,
 					node);
 	if (likely(obj))
@@ -5235,6 +5238,7 @@ void __init skb_init(void)
 						0,
 						SKB_SMALL_HEAD_HEADROOM,
 						NULL);
+	skb_data_buckets = kmem_buckets_create("skb_data", 0, INT_MAX);
 	skb_extensions_init();
 }
 
-- 
2.55.0


^ permalink raw reply	[flat|nested] 9+ messages in thread

end of thread, other threads:[~2026-10-06  9:20 UTC | newest]

Thread overview: 9+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-06  9:20 [PATCH net-next v6 0/8] net: skb: isolate skb data area allocations into a separate bucket Kees Cook
2026-10-06  9:20 ` [PATCH net-next v6 1/8] mm/slab: Mark the kmem_buckets_create() context as a Context: section Kees Cook
2026-10-06  9:20 ` [PATCH net-next v6 2/8] ipc, msg: Account msg_msg allocations with GFP_KERNEL_ACCOUNT again Kees Cook
2026-10-06  9:20 ` [PATCH net-next v6 3/8] mm/slab: Drop the ctor and flags arguments from kmem_buckets_create() Kees Cook
2026-10-06  9:20 ` [PATCH net-next v6 4/8] mm/slab: Give bucket caches the alignment of the caches they mirror Kees Cook
2026-10-06  9:20 ` [PATCH net-next v6 5/8] mm/slab: Add kmem_buckets_destroy() Kees Cook
2026-10-06  9:20 ` [PATCH net-next v6 6/8] mm/slab: Add tests for the existing kmem_buckets behaviour Kees Cook
2026-10-06  9:20 ` [PATCH net-next v6 7/8] mm/slab: Provide kmalloc type fallback for bucket allocations Kees Cook
2026-10-06  9:20 ` [PATCH net-next v6 8/8] net: skb: isolate skb data area allocations into a separate bucket Kees Cook

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®