* [PATCH net-next v5 0/7] net: skb: isolate skb data area allocations into a separate bucket
@ 2026-10-02 23:11 Kees Cook
2026-10-02 23:11 ` [PATCH net-next v5 1/7] mm/slab: Mark the kmem_buckets_create() context as a Context: section Kees Cook
` (6 more replies)
0 siblings, 7 replies; 14+ messages in thread
From: Kees Cook @ 2026-10-02 23:11 UTC (permalink / raw)
To: Vlastimil Babka
Cc: Kees Cook, Pedro Falcato, Harry Yoo (Meta),
David S. Miller, Andrew Morton, Hao Li, Christoph Lameter,
David Rientjes, Roman Gushchin, Johannes Weiner, Michal Hocko,
Shakeel Butt, Muchun Song, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Simon Horman, Jason Xing, Willem de Bruijn,
Mina Almasry, Björn Töpel, Jiayuan Chen,
Kuniyuki Iwashima, linux-kernel, linux-mm, cgroups, netdev,
linux-hardening
Hi!
This gets the buckets able to handle memcg (GFP_KERNEL_ACCOUNT) with
isolation (since it's common due to AF_UNIX), and GFP_DMA with fall back
(since it's rare). It gave me an excuse to build out bucket kunit tests
too, and that (and LLM review) found a couple other issues that needed
fixing too.
The bulk of this is mm/slab, but the final patch is netdev, which Paolo
acked in v4, so I'm hoping this whole series can go via slab?
Thanks!
-Kees
v5:
- 2/7: kmem_buckets_create() and kmem_buckets_create_types() take an
alignment, where kmem_cache_create() has it; 0 keeps v4's behaviour,
the alignment of the kmalloc cache being mirrored. All callers pass 0.
New subject: "mm/slab: Let kmem_buckets_create() take an alignment"
(Harry)
- 4/7: test that an explicit alignment is honored.
- 5/7: export mem_cgroup_kmem_disabled() for KUnit only, fixing a
modular build of slub_kunit (Harry).
- 6/7: commit log shows the per-set cost of indexing rows by
kmalloc_cache_type; test that an explicit alignment reaches the
accounted row too.
- 7/7: pass the new alignment argument.
- Collected tags on 1/7 (Harry, Pedro, Hao) and 7/7 (Paolo).
- v4..v5 diff: https://git.kernel.org/pub/scm/linux/kernel/git/kees/linux.git/diff/?id=dev/v7.3-rc2/skb-buckets/v5&id2=dev/v7.3-rc2/skb-buckets/v4
v4: https://lore.kernel.org/all/20260921075811.too.775-kees@kernel.org/
v3: https://lore.kernel.org/all/20260702170728.168755-1-pfalcato@suse.de/
Kees Cook (6):
mm/slab: Mark the kmem_buckets_create() context as a Context: section
mm/slab: Let kmem_buckets_create() take an alignment
mm/slab: Add kmem_buckets_destroy()
mm/slab: Add tests for the existing kmem_buckets behaviour
mm/slab: Provide kmalloc type fallback for bucket allocations
mm/slab: Let a bucket set handle __GFP_ACCOUNT
Pedro Falcato (1):
net: skb: isolate skb data area allocations into a separate bucket
include/linux/slab.h | 66 ++++++-
mm/slab.h | 44 ++++-
ipc/msgutil.c | 2 +-
lib/tests/slub_kunit.c | 397 +++++++++++++++++++++++++++++++++++++++++
mm/memcontrol.c | 2 +
mm/slab_common.c | 173 +++++++++++++++---
mm/util.c | 2 +-
net/core/skbuff.c | 11 +-
8 files changed, 663 insertions(+), 34 deletions(-)
--
2.34.1
^ permalink raw reply [flat|nested] 14+ messages in thread
* [PATCH net-next v5 1/7] mm/slab: Mark the kmem_buckets_create() context as a Context: section
2026-10-02 23:11 [PATCH net-next v5 0/7] net: skb: isolate skb data area allocations into a separate bucket Kees Cook
@ 2026-10-02 23:11 ` Kees Cook
2026-10-02 23:11 ` [PATCH net-next v5 2/7] mm/slab: Let kmem_buckets_create() take an alignment Kees Cook
` (5 subsequent siblings)
6 siblings, 0 replies; 14+ messages in thread
From: Kees Cook @ 2026-10-02 23:11 UTC (permalink / raw)
To: Vlastimil Babka
Cc: Kees Cook, Harry Yoo, Andrew Morton, Hao Li, Christoph Lameter,
David Rientjes, Roman Gushchin, linux-mm, Pedro Falcato,
Kuniyuki Iwashima, linux-hardening, David S. Miller,
Johannes Weiner, Michal Hocko, Shakeel Butt, Muchun Song,
Eric Dumazet, Jakub Kicinski, Paolo Abeni, Simon Horman,
Jason Xing, Willem de Bruijn, Mina Almasry,
Björn Töpel, Jiayuan Chen, linux-kernel, cgroups,
netdev
kmem_buckets_create() had the same sentence about calling context as
kmem_cache_create(), but lacked the "Context:" prefix, so kernel-doc
rendered in the body instead of as a "context" section.
Give it the missing prefix. Additionally fix the "a interrupt" typo
__kmem_cache_create_args() had.
Fixes: b32801d1255be ("mm/slab: Introduce kmem_buckets_create() and family")
Assisted-by: LLM
Reviewed-by: Harry Yoo (Meta) <harry@kernel.org>
Acked-by: Pedro Falcato <pfalcato@suse.de>
Reviewed-by: Hao Li <hao.li@linux.dev>
Signed-off-by: Kees Cook <kees@kernel.org>
---
mm/slab_common.c | 4 ++--
1 file changed, 2 insertions(+), 2 deletions(-)
diff --git a/mm/slab_common.c b/mm/slab_common.c
index b19ba1b31484..270408ce5a9d 100644
--- a/mm/slab_common.c
+++ b/mm/slab_common.c
@@ -311,7 +311,7 @@ __kmem_cache_alias(const char *name, unsigned int size, slab_flags_t flags,
* &SLAB_TYPESAFE_BY_RCU - Slab page (not individual objects) freeing delayed
* by a grace period - see the full description before using.
*
- * Context: Cannot be called within a interrupt, but can be interrupted.
+ * Context: Cannot be called within an interrupt, but can be interrupted.
*
* Return: a pointer to the cache on success, NULL on failure.
*/
@@ -422,7 +422,7 @@ static struct kmem_cache *kmem_buckets_cache __ro_after_init;
* to/from userspace.
* @ctor: A constructor for the objects, run when new allocations are made.
*
- * Cannot be called within an interrupt, but can be interrupted.
+ * Context: Cannot be called within an interrupt, but can be interrupted.
*
* Return: a pointer to the cache on success, NULL on failure. When
* CONFIG_SLAB_BUCKETS is not enabled, ZERO_SIZE_PTR is returned, and
--
2.34.1
^ permalink raw reply [flat|nested] 14+ messages in thread
* [PATCH net-next v5 2/7] mm/slab: Let kmem_buckets_create() take an alignment
2026-10-02 23:11 [PATCH net-next v5 0/7] net: skb: isolate skb data area allocations into a separate bucket Kees Cook
2026-10-02 23:11 ` [PATCH net-next v5 1/7] mm/slab: Mark the kmem_buckets_create() context as a Context: section Kees Cook
@ 2026-10-02 23:11 ` Kees Cook
2026-10-03 23:13 ` netdev-bot+sashiko
2026-10-02 23:11 ` [PATCH net-next v5 3/7] mm/slab: Add kmem_buckets_destroy() Kees Cook
` (4 subsequent siblings)
6 siblings, 1 reply; 14+ messages in thread
From: Kees Cook @ 2026-10-02 23:11 UTC (permalink / raw)
To: Vlastimil Babka
Cc: Kees Cook, Harry Yoo, Andrew Morton, Hao Li, Christoph Lameter,
David Rientjes, Roman Gushchin, linux-mm, Pedro Falcato,
Kuniyuki Iwashima, linux-hardening, David S. Miller,
Johannes Weiner, Michal Hocko, Shakeel Butt, Muchun Song,
Eric Dumazet, Jakub Kicinski, Paolo Abeni, Simon Horman,
Jason Xing, Willem de Bruijn, Mina Almasry,
Björn Töpel, Jiayuan Chen, linux-kernel, cgroups,
netdev
A bucket set is created with kmem_cache_create_usercopy(..., align = 0),
so calculate_alignment() falls back to arch_slab_minalign(), typically 8
bytes. The general kmalloc caches it stands in for are created through
create_boot_cache(), which starts from ARCH_KMALLOC_MINALIGN and raises
it to the largest power-of-two divisor of the size:
if (flags & SLAB_KMALLOC)
align = max(align, 1U << (ffs(size) - 1));
This is only a problem when slab metadata is enabled with
CONFIG_KASAN=y, CONFIG_SLUB_DEBUG_ON=y, or "slab_debug=...", because
metadata changes the stride size off a power of two, for example:
size 128: bucket align=8 size=224 | kmalloc align=128 size=384
size 512: bucket align=8 size=608 | kmalloc align=512 size=1536
size 2048: bucket align=8 size=2144 | kmalloc align=2048 size=6144
So bucket allocations will fail the IS_ALIGNED(p, ARCH_DMA_MINALIGN)
check, potentially creating problems for non-coherent DMA situation.
As discussed in review, add an alignment argument to
kmem_buckets_create(), where kmem_cache_create() has it. A caller that
needs a particular alignment passes it. With 0, each cache takes the
alignment of the kmalloc cache being mirrored, which is where the size
and the name suffix already come from, so a caller moving from kmalloc()
keeps the alignment it had. Both existing callers pass 0.
Fixes: b32801d1255be ("mm/slab: Introduce kmem_buckets_create() and family")
Assisted-by: LLM
Signed-off-by: Kees Cook <kees@kernel.org>
---
include/linux/slab.h | 3 ++-
ipc/msgutil.c | 2 +-
mm/slab_common.c | 9 +++++++--
mm/util.c | 2 +-
4 files changed, 11 insertions(+), 5 deletions(-)
diff --git a/include/linux/slab.h b/include/linux/slab.h
index cda126def67a..94709e4e14f3 100644
--- a/include/linux/slab.h
+++ b/include/linux/slab.h
@@ -890,7 +890,8 @@ void *kmem_cache_alloc_lru_noprof(struct kmem_cache *s, struct list_lru *lru,
bool kmem_cache_charge(void *objp, gfp_t gfpflags);
void kmem_cache_free(struct kmem_cache *s, void *objp);
-kmem_buckets *kmem_buckets_create(const char *name, slab_flags_t flags,
+kmem_buckets *kmem_buckets_create(const char *name, unsigned int align,
+ slab_flags_t flags,
unsigned int useroffset, unsigned int usersize,
void (*ctor)(void *));
diff --git a/ipc/msgutil.c b/ipc/msgutil.c
index e28f0cecb2ec..b49b8f65a582 100644
--- a/ipc/msgutil.c
+++ b/ipc/msgutil.c
@@ -43,7 +43,7 @@ static kmem_buckets *msg_buckets __ro_after_init;
static int __init init_msg_buckets(void)
{
- msg_buckets = kmem_buckets_create("msg_msg", SLAB_ACCOUNT,
+ msg_buckets = kmem_buckets_create("msg_msg", 0, SLAB_ACCOUNT,
sizeof(struct msg_msg),
DATALEN_MSG, NULL);
diff --git a/mm/slab_common.c b/mm/slab_common.c
index 270408ce5a9d..301f3d4f4ca2 100644
--- a/mm/slab_common.c
+++ b/mm/slab_common.c
@@ -415,6 +415,9 @@ static struct kmem_cache *kmem_buckets_cache __ro_after_init;
* allocations via kmem_buckets_alloc()
* @name: A prefix string which is used in /proc/slabinfo to identify this
* cache. The individual caches with have their sizes as the suffix.
+ * @align: The required alignment for the objects, or 0 to give each cache
+ * the alignment of the kmalloc cache of the same size, as a caller
+ * moving from kmalloc() may depend on.
* @flags: SLAB flags (see kmem_cache_create() for details).
* @useroffset: Starting offset within an allocation that may be copied
* to/from userspace.
@@ -429,7 +432,8 @@ static struct kmem_cache *kmem_buckets_cache __ro_after_init;
* subsequent calls to kmem_buckets_alloc() will fall back to kmalloc().
* (i.e. callers only need to check for NULL on failure.)
*/
-kmem_buckets *kmem_buckets_create(const char *name, slab_flags_t flags,
+kmem_buckets *kmem_buckets_create(const char *name, unsigned int align,
+ slab_flags_t flags,
unsigned int useroffset,
unsigned int usersize,
void (*ctor)(void *))
@@ -487,7 +491,8 @@ kmem_buckets *kmem_buckets_create(const char *name, slab_flags_t flags,
if (WARN_ON(!cache_name))
goto fail;
(*b)[aligned_idx] = kmem_cache_create_usercopy(cache_name, size,
- 0, flags, cache_useroffset,
+ align ?: kmalloc_caches[KMALLOC_NORMAL][idx]->align,
+ flags, cache_useroffset,
cache_usersize, ctor);
kfree(cache_name);
if (WARN_ON(!(*b)[aligned_idx]))
diff --git a/mm/util.c b/mm/util.c
index bf0513d1d3d0..3f04d77ef4b8 100644
--- a/mm/util.c
+++ b/mm/util.c
@@ -199,7 +199,7 @@ static kmem_buckets *user_buckets __ro_after_init;
static int __init init_user_buckets(void)
{
- user_buckets = kmem_buckets_create("memdup_user", 0, 0, INT_MAX, NULL);
+ user_buckets = kmem_buckets_create("memdup_user", 0, 0, 0, INT_MAX, NULL);
return 0;
}
--
2.34.1
^ permalink raw reply [flat|nested] 14+ messages in thread
* [PATCH net-next v5 3/7] mm/slab: Add kmem_buckets_destroy()
2026-10-02 23:11 [PATCH net-next v5 0/7] net: skb: isolate skb data area allocations into a separate bucket Kees Cook
2026-10-02 23:11 ` [PATCH net-next v5 1/7] mm/slab: Mark the kmem_buckets_create() context as a Context: section Kees Cook
2026-10-02 23:11 ` [PATCH net-next v5 2/7] mm/slab: Let kmem_buckets_create() take an alignment Kees Cook
@ 2026-10-02 23:11 ` Kees Cook
2026-10-03 23:13 ` netdev-bot+sashiko
2026-10-02 23:11 ` [PATCH net-next v5 4/7] mm/slab: Add tests for the existing kmem_buckets behaviour Kees Cook
` (3 subsequent siblings)
6 siblings, 1 reply; 14+ messages in thread
From: Kees Cook @ 2026-10-02 23:11 UTC (permalink / raw)
To: Vlastimil Babka
Cc: Kees Cook, Harry Yoo, Andrew Morton, Hao Li, Christoph Lameter,
David Rientjes, Roman Gushchin, linux-mm, Pedro Falcato,
Kuniyuki Iwashima, linux-hardening, David S. Miller,
Johannes Weiner, Michal Hocko, Shakeel Butt, Muchun Song,
Eric Dumazet, Jakub Kicinski, Paolo Abeni, Simon Horman,
Jason Xing, Willem de Bruijn, Mina Almasry,
Björn Töpel, Jiayuan Chen, linux-kernel, cgroups,
netdev
kmem_buckets_create() intentionally had no "destroy" counterpart. Every
caller has lived in core kernel code and creates its set once at boot,
so nothing has needed to take one down. However, KUnit tests may be
built module, so we need it now to support the coming tests.
Some caches have size aliases, so the same pointer is stored at more
then one index, so we have to save it, clear all matching instances, and
then free the saved cache pointer. (This is what the bitmap was tracking
before in the "allocation failed" error path.)
When CONFIG_SLAB_BUCKETS=n the whole body compiles away, matching the
ZERO_SIZE_PTR that kmem_buckets_create() hands back in that configuration.
Link: https://lore.kernel.org/all/20240809073309.2134488-1-kees@kernel.org/
Assisted-by: LLM
Signed-off-by: Kees Cook <kees@kernel.org>
---
include/linux/slab.h | 1 +
mm/slab_common.c | 49 +++++++++++++++++++++++++++++++++++++-------
2 files changed, 43 insertions(+), 7 deletions(-)
diff --git a/include/linux/slab.h b/include/linux/slab.h
index 94709e4e14f3..4e6e74b3a990 100644
--- a/include/linux/slab.h
+++ b/include/linux/slab.h
@@ -894,6 +894,7 @@ kmem_buckets *kmem_buckets_create(const char *name, unsigned int align,
slab_flags_t flags,
unsigned int useroffset, unsigned int usersize,
void (*ctor)(void *));
+void kmem_buckets_destroy(kmem_buckets *bucket);
/*
* Bulk allocation and freeing operations. These are accelerated in an
diff --git a/mm/slab_common.c b/mm/slab_common.c
index 301f3d4f4ca2..d66e5e56a0f1 100644
--- a/mm/slab_common.c
+++ b/mm/slab_common.c
@@ -438,12 +438,9 @@ kmem_buckets *kmem_buckets_create(const char *name, unsigned int align,
unsigned int usersize,
void (*ctor)(void *))
{
- unsigned long mask = 0;
unsigned int idx;
kmem_buckets *b;
- BUILD_BUG_ON(ARRAY_SIZE(kmalloc_caches[KMALLOC_NORMAL]) > BITS_PER_LONG);
-
/*
* When the separate buckets API is not built in, just return
* a non-NULL value for the kmem_buckets pointer, which will be
@@ -497,7 +494,6 @@ kmem_buckets *kmem_buckets_create(const char *name, unsigned int align,
kfree(cache_name);
if (WARN_ON(!(*b)[aligned_idx]))
goto fail;
- set_bit(aligned_idx, &mask);
}
if (idx != aligned_idx)
(*b)[idx] = (*b)[aligned_idx];
@@ -506,14 +502,53 @@ kmem_buckets *kmem_buckets_create(const char *name, unsigned int align,
return b;
fail:
- for_each_set_bit(idx, &mask, ARRAY_SIZE(kmalloc_caches[KMALLOC_NORMAL]))
- kmem_cache_destroy((*b)[idx]);
- kmem_cache_free(kmem_buckets_cache, b);
+ kmem_buckets_destroy(b);
return NULL;
}
EXPORT_SYMBOL(kmem_buckets_create);
+/**
+ * kmem_buckets_destroy - Destroy a set of caches made by kmem_buckets_create()
+ * @bucket: The set to destroy, which may be NULL.
+ *
+ * Destroys each cache in @bucket and then frees @bucket itself. As for
+ * kmem_cache_destroy(), every object allocated from @bucket must have been
+ * freed beforehand, and @bucket must not be used afterwards.
+ *
+ * Context: Process context. May sleep, as kmem_cache_destroy() takes the
+ * slab mutex and can wait on RCU callbacks for each cache.
+ */
+void kmem_buckets_destroy(kmem_buckets *bucket)
+{
+ unsigned int idx, i;
+
+ if (!IS_ENABLED(CONFIG_SLAB_BUCKETS) || ZERO_OR_NULL_PTR(bucket))
+ return;
+
+ for (idx = 0; idx < ARRAY_SIZE(kmalloc_caches[KMALLOC_NORMAL]); idx++) {
+ struct kmem_cache *cache = (*bucket)[idx];
+
+ if (!cache)
+ continue;
+
+ /*
+ * Sizes below arch_slab_minalign() share one cache, which
+ * kmem_buckets_create() then stores at each of their indices.
+ * Drop every reference to it before destroying it, so that no
+ * later pass reads a pointer to a cache that is already gone.
+ */
+ for (i = idx; i < ARRAY_SIZE(kmalloc_caches[KMALLOC_NORMAL]); i++)
+ if ((*bucket)[i] == cache)
+ (*bucket)[i] = NULL;
+
+ kmem_cache_destroy(cache);
+ }
+
+ kmem_cache_free(kmem_buckets_cache, bucket);
+}
+EXPORT_SYMBOL(kmem_buckets_destroy);
+
/*
* For a given kmem_cache, kmem_cache_destroy() should only be called
* once or there will be a use-after-free problem. The actual deletion
--
2.34.1
^ permalink raw reply [flat|nested] 14+ messages in thread
* [PATCH net-next v5 4/7] mm/slab: Add tests for the existing kmem_buckets behaviour
2026-10-02 23:11 [PATCH net-next v5 0/7] net: skb: isolate skb data area allocations into a separate bucket Kees Cook
` (2 preceding siblings ...)
2026-10-02 23:11 ` [PATCH net-next v5 3/7] mm/slab: Add kmem_buckets_destroy() Kees Cook
@ 2026-10-02 23:11 ` Kees Cook
2026-10-03 23:13 ` netdev-bot+sashiko
2026-10-02 23:11 ` [PATCH net-next v5 5/7] mm/slab: Provide kmalloc type fallback for bucket allocations Kees Cook
` (2 subsequent siblings)
6 siblings, 1 reply; 14+ messages in thread
From: Kees Cook @ 2026-10-02 23:11 UTC (permalink / raw)
To: Vlastimil Babka
Cc: Kees Cook, Harry Yoo, Andrew Morton, Hao Li, Christoph Lameter,
David Rientjes, Roman Gushchin, linux-mm, Pedro Falcato,
Kuniyuki Iwashima, linux-hardening, David S. Miller,
Johannes Weiner, Michal Hocko, Shakeel Butt, Muchun Song,
Eric Dumazet, Jakub Kicinski, Paolo Abeni, Simon Horman,
Jason Xing, Willem de Bruijn, Mina Almasry,
Björn Töpel, Jiayuan Chen, linux-kernel, cgroups,
netdev
kmem_buckets has had no test coverage since it was added. Add tests,
including stuff unique to the bucket design:
- A bucket allocation comes from a cache of the set's own, and that
cache carries SLAB_NO_MERGE.
- Each size class is served by the set, including 96 and 192, which are
not powers of two and are filled in from an aligned index by
kmem_buckets_create(). Sizes above KMALLOC_MAX_CACHE_SIZE go to the
page allocator instead, bucket set or not.
- Each cache is aligned like the kmalloc cache it mirrors, or, when the
set was created with an alignment, to that one.
- With CONFIG_SLAB_BUCKETS=n, kmem_buckets_create() still returns a
non-NULL (zero size alloc pointer), so that callers only have to check
for failure, and allocations through it come from the general caches.
- Destroying a set takes its caches down rather than only freeing the
set, which is what a module creating one on each load depends on.
A set freed without its caches would leave the names taken, and the
next load would warn about every one of them.
The tests skip rather than compile out, which is useful for testing
the CONFIG_SLAB_BUCKETS=n behaviors.
Built and tests passing (with expected skips) on ARCH=x86_64 defconfig
with GCC 16.2.0, with CONFIG_SLAB_BUCKETS as y and n.
Assisted-by: LLM
Signed-off-by: Kees Cook <kees@kernel.org>
---
lib/tests/slub_kunit.c | 252 +++++++++++++++++++++++++++++++++++++++++
1 file changed, 252 insertions(+)
diff --git a/lib/tests/slub_kunit.c b/lib/tests/slub_kunit.c
index e3b63f0338d5..3c923a3af825 100644
--- a/lib/tests/slub_kunit.c
+++ b/lib/tests/slub_kunit.c
@@ -1,6 +1,7 @@
// SPDX-License-Identifier: GPL-2.0
#include <kunit/test.h>
#include <kunit/test-bug.h>
+#include <kunit/resource.h>
#include <linux/mm.h>
#include <linux/slab.h>
#include <linux/module.h>
@@ -474,6 +475,251 @@ static int test_init(struct kunit *test)
return 0;
}
+/* Destroy buckets on test exit so a failed KUNIT_ASSERT_*() doesn't leak. */
+KUNIT_DEFINE_ACTION_WRAPPER(destroy_buckets, kmem_buckets_destroy, kmem_buckets *);
+
+#define KUNIT_ASSERT_BUCKETS_CREATED(test, b) \
+ do { \
+ KUNIT_ASSERT_NOT_NULL(test, b); \
+ KUNIT_ASSERT_EQ(test, 0, \
+ kunit_add_action_or_reset(test, \
+ destroy_buckets, b)); \
+ } while (0)
+
+/*
+ * The cache an allocation came from, or NULL if it came from no cache at
+ * all, e.g. a size too big for any of them is served by the page allocator.
+ */
+static struct kmem_cache *cache_of(void *p)
+{
+ struct slab *slab = virt_to_slab(p);
+
+ return slab ? slab->slab_cache : NULL;
+}
+
+/*
+ * A bucket set exists to keep its allocations out of the caches everything
+ * else uses, so check the two things that make that true: they come from a
+ * cache of the set's own, and that cache is never merged into another.
+ */
+static void test_kmem_buckets_isolation(struct kunit *test)
+{
+ struct kmem_cache *bucket_cache, *general_cache;
+ kmem_buckets *b;
+ void *p, *q;
+
+ if (!IS_ENABLED(CONFIG_SLAB_BUCKETS))
+ kunit_skip(test, "needs CONFIG_SLAB_BUCKETS");
+
+ b = kmem_buckets_create("isolated_buckets", 0, 0, 0, INT_MAX, NULL);
+ KUNIT_ASSERT_BUCKETS_CREATED(test, b);
+
+ /*
+ * Free each allocation before asserting on the next one: the cache
+ * outlives its objects, so nothing below needs them, and an assertion
+ * that leaves one behind would make the deferred teardown report a
+ * cache that is still in use.
+ */
+ p = kmem_buckets_alloc(b, 128, GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, p);
+ bucket_cache = cache_of(p);
+ kfree(p);
+ KUNIT_ASSERT_NOT_NULL(test, bucket_cache);
+
+ KUNIT_EXPECT_TRUE_MSG(test, strstarts(bucket_cache->name, "isolated_buckets-"),
+ "expected a bucket cache, got %s", bucket_cache->name);
+
+ /*
+ * Cache merging is on by default, and a bucket cache merged into a
+ * same-sized general one would quietly undo the whole separation.
+ */
+ KUNIT_EXPECT_TRUE(test, bucket_cache->flags & SLAB_NO_MERGE);
+
+ q = kmalloc(128, GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, q);
+ general_cache = cache_of(q);
+ kfree(q);
+ KUNIT_ASSERT_NOT_NULL(test, general_cache);
+
+ KUNIT_EXPECT_PTR_NE(test, bucket_cache, general_cache);
+}
+
+/*
+ * Every size class gets its own cache in the set, including the ones that
+ * are not powers of two and are filled in from an aligned index. Sizes past
+ * the largest cache are served by the page allocator, bucket set or not.
+ */
+static void test_kmem_buckets_sizes(struct kunit *test)
+{
+ static const size_t sizes[] = { 8, 96, 192, 1024, 4096 };
+ struct kmem_cache *c;
+ kmem_buckets *b;
+ void *p;
+ int i;
+
+ if (!IS_ENABLED(CONFIG_SLAB_BUCKETS))
+ kunit_skip(test, "needs CONFIG_SLAB_BUCKETS");
+
+ b = kmem_buckets_create("sized_buckets", 0, 0, 0, INT_MAX, NULL);
+ KUNIT_ASSERT_BUCKETS_CREATED(test, b);
+
+ for (i = 0; i < ARRAY_SIZE(sizes); i++) {
+ p = kmem_buckets_alloc(b, sizes[i], GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, p);
+ c = cache_of(p);
+ kfree(p);
+ KUNIT_ASSERT_NOT_NULL(test, c);
+
+ KUNIT_EXPECT_TRUE_MSG(test, strstarts(c->name, "sized_buckets-"),
+ "size %zu: expected a bucket cache, got %s",
+ sizes[i], c->name);
+ KUNIT_EXPECT_GE(test, c->object_size, sizes[i]);
+ }
+
+ /* Too big for any cache: a folio from the page allocator, not a slab. */
+ p = kmem_buckets_alloc(b, KMALLOC_MAX_CACHE_SIZE + 1, GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, p);
+ c = cache_of(p);
+ kfree(p);
+
+ KUNIT_EXPECT_NULL(test, c);
+}
+
+/*
+ * A bucket cache stands in for a kmalloc cache, so it has to be aligned like
+ * one. The DMA layer decides whether a buffer needs bouncing from its size,
+ * on the grounds that a kmalloc cache of that size is already aligned for
+ * the device, so a weaker alignment here is not something a caller can see
+ * coming. Without slab debugging the size implies the alignment and this
+ * holds either way; with it, only the cache's own alignment does.
+ */
+static void test_kmem_buckets_alignment(struct kunit *test)
+{
+ static const size_t sizes[] = { 128, 512, 2048 };
+ struct kmem_cache *bucket_cache, *general_cache;
+ kmem_buckets *b;
+ void *p;
+ int i;
+
+ if (!IS_ENABLED(CONFIG_SLAB_BUCKETS))
+ kunit_skip(test, "needs CONFIG_SLAB_BUCKETS");
+
+ b = kmem_buckets_create("aligned_buckets", 0, 0, 0, INT_MAX, NULL);
+ KUNIT_ASSERT_BUCKETS_CREATED(test, b);
+
+ for (i = 0; i < ARRAY_SIZE(sizes); i++) {
+ p = kmem_buckets_alloc(b, sizes[i], GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, p);
+ bucket_cache = cache_of(p);
+ KUNIT_EXPECT_TRUE_MSG(test,
+ IS_ALIGNED((unsigned long)p, ARCH_DMA_MINALIGN),
+ "size %zu: object %p is not %d byte aligned",
+ sizes[i], p, (int)ARCH_DMA_MINALIGN);
+ kfree(p);
+
+ p = kmalloc(sizes[i], GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, p);
+ general_cache = cache_of(p);
+ kfree(p);
+
+ KUNIT_ASSERT_NOT_NULL(test, bucket_cache);
+ KUNIT_ASSERT_NOT_NULL(test, general_cache);
+ KUNIT_EXPECT_EQ_MSG(test, bucket_cache->align, general_cache->align,
+ "size %zu: bucket cache aligned to %u, %s to %u",
+ sizes[i], bucket_cache->align,
+ general_cache->name, general_cache->align);
+ }
+}
+
+/*
+ * A set created with an alignment gives every one of its caches that
+ * alignment in place of the kmalloc caches' own. 256 is stronger than
+ * kmalloc's alignment for the 64 and 128 byte caches and weaker than it
+ * for the 512 and 2048 byte ones, so this checks the override both ways.
+ */
+static void test_kmem_buckets_explicit_alignment(struct kunit *test)
+{
+ static const size_t sizes[] = { 64, 128, 512, 2048 };
+ const unsigned int align = 256;
+ struct kmem_cache *c;
+ kmem_buckets *b;
+ void *p;
+ int i;
+
+ if (!IS_ENABLED(CONFIG_SLAB_BUCKETS))
+ kunit_skip(test, "needs CONFIG_SLAB_BUCKETS");
+
+ b = kmem_buckets_create("explicit_buckets", align, 0, 0, INT_MAX, NULL);
+ KUNIT_ASSERT_BUCKETS_CREATED(test, b);
+
+ for (i = 0; i < ARRAY_SIZE(sizes); i++) {
+ p = kmem_buckets_alloc(b, sizes[i], GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, p);
+ c = cache_of(p);
+ KUNIT_EXPECT_TRUE_MSG(test, IS_ALIGNED((unsigned long)p, align),
+ "size %zu: object %p is not %u byte aligned",
+ sizes[i], p, align);
+ kfree(p);
+
+ KUNIT_ASSERT_NOT_NULL(test, c);
+ KUNIT_EXPECT_EQ_MSG(test, c->align, align,
+ "size %zu: bucket cache aligned to %u, not %u",
+ sizes[i], c->align, align);
+ }
+}
+
+/*
+ * With the feature compiled out, kmem_buckets_create() still returns
+ * something non-NULL so that callers only have to check for failure, and
+ * allocations through it work (i.e. come from the general caches).
+ */
+static void test_kmem_buckets_disabled(struct kunit *test)
+{
+ kmem_buckets *b;
+ struct kmem_cache *c;
+ void *p;
+
+ if (IS_ENABLED(CONFIG_SLAB_BUCKETS))
+ kunit_skip(test, "only meaningful without CONFIG_SLAB_BUCKETS");
+
+ b = kmem_buckets_create("disabled_buckets", 0, 0, 0, INT_MAX, NULL);
+ KUNIT_ASSERT_BUCKETS_CREATED(test, b);
+
+ p = kmem_buckets_alloc(b, 128, GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, p);
+ c = cache_of(p);
+ kfree(p);
+ KUNIT_ASSERT_NOT_NULL(test, c);
+
+ KUNIT_EXPECT_TRUE_MSG(test, !strstarts(c->name, "disabled_buckets-"),
+ "expected a general cache, got %s", c->name);
+}
+
+/* Destroying a set has to take its caches down, not just free the set. */
+static void test_kmem_buckets_destroy(struct kunit *test)
+{
+ kmem_buckets *b;
+ void *p;
+
+ if (!IS_ENABLED(CONFIG_SLAB_BUCKETS))
+ kunit_skip(test, "needs CONFIG_SLAB_BUCKETS");
+
+ b = kmem_buckets_create("destroyed_buckets", 0, 0, 0, INT_MAX, NULL);
+ KUNIT_ASSERT_NOT_NULL(test, b);
+
+ /*
+ * Deliberately leaked, as test_leak_destroy() leaks its own: the
+ * teardown below has to find it. kmem_cache_destroy() unlists the
+ * cache either way, so the name is still released.
+ */
+ p = kmem_buckets_alloc(b, 128, GFP_KERNEL);
+ KUNIT_EXPECT_NOT_NULL(test, p);
+
+ kmem_buckets_destroy(b);
+
+ KUNIT_EXPECT_EQ(test, 2, slab_errors);
+}
+
static struct kunit_case test_cases[] = {
KUNIT_CASE(test_clobber_zone),
@@ -495,6 +741,12 @@ static struct kunit_case test_cases[] = {
#if defined(CONFIG_KPROBES) && defined(CONFIG_SMP)
KUNIT_CASE_SLOW(test_kmalloc_nolock_and_friends_kprobe),
#endif
+ KUNIT_CASE(test_kmem_buckets_isolation),
+ KUNIT_CASE(test_kmem_buckets_sizes),
+ KUNIT_CASE(test_kmem_buckets_alignment),
+ KUNIT_CASE(test_kmem_buckets_explicit_alignment),
+ KUNIT_CASE(test_kmem_buckets_disabled),
+ KUNIT_CASE(test_kmem_buckets_destroy),
{}
};
--
2.34.1
^ permalink raw reply [flat|nested] 14+ messages in thread
* [PATCH net-next v5 5/7] mm/slab: Provide kmalloc type fallback for bucket allocations
2026-10-02 23:11 [PATCH net-next v5 0/7] net: skb: isolate skb data area allocations into a separate bucket Kees Cook
` (3 preceding siblings ...)
2026-10-02 23:11 ` [PATCH net-next v5 4/7] mm/slab: Add tests for the existing kmem_buckets behaviour Kees Cook
@ 2026-10-02 23:11 ` Kees Cook
2026-10-03 23:13 ` netdev-bot+sashiko
2026-10-02 23:11 ` [PATCH net-next v5 6/7] mm/slab: Let a bucket set handle __GFP_ACCOUNT Kees Cook
2026-10-02 23:11 ` [PATCH net-next v5 7/7] net: skb: isolate skb data area allocations into a separate bucket Kees Cook
6 siblings, 1 reply; 14+ messages in thread
From: Kees Cook @ 2026-10-02 23:11 UTC (permalink / raw)
To: Vlastimil Babka
Cc: Kees Cook, Harry Yoo, Andrew Morton, Hao Li, Christoph Lameter,
David Rientjes, Roman Gushchin, linux-mm, Pedro Falcato,
Kuniyuki Iwashima, linux-hardening, Johannes Weiner,
Michal Hocko, Shakeel Butt, Muchun Song, cgroups,
David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Simon Horman, Jason Xing, Willem de Bruijn, Mina Almasry,
Björn Töpel, Jiayuan Chen, linux-kernel, netdev
kmem_buckets_create() clones kmalloc_caches[KMALLOC_NORMAL].
kmalloc_slab() figures out the kmalloc type the caller asks for, but
then ignored it whenever a bucket set was in use, returning a normal
cache regardless. This would be a problem if a caller asked for GFP_DMA,
__GFP_ACCOUNT, etc. None of the current users do this, so there is no
problem, but it makes adding new users fragile. For example, skb data[1]
needs to handle GFP_DMA (rarely) and __GFP_ACCOUNT (often).
Send those allocations to the general caches instead so nothing breaks and
regular allocations remain isolated with the bucket. The kmem_bucket_type
enum contains only a single item here, but will be expanded in the next
patch.
The test checks mem_cgroup_kmem_disabled() before expecting an accounted
cache, so export it for KUnit only, which a modular build of the test
needs to link.
Built and tests pass with ARCH=x86_64 defconfig with GCC 16.2.0, with
CONFIG_SLAB_BUCKETS as y and n.
Assisted-by: LLM
Link: https://lore.kernel.org/all/04debe19-bbe8-4b5f-9668-753d1f97832d@redhat.com/ [1]
Signed-off-by: Kees Cook <kees@kernel.org>
---
include/linux/slab.h | 13 ++++++++++
mm/slab.h | 23 ++++++++++++++++--
lib/tests/slub_kunit.c | 54 ++++++++++++++++++++++++++++++++++++++++++
mm/memcontrol.c | 2 ++
4 files changed, 90 insertions(+), 2 deletions(-)
diff --git a/include/linux/slab.h b/include/linux/slab.h
index 4e6e74b3a990..bdd00235d4b0 100644
--- a/include/linux/slab.h
+++ b/include/linux/slab.h
@@ -742,6 +742,19 @@ typedef struct kmem_cache * kmem_buckets[KMALLOC_SHIFT_HIGH + 1];
extern kmem_buckets kmalloc_caches[NR_KMALLOC_TYPES];
+/*
+ * The kmalloc types a bucket set can hold a copy of. This is deliberately not
+ * enum kmalloc_cache_type: the KMALLOC_PARTITION copies are all "normal" to a
+ * bucket set, which already separates what they were there to separate, so
+ * indexing by those would mean up to KMALLOC_PARTITION_CACHES_NR unusable
+ * rows per set. Allocations of any type not listed here are served by the
+ * general caches.
+ */
+enum kmem_bucket_type {
+ KMEM_BUCKET_NORMAL = 0,
+ NR_KMEM_BUCKET_TYPES
+};
+
/*
* Define gfp bits that should not be set for KMALLOC_NORMAL.
*/
diff --git a/mm/slab.h b/mm/slab.h
index 8fd6835e4235..7f1bfee83b92 100644
--- a/mm/slab.h
+++ b/mm/slab.h
@@ -421,6 +421,26 @@ static inline unsigned int size_index_elem(unsigned int bytes)
return (bytes - 1) / 8;
}
+/*
+ * Which set of buckets to use for the given kmalloc_cache_type. If not
+ * handled by the kmem_buckets, fall back to general caches.
+ */
+static inline kmem_buckets *
+kmalloc_choose_bucket(kmem_buckets *bucket, enum kmalloc_cache_type type)
+{
+ enum kmem_bucket_type btype;
+
+ if (!bucket)
+ return &kmalloc_caches[type];
+
+ if (type <= KMALLOC_PARTITION_END)
+ btype = KMEM_BUCKET_NORMAL;
+ else
+ return &kmalloc_caches[type]; /* No set holds a row for it. */
+
+ return &bucket[btype];
+}
+
/*
* Find the kmem_cache structure that serves a given size of
* allocation
@@ -438,8 +458,7 @@ kmalloc_slab(size_t size, kmem_buckets *b, gfp_t flags, kmalloc_token_t token,
if (alloc_flags & SLAB_ALLOC_NO_OBJ_EXT)
type = KMALLOC_NO_OBJ_EXT;
- if (!b)
- b = &kmalloc_caches[type];
+ b = kmalloc_choose_bucket(b, type);
if (size <= 192)
index = kmalloc_size_index[size_index_elem(size)];
else
diff --git a/lib/tests/slub_kunit.c b/lib/tests/slub_kunit.c
index 3c923a3af825..9768de01e6f9 100644
--- a/lib/tests/slub_kunit.c
+++ b/lib/tests/slub_kunit.c
@@ -720,6 +720,58 @@ static void test_kmem_buckets_destroy(struct kunit *test)
KUNIT_EXPECT_EQ(test, 2, slab_errors);
}
+/*
+ * A bucket set holds only the kmalloc types it was created with, so an
+ * allocation that asks for a different one has to come from the general
+ * caches. Check that it does, rather than being served a normal cache that
+ * does not satisfy what the flags asked for.
+ */
+static void test_kmem_buckets_type_fallback(struct kunit *test)
+{
+ struct kmem_cache *c;
+ kmem_buckets *b;
+ void *p;
+
+ if (!IS_ENABLED(CONFIG_SLAB_BUCKETS))
+ kunit_skip(test, "needs CONFIG_SLAB_BUCKETS");
+
+ b = kmem_buckets_create("test_buckets", 0, 0, 0, INT_MAX, NULL);
+ KUNIT_ASSERT_BUCKETS_CREATED(test, b);
+
+ /* A plain allocation stays isolated in the bucket set. */
+ p = kmem_buckets_alloc(b, 128, GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, p);
+ c = cache_of(p);
+ kfree(p);
+ KUNIT_ASSERT_NOT_NULL(test, c);
+
+ KUNIT_EXPECT_TRUE_MSG(test, strstarts(c->name, "test_buckets-"),
+ "expected a bucket cache, got %s", c->name);
+
+ /* One that needs ZONE_DMA cannot, so it falls back. */
+ if (IS_ENABLED(CONFIG_ZONE_DMA)) {
+ p = kmem_buckets_alloc(b, 128, GFP_KERNEL | GFP_DMA);
+ KUNIT_ASSERT_NOT_NULL(test, p);
+ c = cache_of(p);
+ kfree(p);
+ KUNIT_ASSERT_NOT_NULL(test, c);
+
+ KUNIT_EXPECT_TRUE_MSG(test, strstarts(c->name, "dma-kmalloc-"),
+ "expected a DMA cache, got %s", c->name);
+ }
+
+ /* Nor can one that has to be accounted. */
+ if (IS_ENABLED(CONFIG_MEMCG) && !mem_cgroup_kmem_disabled()) {
+ p = kmem_buckets_alloc(b, 128, GFP_KERNEL | __GFP_ACCOUNT);
+ KUNIT_ASSERT_NOT_NULL(test, p);
+ c = virt_to_slab(p)->slab_cache;
+ kfree(p);
+
+ KUNIT_EXPECT_TRUE_MSG(test, strstarts(c->name, "kmalloc-cg-"),
+ "expected an accounted cache, got %s", c->name);
+ }
+}
+
static struct kunit_case test_cases[] = {
KUNIT_CASE(test_clobber_zone),
@@ -747,6 +799,7 @@ static struct kunit_case test_cases[] = {
KUNIT_CASE(test_kmem_buckets_explicit_alignment),
KUNIT_CASE(test_kmem_buckets_disabled),
KUNIT_CASE(test_kmem_buckets_destroy),
+ KUNIT_CASE(test_kmem_buckets_type_fallback),
{}
};
@@ -757,5 +810,6 @@ static struct kunit_suite test_suite = {
};
kunit_test_suite(test_suite);
+MODULE_IMPORT_NS("EXPORTED_FOR_KUNIT_TESTING");
MODULE_DESCRIPTION("Kunit tests for slub allocator");
MODULE_LICENSE("GPL");
diff --git a/mm/memcontrol.c b/mm/memcontrol.c
index 1271d390b617..5ceeb5a0b614 100644
--- a/mm/memcontrol.c
+++ b/mm/memcontrol.c
@@ -62,6 +62,7 @@
#include <linux/seq_buf.h>
#include <linux/sched/isolation.h>
#include <linux/kmemleak.h>
+#include <kunit/visibility.h>
#include "internal.h"
#include "swap.h"
#include "swap_table.h"
@@ -135,6 +136,7 @@ bool mem_cgroup_kmem_disabled(void)
{
return cgroup_memory_nokmem;
}
+EXPORT_SYMBOL_IF_KUNIT(mem_cgroup_kmem_disabled);
static void memcg_uncharge(struct mem_cgroup *memcg, unsigned int nr_pages);
--
2.34.1
^ permalink raw reply [flat|nested] 14+ messages in thread
* [PATCH net-next v5 6/7] mm/slab: Let a bucket set handle __GFP_ACCOUNT
2026-10-02 23:11 [PATCH net-next v5 0/7] net: skb: isolate skb data area allocations into a separate bucket Kees Cook
` (4 preceding siblings ...)
2026-10-02 23:11 ` [PATCH net-next v5 5/7] mm/slab: Provide kmalloc type fallback for bucket allocations Kees Cook
@ 2026-10-02 23:11 ` Kees Cook
2026-10-03 23:13 ` netdev-bot+sashiko
2026-10-02 23:11 ` [PATCH net-next v5 7/7] net: skb: isolate skb data area allocations into a separate bucket Kees Cook
6 siblings, 1 reply; 14+ messages in thread
From: Kees Cook @ 2026-10-02 23:11 UTC (permalink / raw)
To: Vlastimil Babka
Cc: Kees Cook, Harry Yoo, Andrew Morton, Hao Li, Christoph Lameter,
David Rientjes, Roman Gushchin, linux-mm, Pedro Falcato,
Kuniyuki Iwashima, linux-hardening, David S. Miller,
Johannes Weiner, Michal Hocko, Shakeel Butt, Muchun Song,
Eric Dumazet, Jakub Kicinski, Paolo Abeni, Simon Horman,
Jason Xing, Willem de Bruijn, Mina Almasry,
Björn Töpel, Jiayuan Chen, linux-kernel, cgroups,
netdev
A bucket set holds one row of caches, cloned from KMALLOC_NORMAL, and an
allocation of any other kmalloc type falls back to the general caches.
Extend this to handle __GFP_ACCOUNT, so that a single bucket user can
isolate either GFP_KERNEL or GFP_KERNEL_ACCOUNT allocations, as is
needed for skb data, where AF_UNIX uses:
sk->sk_allocation = GFP_KERNEL_ACCOUNT;
The coverage is selected at bucket creation time:
b = kmem_buckets_create_types(name, 0, flags, 0, INT_MAX, NULL,
BIT(KMEM_BUCKET_NORMAL) |
BIT(KMEM_BUCKET_CGROUP));
The prior kmem_buckets_create() function keeps its name and defaults
to only KMEM_BUCKET_NORMAL, leaving existing users as-is.
Only the accounted type is offered. Nothing wants a reclaimable or
no-obj-ext row, and of the twelve places passing GFP_DMA to an skb
allocator, all are rare hardware: b44, b43legacy, prestera, and s390 ctcm.
The choice is made at creation rather than every set getting every type
because the rows, when populated, are not free. Each holds 13 caches, and
a cache is a 1208 byte struct plus an unconditional per-cpu allocation,
a node struct, and an entry in /proc/slabinfo and under /sys/kernel/slab.
The rows are indexed by enum kmem_bucket_type, not enum
kmalloc_cache_type, which counts each KMALLOC_PARTITION copy as a type
of its own. Indexed by kmalloc_cache_type, a set would carry rows it
can never use: on x86_64 with CONFIG_KMALLOC_PARTITION_CACHES=y,
CONFIG_MEMCG=y and CONFIG_ZONE_DMA=y, 20 rows of 112 bytes (2240 bytes)
per set, against the 2 rows (224 bytes) used here.
KMEM_BUCKET_CGROUP collapses to KMEM_BUCKET_NORMAL without CONFIG_MEMCG,
exactly as KMALLOC_CGROUP does, so NR_KMEM_BUCKET_TYPES is 1 there and a
bucket set is the same single row it is today. Where the type is asked for
but the system is not creating caches of it (under "cgroup.memory=nokmem")
the row is aliased to the normal one, as new_kmalloc_cache() does for the
general caches, so those allocations stay isolated rather than falling
back to the general caches.
Built and tests pass (and skip as expected) on ARCH=x86_64 defconfig
with GCC 16.2.0 in all combinations of CONFIG_SLAB_BUCKETS=y/n and
CONFIG_MEMCG=y/n/y+"cgroup.memory=nokmem".
Assisted-by: LLM
Signed-off-by: Kees Cook <kees@kernel.org>
---
include/linux/slab.h | 53 +++++++++++++--
mm/slab.h | 23 ++++++-
lib/tests/slub_kunit.c | 101 ++++++++++++++++++++++++++--
mm/slab_common.c | 149 ++++++++++++++++++++++++++++++++---------
4 files changed, 283 insertions(+), 43 deletions(-)
diff --git a/include/linux/slab.h b/include/linux/slab.h
index bdd00235d4b0..ee359dc90333 100644
--- a/include/linux/slab.h
+++ b/include/linux/slab.h
@@ -752,6 +752,11 @@ extern kmem_buckets kmalloc_caches[NR_KMALLOC_TYPES];
*/
enum kmem_bucket_type {
KMEM_BUCKET_NORMAL = 0,
+#ifdef CONFIG_MEMCG
+ KMEM_BUCKET_CGROUP,
+#else
+ KMEM_BUCKET_CGROUP = KMEM_BUCKET_NORMAL,
+#endif
NR_KMEM_BUCKET_TYPES
};
@@ -903,10 +908,50 @@ void *kmem_cache_alloc_lru_noprof(struct kmem_cache *s, struct list_lru *lru,
bool kmem_cache_charge(void *objp, gfp_t gfpflags);
void kmem_cache_free(struct kmem_cache *s, void *objp);
-kmem_buckets *kmem_buckets_create(const char *name, unsigned int align,
- slab_flags_t flags,
- unsigned int useroffset, unsigned int usersize,
- void (*ctor)(void *));
+kmem_buckets *kmem_buckets_create_types(const char *name, unsigned int align,
+ slab_flags_t flags,
+ unsigned int useroffset, unsigned int usersize,
+ void (*ctor)(void *),
+ unsigned int type_mask);
+
+/**
+ * kmem_buckets_create - Create a set of caches that handle dynamic sized
+ * allocations via kmem_buckets_alloc()
+ * @name: A prefix string which is used in /proc/slabinfo to identify this
+ * cache. The individual caches with have their sizes as the suffix.
+ * @align: The required alignment for the objects, or 0 to give each cache
+ * the alignment of the kmalloc cache of the same size, as a caller
+ * moving from kmalloc() may depend on.
+ * @flags: SLAB flags (see kmem_cache_create() for details).
+ * @useroffset: Starting offset within an allocation that may be copied
+ * to/from userspace.
+ * @usersize: How many bytes, starting at @useroffset, may be copied
+ * to/from userspace.
+ * @ctor: A constructor for the objects, run when new allocations are made.
+ *
+ * Covers KMEM_BUCKET_NORMAL only. Allocations needing another kmalloc type
+ * are served by the general caches, keeping the type they asked for and
+ * losing only the isolation. Use kmem_buckets_create_types() to cover more.
+ *
+ * Context: Cannot be called within an interrupt, but can be interrupted.
+ *
+ * Return: a pointer to the cache on success, NULL on failure. When
+ * CONFIG_SLAB_BUCKETS is not enabled, ZERO_SIZE_PTR is returned, and
+ * subsequent calls to kmem_buckets_alloc() will fall back to kmalloc().
+ * (i.e. callers only need to check for NULL on failure.)
+ */
+static inline kmem_buckets *kmem_buckets_create(const char *name,
+ unsigned int align,
+ slab_flags_t flags,
+ unsigned int useroffset,
+ unsigned int usersize,
+ void (*ctor)(void *))
+{
+ return kmem_buckets_create_types(name, align, flags, useroffset,
+ usersize, ctor,
+ BIT(KMEM_BUCKET_NORMAL));
+}
+
void kmem_buckets_destroy(kmem_buckets *bucket);
/*
diff --git a/mm/slab.h b/mm/slab.h
index 7f1bfee83b92..2af44e09edda 100644
--- a/mm/slab.h
+++ b/mm/slab.h
@@ -435,10 +435,31 @@ kmalloc_choose_bucket(kmem_buckets *bucket, enum kmalloc_cache_type type)
if (type <= KMALLOC_PARTITION_END)
btype = KMEM_BUCKET_NORMAL;
+ else if (IS_ENABLED(CONFIG_MEMCG) && type == KMALLOC_CGROUP)
+ btype = KMEM_BUCKET_CGROUP;
else
return &kmalloc_caches[type]; /* No set holds a row for it. */
- return &bucket[btype];
+ /*
+ * Either this row was created, and holds a cache everywhere the
+ * general caches hold one, or it was never created and holds nothing.
+ * Test with the KMALLOC_SHIFT_LOW which exists in every configuration.
+ */
+ if (likely(bucket[btype][KMALLOC_SHIFT_LOW]))
+ return &bucket[btype];
+
+ /*
+ * A row this set _could_ have held, but was not created with: the type
+ * mask passed to kmem_buckets_create_types() did not cover what its
+ * callers actually tried to allocate. Report the mismatch but still
+ * fall back to the general caches.
+ *
+ * At present, only __GFP_ACCOUNT can be missing.
+ */
+ WARN_ONCE(1,
+ "kmem_buckets: __GFP_ACCOUNT needs BIT(KMEM_BUCKET_CGROUP) in create mask\n");
+
+ return &kmalloc_caches[type];
}
/*
diff --git a/lib/tests/slub_kunit.c b/lib/tests/slub_kunit.c
index 9768de01e6f9..bc5e52f20179 100644
--- a/lib/tests/slub_kunit.c
+++ b/lib/tests/slub_kunit.c
@@ -760,15 +760,104 @@ static void test_kmem_buckets_type_fallback(struct kunit *test)
"expected a DMA cache, got %s", c->name);
}
- /* Nor can one that has to be accounted. */
+ /*
+ * An accounted allocation would fall back too, but a bucket set can
+ * hold that type, so reaching the fallback means the create mask was
+ * wrong and kmalloc_slab() warns. Not exercised here for that reason;
+ * test_kmem_buckets_type_covered() checks the type that is asked for.
+ */
+}
+
+/*
+ * A bucket set created for a kmalloc type keeps those allocations isolated
+ * too, rather than sending them to the general caches. Where nothing creates
+ * accounted caches at all, the row aliases the normal one, so this also
+ * covers tearing down a set whose rows share their caches.
+ */
+static void test_kmem_buckets_type_covered(struct kunit *test)
+{
+ struct kmem_cache *c, *normal_cache;
+ kmem_buckets *b;
+ void *p;
+
+ if (!IS_ENABLED(CONFIG_SLAB_BUCKETS))
+ kunit_skip(test, "needs CONFIG_SLAB_BUCKETS");
+
+ b = kmem_buckets_create_types("covered_buckets", 0, 0, 0, INT_MAX, NULL,
+ BIT(KMEM_BUCKET_NORMAL) |
+ BIT(KMEM_BUCKET_CGROUP));
+ KUNIT_ASSERT_BUCKETS_CREATED(test, b);
+
+ p = kmem_buckets_alloc(b, 128, GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, p);
+ normal_cache = cache_of(p);
+ kfree(p);
+ KUNIT_ASSERT_NOT_NULL(test, normal_cache);
+
+ KUNIT_EXPECT_TRUE_MSG(test, strstarts(normal_cache->name, "covered_buckets-128"),
+ "expected the normal bucket cache, got %s",
+ normal_cache->name);
+
+ /* Accounted, and still in the bucket set rather than kmalloc-cg-*. */
+ p = kmem_buckets_alloc(b, 128, GFP_KERNEL | __GFP_ACCOUNT);
+ KUNIT_ASSERT_NOT_NULL(test, p);
+ c = cache_of(p);
+ kfree(p);
+ KUNIT_ASSERT_NOT_NULL(test, c);
+
if (IS_ENABLED(CONFIG_MEMCG) && !mem_cgroup_kmem_disabled()) {
- p = kmem_buckets_alloc(b, 128, GFP_KERNEL | __GFP_ACCOUNT);
+ KUNIT_EXPECT_TRUE_MSG(test, strstarts(c->name, "covered_buckets-cg-"),
+ "expected the accounted bucket cache, got %s",
+ c->name);
+ KUNIT_EXPECT_TRUE(test, c->flags & SLAB_ACCOUNT);
+ } else {
+ /*
+ * Nothing is creating accounted caches, so the row aliases
+ * the normal one and the allocation lands there -- isolated
+ * still, just not separately accounted.
+ */
+ KUNIT_EXPECT_PTR_EQ(test, c, normal_cache);
+ }
+}
+
+/*
+ * The alignment a set is created with reaches every row it holds, not just
+ * the normal one. 256 is stronger than kmalloc's alignment for 128 byte
+ * objects, so a row built without it would show here.
+ */
+static void test_kmem_buckets_type_covered_alignment(struct kunit *test)
+{
+ static const gfp_t gfps[] = { GFP_KERNEL, GFP_KERNEL | __GFP_ACCOUNT };
+ const unsigned int align = 256;
+ struct kmem_cache *c;
+ kmem_buckets *b;
+ void *p;
+ int i;
+
+ if (!IS_ENABLED(CONFIG_SLAB_BUCKETS))
+ kunit_skip(test, "needs CONFIG_SLAB_BUCKETS");
+
+ b = kmem_buckets_create_types("covered_aligned", align, 0, 0, INT_MAX,
+ NULL, BIT(KMEM_BUCKET_NORMAL) |
+ BIT(KMEM_BUCKET_CGROUP));
+ KUNIT_ASSERT_BUCKETS_CREATED(test, b);
+
+ for (i = 0; i < ARRAY_SIZE(gfps); i++) {
+ p = kmem_buckets_alloc(b, 128, gfps[i]);
KUNIT_ASSERT_NOT_NULL(test, p);
- c = virt_to_slab(p)->slab_cache;
+ c = cache_of(p);
+ KUNIT_EXPECT_TRUE_MSG(test, IS_ALIGNED((unsigned long)p, align),
+ "gfp %pGg: object %p is not %u byte aligned",
+ &gfps[i], p, align);
kfree(p);
- KUNIT_EXPECT_TRUE_MSG(test, strstarts(c->name, "kmalloc-cg-"),
- "expected an accounted cache, got %s", c->name);
+ KUNIT_ASSERT_NOT_NULL(test, c);
+ KUNIT_EXPECT_TRUE_MSG(test, strstarts(c->name, "covered_aligned-"),
+ "gfp %pGg: expected a bucket cache, got %s",
+ &gfps[i], c->name);
+ KUNIT_EXPECT_EQ_MSG(test, c->align, align,
+ "gfp %pGg: %s aligned to %u, not %u",
+ &gfps[i], c->name, c->align, align);
}
}
@@ -800,6 +889,8 @@ static struct kunit_case test_cases[] = {
KUNIT_CASE(test_kmem_buckets_disabled),
KUNIT_CASE(test_kmem_buckets_destroy),
KUNIT_CASE(test_kmem_buckets_type_fallback),
+ KUNIT_CASE(test_kmem_buckets_type_covered),
+ KUNIT_CASE(test_kmem_buckets_type_covered_alignment),
{}
};
diff --git a/mm/slab_common.c b/mm/slab_common.c
index d66e5e56a0f1..c045e2d41ee5 100644
--- a/mm/slab_common.c
+++ b/mm/slab_common.c
@@ -410,9 +410,15 @@ EXPORT_SYMBOL(__kmem_cache_create_args);
static struct kmem_cache *kmem_buckets_cache __ro_after_init;
+static int kmem_buckets_create_row(kmem_buckets *b,
+ enum kmalloc_cache_type type,
+ const char *name, unsigned int align,
+ slab_flags_t flags, unsigned int useroffset,
+ unsigned int usersize, void (*ctor)(void *));
+
/**
- * kmem_buckets_create - Create a set of caches that handle dynamic sized
- * allocations via kmem_buckets_alloc()
+ * kmem_buckets_create_types - Create a set of caches that handle dynamic sized
+ * allocations via kmem_buckets_alloc()
* @name: A prefix string which is used in /proc/slabinfo to identify this
* cache. The individual caches with have their sizes as the suffix.
* @align: The required alignment for the objects, or 0 to give each cache
@@ -424,6 +430,11 @@ static struct kmem_cache *kmem_buckets_cache __ro_after_init;
* @usersize: How many bytes, starting at @useroffset, may be copied
* to/from userspace.
* @ctor: A constructor for the objects, run when new allocations are made.
+ * @type_mask: Which kmalloc types to hold caches for, as a mask of
+ * BIT(KMEM_BUCKET_*). KMEM_BUCKET_NORMAL is always included.
+ * Allocations of a type that is not covered are served by the
+ * general caches instead, so a caller need not know in advance
+ * which types its own callers will ask for.
*
* Context: Cannot be called within an interrupt, but can be interrupted.
*
@@ -432,13 +443,14 @@ static struct kmem_cache *kmem_buckets_cache __ro_after_init;
* subsequent calls to kmem_buckets_alloc() will fall back to kmalloc().
* (i.e. callers only need to check for NULL on failure.)
*/
-kmem_buckets *kmem_buckets_create(const char *name, unsigned int align,
- slab_flags_t flags,
- unsigned int useroffset,
- unsigned int usersize,
- void (*ctor)(void *))
-{
- unsigned int idx;
+kmem_buckets *kmem_buckets_create_types(const char *name, unsigned int align,
+ slab_flags_t flags,
+ unsigned int useroffset,
+ unsigned int usersize,
+ void (*ctor)(void *),
+ unsigned int type_mask)
+{
+ enum kmem_bucket_type btype;
kmem_buckets *b;
/*
@@ -457,20 +469,87 @@ kmem_buckets *kmem_buckets_create(const char *name, unsigned int align,
return NULL;
flags |= SLAB_NO_MERGE;
+ type_mask |= BIT(KMEM_BUCKET_NORMAL);
+
+ for (btype = 0; btype < NR_KMEM_BUCKET_TYPES; btype++) {
+ enum kmalloc_cache_type src = KMALLOC_NORMAL;
+ slab_flags_t type_flags = 0;
+
+ if (!(type_mask & BIT(btype)))
+ continue;
+
+ /*
+ * Under CONFIG_MEMCG=n the two types are the same value, so
+ * the IS_ENABLED() is what keeps the normal row out of here.
+ */
+ if (IS_ENABLED(CONFIG_MEMCG) && btype == KMEM_BUCKET_CGROUP) {
+ /*
+ * Aliasing below reads the normal row, so this loop
+ * must have built it already. That holds only while
+ * the normal type sorts first.
+ */
+ BUILD_BUG_ON(KMEM_BUCKET_CGROUP <= KMEM_BUCKET_NORMAL);
+
+ /*
+ * Nothing anywhere is creating accounted caches, as
+ * with "cgroup.memory=nokmem". Point this row's
+ * entries at the normal row's caches, the way
+ * new_kmalloc_cache() aliases kmalloc_caches[] for
+ * the same reason. Leaving the row empty instead
+ * would send every accounted allocation out of the
+ * set and into the general caches.
+ */
+ if (mem_cgroup_kmem_disabled()) {
+ memcpy(b[btype], b[KMEM_BUCKET_NORMAL],
+ sizeof(b[btype]));
+ continue;
+ }
+
+ type_flags = SLAB_ACCOUNT;
+ src = KMALLOC_CGROUP;
+ }
+
+ if (kmem_buckets_create_row(&b[btype], src, name, align,
+ flags | type_flags, useroffset,
+ usersize, ctor))
+ goto fail;
+ }
+
+ return b;
+
+fail:
+ kmem_buckets_destroy(b);
+
+ return NULL;
+}
+EXPORT_SYMBOL(kmem_buckets_create_types);
+
+/*
+ * Build one row of @b by mirroring the general caches of @type: a cache per
+ * kmalloc size, each named "@name-" followed by that cache's own suffix, so
+ * a row of KMALLOC_CGROUP ("kmalloc-cg-96") gets "@name-cg-96".
+ */
+static int kmem_buckets_create_row(kmem_buckets *b,
+ enum kmalloc_cache_type type,
+ const char *name, unsigned int align,
+ slab_flags_t flags, unsigned int useroffset,
+ unsigned int usersize, void (*ctor)(void *))
+{
+ unsigned int idx;
for (idx = 0; idx < ARRAY_SIZE(kmalloc_caches[KMALLOC_NORMAL]); idx++) {
char *short_size, *cache_name;
unsigned int cache_useroffset, cache_usersize;
unsigned int size, aligned_idx;
- if (!kmalloc_caches[KMALLOC_NORMAL][idx])
+ if (!kmalloc_caches[type][idx])
continue;
- size = kmalloc_caches[KMALLOC_NORMAL][idx]->object_size;
+ size = kmalloc_caches[type][idx]->object_size;
if (!size)
continue;
- short_size = strchr(kmalloc_caches[KMALLOC_NORMAL][idx]->name, '-');
+ short_size = strchr(kmalloc_caches[type][idx]->name, '-');
if (WARN_ON(!short_size))
goto fail;
@@ -488,7 +567,7 @@ kmem_buckets *kmem_buckets_create(const char *name, unsigned int align,
if (WARN_ON(!cache_name))
goto fail;
(*b)[aligned_idx] = kmem_cache_create_usercopy(cache_name, size,
- align ?: kmalloc_caches[KMALLOC_NORMAL][idx]->align,
+ align ?: kmalloc_caches[type][idx]->align,
flags, cache_useroffset,
cache_usersize, ctor);
kfree(cache_name);
@@ -499,14 +578,11 @@ kmem_buckets *kmem_buckets_create(const char *name, unsigned int align,
(*b)[idx] = (*b)[aligned_idx];
}
- return b;
+ return 0;
fail:
- kmem_buckets_destroy(b);
-
- return NULL;
+ return -ENOMEM;
}
-EXPORT_SYMBOL(kmem_buckets_create);
/**
* kmem_buckets_destroy - Destroy a set of caches made by kmem_buckets_create()
@@ -521,28 +597,35 @@ EXPORT_SYMBOL(kmem_buckets_create);
*/
void kmem_buckets_destroy(kmem_buckets *bucket)
{
+ enum kmem_bucket_type btype, t;
unsigned int idx, i;
if (!IS_ENABLED(CONFIG_SLAB_BUCKETS) || ZERO_OR_NULL_PTR(bucket))
return;
- for (idx = 0; idx < ARRAY_SIZE(kmalloc_caches[KMALLOC_NORMAL]); idx++) {
- struct kmem_cache *cache = (*bucket)[idx];
+ for (btype = 0; btype < NR_KMEM_BUCKET_TYPES; btype++) {
+ for (idx = 0; idx < ARRAY_SIZE(bucket[btype]); idx++) {
+ struct kmem_cache *cache = bucket[btype][idx];
- if (!cache)
- continue;
+ if (!cache)
+ continue;
- /*
- * Sizes below arch_slab_minalign() share one cache, which
- * kmem_buckets_create() then stores at each of their indices.
- * Drop every reference to it before destroying it, so that no
- * later pass reads a pointer to a cache that is already gone.
- */
- for (i = idx; i < ARRAY_SIZE(kmalloc_caches[KMALLOC_NORMAL]); i++)
- if ((*bucket)[i] == cache)
- (*bucket)[i] = NULL;
+ /*
+ * A cache is reachable from more than one entry: sizes
+ * below arch_slab_minalign() share one, and a row that
+ * kmem_buckets_create_types() aliased onto the normal
+ * one under "cgroup.memory=nokmem" holds all of them a
+ * second time. Drop every reference before destroying
+ * it, so that no later pass reads a pointer to a cache
+ * that is already gone.
+ */
+ for (t = 0; t < NR_KMEM_BUCKET_TYPES; t++)
+ for (i = 0; i < ARRAY_SIZE(bucket[t]); i++)
+ if (bucket[t][i] == cache)
+ bucket[t][i] = NULL;
- kmem_cache_destroy(cache);
+ kmem_cache_destroy(cache);
+ }
}
kmem_cache_free(kmem_buckets_cache, bucket);
@@ -1092,7 +1175,7 @@ void __init create_kmalloc_caches(void)
if (IS_ENABLED(CONFIG_SLAB_BUCKETS))
kmem_buckets_cache = kmem_cache_create("kmalloc_buckets",
- sizeof(kmem_buckets),
+ sizeof(kmem_buckets) * NR_KMEM_BUCKET_TYPES,
0, SLAB_NO_MERGE, NULL);
}
--
2.34.1
^ permalink raw reply [flat|nested] 14+ messages in thread
* [PATCH net-next v5 7/7] net: skb: isolate skb data area allocations into a separate bucket
2026-10-02 23:11 [PATCH net-next v5 0/7] net: skb: isolate skb data area allocations into a separate bucket Kees Cook
` (5 preceding siblings ...)
2026-10-02 23:11 ` [PATCH net-next v5 6/7] mm/slab: Let a bucket set handle __GFP_ACCOUNT Kees Cook
@ 2026-10-02 23:11 ` Kees Cook
2026-10-03 23:13 ` netdev-bot+sashiko
6 siblings, 1 reply; 14+ messages in thread
From: Kees Cook @ 2026-10-02 23:11 UTC (permalink / raw)
To: Vlastimil Babka
Cc: Kees Cook, Pedro Falcato, David S. Miller, Eric Dumazet,
Jakub Kicinski, Paolo Abeni, Simon Horman, Willem de Bruijn,
Jason Xing, netdev, Kuniyuki Iwashima, linux-hardening,
Harry Yoo (Meta),
Andrew Morton, Hao Li, Christoph Lameter, David Rientjes,
Roman Gushchin, Johannes Weiner, Michal Hocko, Shakeel Butt,
Muchun Song, Mina Almasry, Björn Töpel, Jiayuan Chen,
linux-kernel, linux-mm, cgroups
From: Pedro Falcato <pfalcato@suse.de>
SKB data area allocations (as done from alloc_skb()) use kmalloc().
These allocations can be variably sized and their contents can be more
or less controlled from userspace, which makes them useful for attackers
that want to overwrite a use-after-free'd object from the same kmalloc slab
(which often just requires the sizes to roughly match into the same kmalloc
bucket). [0] is an easy example of an exploit that uses netlink skb
allocation to target another similarly-sized accidentally freed object.
While other mitigations like CONFIG_RANDOM_KMALLOC_CACHES exist, these are
probabilistic. Use the existing kmem buckets API to further isolate these
allocations in a guaranteed fashion, when CONFIG_SLAB_BUCKETS=y.
Ask for the accounted kmalloc type as well as the normal one. AF_UNIX
sets sk_allocation to GFP_KERNEL_ACCOUNT, so without it every AF_UNIX
skb data area would fall back to the general caches, and those are the
ones most worth isolating. GFP_DMA is left to fall back, being passed to
an skb allocator only by rare devices.
Link: https://github.com/google/security-research/blob/master/pocs/linux/kernelctf/CVE-2023-4207_lts_cos_mitigation_2/docs/exploit.md [0]
Reviewed-by: Kees Cook <kees@kernel.org>
Signed-off-by: Pedro Falcato <pfalcato@suse.de>
Acked-by: Paolo Abeni <pabeni@redhat.com>
Signed-off-by: Kees Cook <kees@kernel.org>
---
net/core/skbuff.c | 11 +++++++++--
1 file changed, 9 insertions(+), 2 deletions(-)
diff --git a/net/core/skbuff.c b/net/core/skbuff.c
index 966af3beed94..e0660b2dbc19 100644
--- a/net/core/skbuff.c
+++ b/net/core/skbuff.c
@@ -586,6 +586,8 @@ struct sk_buff *napi_build_skb(void *data, unsigned int frag_size)
}
EXPORT_SYMBOL(napi_build_skb);
+static kmem_buckets *skb_data_buckets __ro_after_init;
+
static void *kmalloc_pfmemalloc(size_t obj_size, gfp_t flags, int node)
{
if (!gfp_pfmemalloc_allowed(flags))
@@ -593,7 +595,8 @@ static void *kmalloc_pfmemalloc(size_t obj_size, gfp_t flags, int node)
if (!obj_size)
return kmem_cache_alloc_node(net_hotdata.skb_small_head_cache,
flags, node);
- return kmalloc_node_track_caller(obj_size, flags, node);
+ return kmem_buckets_alloc_node_track_caller(skb_data_buckets, obj_size,
+ flags, node);
}
/*
@@ -634,7 +637,7 @@ static void *kmalloc_reserve(unsigned int *size, gfp_t flags, int node,
* Try a regular allocation, when that fails and we're not entitled
* to the reserves, fail.
*/
- obj = kmalloc_node_track_caller(obj_size,
+ obj = kmem_buckets_alloc_node_track_caller(skb_data_buckets, obj_size,
flags | __GFP_NOMEMALLOC | __GFP_NOWARN,
node);
if (likely(obj))
@@ -5235,6 +5238,10 @@ void __init skb_init(void)
0,
SKB_SMALL_HEAD_HEADROOM,
NULL);
+ skb_data_buckets = kmem_buckets_create_types("skb_data", 0, SLAB_PANIC,
+ 0, INT_MAX, NULL,
+ BIT(KMEM_BUCKET_NORMAL) |
+ BIT(KMEM_BUCKET_CGROUP));
skb_extensions_init();
}
--
2.34.1
^ permalink raw reply [flat|nested] 14+ messages in thread
* Re: [PATCH net-next v5 2/7] mm/slab: Let kmem_buckets_create() take an alignment
2026-10-02 23:11 ` [PATCH net-next v5 2/7] mm/slab: Let kmem_buckets_create() take an alignment Kees Cook
@ 2026-10-03 23:13 ` netdev-bot+sashiko
0 siblings, 0 replies; 14+ messages in thread
From: netdev-bot+sashiko @ 2026-10-03 23:13 UTC (permalink / raw)
To: kees
Cc: vbabka, harry, akpm, hao.li, cl, rientjes, roman.gushchin,
linux-mm, pfalcato, kuniyu, linux-hardening, davem, hannes,
mhocko, shakeel.butt, muchun.song, edumazet, kuba, pabeni, horms,
kerneljasonxing, willemb, almasrymina, bjorn, jiayuan.chen,
linux-kernel, cgroups, netdev
Thank you for your contribution! Sashiko AI review found 2 potential
issue(s) to consider.
Critical: 0 · High: 0 · Medium: 1 · Low: 1
If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).
- [Medium] When kmem_buckets_create() gets a non-zero @align, it uses
that value as the cache alignment instead of the alignment of the…
- [Low] The new kernel-doc for kmem_buckets_create() (mm/slab_common.c
lines 418-420, copied to include/linux/slab.h and…
--
mm/slab: Let kmem_buckets_create() take an alignment
This commit adds an align argument to kmem_buckets_create(). When align is
0, each bucket cache takes the alignment of the kmalloc cache of the same
size, so bucket allocations match kmalloc() when slab metadata is enabled.
Both existing callers now pass 0.
> diff --git a/mm/slab_common.c b/mm/slab_common.c
> index b93fdb5cb0219..71926f726a07a 100644
> --- a/mm/slab_common.c
> +++ b/mm/slab_common.c
> @@ -415,6 +415,9 @@ static struct kmem_cache *kmem_buckets_cache __ro_after_init;
> * allocations via kmem_buckets_alloc()
> * @name: A prefix string which is used in /proc/slabinfo to identify this
> * cache. The individual caches with have their sizes as the suffix.
> + * @align: The required alignment for the objects, or 0 to give each cache
> + * the alignment of the kmalloc cache of the same size, as a caller
> + * moving from kmalloc() may depend on.
[Severity: Low]
The new kernel-doc calls @align "The required alignment for the objects".
Does every path covered by this doc honour it?
With CONFIG_SLAB_BUCKETS=n, kmem_buckets_create() returns without looking
at align:
if (!IS_ENABLED(CONFIG_SLAB_BUCKETS))
return ZERO_SIZE_PTR;
PASS_BUCKET_PARAM(_b) is also NULL in that config, so allocations are served
from the general caches.
When set creation fails, the doc says callers can keep using the NULL
result and that allocations "will fall back to kmalloc()". The caller in
ipc/msgutil.c:init_msg_buckets() never checks the return value.
In both cases, kmalloc_slab() does this:
if (!b)
b = &kmalloc_caches[type];
The object then gets only kmalloc's natural alignment for that size. Take
align=256 with 64-byte objects, which is the case in the series' kunit test
(the test skips itself when !CONFIG_SLAB_BUCKETS). That caller would get
less-aligned memory on some configs and see no warning.
Later in the series, allocations of kmalloc types the set does not cover
also go to the general caches and lose align in the same way.
Should the kernel-doc say that a non-zero align only applies when the set
was actually created? Or should these fallback paths honour it?
> * @flags: SLAB flags (see kmem_cache_create() for details).
> * @useroffset: Starting offset within an allocation that may be copied
> * to/from userspace.
[ ... ]
> @@ -487,7 +491,8 @@ kmem_buckets *kmem_buckets_create(const char *name, slab_flags_t flags,
> if (WARN_ON(!cache_name))
> goto fail;
> (*b)[aligned_idx] = kmem_cache_create_usercopy(cache_name, size,
> - 0, flags, cache_useroffset,
> + align ?: kmalloc_caches[KMALLOC_NORMAL][idx]->align,
[Severity: Medium]
A non-zero align replaces the alignment of the mirrored kmalloc cache
instead of being combined with it. Below this, calculate_alignment() only
raises it to arch_slab_minalign():
align = max(align, arch_slab_minalign());
Nothing raises it to ARCH_KMALLOC_MINALIGN, ARCH_DMA_MINALIGN, or the
power-of-two size alignment that create_boot_cache() gives kmalloc caches.
Could a caller that passes a small non-zero align end up with weaker
alignment than plain kmalloc()?
kmem_buckets_alloc() goes through __kmalloc_node_noprof(), which is
declared __assume_kmalloc_alignment. Some architectures leave
ARCH_KMALLOC_MINALIGN at its default of ARCH_DMA_MINALIGN. On those, an
explicit align below that value, with KASAN, SLUB_DEBUG_ON or slab_debug=
enabled, would break the compiler's alignment assumption.
dma_kmalloc_size_aligned() also decides whether to bounce from the size
alone:
/*
* Larger kmalloc() sizes are guaranteed to be aligned to
* ARCH_DMA_MINALIGN.
*/
if (size >= 2 * ARCH_DMA_MINALIGN ||
For example, take align=64 on arm64 (ARCH_DMA_MINALIGN=128) with slab_debug
enabled. The 2048-byte objects would be only 64-byte aligned, and swiotlb
would never bounce them. Is that the same non-coherent DMA problem this
patch fixes for align == 0, just reached through a non-zero align?
Every in-tree caller passes 0 through the end of the series. The only
non-zero caller is the kunit test with align=256, so nothing triggers this
today. The comment in test_kmem_buckets_explicit_alignment says it "checks
the override both ways". That suggests alignment weaker than kmalloc's is
intended.
The same expression appears later in the series in
kmem_buckets_create_row(), as align ?: kmalloc_caches[type][idx]->align.
Would max(align, kmalloc_caches[KMALLOC_NORMAL][idx]->align) keep the
kmalloc guarantees that bucket users and the DMA layer rely on? Clamping to
at least ARCH_KMALLOC_MINALIGN might also work.
> + flags, cache_useroffset,
> cache_usersize, ctor);
> kfree(cache_name);
> if (WARN_ON(!(*b)[aligned_idx]))
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20261002231120.late.500-kees%40kernel.org
^ permalink raw reply [flat|nested] 14+ messages in thread
* Re: [PATCH net-next v5 3/7] mm/slab: Add kmem_buckets_destroy()
2026-10-02 23:11 ` [PATCH net-next v5 3/7] mm/slab: Add kmem_buckets_destroy() Kees Cook
@ 2026-10-03 23:13 ` netdev-bot+sashiko
0 siblings, 0 replies; 14+ messages in thread
From: netdev-bot+sashiko @ 2026-10-03 23:13 UTC (permalink / raw)
To: kees
Cc: vbabka, harry, akpm, hao.li, cl, rientjes, roman.gushchin,
linux-mm, pfalcato, kuniyu, linux-hardening, davem, hannes,
mhocko, shakeel.butt, muchun.song, edumazet, kuba, pabeni, horms,
kerneljasonxing, willemb, almasrymina, bjorn, jiayuan.chen,
linux-kernel, cgroups, netdev
Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.
Low priority / nit picks for your consideration:
- [Low] The new comment in kmem_buckets_destroy() (mm/slab_common.c)
says: "Sizes below arch_slab_minalign() share one cache, which…
--
mm/slab: Add kmem_buckets_destroy()
This adds kmem_buckets_destroy() as a counterpart to kmem_buckets_create(),
so that KUnit tests built as modules can tear down a bucket set. The
create path's bitmap-based error unwinding is replaced with a call to the
new helper. The helper clears every aliased pointer to a cache before
destroying it.
Link: https://lore.kernel.org/all/20240809073309.2134488-1-kees@kernel.org/
> diff --git a/mm/slab_common.c b/mm/slab_common.c
> index 71926f726a07a..bdc34a5e6d6f5 100644
> --- a/mm/slab_common.c
> +++ b/mm/slab_common.c
[ ... ]
> +void kmem_buckets_destroy(kmem_buckets *bucket)
> +{
> + unsigned int idx, i;
> +
> + if (!IS_ENABLED(CONFIG_SLAB_BUCKETS) || ZERO_OR_NULL_PTR(bucket))
> + return;
> +
> + for (idx = 0; idx < ARRAY_SIZE(kmalloc_caches[KMALLOC_NORMAL]); idx++) {
> + struct kmem_cache *cache = (*bucket)[idx];
> +
> + if (!cache)
> + continue;
> +
> + /*
> + * Sizes below arch_slab_minalign() share one cache, which
> + * kmem_buckets_create() then stores at each of their indices.
[Severity: Low]
This isn't a bug, but is this comment accurate about where the aliasing
comes from?
The aliasing in kmalloc_caches[] is set up by new_kmalloc_cache(). It uses
__kmalloc_minalign(), and that is not arch_slab_minalign() on its own:
mm/slab_common.c:__kmalloc_minalign() {
unsigned int minalign = dma_get_cache_alignment();
...
return max(minalign, arch_slab_minalign());
}
new_kmalloc_cache() then rounds every size up to that alignment:
if (minalign > ARCH_KMALLOC_MINALIGN) {
aligned_size = ALIGN(aligned_size, minalign);
aligned_idx = __kmalloc_index(aligned_size, false);
}
This aliases sizes that are not multiples of minalign, even when they are
not below it. For example, kmalloc-96 maps to the 128 cache when minalign
is 64, and kmalloc-192 maps to the 256 cache when minalign is 128.
kmem_buckets_create() reads object_size from the already-aliased
kmalloc_caches[KMALLOC_NORMAL][idx], so the same aliasing ends up in the
bucket set.
The code itself is fine, because the inner loop clears every matching
pointer whatever caused the alias. Could the comment say that sizes are
rounded up to the kmalloc minimum alignment, rather than "sizes below
arch_slab_minalign()"?
The same wording also survives in the later commit "mm/slab: Let a
bucket set handle __GFP_ACCOUNT". That commit rewords this comment but
keeps "sizes below arch_slab_minalign() share one".
> + * Drop every reference to it before destroying it, so that no
> + * later pass reads a pointer to a cache that is already gone.
> + */
> + for (i = idx; i < ARRAY_SIZE(kmalloc_caches[KMALLOC_NORMAL]); i++)
> + if ((*bucket)[i] == cache)
> + (*bucket)[i] = NULL;
> +
> + kmem_cache_destroy(cache);
> + }
[ ... ]
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20261002231120.late.500-kees%40kernel.org
^ permalink raw reply [flat|nested] 14+ messages in thread
* Re: [PATCH net-next v5 4/7] mm/slab: Add tests for the existing kmem_buckets behaviour
2026-10-02 23:11 ` [PATCH net-next v5 4/7] mm/slab: Add tests for the existing kmem_buckets behaviour Kees Cook
@ 2026-10-03 23:13 ` netdev-bot+sashiko
0 siblings, 0 replies; 14+ messages in thread
From: netdev-bot+sashiko @ 2026-10-03 23:13 UTC (permalink / raw)
To: kees
Cc: vbabka, harry, akpm, hao.li, cl, rientjes, roman.gushchin,
linux-mm, pfalcato, kuniyu, linux-hardening, davem, hannes,
mhocko, shakeel.butt, muchun.song, edumazet, kuba, pabeni, horms,
kerneljasonxing, willemb, almasrymina, bjorn, jiayuan.chen,
linux-kernel, cgroups, netdev
Thank you for your contribution! Sashiko AI review found 2 potential
issue(s) to consider.
Critical: 0 · High: 0 · Medium: 1 · Low: 1
If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).
- [Medium] test_kmem_buckets_destroy() (lib/tests/slub_kunit.c) can fail
at random on CONFIG_KFENCE=y kernels.
- [Low] test_kmem_buckets_sizes() (lib/tests/slub_kunit.c) checks less
than its comment and the commit message say it does.
--
mm/slab: Add tests for the existing kmem_buckets behaviour
This adds KUnit tests to lib/tests/slub_kunit.c for kmem_buckets. They
cover cache isolation, size class routing, default and explicit cache
alignment, the CONFIG_SLAB_BUCKETS=n fallback, and teardown of a set's
caches by kmem_buckets_destroy().
> diff --git a/lib/tests/slub_kunit.c b/lib/tests/slub_kunit.c
> index e3b63f0338d5f..3c923a3af825b 100644
> --- a/lib/tests/slub_kunit.c
> +++ b/lib/tests/slub_kunit.c
[ ... ]
> @@ -474,6 +475,251 @@ static int test_init(struct kunit *test)
[ ... ]
> +/*
> + * Every size class gets its own cache in the set, including the ones that
> + * are not powers of two and are filled in from an aligned index. Sizes past
> + * the largest cache are served by the page allocator, bucket set or not.
> + */
> +static void test_kmem_buckets_sizes(struct kunit *test)
> +{
> + static const size_t sizes[] = { 8, 96, 192, 1024, 4096 };
[ ... ]
> + for (i = 0; i < ARRAY_SIZE(sizes); i++) {
> + p = kmem_buckets_alloc(b, sizes[i], GFP_KERNEL);
> + KUNIT_ASSERT_NOT_NULL(test, p);
> + c = cache_of(p);
> + kfree(p);
> + KUNIT_ASSERT_NOT_NULL(test, c);
> +
> + KUNIT_EXPECT_TRUE_MSG(test, strstarts(c->name, "sized_buckets-"),
> + "size %zu: expected a bucket cache, got %s",
> + sizes[i], c->name);
> + KUNIT_EXPECT_GE(test, c->object_size, sizes[i]);
[Severity: Low]
Is this check strong enough to back up the comment above and the commit
message? The commit message says:
Each size class is served by the set, including 96 and 192, which are
not powers of two and are filled in from an aligned index by
kmem_buckets_create().
The test only checks the name prefix and uses >= on object_size. It would
still pass if a size went to a larger bucket cache, for example 8 bytes
ending up in sized_buckets-8k.
On x86_64, __kmalloc_minalign() returns 8, which is the same as
ARCH_KMALLOC_MINALIGN. So new_kmalloc_cache() aliases nothing, and this
branch in kmem_buckets_create() never runs there:
mm/slab_common.c:kmem_buckets_create() {
...
if (idx != aligned_idx)
(*b)[idx] = (*b)[aligned_idx];
...
}
On configurations where minalign is larger than ARCH_KMALLOC_MINALIGN
(for example arm64 without swiotlb), sizes such as 8 and 96 share a cache
with a larger class. "Every size class gets its own cache" doesn't hold
there, and the >= check hides that.
Would it be tighter to compare c->object_size against
kmalloc_size_roundup(sizes[i]), or against the object_size of the
matching general kmalloc cache? The test also looks unchanged at the end
of the series.
> + }
[ ... ]
> +/* Destroying a set has to take its caches down, not just free the set. */
> +static void test_kmem_buckets_destroy(struct kunit *test)
> +{
[ ... ]
> + b = kmem_buckets_create("destroyed_buckets", 0, 0, 0, INT_MAX, NULL);
> + KUNIT_ASSERT_NOT_NULL(test, b);
> +
> + /*
> + * Deliberately leaked, as test_leak_destroy() leaks its own: the
> + * teardown below has to find it. kmem_cache_destroy() unlists the
> + * cache either way, so the name is still released.
> + */
> + p = kmem_buckets_alloc(b, 128, GFP_KERNEL);
> + KUNIT_EXPECT_NOT_NULL(test, p);
> +
> + kmem_buckets_destroy(b);
> +
> + KUNIT_EXPECT_EQ(test, 2, slab_errors);
[Severity: Medium]
Can this fail at random on CONFIG_KFENCE=y kernels?
kmem_buckets_create() only adds SLAB_NO_MERGE to the bucket caches.
Nothing sets SLAB_SKIP_KFENCE, so slab_alloc_node()->kfence_alloc() can
serve this 128 byte allocation from the KFENCE pool.
In that case the leaked object is not in any slab on the node partial
list:
kmem_buckets_destroy()
kmem_cache_destroy()
__kmem_cache_shutdown()
free_partial() <- no in-use slab, so list_slab_objects()
is never called and returns 0
kmem_cache_release()
kfence_shutdown_cache() <- object becomes a zombie, no kunit error
slab_errors would then stay at 0, and the check against 2 would fail.
Whether it fails depends only on the KFENCE sample timer.
The existing test_leak_destroy() avoids this by going through
test_kmem_cache_create(), which does:
s->flags |= SLAB_SKIP_KFENCE;
Could this test set SLAB_SKIP_KFENCE on the 128 byte bucket cache before
allocating? Another option is to skip or retry when is_kfence_address(p)
is true. This also looks unchanged at the end of the series.
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20261002231120.late.500-kees%40kernel.org
^ permalink raw reply [flat|nested] 14+ messages in thread
* Re: [PATCH net-next v5 5/7] mm/slab: Provide kmalloc type fallback for bucket allocations
2026-10-02 23:11 ` [PATCH net-next v5 5/7] mm/slab: Provide kmalloc type fallback for bucket allocations Kees Cook
@ 2026-10-03 23:13 ` netdev-bot+sashiko
0 siblings, 0 replies; 14+ messages in thread
From: netdev-bot+sashiko @ 2026-10-03 23:13 UTC (permalink / raw)
To: kees
Cc: vbabka, harry, akpm, hao.li, cl, rientjes, roman.gushchin,
linux-mm, pfalcato, kuniyu, linux-hardening, davem, hannes,
mhocko, shakeel.butt, muchun.song, edumazet, kuba, pabeni, horms,
kerneljasonxing, willemb, almasrymina, bjorn, jiayuan.chen,
linux-kernel, cgroups, netdev
Thank you for your contribution! Sashiko AI review found 1 potential
issue(s) to consider.
Critical: 0 · High: 0 · Medium: 1 · Low: 0
If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).
- [Medium] kmem_buckets_create() takes @ctor, @align,
@useroffset/@usersize and @flags (e.g. SLAB_ACCOUNT), and applies them…
--
mm/slab: Provide kmalloc type fallback for bucket allocations
This patch changes kmalloc_slab() so that a bucketed allocation whose
kmalloc type the bucket set doesn't hold (DMA, reclaimable, accounted,
and so on) goes to the general kmalloc caches. Before, it was always
served from the bucket's normal caches. It also adds a KUnit test for
the fallback and exports mem_cgroup_kmem_disabled() for KUnit.
Link: https://lore.kernel.org/all/04debe19-bbe8-4b5f-9668-753d1f97832d@redhat.com/ [1]
> diff --git a/mm/slab.h b/mm/slab.h
> index 8fd6835e42352..7f1bfee83b92e 100644
> --- a/mm/slab.h
> +++ b/mm/slab.h
> @@ -421,6 +421,26 @@ static inline unsigned int size_index_elem(unsigned int bytes)
> return (bytes - 1) / 8;
> }
>
> +/*
> + * Which set of buckets to use for the given kmalloc_cache_type. If not
> + * handled by the kmem_buckets, fall back to general caches.
> + */
> +static inline kmem_buckets *
> +kmalloc_choose_bucket(kmem_buckets *bucket, enum kmalloc_cache_type type)
> +{
> + enum kmem_bucket_type btype;
> +
> + if (!bucket)
> + return &kmalloc_caches[type];
> +
> + if (type <= KMALLOC_PARTITION_END)
> + btype = KMEM_BUCKET_NORMAL;
> + else
> + return &kmalloc_caches[type]; /* No set holds a row for it. */
[Severity: Medium]
If a bucket set is in use and the type is above KMALLOC_PARTITION_END,
this returns the general kmalloc caches. That happens for GFP_DMA with
CONFIG_ZONE_DMA, __GFP_RECLAIMABLE, and __GFP_ACCOUNT with CONFIG_MEMCG.
Do the parameters passed to kmem_buckets_create() still apply in that
case?
kmem_buckets_create() applies align, flags, useroffset/usersize and ctor
only to the bucket's own caches:
mm/slab_common.c:kmem_buckets_create() {
...
(*b)[aligned_idx] = kmem_cache_create_usercopy(cache_name, size,
align ?: kmalloc_caches[KMALLOC_NORMAL][idx]->align,
flags, cache_useroffset,
cache_usersize, ctor);
...
}
The general caches have no ctor, and their usercopy window covers the
whole object:
mm/slab_common.c:create_kmalloc_cache() {
...
create_boot_cache(s, name, size, flags | SLAB_KMALLOC, 0, size);
...
}
With the fallback, it looks like four things are lost:
- the ctor never runs
- the requested alignment is not honoured
- the hardened usercopy window widens to the whole object
- creator flags such as SLAB_ACCOUNT are dropped
For the SLAB_ACCOUNT case, ipc/msgutil.c creates msg_msg this way:
msg_buckets = kmem_buckets_create("msg_msg", 0, SLAB_ACCOUNT,
sizeof(struct msg_msg),
DATALEN_MSG, NULL);
Suppose a bucket like this gets GFP_DMA or __GFP_RECLAIMABLE without
__GFP_ACCOUNT. The allocation would land in dma-kmalloc-N or
kmalloc-rcl-N, and neither has SLAB_ACCOUNT. memcg_slab_post_alloc_hook()
then returns early:
if (likely(!(flags & __GFP_ACCOUNT) && !(s->flags & SLAB_ACCOUNT)))
return true;
Wouldn't the memcg charge be skipped here?
Before this patch, every bucketed allocation got these parameters.
kmem_buckets_create() doesn't reject or warn about combinations that the
fallback can't honour.
No in-tree caller hits this today. memdup_user uses GFP_USER and msg_msg
uses GFP_KERNEL. skb data uses ctor=NULL, align=0 and a full usercopy
window, so it loses only isolation on its rare GFP_DMA path. The API is
exported, though.
A later patch in the series doesn't change this either. After "mm/slab:
Let a bucket set handle __GFP_ACCOUNT", the kmem_buckets_create()
kernel-doc in include/linux/slab.h says:
* Covers KMEM_BUCKET_NORMAL only. Allocations needing another kmalloc type
* are served by the general caches, keeping the type they asked for and
* losing only the isolation.
Is "losing only the isolation" accurate, given that the ctor, alignment,
usercopy restriction and creator flags are lost as well? The existing
kernel-doc in mm/slab_common.c also says the ctor is "run when new
allocations are made", with no exception for fallback types.
kmem_buckets_create_types() has the same fallback for any type it doesn't
cover.
Should kmem_buckets_create() reject or warn about the parameter
combinations the fallback can't satisfy? Or should the documentation
describe what is lost?
> +
> + return &bucket[btype];
> +}
> +
[ ... ]
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20261002231120.late.500-kees%40kernel.org
^ permalink raw reply [flat|nested] 14+ messages in thread
* Re: [PATCH net-next v5 6/7] mm/slab: Let a bucket set handle __GFP_ACCOUNT
2026-10-02 23:11 ` [PATCH net-next v5 6/7] mm/slab: Let a bucket set handle __GFP_ACCOUNT Kees Cook
@ 2026-10-03 23:13 ` netdev-bot+sashiko
0 siblings, 0 replies; 14+ messages in thread
From: netdev-bot+sashiko @ 2026-10-03 23:13 UTC (permalink / raw)
To: kees
Cc: vbabka, harry, akpm, hao.li, cl, rientjes, roman.gushchin,
linux-mm, pfalcato, kuniyu, linux-hardening, davem, hannes,
mhocko, shakeel.butt, muchun.song, edumazet, kuba, pabeni, horms,
kerneljasonxing, willemb, almasrymina, bjorn, jiayuan.chen,
linux-kernel, cgroups, netdev
Thank you for your contribution! Sashiko AI review found 1 potential
issue(s) to consider.
Critical: 0 · High: 0 · Medium: 1 · Low: 0
If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).
- [Medium] The patch adds a WARN_ONCE() in kmalloc_choose_bucket()
(mm/slab.h).
--
mm/slab: Let a bucket set handle __GFP_ACCOUNT
This patch adds an optional accounted row (KMEM_BUCKET_CGROUP) to a
kmem_buckets set. The row is selected through the type mask of a new
kmem_buckets_create_types(), so __GFP_ACCOUNT allocations can stay
isolated in the set. kmem_buckets_create() becomes an inline wrapper that
builds only the normal row, and under cgroup.memory=nokmem the accounted
row aliases the normal one.
> diff --git a/mm/slab.h b/mm/slab.h
> index 7f1bfee83b92e..2af44e09edda1 100644
> --- a/mm/slab.h
> +++ b/mm/slab.h
> @@ -435,10 +435,31 @@ kmalloc_choose_bucket(kmem_buckets *bucket, enum kmalloc_cache_type type)
>
> if (type <= KMALLOC_PARTITION_END)
> btype = KMEM_BUCKET_NORMAL;
> + else if (IS_ENABLED(CONFIG_MEMCG) && type == KMALLOC_CGROUP)
> + btype = KMEM_BUCKET_CGROUP;
> else
> return &kmalloc_caches[type]; /* No set holds a row for it. */
>
> - return &bucket[btype];
> + /*
> + * Either this row was created, and holds a cache everywhere the
> + * general caches hold one, or it was never created and holds nothing.
> + * Test with the KMALLOC_SHIFT_LOW which exists in every configuration.
> + */
> + if (likely(bucket[btype][KMALLOC_SHIFT_LOW]))
> + return &bucket[btype];
> +
> + /*
> + * A row this set _could_ have held, but was not created with: the type
> + * mask passed to kmem_buckets_create_types() did not cover what its
> + * callers actually tried to allocate. Report the mismatch but still
> + * fall back to the general caches.
> + *
> + * At present, only __GFP_ACCOUNT can be missing.
> + */
> + WARN_ONCE(1,
> + "kmem_buckets: __GFP_ACCOUNT needs BIT(KMEM_BUCKET_CGROUP) in create mask\n");
[Severity: Medium]
Can a set built with plain kmem_buckets_create() reach this WARN_ONCE()?
That wrapper passes only BIT(KMEM_BUCKET_NORMAL), so the
KMEM_BUCKET_CGROUP row stays zeroed. With CONFIG_MEMCG=y, take this call:
kmem_buckets_alloc(b, size, GFP_KERNEL | __GFP_ACCOUNT);
kmalloc_type() returns KMALLOC_CGROUP for it, including under
cgroup.memory=nokmem. kmalloc_choose_bucket() then maps that to
KMEM_BUCKET_CGROUP, finds bucket[KMEM_BUCKET_CGROUP][KMALLOC_SHIFT_LOW]
is NULL, and warns.
Before this patch, that allocation fell back to
kmalloc_caches[KMALLOC_CGROUP] without a warning, and
test_kmem_buckets_type_fallback() checked for a "kmalloc-cg-" cache. In
this patch that check is replaced by a comment saying the fallback warns.
The documentation added in this patch still seems to describe the old
behaviour. The kernel-doc for @type_mask in kmem_buckets_create_types()
says:
Allocations of a type that is not covered are served by the
general caches instead, so a caller need not know in advance
which types its own callers will ask for.
The kernel-doc for kmem_buckets_create() in include/linux/slab.h says:
Covers KMEM_BUCKET_NORMAL only. Allocations needing another kmalloc type
are served by the general caches, keeping the type they asked for and
losing only the isolation.
The commit message also says kmem_buckets_create() keeps its name,
"leaving existing users as-is", and does not mention the new warning.
No in-tree caller appears to hit this, and that is still true at the end
of the series. However, kmem_buckets_create_types() is exported, and a
caller that follows the documentation would get a WARN splat. With
panic_on_warn set, that becomes a panic.
Should the WARN_ONCE() be dropped to match the documented contract? If
the warning is intended, could the commit message and both kernel-doc
blocks say that __GFP_ACCOUNT allocations need BIT(KMEM_BUCKET_CGROUP)
and warn otherwise?
> +
> + return &kmalloc_caches[type];
> }
[ ... ]
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20261002231120.late.500-kees%40kernel.org
^ permalink raw reply [flat|nested] 14+ messages in thread
* Re: [PATCH net-next v5 7/7] net: skb: isolate skb data area allocations into a separate bucket
2026-10-02 23:11 ` [PATCH net-next v5 7/7] net: skb: isolate skb data area allocations into a separate bucket Kees Cook
@ 2026-10-03 23:13 ` netdev-bot+sashiko
0 siblings, 0 replies; 14+ messages in thread
From: netdev-bot+sashiko @ 2026-10-03 23:13 UTC (permalink / raw)
To: kees
Cc: vbabka, harry, akpm, hao.li, cl, rientjes, roman.gushchin,
linux-mm, pfalcato, kuniyu, linux-hardening, davem, hannes,
mhocko, shakeel.butt, muchun.song, edumazet, kuba, pabeni, horms,
kerneljasonxing, willemb, almasrymina, bjorn, jiayuan.chen,
linux-kernel, cgroups, netdev
Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.
Low priority / nit picks for your consideration:
- [Low] The block comment above kmalloc_reserve() in net/core/skbuff.c
(line 603) still says "kmalloc_reserve is a wrapper around…
--
net: skb: isolate skb data area allocations into a separate bucket
This patch moves skb data area allocations in kmalloc_reserve() and
kmalloc_pfmemalloc() off the general kmalloc caches and onto a dedicated
kmem_buckets set named "skb_data". The set is created in skb_init() with
both the normal and the accounted kmalloc types, so GFP_KERNEL_ACCOUNT
users such as AF_UNIX are isolated as well.
Link: https://github.com/google/security-research/blob/master/pocs/linux/kernelctf/CVE-2023-4207_lts_cos_mitigation_2/docs/exploit.md [0]
> diff --git a/net/core/skbuff.c b/net/core/skbuff.c
> index 4aea06d5167d1..a83689661fc28 100644
> --- a/net/core/skbuff.c
> +++ b/net/core/skbuff.c
[ ... ]
> @@ -593,7 +595,8 @@ static void *kmalloc_pfmemalloc(size_t obj_size, gfp_t flags, int node)
> if (!obj_size)
> return kmem_cache_alloc_node(net_hotdata.skb_small_head_cache,
> flags, node);
> - return kmalloc_node_track_caller(obj_size, flags, node);
> + return kmem_buckets_alloc_node_track_caller(skb_data_buckets, obj_size,
> + flags, node);
> }
>
> /*
[Severity: Low]
This isn't a bug, but the block comment that follows here, above
kmalloc_reserve() in net/core/skbuff.c, still says:
* kmalloc_reserve is a wrapper around kmalloc_node_track_caller that tells
* the caller if emergency pfmemalloc reserves are being used. ...
After this patch, neither kmalloc_reserve() nor kmalloc_pfmemalloc() calls
kmalloc_node_track_caller(). Both now use
kmem_buckets_alloc_node_track_caller(skb_data_buckets, ...).
Should the comment be updated to name the bucket allocator? As it stands,
it points readers at the general kmalloc caches instead of the skb_data
caches.
> @@ -634,7 +637,7 @@ static void *kmalloc_reserve(unsigned int *size, gfp_t flags, int node,
> * Try a regular allocation, when that fails and we're not entitled
> * to the reserves, fail.
> */
> - obj = kmalloc_node_track_caller(obj_size,
> + obj = kmem_buckets_alloc_node_track_caller(skb_data_buckets, obj_size,
> flags | __GFP_NOMEMALLOC | __GFP_NOWARN,
> node);
> if (likely(obj))
[ ... ]
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20261002231120.late.500-kees%40kernel.org
^ permalink raw reply [flat|nested] 14+ messages in thread
end of thread, other threads:[~2026-10-03 23:13 UTC | newest]
Thread overview: 14+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-02 23:11 [PATCH net-next v5 0/7] net: skb: isolate skb data area allocations into a separate bucket Kees Cook
2026-10-02 23:11 ` [PATCH net-next v5 1/7] mm/slab: Mark the kmem_buckets_create() context as a Context: section Kees Cook
2026-10-02 23:11 ` [PATCH net-next v5 2/7] mm/slab: Let kmem_buckets_create() take an alignment Kees Cook
2026-10-03 23:13 ` netdev-bot+sashiko
2026-10-02 23:11 ` [PATCH net-next v5 3/7] mm/slab: Add kmem_buckets_destroy() Kees Cook
2026-10-03 23:13 ` netdev-bot+sashiko
2026-10-02 23:11 ` [PATCH net-next v5 4/7] mm/slab: Add tests for the existing kmem_buckets behaviour Kees Cook
2026-10-03 23:13 ` netdev-bot+sashiko
2026-10-02 23:11 ` [PATCH net-next v5 5/7] mm/slab: Provide kmalloc type fallback for bucket allocations Kees Cook
2026-10-03 23:13 ` netdev-bot+sashiko
2026-10-02 23:11 ` [PATCH net-next v5 6/7] mm/slab: Let a bucket set handle __GFP_ACCOUNT Kees Cook
2026-10-03 23:13 ` netdev-bot+sashiko
2026-10-02 23:11 ` [PATCH net-next v5 7/7] net: skb: isolate skb data area allocations into a separate bucket Kees Cook
2026-10-03 23:13 ` netdev-bot+sashiko
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®