* [PATCH 0/1] rcu: drain kfree_rcu sheaves from the userspace barrier hook
@ 2026-09-10 10:11 Matthias Goergens
2026-09-10 10:11 ` [PATCH 1/1] " Matthias Goergens
2026-09-10 17:00 ` [PATCH v2 0/1] rcu: make userspace barrier hook drain kvfree_rcu work Matthias Goergens
0 siblings, 2 replies; 10+ messages in thread
From: Matthias Goergens @ 2026-09-10 10:11 UTC (permalink / raw)
To: paulmck, frederic, neeraj.upadhyay, joelagnelf, josh, boqun, urezki
Cc: rostedt, mathieu.desnoyers, jiangshanlai, qiang.zhang, corbet,
skhan, rdunlap, harry, surenb, vbabka, rcu, linux-doc,
linux-kernel
The rcutree.do_rcu_barrier test hook currently waits for ordinary RCU
callbacks, but it can return while kfree_rcu() still retains an object in a
partial per-CPU sheaf. This defeats the hook's purpose of preventing deferred
frees from one test spilling into the next.
The patch drains kfree_rcu sheaves and batches before retaining the hook's
explicit ordinary rcu_barrier(). Four counterbalanced fresh-VM pairs with the
full validation fixture reported 60 -> 60 active objects on the unpatched
kernel and 60 -> 59 on the patched kernel. An ordinary-callback regression
test passed on both kernels.
The primary reproducer below removes that separate regression machinery. One
fresh control/treatment pair with this exact 41-line source reproduced the
same 60 -> 60 versus 60 -> 59 split; both cells reached TEST SUCCESS with no
problem-class kernel records.
Save the source as rcu_barrier_sheaf_repro.c and create a Makefile containing:
obj-m := rcu_barrier_sheaf_repro.o
Build it with:
make -C /lib/modules/$(uname -r)/build M="$PWD" modules
Then, as root on a disposable test kernel:
insmod rcu_barrier_sheaf_repro.ko
awk '$1 == "rcu_barrier_sheaf_repro" { print $2 }' /proc/slabinfo
cat /sys/kernel/slab/rcu_barrier_sheaf_repro/sheaf_capacity
echo 1 > /sys/module/rcutree/parameters/do_rcu_barrier
awk '$1 == "rcu_barrier_sheaf_repro" { print $2 }' /proc/slabinfo
rmmod rcu_barrier_sheaf_repro
The first and second slabinfo readings are 60 and 60 without the patch, and
60 and 59 with it. kmem_cache_destroy() performs per-cache deferred-free
cleanup when the module is removed, after the measurement.
// SPDX-License-Identifier: GPL-2.0
#include <linux/init.h>
#include <linux/module.h>
#include <linux/rcupdate.h>
#include <linux/slab.h>
struct repro_object {
struct rcu_head rcu;
unsigned long payload;
};
static struct kmem_cache *repro_cache;
static int __init rcu_barrier_sheaf_repro_init(void)
{
struct repro_object *object;
repro_cache = kmem_cache_create("rcu_barrier_sheaf_repro",
sizeof(*object), 0, SLAB_NO_MERGE, NULL);
if (!repro_cache)
return -ENOMEM;
object = kmem_cache_alloc(repro_cache, GFP_KERNEL);
if (!object) {
kmem_cache_destroy(repro_cache);
return -ENOMEM;
}
kfree_rcu(object, rcu);
return 0;
}
static void __exit rcu_barrier_sheaf_repro_exit(void)
{
kmem_cache_destroy(repro_cache);
}
module_init(rcu_barrier_sheaf_repro_init);
module_exit(rcu_barrier_sheaf_repro_exit);
MODULE_LICENSE("GPL");
MODULE_DESCRIPTION("Reproduce incomplete rcutree.do_rcu_barrier drains");
Matthias Goergens (1):
rcu: drain kfree_rcu sheaves from the userspace barrier hook
.../admin-guide/kernel-parameters.txt | 7 ++---
kernel/rcu/tree.c | 27 ++++++++++++-------
2 files changed, 21 insertions(+), 13 deletions(-)
base-commit: 50d05c7c76c96b90462f24debacca971d2e86713
--
2.55.0
^ permalink raw reply [flat|nested] 10+ messages in thread* [PATCH 1/1] rcu: drain kfree_rcu sheaves from the userspace barrier hook 2026-09-10 10:11 [PATCH 0/1] rcu: drain kfree_rcu sheaves from the userspace barrier hook Matthias Goergens @ 2026-09-10 10:11 ` Matthias Goergens 2026-09-10 12:03 ` Harry Yoo 2026-09-10 17:00 ` [PATCH v2 0/1] rcu: make userspace barrier hook drain kvfree_rcu work Matthias Goergens 1 sibling, 1 reply; 10+ messages in thread From: Matthias Goergens @ 2026-09-10 10:11 UTC (permalink / raw) To: paulmck, frederic, neeraj.upadhyay, joelagnelf, josh, boqun, urezki Cc: rostedt, mathieu.desnoyers, jiangshanlai, qiang.zhang, corbet, skhan, rdunlap, harry, surenb, vbabka, rcu, linux-doc, linux-kernel The rcutree.do_rcu_barrier test hook is intended to prevent deferred RCU callbacks from one stress test spilling into the next. Since kfree_rcu() sheaves were added, an object can remain deferred without appearing on an ordinary RCU callback list. rcu_barrier() therefore no longer fulfils the hook's stated purpose by itself. Drain kfree_rcu sheaves and kvfree_rcu batches before completing the ordinary RCU barrier. Keep the explicit rcu_barrier() because the hook's original ordinary-callback contract should not depend on the current, undocumented fact that kvfree_rcu_barrier() includes one internally. Keep the existing throttling because this remains a deliberately expensive test-only action. Do not coalesce requests based on the ordinary rcu_barrier() sequence: an unrelated ordinary barrier does not prove that sheaves were drained. A reproducer creates a private SLAB_NO_MERGE cache whose first allocation populates a 60-object slab. It queues that object with kfree_rcu(), invokes the hook, and reads the active-object count from /proc/slabinfo. In four fresh VM pairs, the parent retained the object (60 to 60). The patched hook drained it (60 to 59). Fixes: ec66e0d59952 ("slab: add sheaf support for batching kfree_rcu() operations") Signed-off-by: Matthias Goergens <matthias.goergens@gmail.com> --- .../admin-guide/kernel-parameters.txt | 7 ++--- kernel/rcu/tree.c | 27 ++++++++++++------- 2 files changed, 21 insertions(+), 13 deletions(-) diff --git a/Documentation/admin-guide/kernel-parameters.txt b/Documentation/admin-guide/kernel-parameters.txt index 68647ff4bdd2..244a53166249 100644 --- a/Documentation/admin-guide/kernel-parameters.txt +++ b/Documentation/admin-guide/kernel-parameters.txt @@ -5699,9 +5699,10 @@ Kernel parameters there is an ongoing too-long CSD-lock wait. rcutree.do_rcu_barrier= [KNL] - Request a call to rcu_barrier(). This is - throttled so that userspace tests can safely - hammer on the sysfs variable if they so choose. + Request that deferred kfree_rcu() objects and + ordinary call_rcu() callbacks be drained. This is + throttled so that userspace tests can safely hammer + on the sysfs variable if they so choose. If triggered before the RCU grace-period machinery is fully active, this will error out with EAGAIN. diff --git a/kernel/rcu/tree.c b/kernel/rcu/tree.c index 96848fc1f02b..014e28ec3bd3 100644 --- a/kernel/rcu/tree.c +++ b/kernel/rcu/tree.c @@ -3989,12 +3989,12 @@ EXPORT_SYMBOL_GPL(rcu_barrier); static unsigned long rcu_barrier_last_throttle; /** - * rcu_barrier_throttled - Do rcu_barrier(), but limit to one per second + * rcu_barrier_throttled - Drain deferred RCU frees, but rate-limit starts * - * This can be thought of as guard rails around rcu_barrier() that - * permits unrestricted userspace use, at least assuming the hardware's - * try_cmpxchg() is robust. There will be at most one call per second to - * rcu_barrier() system-wide from use of this function, which means that + * This can be thought of as guard rails around the deferred-free barriers + * that permit unrestricted userspace use, at least assuming the hardware's + * try_cmpxchg() is robust. There will be at most one drain operation started + * per sixteenth of a second from use of this function, which means that * callers might needlessly wait a second or three. * * This is intended for use by test suites to avoid OOM by flushing RCU @@ -4011,18 +4011,25 @@ static void rcu_barrier_throttled(void) { unsigned long j = jiffies; unsigned long old = READ_ONCE(rcu_barrier_last_throttle); - unsigned long s = rcu_seq_snap(&rcu_state.barrier_sequence); while (time_in_range(j, old, old + HZ / 16) || !try_cmpxchg(&rcu_barrier_last_throttle, &old, j)) { schedule_timeout_idle(HZ / 16); - if (rcu_seq_done(&rcu_state.barrier_sequence, s)) { - smp_mb(); /* caller's subsequent code after above check. */ - return; - } j = jiffies; old = READ_ONCE(rcu_barrier_last_throttle); } + /* + * kfree_rcu() can retain objects outside the ordinary callback lists in + * per-CPU SLUB sheaves and kvfree_rcu batches. Test suites use this hook + * to prevent deferred frees from spilling into the following test, so + * drain those queues as well as ordinary call_rcu() callbacks. + * + * kvfree_rcu_barrier() currently includes an ordinary barrier, but that + * is not part of its documented API. Keep the explicit rcu_barrier() so + * this hook's original contract does not depend on slab implementation + * details. + */ + kvfree_rcu_barrier(); rcu_barrier(); } -- 2.55.0 ^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [PATCH 1/1] rcu: drain kfree_rcu sheaves from the userspace barrier hook 2026-09-10 10:11 ` [PATCH 1/1] " Matthias Goergens @ 2026-09-10 12:03 ` Harry Yoo 2026-09-10 13:46 ` Matthias Goergens 0 siblings, 1 reply; 10+ messages in thread From: Harry Yoo @ 2026-09-10 12:03 UTC (permalink / raw) To: Matthias Goergens Cc: paulmck, frederic, neeraj.upadhyay, joelagnelf, josh, boqun, urezki, rostedt, mathieu.desnoyers, jiangshanlai, qiang.zhang, corbet, skhan, rdunlap, surenb, vbabka, rcu, linux-doc, linux-kernel On Thu, Sep 10, 2026 at 06:11:12PM +0800, Matthias Goergens wrote: > The rcutree.do_rcu_barrier test hook is intended to prevent deferred RCU > callbacks from one stress test spilling into the next. Since kfree_rcu() > sheaves were added, an object can remain deferred without appearing on an > ordinary RCU callback list. rcu_barrier() therefore no longer fulfils the > hook's stated purpose by itself. > > Drain kfree_rcu sheaves and kvfree_rcu batches before completing the > ordinary RCU barrier. I'm convinced that the behavior of "do rcu_barrier()" knob to imply a kvfree_rcu_barrier() is the right fix. And also I wonder what's the user-facing problem you are trying to fix. You can't pile up unbounded amount of objects via kvfree_rcu() to cause an OOM during the userspace tests? How did you discover the problem? > Keep the explicit rcu_barrier() because the hook's > original ordinary-callback contract should not depend on the current, > undocumented fact that kvfree_rcu_barrier() includes one internally. > > Keep the existing throttling because this remains a deliberately expensive > test-only action. Do not coalesce requests based on the ordinary > rcu_barrier() sequence: an unrelated ordinary barrier does not prove that > sheaves were drained. > > A reproducer creates a private SLAB_NO_MERGE cache whose first allocation > populates a 60-object slab. It queues that object with kfree_rcu(), invokes > the hook, and reads the active-object count from /proc/slabinfo. In four > fresh VM pairs, the parent retained the object (60 to 60). The patched hook > drained it (60 to 59). > > Fixes: ec66e0d59952 ("slab: add sheaf support for batching kfree_rcu() operations") This Fixes: commit is incorrect because the problem was introduced by kvfree_rcu(), not sheaves. > Signed-off-by: Matthias Goergens <matthias.goergens@gmail.com> > --- > .../admin-guide/kernel-parameters.txt | 7 ++--- > kernel/rcu/tree.c | 27 ++++++++++++------- > 2 files changed, 21 insertions(+), 13 deletions(-) > -- Cheers, Harry / Hyeonggon ^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [PATCH 1/1] rcu: drain kfree_rcu sheaves from the userspace barrier hook 2026-09-10 12:03 ` Harry Yoo @ 2026-09-10 13:46 ` Matthias Goergens 2026-09-10 15:37 ` Harry Yoo 0 siblings, 1 reply; 10+ messages in thread From: Matthias Goergens @ 2026-09-10 13:46 UTC (permalink / raw) To: Harry Yoo Cc: Matthias Goergens, paulmck, frederic, neeraj.upadhyay, joelagnelf, josh, boqun, urezki, rostedt, mathieu.desnoyers, jiangshanlai, qiang.zhang, corbet, skhan, rdunlap, surenb, vbabka, rcu, linux-doc, linux-kernel Thanks. I ran into this while testing bcachefs performance changes. The bcachefs ktest end checks write `do_rcu_barrier` before reading `/proc/allocinfo`, with the expectation that allocations still reported afterwards are leaks. Small objects released with `kfree_rcu()` remained visible after repeated writes to the hook and 20 seconds of waiting, so otherwise clean tests failed their leak check. Strictly, that ktest is assuming a stronger contract than the hook currently documents: `do_rcu_barrier` promises an ordinary `rcu_barrier()`, not a complete drain of objects still held in `kfree_rcu()` batching. I nevertheless think the stronger behaviour is useful for this test-only quiescence hook, because it lets allocation-leak checks reliably separate deferred frees from genuine leaks. I followed those allocations across repeated filesystem lifecycles. Their number eventually fell when an RCU sheaf filled, so I have no evidence that this path grows without bound or causes OOM. The problem I observed is limited to test isolation: the hook can leave deferred frees behind and make them look like leaks. With the proposed change, the same unmodified bcachefs workload passed the leak check. After tracing that behaviour, I wrote the private-cache module in the cover letter to reproduce it without bcachefs. I can publish the original bcachefs workload and results if anyone is interested. I also found that this exact follow-on was discussed when `kvfree_rcu_barrier()` was added in 2024: Paul proposed calling it from `rcu_barrier_throttled()` for clean userspace benchmark baselines, and Uladzislau agreed that adding it and documenting both operations was safest: https://lore.kernel.org/all/20240820155935.1167988-1-urezki@gmail.com/ You are right about the `Fixes:` tag. `kvfree_rcu()` batching predates `do_rcu_barrier`, and the existing interface does what it documents, so the later sheaf commit is not the right introduction point. I will omit the `Fixes:` tag in v2 and present this as a strengthening of the test interface. -- Matthias ^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [PATCH 1/1] rcu: drain kfree_rcu sheaves from the userspace barrier hook 2026-09-10 13:46 ` Matthias Goergens @ 2026-09-10 15:37 ` Harry Yoo 0 siblings, 0 replies; 10+ messages in thread From: Harry Yoo @ 2026-09-10 15:37 UTC (permalink / raw) To: Matthias Goergens Cc: paulmck, frederic, neeraj.upadhyay, joelagnelf, josh, boqun, urezki, rostedt, mathieu.desnoyers, jiangshanlai, qiang.zhang, corbet, skhan, rdunlap, surenb, vbabka, rcu, linux-doc, linux-kernel Hi Matthias, thanks for reply. Would you please wraparound the text when writing a reply? Each line becomes way too long on text based editors :-) Anyway, moving on to the topic... On Thu, Sep 10, 2026 at 09:46:40PM +0800, Matthias Goergens wrote: > Thanks. I ran into this while testing bcachefs performance changes. > The bcachefs ktest end checks write `do_rcu_barrier` before reading > `/proc/allocinfo`, with the expectation that allocations still > reported afterwards are leaks. Small objects released with > `kfree_rcu()` remained visible after repeated writes to the hook and > 20 seconds of waiting, so otherwise clean tests failed their leak > check. Thanks. Some background like this would be nice to be covered in the cover letter or commit message :-) > Strictly, that ktest is assuming a stronger contract than the hook > currently documents: `do_rcu_barrier` promises an ordinary > `rcu_barrier()`, not a complete drain of objects still held in > `kfree_rcu()` batching. I nevertheless think the stronger behaviour > is useful for this test-only quiescence hook, because it lets > allocation-leak checks reliably separate deferred frees from genuine > leaks. Ack. > I followed those allocations across repeated filesystem lifecycles. > Their number eventually fell when an RCU sheaf filled, so I have no > evidence that this path grows without bound or causes OOM. Ah, the reason I mentioned unbounded amount of objects was because the commit 16128b1f8c823438dc that introduced the knob explains what can go wrong (e.g OOMs during the test) without the rcutree.do_rcu_barrier. > The problem I observed is limited to test isolation: the hook can > leave deferred frees behind and make them look like leaks. With the > proposed change, the same unmodified bcachefs workload passed the > leak check. It'd be more convincing if this part is included in the changelog ;-) > I also found that this exact follow-on was discussed when > `kvfree_rcu_barrier()` was added in 2024: Paul proposed calling it > from `rcu_barrier_throttled()` for clean userspace benchmark > baselines, and Uladzislau agreed that adding it and documenting both > operations was safest: > https://lore.kernel.org/all/20240820155935.1167988-1-urezki@gmail.com/ Didn't notice that, and makes sense given the purpose of the knob. > I will omit the `Fixes:` tag in v2 and present this as a strengthening > of the test interface. Makes sense to me. -- Cheers, Harry / Hyeonggon ^ permalink raw reply [flat|nested] 10+ messages in thread
* [PATCH v2 0/1] rcu: make userspace barrier hook drain kvfree_rcu work 2026-09-10 10:11 [PATCH 0/1] rcu: drain kfree_rcu sheaves from the userspace barrier hook Matthias Goergens 2026-09-10 10:11 ` [PATCH 1/1] " Matthias Goergens @ 2026-09-10 17:00 ` Matthias Goergens 2026-09-10 17:00 ` [PATCH v2 1/1] " Matthias Goergens 1 sibling, 1 reply; 10+ messages in thread From: Matthias Goergens @ 2026-09-10 17:00 UTC (permalink / raw) To: paulmck, urezki, harry Cc: frederic, neeraj.upadhyay, joelagnelf, josh, boqun, rostedt, mathieu.desnoyers, jiangshanlai, qiang.zhang, corbet, skhan, rdunlap, surenb, vbabka, rcu, linux-doc, linux-kernel The rcutree.do_rcu_barrier hook currently waits for ordinary RCU callbacks, but objects may still be retained in kfree_rcu() batching or a partial per-CPU SLUB sheaf. This is consistent with the hook's documented rcu_barrier() operation, but incomplete for its intended use as a boundary between userspace tests. The immediate trigger was a false allocation-leak failure in the bcachefs ktest suite while testing performance changes. Its end check writes the hook before reading /proc/allocinfo, assuming a complete deferred-free drain. Small objects remained visible after repeated hook writes and 20 seconds of waiting, so otherwise clean tests failed their leak check. Changing the hook to drain kvfree_rcu() work let the same unmodified bcachefs workload pass its allocation check. All eight checkpoints in one VM, after 50 through 400 option changes, reported zero retained reconcile_scan objects. The retained population on the original kernel eventually fell as a sheaf filled; there is no evidence here of unbounded growth or OOM. Calling kvfree_rcu_barrier() from rcu_barrier_throttled() was proposed and agreed during review of the former API in 2024, specifically to restore a clean baseline between userspace benchmark runs: https://lore.kernel.org/all/20240820155935.1167988-1-urezki@gmail.com/ This patch implements that follow-up and documents the expanded hook. It also removes the old ordinary-barrier completion shortcut: an unrelated rcu_barrier() does not establish that kvfree_rcu() work was drained. Four counterbalanced fresh-VM pairs with the full private-cache fixture reported 60 to 60 active objects on the unpatched kernel and 60 to 59 on the patched kernel. A separate ordinary-callback regression test passed on both kernels. The simplified reproducer below removes that separate regression machinery. One additional fresh control/treatment pair with this exact 41-line source confirmed the same 60 to 60 versus 60 to 59 split. These counts reflect the slab layout in the tested configuration. Save the source as rcu_barrier_sheaf_repro.c and create a Makefile containing: obj-m := rcu_barrier_sheaf_repro.o Build it with: make -C /lib/modules/$(uname -r)/build M="$PWD" modules Then, as root on a disposable test kernel: insmod rcu_barrier_sheaf_repro.ko awk '$1 == "rcu_barrier_sheaf_repro" { print $2 }' /proc/slabinfo cat /sys/kernel/slab/rcu_barrier_sheaf_repro/sheaf_capacity echo 1 > /sys/module/rcutree/parameters/do_rcu_barrier awk '$1 == "rcu_barrier_sheaf_repro" { print $2 }' /proc/slabinfo rmmod rcu_barrier_sheaf_repro The first and second slabinfo readings are 60 and 60 without the patch, and 60 and 59 with it. kmem_cache_destroy() performs per-cache deferred-free cleanup when the module is removed, after the measurement. // SPDX-License-Identifier: GPL-2.0 #include <linux/init.h> #include <linux/module.h> #include <linux/rcupdate.h> #include <linux/slab.h> struct repro_object { struct rcu_head rcu; unsigned long payload; }; static struct kmem_cache *repro_cache; static int __init rcu_barrier_sheaf_repro_init(void) { struct repro_object *object; repro_cache = kmem_cache_create("rcu_barrier_sheaf_repro", sizeof(*object), 0, SLAB_NO_MERGE, NULL); if (!repro_cache) return -ENOMEM; object = kmem_cache_alloc(repro_cache, GFP_KERNEL); if (!object) { kmem_cache_destroy(repro_cache); return -ENOMEM; } kfree_rcu(object, rcu); return 0; } static void __exit rcu_barrier_sheaf_repro_exit(void) { kmem_cache_destroy(repro_cache); } module_init(rcu_barrier_sheaf_repro_init); module_exit(rcu_barrier_sheaf_repro_exit); MODULE_LICENSE("GPL"); MODULE_DESCRIPTION("Reproduce incomplete rcutree.do_rcu_barrier drains"); --- Changes since v1: - Add the motivating bcachefs failure and the successful unmodified workload result to both the cover letter and commit message. - Drop the incorrect sheaf Fixes: tag and regression framing; describe this as a strengthening of the existing test interface. - Credit the agreed 2024 proposal for this extension. - Broaden the subject and changelog from sheaves to kvfree_rcu work. - Hard-wrap the prose for text-based mail readers. The code diff is unchanged from v1. The results above are the existing validation results; no new kernel tests were run for this prose revision. v1: https://lore.kernel.org/all/20260910101112.1648978-1-matthias.goergens@gmail.com/ Matthias Goergens (1): rcu: make userspace barrier hook drain kvfree_rcu work .../admin-guide/kernel-parameters.txt | 7 ++--- kernel/rcu/tree.c | 27 ++++++++++++------- 2 files changed, 21 insertions(+), 13 deletions(-) base-commit: 50d05c7c76c96b90462f24debacca971d2e86713 -- 2.55.0 ^ permalink raw reply [flat|nested] 10+ messages in thread
* [PATCH v2 1/1] rcu: make userspace barrier hook drain kvfree_rcu work 2026-09-10 17:00 ` [PATCH v2 0/1] rcu: make userspace barrier hook drain kvfree_rcu work Matthias Goergens @ 2026-09-10 17:00 ` Matthias Goergens 2026-09-10 17:51 ` Paul E. McKenney 0 siblings, 1 reply; 10+ messages in thread From: Matthias Goergens @ 2026-09-10 17:00 UTC (permalink / raw) To: paulmck, urezki, harry Cc: frederic, neeraj.upadhyay, joelagnelf, josh, boqun, rostedt, mathieu.desnoyers, jiangshanlai, qiang.zhang, corbet, skhan, rdunlap, surenb, vbabka, rcu, linux-doc, linux-kernel The bcachefs ktest allocation-leak check writes rcutree.do_rcu_barrier before reading /proc/allocinfo. While testing bcachefs performance changes, small objects released with kfree_rcu() remained visible after repeated writes to the hook and 20 seconds of waiting, causing otherwise clean tests to fail their leak check. The test assumes a stronger contract than the hook currently documents: rcu_barrier() waits for ordinary callbacks, but does not flush objects still held in kfree_rcu() batching or per-CPU SLUB sheaves. The retained population eventually fell as a sheaf filled; there is no evidence here of unbounded growth or OOM. Changing the hook to drain kvfree_rcu() work let the same unmodified bcachefs workload pass its allocation check. All eight checkpoints in one VM, after 50 through 400 option changes, reported zero retained reconcile_scan objects. This motivated the separate private-cache reproducer used to isolate the incomplete drain from bcachefs. Calling kvfree_rcu_barrier() from rcu_barrier_throttled() was proposed when the former API was added in 2024, to restore a clean baseline between userspace benchmark runs. The discussion concluded that keeping the existing hook name, adding the second operation and documenting both was the safest compatibility choice, but the follow-up was not added. Add that drain and document the stronger test interface. Keep the explicit rcu_barrier() so the hook's ordinary-callback contract does not depend on kvfree_rcu_barrier() reaching an ordinary barrier internally. Do not reuse the ordinary rcu_barrier() sequence as an early-completion check while throttling: an unrelated ordinary barrier does not establish that kvfree_rcu() work was drained. Retain the existing start-rate limit. Four fresh VM pairs with the full private-cache fixture retained the queued object without the patch (60 to 60 active objects) and drained it with the patch (60 to 59). A separate ordinary-callback regression test passed on both kernels. Link: https://lore.kernel.org/all/20240820155935.1167988-1-urezki@gmail.com/ Signed-off-by: Matthias Goergens <matthias.goergens@gmail.com> --- .../admin-guide/kernel-parameters.txt | 7 ++--- kernel/rcu/tree.c | 27 ++++++++++++------- 2 files changed, 21 insertions(+), 13 deletions(-) diff --git a/Documentation/admin-guide/kernel-parameters.txt b/Documentation/admin-guide/kernel-parameters.txt index 68647ff4bdd2..244a53166249 100644 --- a/Documentation/admin-guide/kernel-parameters.txt +++ b/Documentation/admin-guide/kernel-parameters.txt @@ -5699,9 +5699,10 @@ Kernel parameters there is an ongoing too-long CSD-lock wait. rcutree.do_rcu_barrier= [KNL] - Request a call to rcu_barrier(). This is - throttled so that userspace tests can safely - hammer on the sysfs variable if they so choose. + Request that deferred kfree_rcu() objects and + ordinary call_rcu() callbacks be drained. This is + throttled so that userspace tests can safely hammer + on the sysfs variable if they so choose. If triggered before the RCU grace-period machinery is fully active, this will error out with EAGAIN. diff --git a/kernel/rcu/tree.c b/kernel/rcu/tree.c index 96848fc1f02b..014e28ec3bd3 100644 --- a/kernel/rcu/tree.c +++ b/kernel/rcu/tree.c @@ -3989,12 +3989,12 @@ EXPORT_SYMBOL_GPL(rcu_barrier); static unsigned long rcu_barrier_last_throttle; /** - * rcu_barrier_throttled - Do rcu_barrier(), but limit to one per second + * rcu_barrier_throttled - Drain deferred RCU frees, but rate-limit starts * - * This can be thought of as guard rails around rcu_barrier() that - * permits unrestricted userspace use, at least assuming the hardware's - * try_cmpxchg() is robust. There will be at most one call per second to - * rcu_barrier() system-wide from use of this function, which means that + * This can be thought of as guard rails around the deferred-free barriers + * that permit unrestricted userspace use, at least assuming the hardware's + * try_cmpxchg() is robust. There will be at most one drain operation started + * per sixteenth of a second from use of this function, which means that * callers might needlessly wait a second or three. * * This is intended for use by test suites to avoid OOM by flushing RCU @@ -4011,18 +4011,25 @@ static void rcu_barrier_throttled(void) { unsigned long j = jiffies; unsigned long old = READ_ONCE(rcu_barrier_last_throttle); - unsigned long s = rcu_seq_snap(&rcu_state.barrier_sequence); while (time_in_range(j, old, old + HZ / 16) || !try_cmpxchg(&rcu_barrier_last_throttle, &old, j)) { schedule_timeout_idle(HZ / 16); - if (rcu_seq_done(&rcu_state.barrier_sequence, s)) { - smp_mb(); /* caller's subsequent code after above check. */ - return; - } j = jiffies; old = READ_ONCE(rcu_barrier_last_throttle); } + /* + * kfree_rcu() can retain objects outside the ordinary callback lists in + * per-CPU SLUB sheaves and kvfree_rcu batches. Test suites use this hook + * to prevent deferred frees from spilling into the following test, so + * drain those queues as well as ordinary call_rcu() callbacks. + * + * kvfree_rcu_barrier() currently includes an ordinary barrier, but that + * is not part of its documented API. Keep the explicit rcu_barrier() so + * this hook's original contract does not depend on slab implementation + * details. + */ + kvfree_rcu_barrier(); rcu_barrier(); } -- 2.55.0 ^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [PATCH v2 1/1] rcu: make userspace barrier hook drain kvfree_rcu work 2026-09-10 17:00 ` [PATCH v2 1/1] " Matthias Goergens @ 2026-09-10 17:51 ` Paul E. McKenney 2026-09-11 3:40 ` [PATCH v3 0/1] " Matthias Goergens 0 siblings, 1 reply; 10+ messages in thread From: Paul E. McKenney @ 2026-09-10 17:51 UTC (permalink / raw) To: Matthias Goergens Cc: urezki, harry, frederic, neeraj.upadhyay, joelagnelf, josh, boqun, rostedt, mathieu.desnoyers, jiangshanlai, qiang.zhang, corbet, skhan, rdunlap, surenb, vbabka, rcu, linux-doc, linux-kernel On Fri, Sep 11, 2026 at 01:00:40AM +0800, Matthias Goergens wrote: > The bcachefs ktest allocation-leak check writes rcutree.do_rcu_barrier > before reading /proc/allocinfo. While testing bcachefs performance > changes, small objects released with kfree_rcu() remained visible after > repeated writes to the hook and 20 seconds of waiting, causing otherwise > clean tests to fail their leak check. > > The test assumes a stronger contract than the hook currently documents: > rcu_barrier() waits for ordinary callbacks, but does not flush objects > still held in kfree_rcu() batching or per-CPU SLUB sheaves. The retained > population eventually fell as a sheaf filled; there is no evidence here > of unbounded growth or OOM. > > Changing the hook to drain kvfree_rcu() work let the same unmodified > bcachefs workload pass its allocation check. All eight checkpoints in > one VM, after 50 through 400 option changes, reported zero retained > reconcile_scan objects. This motivated the separate private-cache > reproducer used to isolate the incomplete drain from bcachefs. > > Calling kvfree_rcu_barrier() from rcu_barrier_throttled() was proposed > when the former API was added in 2024, to restore a clean baseline > between userspace benchmark runs. The discussion concluded that keeping > the existing hook name, adding the second operation and documenting both > was the safest compatibility choice, but the follow-up was not added. > > Add that drain and document the stronger test interface. Keep the > explicit rcu_barrier() so the hook's ordinary-callback contract does not > depend on kvfree_rcu_barrier() reaching an ordinary barrier internally. > > Do not reuse the ordinary rcu_barrier() sequence as an early-completion > check while throttling: an unrelated ordinary barrier does not establish > that kvfree_rcu() work was drained. Retain the existing start-rate limit. > > Four fresh VM pairs with the full private-cache fixture retained the > queued object without the patch (60 to 60 active objects) and drained it > with the patch (60 to 59). A separate ordinary-callback regression test > passed on both kernels. > > Link: https://lore.kernel.org/all/20240820155935.1167988-1-urezki@gmail.com/ > Signed-off-by: Matthias Goergens <matthias.goergens@gmail.com> > --- > .../admin-guide/kernel-parameters.txt | 7 ++--- > kernel/rcu/tree.c | 27 ++++++++++++------- > 2 files changed, 21 insertions(+), 13 deletions(-) > > diff --git a/Documentation/admin-guide/kernel-parameters.txt b/Documentation/admin-guide/kernel-parameters.txt > index 68647ff4bdd2..244a53166249 100644 > --- a/Documentation/admin-guide/kernel-parameters.txt > +++ b/Documentation/admin-guide/kernel-parameters.txt > @@ -5699,9 +5699,10 @@ Kernel parameters > there is an ongoing too-long CSD-lock wait. > > rcutree.do_rcu_barrier= [KNL] > - Request a call to rcu_barrier(). This is > - throttled so that userspace tests can safely > - hammer on the sysfs variable if they so choose. > + Request that deferred kfree_rcu() objects and > + ordinary call_rcu() callbacks be drained. This is > + throttled so that userspace tests can safely hammer > + on the sysfs variable if they so choose. > If triggered before the RCU grace-period machinery > is fully active, this will error out with EAGAIN. > > diff --git a/kernel/rcu/tree.c b/kernel/rcu/tree.c > index 96848fc1f02b..014e28ec3bd3 100644 > --- a/kernel/rcu/tree.c > +++ b/kernel/rcu/tree.c > @@ -3989,12 +3989,12 @@ EXPORT_SYMBOL_GPL(rcu_barrier); > static unsigned long rcu_barrier_last_throttle; > > /** > - * rcu_barrier_throttled - Do rcu_barrier(), but limit to one per second > + * rcu_barrier_throttled - Drain deferred RCU frees, but rate-limit starts > * > - * This can be thought of as guard rails around rcu_barrier() that > - * permits unrestricted userspace use, at least assuming the hardware's > - * try_cmpxchg() is robust. There will be at most one call per second to > - * rcu_barrier() system-wide from use of this function, which means that > + * This can be thought of as guard rails around the deferred-free barriers > + * that permit unrestricted userspace use, at least assuming the hardware's > + * try_cmpxchg() is robust. There will be at most one drain operation started > + * per sixteenth of a second from use of this function, which means that > * callers might needlessly wait a second or three. > * > * This is intended for use by test suites to avoid OOM by flushing RCU > @@ -4011,18 +4011,25 @@ static void rcu_barrier_throttled(void) > { > unsigned long j = jiffies; > unsigned long old = READ_ONCE(rcu_barrier_last_throttle); > - unsigned long s = rcu_seq_snap(&rcu_state.barrier_sequence); > > while (time_in_range(j, old, old + HZ / 16) || > !try_cmpxchg(&rcu_barrier_last_throttle, &old, j)) { > schedule_timeout_idle(HZ / 16); > - if (rcu_seq_done(&rcu_state.barrier_sequence, s)) { > - smp_mb(); /* caller's subsequent code after above check. */ > - return; Don't we still want to skip the rcu_barrier() in this case? Or am I missing something subtle here? Thanx, Paul > - } > j = jiffies; > old = READ_ONCE(rcu_barrier_last_throttle); > } > + /* > + * kfree_rcu() can retain objects outside the ordinary callback lists in > + * per-CPU SLUB sheaves and kvfree_rcu batches. Test suites use this hook > + * to prevent deferred frees from spilling into the following test, so > + * drain those queues as well as ordinary call_rcu() callbacks. > + * > + * kvfree_rcu_barrier() currently includes an ordinary barrier, but that > + * is not part of its documented API. Keep the explicit rcu_barrier() so > + * this hook's original contract does not depend on slab implementation > + * details. > + */ > + kvfree_rcu_barrier(); > rcu_barrier(); > } > > -- > 2.55.0 > ^ permalink raw reply [flat|nested] 10+ messages in thread
* [PATCH v3 0/1] rcu: make userspace barrier hook drain kvfree_rcu work 2026-09-10 17:51 ` Paul E. McKenney @ 2026-09-11 3:40 ` Matthias Goergens 2026-09-11 3:40 ` [PATCH v3 1/1] " Matthias Goergens 0 siblings, 1 reply; 10+ messages in thread From: Matthias Goergens @ 2026-09-11 3:40 UTC (permalink / raw) To: paulmck, urezki, harry Cc: frederic, neeraj.upadhyay, joelagnelf, josh, boqun, rostedt, mathieu.desnoyers, jiangshanlai, qiang.zhang, corbet, skhan, rdunlap, surenb, vbabka, rcu, linux-doc, linux-kernel Extend rcutree.do_rcu_barrier to drain kfree_rcu() batches and per-CPU SLUB sheaves as well as ordinary callbacks. This strengthens the documented test interface, implementing the follow-up proposed when kvfree_rcu_barrier() was added in 2024: https://lore.kernel.org/all/20240820155935.1167988-1-urezki@gmail.com/ While testing bcachefs performance changes, ktest's /proc/allocinfo check falsely reported leaks despite repeated hook writes and a 20-second wait. An earlier prototype let the same unmodified workload pass: eight checkpoints in one VM, after 50 through 400 option changes, found zero retained reconcile_scan objects. The original kernel's retained count eventually fell as a sheaf filled; no unbounded growth or OOM was observed. Following Paul's v2 review, retain the entry sequence snapshot and check it after the unconditional kvfree_rcu() drain. If complete, retain smp_mb() and skip the final ordinary barrier; otherwise, call rcu_barrier(). An unrelated ordinary barrier cannot justify skipping the deferred-free drain. The start-rate throttle remains unconditional. The documentation now scopes completion to work queued before the request, without preventing new concurrent work. This can avoid an extra barrier operation: v2's trailing rcu_barrier() takes a fresh snapshot and need not reuse the internal barrier. No elapsed-time improvement has been measured. Exact-v3 builds and interface smoke checks passed in four fresh two-vCPU TREE_RCU VMs: two with kvfree batching and two with SLUB_TINY (unbatched). Each completed three true requests, remained idle after false, and rejected invalid input with EINVAL. Full kernel logs showed no WARN, oops, panic or RCU-stall diagnostics. The v1/v2 private-cache completion and ordinary-callback regression tests have not been rerun on v3. These smoke tests neither prove concurrency/weak-memory correctness nor force the guarded fallback. v2: https://lore.kernel.org/all/20260910170040.344864-1-matthias.goergens@gmail.com/ Matthias Goergens (1): rcu: make userspace barrier hook drain kvfree_rcu work .../admin-guide/kernel-parameters.txt | 9 ++++-- kernel/rcu/tree.c | 30 ++++++++++++------- 2 files changed, 26 insertions(+), 13 deletions(-) -- 2.55.0 ^ permalink raw reply [flat|nested] 10+ messages in thread
* [PATCH v3 1/1] rcu: make userspace barrier hook drain kvfree_rcu work 2026-09-11 3:40 ` [PATCH v3 0/1] " Matthias Goergens @ 2026-09-11 3:40 ` Matthias Goergens 0 siblings, 0 replies; 10+ messages in thread From: Matthias Goergens @ 2026-09-11 3:40 UTC (permalink / raw) To: paulmck, urezki, harry Cc: frederic, neeraj.upadhyay, joelagnelf, josh, boqun, rostedt, mathieu.desnoyers, jiangshanlai, qiang.zhang, corbet, skhan, rdunlap, surenb, vbabka, rcu, linux-doc, linux-kernel The bcachefs ktest allocation-leak check writes rcutree.do_rcu_barrier before reading /proc/allocinfo. While testing bcachefs performance changes, small objects released with kfree_rcu() remained visible after repeated writes to the hook and 20 seconds of waiting, causing otherwise clean tests to fail their leak check. The test assumes a stronger contract than the hook currently documents: rcu_barrier() waits for ordinary callbacks, but does not flush objects still held in kfree_rcu() batching or per-CPU SLUB sheaves. The retained population eventually fell as a sheaf filled; there is no evidence here of unbounded growth or OOM. Changing the hook to drain kvfree_rcu() work let the same unmodified bcachefs workload pass its allocation check. All eight checkpoints in one VM, after 50 through 400 option changes, reported zero retained reconcile_scan objects. This motivated the separate private-cache test used to isolate the incomplete drain from bcachefs. Calling kvfree_rcu_barrier() from rcu_barrier_throttled() was proposed when kvfree_rcu_barrier() was added in 2024, to restore a clean baseline between userspace benchmark runs. The discussion concluded that keeping the existing hook name, adding the second operation and documenting both was the safest compatibility choice, but the follow-up was not added. Add that drain and document the stronger test interface. Always retain the existing start-rate limit and perform the kvfree_rcu() drain: an unrelated ordinary barrier does not establish that this work completed. Retain the entry ordinary-barrier sequence snapshot. After draining, skip the final ordinary barrier only if that snapshot is complete, preserving the memory barrier on the completion path. Otherwise, invoke rcu_barrier() explicitly. This keeps the ordinary-callback guarantee independent of whether kvfree_rcu_barrier() embeds an ordinary barrier. Clarify that the documented completion guarantee covers work queued before the request, without preventing new work from being queued. Earlier validation of the unconditional-drain version used four fresh VM pairs with a private-cache fixture: controls retained the queued object (60 to 60 active objects), and treatments drained it (60 to 59). An ordinary-callback test passed on both kernels. Those runs predated the guarded skip and do not validate that change. No elapsed-time improvement is claimed. Link: https://lore.kernel.org/all/20240820155935.1167988-1-urezki@gmail.com/ Signed-off-by: Matthias Goergens <matthias.goergens@gmail.com> --- .../admin-guide/kernel-parameters.txt | 9 ++++-- kernel/rcu/tree.c | 30 ++++++++++++------- 2 files changed, 26 insertions(+), 13 deletions(-) diff --git a/Documentation/admin-guide/kernel-parameters.txt b/Documentation/admin-guide/kernel-parameters.txt index 68647ff4bdd2..914b65ae9413 100644 --- a/Documentation/admin-guide/kernel-parameters.txt +++ b/Documentation/admin-guide/kernel-parameters.txt @@ -5699,9 +5699,12 @@ Kernel parameters there is an ongoing too-long CSD-lock wait. rcutree.do_rcu_barrier= [KNL] - Request a call to rcu_barrier(). This is - throttled so that userspace tests can safely - hammer on the sysfs variable if they so choose. + Wait for deferred kfree_rcu() frees and ordinary + call_rcu() callbacks queued before this request to + complete. This does not prevent new work from being + queued concurrently. Requests are throttled so that + userspace tests can safely hammer on the sysfs + variable if they so choose. If triggered before the RCU grace-period machinery is fully active, this will error out with EAGAIN. diff --git a/kernel/rcu/tree.c b/kernel/rcu/tree.c index 96848fc1f02b..93b71682306c 100644 --- a/kernel/rcu/tree.c +++ b/kernel/rcu/tree.c @@ -3989,12 +3989,12 @@ EXPORT_SYMBOL_GPL(rcu_barrier); static unsigned long rcu_barrier_last_throttle; /** - * rcu_barrier_throttled - Do rcu_barrier(), but limit to one per second + * rcu_barrier_throttled - Drain deferred RCU frees, but rate-limit starts * - * This can be thought of as guard rails around rcu_barrier() that - * permits unrestricted userspace use, at least assuming the hardware's - * try_cmpxchg() is robust. There will be at most one call per second to - * rcu_barrier() system-wide from use of this function, which means that + * This can be thought of as guard rails around the deferred-free barriers + * that permit unrestricted userspace use, at least assuming the hardware's + * try_cmpxchg() is robust. There will be at most one drain operation started + * per sixteenth of a second from use of this function, which means that * callers might needlessly wait a second or three. * * This is intended for use by test suites to avoid OOM by flushing RCU @@ -4016,14 +4016,24 @@ static void rcu_barrier_throttled(void) while (time_in_range(j, old, old + HZ / 16) || !try_cmpxchg(&rcu_barrier_last_throttle, &old, j)) { schedule_timeout_idle(HZ / 16); - if (rcu_seq_done(&rcu_state.barrier_sequence, s)) { - smp_mb(); /* caller's subsequent code after above check. */ - return; - } j = jiffies; old = READ_ONCE(rcu_barrier_last_throttle); } - rcu_barrier(); + /* + * kfree_rcu() can retain objects outside the ordinary callback lists in + * per-CPU SLUB sheaves and kvfree_rcu batches. Always drain those queues: + * an ordinary barrier does not establish that this work was drained. + */ + kvfree_rcu_barrier(); + /* + * A completed barrier can still cover ordinary callbacks queued before + * our entry snapshot. Otherwise, retain an explicit ordinary barrier + * without depending on the implementation of kvfree_rcu_barrier(). + */ + if (rcu_seq_done(&rcu_state.barrier_sequence, s)) + smp_mb(); /* caller's subsequent code after above check. */ + else + rcu_barrier(); } /* -- 2.55.0 ^ permalink raw reply [flat|nested] 10+ messages in thread
end of thread, other threads:[~2026-09-11 3:41 UTC | newest] Thread overview: 10+ messages (download: mbox.gz / follow: Atom feed) -- links below jump to the message on this page -- 2026-09-10 10:11 [PATCH 0/1] rcu: drain kfree_rcu sheaves from the userspace barrier hook Matthias Goergens 2026-09-10 10:11 ` [PATCH 1/1] " Matthias Goergens 2026-09-10 12:03 ` Harry Yoo 2026-09-10 13:46 ` Matthias Goergens 2026-09-10 15:37 ` Harry Yoo 2026-09-10 17:00 ` [PATCH v2 0/1] rcu: make userspace barrier hook drain kvfree_rcu work Matthias Goergens 2026-09-10 17:00 ` [PATCH v2 1/1] " Matthias Goergens 2026-09-10 17:51 ` Paul E. McKenney 2026-09-11 3:40 ` [PATCH v3 0/1] " Matthias Goergens 2026-09-11 3:40 ` [PATCH v3 1/1] " Matthias Goergens
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®