* [PATCH v2] mm/slab_common: fix shrink budget underflow in kfree_rcu_shrink_scan
@ 2026-08-26 7:56 Longlong Xia
2026-08-26 8:35 ` Hao Li
2026-08-30 14:22 ` Harry Yoo
0 siblings, 2 replies; 3+ messages in thread
From: Longlong Xia @ 2026-08-26 7:56 UTC (permalink / raw)
To: vbabka, harry, akpm
Cc: hao.li, cl, rientjes, roman.gushchin, linux-mm, linux-kernel,
paulmck, rcu, xialonglong
From: Longlong Xia <xialonglong@kylinos.cn>
The kfree_rcu shrinker decremented sc->nr_to_scan (unsigned long)
and then tested the result with <= 0. When a single CPU's object
count exceeds the remaining budget, the subtraction wraps to a large
positive value and the <= 0 comparison, which is equivalent to == 0
for an unsigned type, never fires again. The scan loop then iterates
through every possible CPU instead of honouring the reclaim budget.
Accumulate into freed and stop once freed >= nr_to_scan. The shrinker
core treats nr_to_scan as input-only, so dropping the decrement is
safe; freed becomes unsigned long to match the return type.
Suggested-by: Hao Li <hao.li@linux.dev>
Assisted-by: Codex:gpt-5.6-sol
Signed-off-by: Longlong Xia <xialonglong@kylinos.cn>
---
Changes in v2:
- Rework per suggestion from Hao Li: accumulate into freed directly,
compare freed >= nr_to_scan instead of decrementing nr_to_scan, and
drop the per-CPU count local; promote freed to unsigned long.
Link: https://lore.kernel.org/all/20260824091838.1692153-1-xialonglong2025@163.com/
---
mm/slab_common.c | 13 +++++--------
1 file changed, 5 insertions(+), 8 deletions(-)
diff --git a/mm/slab_common.c b/mm/slab_common.c
index 657fd75776ea..e227c2ef2a4e 100644
--- a/mm/slab_common.c
+++ b/mm/slab_common.c
@@ -2162,20 +2162,17 @@ kfree_rcu_shrink_count(struct shrinker *shrink, struct shrink_control *sc)
static unsigned long
kfree_rcu_shrink_scan(struct shrinker *shrink, struct shrink_control *sc)
{
- int cpu, freed = 0;
+ int cpu;
+ unsigned long freed = 0;
for_each_possible_cpu(cpu) {
- int count;
struct kfree_rcu_cpu *krcp = per_cpu_ptr(&krc, cpu);
- count = krc_count(krcp);
- count += drain_page_cache(krcp);
+ freed += krc_count(krcp);
+ freed += drain_page_cache(krcp);
kfree_rcu_monitor(&krcp->monitor_work.work);
- sc->nr_to_scan -= count;
- freed += count;
-
- if (sc->nr_to_scan <= 0)
+ if (freed >= sc->nr_to_scan)
break;
}
--
2.43.0
^ permalink raw reply [flat|nested] 3+ messages in thread
* Re: [PATCH v2] mm/slab_common: fix shrink budget underflow in kfree_rcu_shrink_scan
2026-08-26 7:56 [PATCH v2] mm/slab_common: fix shrink budget underflow in kfree_rcu_shrink_scan Longlong Xia
@ 2026-08-26 8:35 ` Hao Li
2026-08-30 14:22 ` Harry Yoo
1 sibling, 0 replies; 3+ messages in thread
From: Hao Li @ 2026-08-26 8:35 UTC (permalink / raw)
To: Longlong Xia
Cc: vbabka, harry, akpm, cl, rientjes, roman.gushchin, linux-mm,
linux-kernel, paulmck, rcu, xialonglong
On Wed, Aug 26, 2026 at 03:56:53PM +0800, Longlong Xia wrote:
> From: Longlong Xia <xialonglong@kylinos.cn>
>
> The kfree_rcu shrinker decremented sc->nr_to_scan (unsigned long)
> and then tested the result with <= 0. When a single CPU's object
> count exceeds the remaining budget, the subtraction wraps to a large
> positive value and the <= 0 comparison, which is equivalent to == 0
> for an unsigned type, never fires again. The scan loop then iterates
> through every possible CPU instead of honouring the reclaim budget.
>
> Accumulate into freed and stop once freed >= nr_to_scan. The shrinker
> core treats nr_to_scan as input-only, so dropping the decrement is
> safe; freed becomes unsigned long to match the return type.
>
> Suggested-by: Hao Li <hao.li@linux.dev>
> Assisted-by: Codex:gpt-5.6-sol
> Signed-off-by: Longlong Xia <xialonglong@kylinos.cn>
> ---
> Changes in v2:
> - Rework per suggestion from Hao Li: accumulate into freed directly,
> compare freed >= nr_to_scan instead of decrementing nr_to_scan, and
> drop the per-CPU count local; promote freed to unsigned long.
>
> Link: https://lore.kernel.org/all/20260824091838.1692153-1-xialonglong2025@163.com/
> ---
> mm/slab_common.c | 13 +++++--------
> 1 file changed, 5 insertions(+), 8 deletions(-)
>
Looks good to me. Thanks.
Reviewed-by: Hao Li <hao.li@linux.dev>
> diff --git a/mm/slab_common.c b/mm/slab_common.c
> index 657fd75776ea..e227c2ef2a4e 100644
> --- a/mm/slab_common.c
> +++ b/mm/slab_common.c
> @@ -2162,20 +2162,17 @@ kfree_rcu_shrink_count(struct shrinker *shrink, struct shrink_control *sc)
> static unsigned long
> kfree_rcu_shrink_scan(struct shrinker *shrink, struct shrink_control *sc)
> {
> - int cpu, freed = 0;
> + int cpu;
> + unsigned long freed = 0;
>
> for_each_possible_cpu(cpu) {
> - int count;
> struct kfree_rcu_cpu *krcp = per_cpu_ptr(&krc, cpu);
>
> - count = krc_count(krcp);
> - count += drain_page_cache(krcp);
> + freed += krc_count(krcp);
> + freed += drain_page_cache(krcp);
> kfree_rcu_monitor(&krcp->monitor_work.work);
>
> - sc->nr_to_scan -= count;
> - freed += count;
> -
> - if (sc->nr_to_scan <= 0)
> + if (freed >= sc->nr_to_scan)
> break;
> }
>
> --
> 2.43.0
>
--
Thanks,
Hao
^ permalink raw reply [flat|nested] 3+ messages in thread
* Re: [PATCH v2] mm/slab_common: fix shrink budget underflow in kfree_rcu_shrink_scan
2026-08-26 7:56 [PATCH v2] mm/slab_common: fix shrink budget underflow in kfree_rcu_shrink_scan Longlong Xia
2026-08-26 8:35 ` Hao Li
@ 2026-08-30 14:22 ` Harry Yoo
1 sibling, 0 replies; 3+ messages in thread
From: Harry Yoo @ 2026-08-30 14:22 UTC (permalink / raw)
To: Longlong Xia
Cc: vbabka, akpm, hao.li, cl, rientjes, roman.gushchin, linux-mm,
linux-kernel, paulmck, rcu, xialonglong
On Wed, Aug 26, 2026 at 03:56:53PM +0800, Longlong Xia wrote:
> From: Longlong Xia <xialonglong@kylinos.cn>
Hi Longlong, I have a few questions.
> The kfree_rcu shrinker decremented sc->nr_to_scan (unsigned long)
> and then tested the result with <= 0. When a single CPU's object
> count exceeds the remaining budget, the subtraction wraps to a large
> positive value and the <= 0 comparison, which is equivalent to == 0
> for an unsigned type, never fires again.
Since when (which commit) has it been broken?
If it has been undiscovered for a very long time, why is that so?
And how did you discover this?
> The scan loop then iterates
> through every possible CPU instead of honouring the reclaim budget.
Did you confirm this actually does happen? If so, how often does the
kernel end up iterating through every possible CPUs, very rarely or
almost always?
Would this affect the kernel's reclamation behavior in some way?
> Accumulate into freed and stop once freed >= nr_to_scan. The shrinker
> core treats nr_to_scan as input-only, so dropping the decrement is
> safe; freed becomes unsigned long to match the return type.
>
> Suggested-by: Hao Li <hao.li@linux.dev>
> Assisted-by: Codex:gpt-5.6-sol
> Signed-off-by: Longlong Xia <xialonglong@kylinos.cn>
> ---
> Changes in v2:
> - Rework per suggestion from Hao Li: accumulate into freed directly,
> compare freed >= nr_to_scan instead of decrementing nr_to_scan, and
> drop the per-CPU count local; promote freed to unsigned long.
>
> Link: https://lore.kernel.org/all/20260824091838.1692153-1-xialonglong2025@163.com/
> ---
> mm/slab_common.c | 13 +++++--------
> 1 file changed, 5 insertions(+), 8 deletions(-)
>
> diff --git a/mm/slab_common.c b/mm/slab_common.c
> index 657fd75776ea..e227c2ef2a4e 100644
> --- a/mm/slab_common.c
> +++ b/mm/slab_common.c
> @@ -2162,20 +2162,17 @@ kfree_rcu_shrink_count(struct shrinker *shrink, struct shrink_control *sc)
> static unsigned long
> kfree_rcu_shrink_scan(struct shrinker *shrink, struct shrink_control *sc)
> {
> - int cpu, freed = 0;
> + int cpu;
> + unsigned long freed = 0;
>
> for_each_possible_cpu(cpu) {
> - int count;
> struct kfree_rcu_cpu *krcp = per_cpu_ptr(&krc, cpu);
>
> - count = krc_count(krcp);
> - count += drain_page_cache(krcp);
> + freed += krc_count(krcp);
> + freed += drain_page_cache(krcp);
> kfree_rcu_monitor(&krcp->monitor_work.work);
>
> - sc->nr_to_scan -= count;
> - freed += count;
> -
> - if (sc->nr_to_scan <= 0)
> + if (freed >= sc->nr_to_scan)
> break;
> }
>
> --
> 2.43.0
--
Cheers,
Harry / Hyeonggon
^ permalink raw reply [flat|nested] 3+ messages in thread
end of thread, other threads:[~2026-08-30 14:23 UTC | newest]
Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-08-26 7:56 [PATCH v2] mm/slab_common: fix shrink budget underflow in kfree_rcu_shrink_scan Longlong Xia
2026-08-26 8:35 ` Hao Li
2026-08-30 14:22 ` Harry Yoo
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®