mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Longlong Xia <xialonglong2025@163.com>
To: Harry Yoo <harry@kernel.org>
Cc: vbabka@kernel.org, akpm@linux-foundation.org, hao.li@linux.dev,
	cl@gentwo.org, rientjes@google.com, roman.gushchin@linux.dev,
	linux-mm@kvack.org, linux-kernel@vger.kernel.org,
	paulmck@kernel.org, rcu@vger.kernel.org, xialonglong@kylinos.cn
Subject: Re: [PATCH v2] mm/slab_common: fix shrink budget underflow in kfree_rcu_shrink_scan
Date: Thu, 3 Sep 2026 17:13:15 +0800	[thread overview]
Message-ID: <baa22f03-e553-46cc-8c35-436363546605@163.com> (raw)
In-Reply-To: <apQsbGJDPPyXcn8R@dev>

Thanks for the reply.

在 2026/8/30 22:22, Harry Yoo 写道:
> On Wed, Aug 26, 2026 at 03:56:53PM +0800, Longlong Xia wrote:
>> From: Longlong Xia <xialonglong@kylinos.cn>
> Hi Longlong, I have a few questions.
>
>> The kfree_rcu shrinker decremented sc->nr_to_scan (unsigned long)
>> and then tested the result with <= 0. When a single CPU's object
>> count exceeds the remaining budget, the subtraction wraps to a large
>> positive value and the <= 0 comparison, which is equivalent to == 0
>> for an unsigned type, never fires again.
> Since when (which commit) has it been broken?
> If it has been undiscovered for a very long time, why is that so?

Since the shrinker was first added: 9154244c1ab6 ("rcu/tree: Add a
shrinker to prevent OOM due to kfree_rcu() batching", v5.8);

>
> And how did you discover this?

While reviewing the kvfree_rcu() batching paths with the help of AI
tooling (hence the Assisted-by: tag in the patch), which flagged the
unsigned subtraction feeding a "<= 0" test.

>> The scan loop then iterates
>> through every possible CPU instead of honouring the reclaim budget.
> Did you confirm this actually does happen? If so, how often does the
> kernel end up iterating through every possible CPUs, very rarely or
> almost always?

Yes, in two ways.


Experiment 1 -- deterministic, instrumented loop.  Per-CPU pr_info()
plus a test entry invoking kfree_rcu_shrink_scan() with a controlled
budget; 4-CPU QEMU guest, 400 objects queued on CPU0, budget 50:
   buggy:   cpu=0 count=400 nr_to_scan_after_sub=18446744073709551266
                 visits all 4 CPUs
   patched: cpu=0 count=400 nr_to_scan=50 BREAK

                stops at CPU0


Experiment 2 -- natural trigger, no kernel modification at all
(kprobes only).  Setup: 1 GiB QEMU guest, 4 possible CPUs, no swap.

   1. Observation: three kprobes.  Scan entry/exit record nr_to_scan
      and the return value; each kfree_rcu_monitor() hit inside a
      scan is one CPU the loop processed.

   2. Memory pressure: ~650 MiB of anonymous memory.

   3. krc backlog: loaded the in-tree test_vmalloc module ,

      which repeatedly vmalloc()s one page and frees it via kvfree_rcu();

     10 million iterations over ~70s kept the per-CPU queues populated.


Result within ~70s: 262 natural invocations, all from kswapd0;
171 of the 261 scans that found a backlog (66%) overshot the budget
(nr_to_scan = 128, freed up to 4479) and visited all 4 possible CPUs.


>
> Would this affect the kernel's reclamation behavior in some way?

The break was meant to cap the scan at the reclaim budget; once
the subtraction wraps, the cap is gone and one scan processes
every possible CPU, draining the whole backlog instead of a
budget-sized slice.


Thanks,

Longlong

>> Accumulate into freed and stop once freed >= nr_to_scan. The shrinker
>> core treats nr_to_scan as input-only, so dropping the decrement is
>> safe; freed becomes unsigned long to match the return type.
>>
>> Suggested-by: Hao Li <hao.li@linux.dev>
>> Assisted-by: Codex:gpt-5.6-sol
>> Signed-off-by: Longlong Xia <xialonglong@kylinos.cn>
>> ---
>> Changes in v2:
>> - Rework per suggestion from Hao Li: accumulate into freed directly,
>>    compare freed >= nr_to_scan instead of decrementing nr_to_scan, and
>>    drop the per-CPU count local; promote freed to unsigned long.
>>
>> Link: https://lore.kernel.org/all/20260824091838.1692153-1-xialonglong2025@163.com/
>> ---
>>   mm/slab_common.c | 13 +++++--------
>>   1 file changed, 5 insertions(+), 8 deletions(-)
>>
>> diff --git a/mm/slab_common.c b/mm/slab_common.c
>> index 657fd75776ea..e227c2ef2a4e 100644
>> --- a/mm/slab_common.c
>> +++ b/mm/slab_common.c
>> @@ -2162,20 +2162,17 @@ kfree_rcu_shrink_count(struct shrinker *shrink, struct shrink_control *sc)
>>   static unsigned long
>>   kfree_rcu_shrink_scan(struct shrinker *shrink, struct shrink_control *sc)
>>   {
>> -	int cpu, freed = 0;
>> +	int cpu;
>> +	unsigned long freed = 0;
>>   
>>   	for_each_possible_cpu(cpu) {
>> -		int count;
>>   		struct kfree_rcu_cpu *krcp = per_cpu_ptr(&krc, cpu);
>>   
>> -		count = krc_count(krcp);
>> -		count += drain_page_cache(krcp);
>> +		freed += krc_count(krcp);
>> +		freed += drain_page_cache(krcp);
>>   		kfree_rcu_monitor(&krcp->monitor_work.work);
>>   
>> -		sc->nr_to_scan -= count;
>> -		freed += count;
>> -
>> -		if (sc->nr_to_scan <= 0)
>> +		if (freed >= sc->nr_to_scan)
>>   			break;
>>   	}
>>   
>> -- 
>> 2.43.0


  reply	other threads:[~2026-09-03  9:14 UTC|newest]

Thread overview: 5+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-26  7:56 Longlong Xia
2026-08-26  8:35 ` Hao Li
2026-08-30 14:22 ` Harry Yoo
2026-09-03  9:13   ` Longlong Xia [this message]
2026-09-03 10:56     ` Harry Yoo

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=baa22f03-e553-46cc-8c35-436363546605@163.com \
    --to=xialonglong2025@163.com \
    --cc=akpm@linux-foundation.org \
    --cc=cl@gentwo.org \
    --cc=hao.li@linux.dev \
    --cc=harry@kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=paulmck@kernel.org \
    --cc=rcu@vger.kernel.org \
    --cc=rientjes@google.com \
    --cc=roman.gushchin@linux.dev \
    --cc=vbabka@kernel.org \
    --cc=xialonglong@kylinos.cn \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®