mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Ridong Chen <ridong.chen@linux.dev>
To: Barry Song <baohua@kernel.org>
Cc: Andrew Morton <akpm@linux-foundation.org>,
	Johannes Weiner <hannes@cmpxchg.org>,
	Kairui Song <kasong@tencent.com>, Qi Zheng <qi.zheng@linux.dev>,
	Shakeel Butt <shakeel.butt@linux.dev>,
	Axel Rasmussen <axelrasmussen@google.com>,
	Yuanchu Xie <yuanchu@google.com>, Wei Xu <weixugc@google.com>,
	Baoquan He <baoquan.he@linux.dev>,
	Baolin Wang <baolin.wang@linux.alibaba.com>,
	David Hildenbrand <david@kernel.org>,
	Michal Hocko <mhocko@kernel.org>,
	Lorenzo Stoakes <ljs@kernel.org>,
	"open list:MEMORY MANAGEMENT - MGLRU (MULTI-GEN LRU)"
	<linux-mm@kvack.org>,
	linux-kernel@vger.kernel.org, Ridong Chen <chenridong@xiaomi.com>
Subject: Re: [PATCH mm-new] mm: vmscan: put rotation-missed folios at the LRU tail
Date: Sun, 20 Sep 2026 19:12:08 +0800	[thread overview]
Message-ID: <63de7ffe-8410-4e88-9e09-02f2cc386539@linux.dev> (raw)
In-Reply-To: <CAGsJ_4xmoGWULDVNGj0LzRFFfxM4VVGg4Bx5Pz5cUfXx9uKrtw@mail.gmail.com>



On 9/20/2026 5:21 PM, Barry Song wrote:
> On Sun, Sep 20, 2026 at 12:30 PM Ridong Chen <ridong.chen@linux.dev> wrote:
>>
>> From: Ridong Chen <chenridong@xiaomi.com>
>>
>> The page reclaim isolates a batch of folios from the tail of an LRU list
>> and works on them one by one.  For a suitable swap-backed folio on an
>> async swap device, it queues the folio for writeback and, after finishing
>> the batch, puts the folio back to the head of the original LRU list.
>>
>> Meanwhile the page writeback flushes the queued folios in its own,
>> independent batches.  For each folio it writes back it calls
>> folio_rotate_reclaimable(), which tries to rotate the folio to the LRU
>> tail.  But folio_rotate_reclaimable() only takes effect once the folio has
>> been put back by reclaim.  If the async swap device is fast enough, the
>> writeback can complete a folio while reclaim is still working on the rest
>> of the batch that contains it.  In that case the folio stays near the head
>> and reclaim will not revisit it before wrapping around, causing a cold/hot
>> inversion: a clean, written-back folio that should be a prime reclaim
>> candidate is kept ahead of hotter folios.
>>
>> commit 359a5e1416ca ("mm: multi-gen LRU: retry folios written back while
>> isolated") addressed this for MGLRU only.  The traditional active/inactive
>> LRU has the same problem, reported at [1].  A reproducer is available at
>> [2].
>>
>> Rather than re-reclaiming those folios (which would drop the swap cache
>> that may still be useful for a future hit [4]), restore the rotation that
>> was missed: when move_folios_to_lru() puts a folio back, add it to the LRU
>> tail if it looks like it missed folio_rotate_reclaimable() (inactive, not
>> mapped, not dirty and not under writeback).  A new do_rotate parameter
>> gates this so it only applies on the reclaim put-back path
>> (shrink_inactive_list()), not on shrink_active_list() where the list order
>> is already deliberate.  This approach was suggested by Barry Song [3].
>>
>> Only the traditional LRU is handled here.  MGLRU already retries such
>> folios via its own clean-list retry pass in evict_folios(), so it is left
>> unchanged.  The same do_rotate scheme could later replace that retry pass
>> to unify both LRUs, which is left for a follow-up.
>>
>> [1] https://lore.kernel.org/linux-kernel/20241010081802.290893-1-chenridong@huaweicloud.com/
>> [2] https://lore.kernel.org/lkml/46037a37-4cf6-448e-a94b-30a4d16e8814@linux.dev/
>> [3] https://lore.kernel.org/lkml/CAGsJ_4zwP3_+EYY5Ug9EJ+yD1UdxsBSGr25u8s1K3u_i7LH3Zg@mail.gmail.com/
>> [4] https://lore.kernel.org/linux-mm/20260911121341.178028-1-alex@ghiti.fr/
>>
>> Suggested-by: Barry Song <baohua@kernel.org>
>> Signed-off-by: Ridong Chen <chenridong@xiaomi.com>
>> ---
>>   mm/vmscan.c | 21 +++++++++++++++------
>>   1 file changed, 15 insertions(+), 6 deletions(-)
>>
>> diff --git a/mm/vmscan.c b/mm/vmscan.c
>> index e200ce3eb056..026b844c681f 100644
>> --- a/mm/vmscan.c
>> +++ b/mm/vmscan.c
>> @@ -1971,7 +1971,7 @@ static bool too_many_isolated(struct pglist_data *pgdat, int file,
>>    *
>>    * Note: The caller must not hold any lruvec lock.
>>    */
>> -static unsigned int move_folios_to_lru(struct list_head *list)
>> +static unsigned int move_folios_to_lru(struct list_head *list, bool do_rotate)
>>   {
>>          int nr_pages, nr_moved = 0;
>>          struct lruvec *lruvec = NULL;
>> @@ -2018,7 +2018,16 @@ static unsigned int move_folios_to_lru(struct list_head *list)
>>                          continue;
>>                  }
>>
>> -               lruvec_add_folio(lruvec, folio);
>> +               /*
>> +                * Put folios that may have missed folio_rotate_reclaimable()
>> +                * at the tail to avoid cold/hot inversion.
>> +                */
>> +               if (do_rotate && !folio_test_active(folio) && !folio_mapped(folio) &&
>> +                   !folio_test_dirty(folio) && !folio_test_writeback(folio))
>> +                       lruvec_add_folio_tail(lruvec, folio);
> 
> sashiko says:
> 
> "Could this heuristic indiscriminately place un-reclaimable clean folios at
> the LRU tail, ensuring they are immediately re-scanned in an infinite loop?
> If shrink_inactive_list() isolates a clean, unmapped file folio with an
> elevated refcount (such as from a concurrent GUP or speculative lookup),
> it will fail __remove_mapping() and be placed on ret_folios.
> Because this pinned folio matches the !active, !mapped, !dirty, and
> !writeback checks, it is placed at the tail of the inactive LRU. The very
> next isolate_lru_folios() pull from the tail will immediately isolate this
> exact same folio again, creating an infinite loop that prevents any other
> folios from being reclaimed and causes kswapd to spin at 100% CPU."
> 
> This seems to be a valid concern. For a clean and unmapped folio, we
> may still fail in `__remove_mapping()` if it is pinned somewhere. So
> perhaps we can add a refcount check when fixing the rotation.
> 

Thanks, I will update.

-- 
Best regards
Ridong


      reply	other threads:[~2026-09-20 11:12 UTC|newest]

Thread overview: 4+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-20  4:30 Ridong Chen
2026-09-20  6:10 ` Ridong Chen
2026-09-20  9:21 ` Barry Song
2026-09-20 11:12   ` Ridong Chen [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=63de7ffe-8410-4e88-9e09-02f2cc386539@linux.dev \
    --to=ridong.chen@linux.dev \
    --cc=akpm@linux-foundation.org \
    --cc=axelrasmussen@google.com \
    --cc=baohua@kernel.org \
    --cc=baolin.wang@linux.alibaba.com \
    --cc=baoquan.he@linux.dev \
    --cc=chenridong@xiaomi.com \
    --cc=david@kernel.org \
    --cc=hannes@cmpxchg.org \
    --cc=kasong@tencent.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=mhocko@kernel.org \
    --cc=qi.zheng@linux.dev \
    --cc=shakeel.butt@linux.dev \
    --cc=weixugc@google.com \
    --cc=yuanchu@google.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®