From: Ridong Chen <ridong.chen@linux.dev>
To: Andrew Morton <akpm@linux-foundation.org>,
Johannes Weiner <hannes@cmpxchg.org>
Cc: David Hildenbrand <david@kernel.org>,
Michal Hocko <mhocko@kernel.org>, Qi Zheng <qi.zheng@linux.dev>,
Shakeel Butt <shakeel.butt@linux.dev>,
Lorenzo Stoakes <ljs@kernel.org>,
Kairui Song <kasong@tencent.com>, Barry Song <baohua@kernel.org>,
Axel Rasmussen <axelrasmussen@google.com>,
Yuanchu Xie <yuanchu@google.com>, Wei Xu <weixugc@google.com>,
Baoquan He <baoquan.he@linux.dev>,
Baolin Wang <baolin.wang@linux.alibaba.com>,
linux-mm@kvack.org, linux-kernel@vger.kernel.org,
Ridong Chen <chenridong@xiaomi.com>
Subject: Re: [PATCH] mm: vmscan: put rotation-missed folios at the LRU tail
Date: Sun, 20 Sep 2026 21:28:31 +0800 [thread overview]
Message-ID: <1d61ee17-edbb-4491-a5e2-32aa2779d839@linux.dev> (raw)
In-Reply-To: <20260920132030.3368138-1-ridong.chen@linux.dev>
On 9/20/2026 9:20 PM, Ridong Chen wrote:
> From: Ridong Chen <chenridong@xiaomi.com>
>
> The page reclaim isolates a batch of folios from the tail of an LRU list
> and works on them one by one. For a suitable swap-backed folio on an
> async swap device, it queues the folio for writeback and, after finishing
> the batch, puts the folio back to the head of the original LRU list.
>
> Meanwhile the page writeback flushes the queued folios in its own,
> independent batches. For each folio it writes back it calls
> folio_rotate_reclaimable(), which tries to rotate the folio to the LRU
> tail. But folio_rotate_reclaimable() only takes effect once the folio has
> been put back by reclaim. If the async swap device is fast enough, the
> writeback can complete a folio while reclaim is still working on the rest
> of the batch that contains it. In that case the folio stays near the head
> and reclaim will not revisit it before wrapping around, causing a cold/hot
> inversion: a clean, written-back folio that should be a prime reclaim
> candidate is kept ahead of hotter folios.
>
> commit 359a5e1416ca ("mm: multi-gen LRU: retry folios written back while
> isolated") addressed this for MGLRU only. The traditional active/inactive
> LRU has the same problem, reported at [1]. A reproducer is available at
> [2].
>
> Rather than re-reclaiming those folios (which would drop the swap cache
> that may still be useful for a future hit [4]), restore the rotation that
> was missed: when move_folios_to_lru() puts a folio back, add it to the LRU
> tail if it looks like it missed folio_rotate_reclaimable() (inactive, not
> mapped, not dirty and not under writeback). A referenced folio is left at
> the head so it still gets a second chance, and a folio with an unexpected
> reference (e.g. a GUP or speculative pin) is left at the head because it
> cannot be reclaimed yet anyway. A new do_rotate parameter gates this so
> it only applies on the reclaim put-back path (shrink_inactive_list()), not
> on shrink_active_list() where the list order is already deliberate. This
> approach was suggested by Barry Song [3].
>
> Only the traditional LRU is handled here. MGLRU already retries such
> folios via its own clean-list retry pass in evict_folios(), so it is left
> unchanged. The same do_rotate scheme could later replace that retry pass
> to unify both LRUs, which is left for a follow-up.
>
> Test result with [2]:
>
> Without patch:
> cat memory.usage_in_bytes
> 1073700864
> cat memory.memsw.usage_in_bytes
> 1413124096
>
> free -h
> total used free
> Mem: 1.6Gi 1.2Gi 299Mi
> Swap: 1.0Gi 678Mi 346Mi
>
> With patch:
> cat memory.usage_in_bytes
> 1071140864
> cat memory.memsw.usage_in_bytes
> 1413423104
>
> free -h
> total used free
> Mem: 1.6Gi 1.2Gi 322Mi
> Swap: 1.0Gi 328Mi 695Mi
>
> After applying the patch, the difference between
> memory.memsw.usage_in_bytes and memory.usage_in_bytes is close to the swap
> "used" value reported by 'free -h'.
>
> [1] https://lore.kernel.org/linux-kernel/20241010081802.290893-1-chenridong@huaweicloud.com/
> [2] https://lore.kernel.org/lkml/46037a37-4cf6-448e-a94b-30a4d16e8814@linux.dev/
> [3] https://lore.kernel.org/lkml/CAGsJ_4zwP3_+EYY5Ug9EJ+yD1UdxsBSGr25u8s1K3u_i7LH3Zg@mail.gmail.com/
> [4] https://lore.kernel.org/linux-mm/20260911121341.178028-1-alex@ghiti.fr/
>
> Suggested-by: Barry Song <baohua@kernel.org>
> Signed-off-by: Ridong Chen <chenridong@xiaomi.com>
> ---
>
> v1 -> v2:
> - skip referenced folios (FOLIOREF_KEEP) and unexpectedly pinned folios
> when rotating to the LRU tail.
> - add test result to the commit message.
>
> mm/vmscan.c | 24 ++++++++++++++++++------
> 1 file changed, 18 insertions(+), 6 deletions(-)
>
> diff --git a/mm/vmscan.c b/mm/vmscan.c
> index e200ce3eb056..91295070ca33 100644
> --- a/mm/vmscan.c
> +++ b/mm/vmscan.c
> @@ -1971,7 +1971,7 @@ static bool too_many_isolated(struct pglist_data *pgdat, int file,
> *
> * Note: The caller must not hold any lruvec lock.
> */
> -static unsigned int move_folios_to_lru(struct list_head *list)
> +static unsigned int move_folios_to_lru(struct list_head *list, bool do_rotate)
> {
> int nr_pages, nr_moved = 0;
> struct lruvec *lruvec = NULL;
> @@ -2018,7 +2018,19 @@ static unsigned int move_folios_to_lru(struct list_head *list)
> continue;
> }
>
> - lruvec_add_folio(lruvec, folio);
> + /*
> + * Put clean, unreferenced and unpinned folios that may have
> + * missed folio_rotate_reclaimable() at the tail to avoid
> + * cold/hot inversion.
> + */
> + if (do_rotate && !folio_test_active(folio) && !folio_mapped(folio) &&
> + !folio_test_dirty(folio) && !folio_test_writeback(folio) &&
> + !folio_test_referenced(folio) &&
> + folio_ref_count(folio) == folio_expected_ref_count(folio))
> + lruvec_add_folio_tail(lruvec, folio);
> + else
> + lruvec_add_folio(lruvec, folio);
> +
> nr_pages = folio_nr_pages(folio);
> nr_moved += nr_pages;
> if (folio_test_active(folio))
> @@ -2135,7 +2147,7 @@ static unsigned long shrink_inactive_list(unsigned long nr_to_scan,
> nr_reclaimed = shrink_folio_list(&folio_list, pgdat, sc, &stat, false,
> lruvec_memcg(lruvec));
>
> - move_folios_to_lru(&folio_list);
> + move_folios_to_lru(&folio_list, true);
>
> mod_lruvec_state(lruvec, PGDEMOTE_KSWAPD + reclaimer_offset(sc),
> stat.nr_demoted);
> @@ -2246,8 +2258,8 @@ static void shrink_active_list(unsigned long nr_to_scan,
> /*
> * Move folios back to the lru list.
> */
> - nr_activate = move_folios_to_lru(&l_active);
> - nr_deactivate = move_folios_to_lru(&l_inactive);
> + nr_activate = move_folios_to_lru(&l_active, false);
> + nr_deactivate = move_folios_to_lru(&l_inactive, false);
>
> count_vm_events(PGDEACTIVATE, nr_deactivate);
> count_memcg_events(lruvec_memcg(lruvec), PGDEACTIVATE, nr_deactivate);
> @@ -5115,7 +5127,7 @@ static int evict_folios(unsigned long nr_to_scan, struct lruvec *lruvec,
> folio_set_active(folio);
> }
>
> - move_folios_to_lru(&list);
> + move_folios_to_lru(&list, false);
>
> walk = current->reclaim_state->mm_walk;
> if (walk && walk->batched) {
Sorry for sending this patch without a version tag. I've resent a new one;
please ignore this patch.
https://lore.kernel.org/lkml/20260920132519.3369946-1-ridong.chen@linux.dev/
Thanks.
--
Best regards
Ridong
prev parent reply other threads:[~2026-09-20 13:28 UTC|newest]
Thread overview: 2+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-20 13:20 Ridong Chen
2026-09-20 13:28 ` Ridong Chen [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=1d61ee17-edbb-4491-a5e2-32aa2779d839@linux.dev \
--to=ridong.chen@linux.dev \
--cc=akpm@linux-foundation.org \
--cc=axelrasmussen@google.com \
--cc=baohua@kernel.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=baoquan.he@linux.dev \
--cc=chenridong@xiaomi.com \
--cc=david@kernel.org \
--cc=hannes@cmpxchg.org \
--cc=kasong@tencent.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=mhocko@kernel.org \
--cc=qi.zheng@linux.dev \
--cc=shakeel.butt@linux.dev \
--cc=weixugc@google.com \
--cc=yuanchu@google.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®