From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mta1.migadu.com (out-41.mta1.migadu.com [95.215.58.41]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 995AD368D4B for ; Sun, 20 Sep 2026 06:10:53 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=95.215.58.41 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789884656; cv=none; b=qK5P7sBn3sZfJB889VQei4/v6mgTz56TLEpBuExUSY1M6M1i89uniUxVPr3qyYlBbxRKqscgfJ7TuzC2MVWHJ/BUGgDLaF5snrno9MNhSClCpfFaj6lTXAsrRR8E2kFK9MFnOuaCoCN9M4NCn7PmWgY++iExnQPEvZ0qA/ySMQ8= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789884656; c=relaxed/simple; bh=XBXICcDKSuY/lMhCIyKc8ruxZp1iN0ZmzDia5GycYgI=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=Dip9lvuHoblF7YeWlJpvYmz73GJaPqJNqRlYwamoR346b5MaO+stLXmFN4i033K+2V2njrrrCsnlAMjh+tBr0GSGeJKei8ZkfYZIezhuO/9tcgLUr3IkJeznJMi2t93SqJnqIuUBsm51sXerfzLlWhlS9MZHvGIu/ndF/Gkeg6c= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=i/diRS5+; arc=none smtp.client-ip=95.215.58.41 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="i/diRS5+" X-Envelope-To: linux-kernel@vger.kernel.org DKIM-Signature: a=rsa-sha256; bh=XBXICcDKSuY/lMhCIyKc8ruxZp1iN0ZmzDia5GycYgI=; c=simple/simple; d=linux.dev; h=from:to:subject:date:message-id:mime-version:content-type; s=key1; t=1789884651; v=1; x=1790489451; b=i/diRS5+SpyBLvfAI2dGtvOfQwq7rIdgK3496V51UaoB9d5LCaVdQAMbDnWllAyaSDcldNoR 1POvFB3sU2IW3rJ5GHqBkmvhf1zCpED78HZfRajI826yiAieXkl+rUNpNOrdqzhE+WNc0eBmgqP dpKegz0UzQrjg4lV14q8yHbk= X-Envelope-To: linux-kernel@vger.kernel.org Received: by smtp.migadu.com with ESMTPS id 0e8465e20e73e58e; Sun, 20 Sep 2026 06:10:51 +0000 X-Mizu-Trace-ID: 0e8465e20e73e58e X-Migadu-Flow: FLOW_OUT Message-ID: <00212347-d64f-48e8-bbda-a4613fe492ae@linux.dev> Date: Sun, 20 Sep 2026 14:10:44 +0800 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH mm-new] mm: vmscan: put rotation-missed folios at the LRU tail To: Andrew Morton , Johannes Weiner Cc: Kairui Song , Qi Zheng , Shakeel Butt , Barry Song , Axel Rasmussen , Yuanchu Xie , Wei Xu , Baoquan He , Baolin Wang , David Hildenbrand , Michal Hocko , Lorenzo Stoakes , "open list:MEMORY MANAGEMENT - MGLRU (MULTI-GEN LRU)" , linux-kernel@vger.kernel.org, Ridong Chen References: <20260920043015.3126317-1-ridong.chen@linux.dev> From: Ridong Chen In-Reply-To: <20260920043015.3126317-1-ridong.chen@linux.dev> Content-Type: text/plain; charset=UTF-8; format=flowed Content-Transfer-Encoding: 7bit On 9/20/2026 12:30 PM, Ridong Chen wrote: > From: Ridong Chen > > The page reclaim isolates a batch of folios from the tail of an LRU list > and works on them one by one. For a suitable swap-backed folio on an > async swap device, it queues the folio for writeback and, after finishing > the batch, puts the folio back to the head of the original LRU list. > > Meanwhile the page writeback flushes the queued folios in its own, > independent batches. For each folio it writes back it calls > folio_rotate_reclaimable(), which tries to rotate the folio to the LRU > tail. But folio_rotate_reclaimable() only takes effect once the folio has > been put back by reclaim. If the async swap device is fast enough, the > writeback can complete a folio while reclaim is still working on the rest > of the batch that contains it. In that case the folio stays near the head > and reclaim will not revisit it before wrapping around, causing a cold/hot > inversion: a clean, written-back folio that should be a prime reclaim > candidate is kept ahead of hotter folios. > > commit 359a5e1416ca ("mm: multi-gen LRU: retry folios written back while > isolated") addressed this for MGLRU only. The traditional active/inactive > LRU has the same problem, reported at [1]. A reproducer is available at > [2]. > > Rather than re-reclaiming those folios (which would drop the swap cache > that may still be useful for a future hit [4]), restore the rotation that > was missed: when move_folios_to_lru() puts a folio back, add it to the LRU > tail if it looks like it missed folio_rotate_reclaimable() (inactive, not > mapped, not dirty and not under writeback). A new do_rotate parameter > gates this so it only applies on the reclaim put-back path > (shrink_inactive_list()), not on shrink_active_list() where the list order > is already deliberate. This approach was suggested by Barry Song [3]. > Add test result: ``` Without patch: cat memory.usage_in_bytes 1073700864 cat memory.memsw.usage_in_bytes 1413124096 free -h total used free Mem: 1.6Gi 1.2Gi 299Mi Swap: 1.0Gi 678Mi 346Mi With patch: cat memory.usage_in_bytes 1071140864 bytes cat memory.memsw.usage_in_bytes 1413423104 bytes free -h total used free Mem: 1.6Gi 1.2Gi 322Mi Swap: 1.0Gi 328Mi 695Mi ``` After applying the patch, the difference between memory.memsw.usage_in_bytes and memory.usage_in_bytes is close to the swap "used" value reported by 'free -h;. This yields the same result as the 'retry' approach in v9 [1]. [1] https://lore.kernel.org/lkml/20260916123846.1242548-1-ridong.chen@linux.dev/ > Only the traditional LRU is handled here. MGLRU already retries such > folios via its own clean-list retry pass in evict_folios(), so it is left > unchanged. The same do_rotate scheme could later replace that retry pass > to unify both LRUs, which is left for a follow-up. > > [1] https://lore.kernel.org/linux-kernel/20241010081802.290893-1-chenridong@huaweicloud.com/ > [2] https://lore.kernel.org/lkml/46037a37-4cf6-448e-a94b-30a4d16e8814@linux.dev/ > [3] https://lore.kernel.org/lkml/CAGsJ_4zwP3_+EYY5Ug9EJ+yD1UdxsBSGr25u8s1K3u_i7LH3Zg@mail.gmail.com/ > [4] https://lore.kernel.org/linux-mm/20260911121341.178028-1-alex@ghiti.fr/ > > Suggested-by: Barry Song > Signed-off-by: Ridong Chen > --- > mm/vmscan.c | 21 +++++++++++++++------ > 1 file changed, 15 insertions(+), 6 deletions(-) > > diff --git a/mm/vmscan.c b/mm/vmscan.c > index e200ce3eb056..026b844c681f 100644 > --- a/mm/vmscan.c > +++ b/mm/vmscan.c > @@ -1971,7 +1971,7 @@ static bool too_many_isolated(struct pglist_data *pgdat, int file, > * > * Note: The caller must not hold any lruvec lock. > */ > -static unsigned int move_folios_to_lru(struct list_head *list) > +static unsigned int move_folios_to_lru(struct list_head *list, bool do_rotate) > { > int nr_pages, nr_moved = 0; > struct lruvec *lruvec = NULL; > @@ -2018,7 +2018,16 @@ static unsigned int move_folios_to_lru(struct list_head *list) > continue; > } > > - lruvec_add_folio(lruvec, folio); > + /* > + * Put folios that may have missed folio_rotate_reclaimable() > + * at the tail to avoid cold/hot inversion. > + */ > + if (do_rotate && !folio_test_active(folio) && !folio_mapped(folio) && > + !folio_test_dirty(folio) && !folio_test_writeback(folio)) > + lruvec_add_folio_tail(lruvec, folio); > + else > + lruvec_add_folio(lruvec, folio); > + > nr_pages = folio_nr_pages(folio); > nr_moved += nr_pages; > if (folio_test_active(folio)) > @@ -2135,7 +2144,7 @@ static unsigned long shrink_inactive_list(unsigned long nr_to_scan, > nr_reclaimed = shrink_folio_list(&folio_list, pgdat, sc, &stat, false, > lruvec_memcg(lruvec)); > > - move_folios_to_lru(&folio_list); > + move_folios_to_lru(&folio_list, true); > > mod_lruvec_state(lruvec, PGDEMOTE_KSWAPD + reclaimer_offset(sc), > stat.nr_demoted); > @@ -2246,8 +2255,8 @@ static void shrink_active_list(unsigned long nr_to_scan, > /* > * Move folios back to the lru list. > */ > - nr_activate = move_folios_to_lru(&l_active); > - nr_deactivate = move_folios_to_lru(&l_inactive); > + nr_activate = move_folios_to_lru(&l_active, false); > + nr_deactivate = move_folios_to_lru(&l_inactive, false); > > count_vm_events(PGDEACTIVATE, nr_deactivate); > count_memcg_events(lruvec_memcg(lruvec), PGDEACTIVATE, nr_deactivate); > @@ -5115,7 +5124,7 @@ static int evict_folios(unsigned long nr_to_scan, struct lruvec *lruvec, > folio_set_active(folio); > } > > - move_folios_to_lru(&list); > + move_folios_to_lru(&list, false); > > walk = current->reclaim_state->mm_walk; > if (walk && walk->batched) { -- Best regards Ridong