From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from out30-132.freemail.mail.aliyun.com (out30-132.freemail.mail.aliyun.com [115.124.30.132]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E1D2D3E0C66 for ; Tue, 31 Mar 2026 09:36:36 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=115.124.30.132 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1774949800; cv=none; b=HNR/j+sOWHAIoRlMFSAuur1R9GqjYz0M2JhI2GRHe6HW/wnWpRFWHG6COLfxmdF9inT9Mlz/upjcbTcIO83zQcxjP3A3Cl+ti92b0Z8B4T/wdhvthuZSOyMiN7tSoV6Zw1TSLAqNJhChoXAiI3fHex7rZRx+MRMIM1enlPCTFW8= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1774949800; c=relaxed/simple; bh=zUM18059upuSoEjpVBJxuoDi6Ah4shPFirFprKA3YD8=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=hCco7WuKaOJNYj+ksN1M+DoCqBjwL88pUD8N7H+0UKD/clDTe3nguAW7+XNeAy3/qui91S4O79tYgr5Z6A599iGwP83mvW4RcZCHcIFQuyumJxy4xCGIIM/W0qJ0HzET4Er+Pa3ZaliYsCHi4k3MTO8Di6/V+WEvE876WboBNEM= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.alibaba.com; spf=pass smtp.mailfrom=linux.alibaba.com; dkim=pass (1024-bit key) header.d=linux.alibaba.com header.i=@linux.alibaba.com header.b=rNYSbd60; arc=none smtp.client-ip=115.124.30.132 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.alibaba.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.alibaba.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.alibaba.com header.i=@linux.alibaba.com header.b="rNYSbd60" DKIM-Signature:v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux.alibaba.com; s=default; t=1774949789; h=Message-ID:Date:MIME-Version:Subject:To:From:Content-Type; bh=VkJVBs8ELvWsSuKQ+1IMOp2IRxZTSWkSwM2U/2pFHug=; b=rNYSbd605J+ap3j7WIWD+yFtM8SQmRIg/UDg2dOzLhJO2LDxA7hEAB725NTONWaVEbEeh3zG6CV4AlVkhFBuyZDdl3k74M90/FnlJjorGCKmhREQUbKtxvKAfEWVGMcXsGpi5FYxiGItPX8+mDC/3Z3UK5HP0DLvC2ipvtkTMtg= X-Alimail-AntiSpam:AC=PASS;BC=-1|-1;BR=01201311R501e4;CH=green;DM=||false|;DS=||;FP=0|-1|-1|-1|0|-1|-1|-1;HT=maildocker-contentspam033037009110;MF=baolin.wang@linux.alibaba.com;NM=1;PH=DS;RN=26;SR=0;TI=SMTPD_---0X03yfwB_1774949786; Received: from 30.74.144.129(mailfrom:baolin.wang@linux.alibaba.com fp:SMTPD_---0X03yfwB_1774949786 cluster:ay36) by smtp.aliyun-inc.com; Tue, 31 Mar 2026 17:36:27 +0800 Message-ID: <522f4898-78c3-453f-8367-29327e29290e@linux.alibaba.com> Date: Tue, 31 Mar 2026 17:36:26 +0800 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH v2 12/12] mm/vmscan: unify writeback reclaim statistic and throttling To: Kairui Song Cc: kasong@tencent.com, linux-mm@kvack.org, Andrew Morton , Axel Rasmussen , Yuanchu Xie , Wei Xu , Johannes Weiner , David Hildenbrand , Michal Hocko , Qi Zheng , Shakeel Butt , Lorenzo Stoakes , Barry Song , David Stevens , Chen Ridong , Leno Hou , Yafang Shao , Yu Zhao , Zicheng Wang , Kalesh Singh , Suren Baghdasaryan , Chris Li , Vernon Yang , linux-kernel@vger.kernel.org, Qi Zheng References: <20260329-mglru-reclaim-v2-0-b53a3678513c@tencent.com> <20260329-mglru-reclaim-v2-12-b53a3678513c@tencent.com> <052ae271-509c-42c3-877e-ac8822b314e5@linux.alibaba.com> From: Baolin Wang In-Reply-To: Content-Type: text/plain; charset=UTF-8; format=flowed Content-Transfer-Encoding: 8bit On 3/31/26 5:29 PM, Kairui Song wrote: > On Tue, Mar 31, 2026 at 05:24:39PM +0800, Baolin Wang wrote: >> >> >> On 3/29/26 3:52 AM, Kairui Song via B4 Relay wrote: >>> From: Kairui Song >>> >>> Currently MGLRU and non-MGLRU handle the reclaim statistic and >>> writeback handling very differently, especially throttling. >>> Basically MGLRU just ignored the throttling part. >>> >>> Let's just unify this part, use a helper to deduplicate the code >>> so both setups will share the same behavior. Also remove the >>> folio_clear_reclaim in isolate_folio which was actively invalidating >>> the congestion control. PG_reclaim is now handled by shrink_folio_list, >>> keeping it in isolate_folio is not helpful. >>> >>> Test using following reproducer using bash: >>> >>> echo "Setup a slow device using dm delay" >>> dd if=/dev/zero of=/var/tmp/backing bs=1M count=2048 >>> LOOP=$(losetup --show -f /var/tmp/backing) >>> mkfs.ext4 -q $LOOP >>> echo "0 $(blockdev --getsz $LOOP) delay $LOOP 0 0 $LOOP 0 1000" | \ >>> dmsetup create slow_dev >>> mkdir -p /mnt/slow && mount /dev/mapper/slow_dev /mnt/slow >>> >>> echo "Start writeback pressure" >>> sync && echo 3 > /proc/sys/vm/drop_caches >>> mkdir /sys/fs/cgroup/test_wb >>> echo 128M > /sys/fs/cgroup/test_wb/memory.max >>> (echo $BASHPID > /sys/fs/cgroup/test_wb/cgroup.procs && \ >>> dd if=/dev/zero of=/mnt/slow/testfile bs=1M count=192) >>> >>> echo "Clean up" >>> echo "0 $(blockdev --getsz $LOOP) error" | dmsetup load slow_dev >>> dmsetup resume slow_dev >>> umount -l /mnt/slow && sync >>> dmsetup remove slow_dev >>> >>> Before this commit, `dd` will get OOM killed immediately if >>> MGLRU is enabled. Classic LRU is fine. >>> >>> After this commit, congestion control is now effective and no more >>> spin on LRU or premature OOM. >>> >>> Stress test on other workloads also looking good. >>> >>> Suggested-by: Chen Ridong >>> Signed-off-by: Kairui Song >>> --- >>> mm/vmscan.c | 93 +++++++++++++++++++++++++++---------------------------------- >>> 1 file changed, 41 insertions(+), 52 deletions(-) >>> >>> diff --git a/mm/vmscan.c b/mm/vmscan.c >>> index 1783da54ada1..83c8fdf8fdc4 100644 >>> --- a/mm/vmscan.c >>> +++ b/mm/vmscan.c >>> @@ -1942,6 +1942,44 @@ static int current_may_throttle(void) >>> return !(current->flags & PF_LOCAL_THROTTLE); >>> } >>> +static void handle_reclaim_writeback(unsigned long nr_taken, >>> + struct pglist_data *pgdat, >>> + struct scan_control *sc, >>> + struct reclaim_stat *stat) >>> +{ >>> + /* >>> + * If dirty folios are scanned that are not queued for IO, it >>> + * implies that flushers are not doing their job. This can >>> + * happen when memory pressure pushes dirty folios to the end of >>> + * the LRU before the dirty limits are breached and the dirty >>> + * data has expired. It can also happen when the proportion of >>> + * dirty folios grows not through writes but through memory >>> + * pressure reclaiming all the clean cache. And in some cases, >>> + * the flushers simply cannot keep up with the allocation >>> + * rate. Nudge the flusher threads in case they are asleep. >>> + */ >>> + if (stat->nr_unqueued_dirty == nr_taken && nr_taken) { >>> + wakeup_flusher_threads(WB_REASON_VMSCAN); >>> + /* >>> + * For cgroupv1 dirty throttling is achieved by waking up >>> + * the kernel flusher here and later waiting on folios >>> + * which are in writeback to finish (see shrink_folio_list()). >>> + * >>> + * Flusher may not be able to issue writeback quickly >>> + * enough for cgroupv1 writeback throttling to work >>> + * on a large system. >>> + */ >>> + if (!writeback_throttling_sane(sc)) >>> + reclaim_throttle(pgdat, VMSCAN_THROTTLE_WRITEBACK); >>> + } >>> + >>> + sc->nr.dirty += stat->nr_dirty; >>> + sc->nr.congested += stat->nr_congested; >>> + sc->nr.writeback += stat->nr_writeback; >>> + sc->nr.immediate += stat->nr_immediate; >>> + sc->nr.taken += nr_taken; >>> +} >>> + >>> /* >>> * shrink_inactive_list() is a helper for shrink_node(). It returns the number >>> * of reclaimed pages >>> @@ -2005,39 +2043,7 @@ static unsigned long shrink_inactive_list(unsigned long nr_to_scan, >>> lruvec_lock_irq(lruvec); >>> lru_note_cost_unlock_irq(lruvec, file, stat.nr_pageout, >>> nr_scanned - nr_reclaimed); >>> - >>> - /* >>> - * If dirty folios are scanned that are not queued for IO, it >>> - * implies that flushers are not doing their job. This can >>> - * happen when memory pressure pushes dirty folios to the end of >>> - * the LRU before the dirty limits are breached and the dirty >>> - * data has expired. It can also happen when the proportion of >>> - * dirty folios grows not through writes but through memory >>> - * pressure reclaiming all the clean cache. And in some cases, >>> - * the flushers simply cannot keep up with the allocation >>> - * rate. Nudge the flusher threads in case they are asleep. >>> - */ >>> - if (stat.nr_unqueued_dirty == nr_taken) { >>> - wakeup_flusher_threads(WB_REASON_VMSCAN); >>> - /* >>> - * For cgroupv1 dirty throttling is achieved by waking up >>> - * the kernel flusher here and later waiting on folios >>> - * which are in writeback to finish (see shrink_folio_list()). >>> - * >>> - * Flusher may not be able to issue writeback quickly >>> - * enough for cgroupv1 writeback throttling to work >>> - * on a large system. >>> - */ >>> - if (!writeback_throttling_sane(sc)) >>> - reclaim_throttle(pgdat, VMSCAN_THROTTLE_WRITEBACK); >>> - } >>> - >>> - sc->nr.dirty += stat.nr_dirty; >>> - sc->nr.congested += stat.nr_congested; >>> - sc->nr.writeback += stat.nr_writeback; >>> - sc->nr.immediate += stat.nr_immediate; >>> - sc->nr.taken += nr_taken; >>> - >>> + handle_reclaim_writeback(nr_taken, pgdat, sc, &stat); >>> trace_mm_vmscan_lru_shrink_inactive(pgdat->node_id, >>> nr_scanned, nr_reclaimed, &stat, sc->priority, file); >>> return nr_reclaimed; >>> @@ -4651,9 +4657,6 @@ static bool isolate_folio(struct lruvec *lruvec, struct folio *folio, struct sca >>> if (!folio_test_referenced(folio)) >>> set_mask_bits(&folio->flags.f, LRU_REFS_MASK, 0); >>> - /* for shrink_folio_list() */ >>> - folio_clear_reclaim(folio); >> >> IMO, Moving this change into patch 8 would make more sense. Otherwise LGTM. > > Thanks for the review! I made it a separate patch so we can better > identify which part had the performance gain, and patch 8 can keep > the review by. Patch 8 is still good without this, a few counters > are updated with no user, kind of wasted but that's harmless. I’m not referring to all the above changes. What I mean is that the 'folio_clear_reclaim' removal should belong to patch 8. Since shrink_folio_list() in patch 8 will handle the writeback logic, folio_clear_reclaim() should also be removed in the same patch.