* [PATCH mm-unstable v4 1/2] mm/vmscan: fix missing NR_ISOLATED counter update in MGLRU reclaim path
2026-08-20 2:48 [PATCH mm-unstable v4 0/2] mm/vmscan: fix NR_ISOLATED accounting and throttling for MGLRU Hui Zhu
@ 2026-08-20 2:48 ` Hui Zhu
2026-08-20 2:48 ` [PATCH mm-unstable v4 2/2] mm/vmscan: apply too_many_isolated() throttling to MGLRU eviction Hui Zhu
2026-08-20 9:03 ` [PATCH mm-unstable v4 0/2] mm/vmscan: fix NR_ISOLATED accounting and throttling for MGLRU Kairui Song
2 siblings, 0 replies; 4+ messages in thread
From: Hui Zhu @ 2026-08-20 2:48 UTC (permalink / raw)
To: Andrew Morton, Kairui Song, Qi Zheng, Shakeel Butt, Barry Song,
Axel Rasmussen, Yuanchu Xie, Wei Xu, Johannes Weiner,
David Hildenbrand, Michal Hocko, Lorenzo Stoakes, Baolin Wang,
linux-mm, linux-kernel
Cc: Hui Zhu
From: Hui Zhu <zhuhui@kylinos.cn>
MGLRU evict_folios() isolates folios from the LRU without updating
the NR_ISOLATED_ANON/FILE counters, unlike the legacy
shrink_inactive_list() path. This causes compaction's
too_many_isolated() check to under-count isolated pages when MGLRU
reclaim is active.
Add NR_ISOLATED counter updates in evict_folios(): increment after
isolate_folios() and decrement after all retry passes complete, using
the existing nr_isolated which holds the original isolated count.
Signed-off-by: Hui Zhu <zhuhui@kylinos.cn>
Reviewed-by: Baolin Wang <baolin.wang@linux.alibaba.com>
Reviewed-by: Barry Song <baohua@kernel.org>
---
mm/vmscan.c | 5 +++++
1 file changed, 5 insertions(+)
diff --git a/mm/vmscan.c b/mm/vmscan.c
index c1404a59523d..98226bb021f3 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -4892,6 +4892,9 @@ static int evict_folios(unsigned long nr_to_scan, struct lruvec *lruvec,
scanned = isolate_folios(nr_to_scan, lruvec, sc, swappiness,
&list, &isolated, &type, &type_scanned);
nr_isolated = isolated;
+ if (nr_isolated)
+ __mod_node_page_state(pgdat, NR_ISOLATED_ANON + type,
+ nr_isolated);
/* Scanning may have emptied the oldest gen, flush it */
if (scanned)
@@ -4954,6 +4957,8 @@ static int evict_folios(unsigned long nr_to_scan, struct lruvec *lruvec,
goto retry;
}
+ mod_node_page_state(pgdat, NR_ISOLATED_ANON + type, -nr_isolated);
+
if (nr_isolated > total_reclaimed)
mod_lruvec_state(lruvec, PGROTATE_ANON + type,
nr_isolated - total_reclaimed);
--
2.53.0
^ permalink raw reply [flat|nested] 4+ messages in thread
* [PATCH mm-unstable v4 2/2] mm/vmscan: apply too_many_isolated() throttling to MGLRU eviction
2026-08-20 2:48 [PATCH mm-unstable v4 0/2] mm/vmscan: fix NR_ISOLATED accounting and throttling for MGLRU Hui Zhu
2026-08-20 2:48 ` [PATCH mm-unstable v4 1/2] mm/vmscan: fix missing NR_ISOLATED counter update in MGLRU reclaim path Hui Zhu
@ 2026-08-20 2:48 ` Hui Zhu
2026-08-20 9:03 ` [PATCH mm-unstable v4 0/2] mm/vmscan: fix NR_ISOLATED accounting and throttling for MGLRU Kairui Song
2 siblings, 0 replies; 4+ messages in thread
From: Hui Zhu @ 2026-08-20 2:48 UTC (permalink / raw)
To: Andrew Morton, Kairui Song, Qi Zheng, Shakeel Butt, Barry Song,
Axel Rasmussen, Yuanchu Xie, Wei Xu, Johannes Weiner,
David Hildenbrand, Michal Hocko, Lorenzo Stoakes, Baolin Wang,
linux-mm, linux-kernel
Cc: Hui Zhu
From: Hui Zhu <zhuhui@kylinos.cn>
The legacy path throttles direct reclaim in shrink_inactive_list() when
too many isolated folios pile up, but MGLRU's evict_folios() isolates
folios without this check, which can lead to unnecessary swapping,
thrashing and OOM.
With the NR_ISOLATED counters now updated in evict_folios(), add
throttle_evictable_types() and call it from evict_folios(), before the
lruvec lock is taken since throttling sleeps. The legacy path is left
untouched.
Unlike the legacy path, where the LRU list to isolate from is known
before isolation, isolate_folios() picks the type to scan from the
refault feedback and may fall back to the other one. Therefore, instead
of throttling on a single type, throttle_evictable_types() collects the
evictable types that do not have too many isolated folios, and only
sleeps when all of them do: like the legacy path, it waits once for
concurrent reclaimers to put their isolated folios back, and gives up
if that makes no progress.
The resulting mask of the types that are not over-isolated is passed to
isolate_folios(), which restricts both its initial choice and its
fallback to it. This way a type that is merely over-isolated never
blocks the reclaim of the other type, and isolation never lands on a
throttled type.
If a fatal signal is pending, fake reclaim progress the same way the
legacy path does, so the dying task exits reclaim quickly instead of
being held in the throttle.
Signed-off-by: Hui Zhu <zhuhui@kylinos.cn>
---
mm/vmscan.c | 85 +++++++++++++++++++++++++++++++++++++++++++++++++----
1 file changed, 80 insertions(+), 5 deletions(-)
diff --git a/mm/vmscan.c b/mm/vmscan.c
index 98226bb021f3..693dc91a1695 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -4833,15 +4833,73 @@ static int get_type_to_scan(struct lruvec *lruvec, int swappiness)
return positive_ctrl_err(&sp, &pv);
}
+/*
+ * Unlike the legacy path, where the LRU list to isolate from is known
+ * before isolation, isolate_folios() picks the type from the refault
+ * feedback and may fall back to the other one. Therefore, instead of
+ * throttling on a single type, collect the evictable types that do not
+ * have too many isolated folios, and only sleep when all of them do.
+ *
+ * Returns the mask of the types isolate_folios() may isolate from, or
+ * 0 if reclaim should stop. Also sets @fatal to tell the caller that
+ * the task received a fatal signal while waiting.
+ */
+static unsigned int throttle_evictable_types(struct pglist_data *pgdat,
+ int swappiness,
+ struct scan_control *sc,
+ bool *fatal)
+{
+ unsigned int allowed;
+ bool stalled = false;
+ int i;
+
+ *fatal = false;
+
+ for (;;) {
+ allowed = 0;
+ for_each_evictable_type(i, swappiness) {
+ if (!too_many_isolated(pgdat, i, sc))
+ allowed |= BIT(i);
+ }
+
+ if (allowed)
+ return allowed;
+
+ /*
+ * All evictable types are over-isolated. Like the legacy
+ * path, wait once for concurrent reclaimers to put their
+ * isolated folios back; give up if that makes no progress.
+ */
+ if (stalled)
+ return 0;
+
+ stalled = true;
+ reclaim_throttle(pgdat, VMSCAN_THROTTLE_ISOLATED);
+
+ /* We are about to die and free our memory. Return now. */
+ if (fatal_signal_pending(current)) {
+ *fatal = true;
+ return 0;
+ }
+ }
+}
+
static int isolate_folios(unsigned long nr_to_scan, struct lruvec *lruvec,
struct scan_control *sc, int swappiness,
- struct list_head *list, int *isolated,
- int *isolate_type, int *isolate_scanned)
+ unsigned int allowed, struct list_head *list,
+ int *isolated, int *isolate_type, int *isolate_scanned)
{
int i;
int total_scanned = 0;
int type = get_type_to_scan(lruvec, swappiness);
+ /*
+ * The preferred type may have been excluded by
+ * throttle_evictable_types(); start from the other one.
+ */
+ if (!(allowed & BIT(type)))
+ type = !type;
+
for_each_evictable_type(i, swappiness) {
int scanned;
int tier = get_tier_idx(lruvec, type);
@@ -4858,9 +4916,10 @@ static int isolate_folios(unsigned long nr_to_scan, struct lruvec *lruvec,
/*
* If scanned > 0 and isolated == 0, avoid falling back to the
* other type, as this type remains sufficient. Falling back
- * too readily can disrupt the positive_ctrl_err() bias.
+ * too readily can disrupt the positive_ctrl_err() bias. Only
+ * fall back to a type that is not throttled.
*/
- if (!scanned)
+ if (!scanned && (allowed & BIT(!type)))
type = !type;
}
@@ -4883,13 +4942,29 @@ static int evict_folios(unsigned long nr_to_scan, struct lruvec *lruvec,
bool skip_retry = false;
struct mem_cgroup *memcg = lruvec_memcg(lruvec);
struct pglist_data *pgdat = lruvec_pgdat(lruvec);
+ unsigned int allowed;
+ bool fatal;
+
+ allowed = throttle_evictable_types(pgdat, swappiness, sc, &fatal);
+ if (!allowed) {
+ /*
+ * We are about to die and free our memory. Like the legacy
+ * path, pretend some pages were reclaimed so reclaim
+ * unwinds quickly instead of looping back into the
+ * throttle.
+ */
+ if (fatal)
+ sc->nr_reclaimed += SWAP_CLUSTER_MAX;
+
+ return 0;
+ }
lruvec_lock_irq(lruvec);
/* In case folio deletion left empty old gens, flush them */
try_to_inc_min_seq(lruvec, swappiness);
- scanned = isolate_folios(nr_to_scan, lruvec, sc, swappiness,
+ scanned = isolate_folios(nr_to_scan, lruvec, sc, swappiness, allowed,
&list, &isolated, &type, &type_scanned);
nr_isolated = isolated;
if (nr_isolated)
--
2.53.0
^ permalink raw reply [flat|nested] 4+ messages in thread* Re: [PATCH mm-unstable v4 0/2] mm/vmscan: fix NR_ISOLATED accounting and throttling for MGLRU
2026-08-20 2:48 [PATCH mm-unstable v4 0/2] mm/vmscan: fix NR_ISOLATED accounting and throttling for MGLRU Hui Zhu
2026-08-20 2:48 ` [PATCH mm-unstable v4 1/2] mm/vmscan: fix missing NR_ISOLATED counter update in MGLRU reclaim path Hui Zhu
2026-08-20 2:48 ` [PATCH mm-unstable v4 2/2] mm/vmscan: apply too_many_isolated() throttling to MGLRU eviction Hui Zhu
@ 2026-08-20 9:03 ` Kairui Song
2 siblings, 0 replies; 4+ messages in thread
From: Kairui Song @ 2026-08-20 9:03 UTC (permalink / raw)
To: Hui Zhu
Cc: Andrew Morton, Qi Zheng, Shakeel Butt, Barry Song,
Axel Rasmussen, Yuanchu Xie, Wei Xu, Johannes Weiner,
David Hildenbrand, Michal Hocko, Lorenzo Stoakes, Baolin Wang,
linux-mm, linux-kernel, Hui Zhu
On Thu, Aug 20, 2026 at 10:49 AM Hui Zhu <hui.zhu@linux.dev> wrote:
>
> From: Hui Zhu <zhuhui@kylinos.cn>
Hello, thanks for the update!
>
> The legacy reclaim path updates the NR_ISOLATED_ANON/FILE node
> counters around isolation and throttles direct reclaimers via
> too_many_isolated() when isolated folios pile up. The MGLRU eviction
> path does neither: evict_folios() isolates folios without touching
> the counters and never consults too_many_isolated().
>
> Patch 1 updates NR_ISOLATED_ANON/FILE around isolation in
> evict_folios(), reusing the existing nr_isolated. Without this the
> counters stay at zero while MGLRU reclaim is active, so compaction's
> too_many_isolated() cannot see the pages MGLRU has isolated.
>
> Patch 2 adds throttle_evictable_types() and calls it from
> evict_folios(), before the lruvec lock is taken since throttling
> sleeps, leaving the legacy path untouched. The MGLRU check differs
> from the legacy per-list one in shrink_inactive_list() because
> isolate_folios() picks the type to scan from the refault feedback
> and may fall back to the other one: it computes the set of evictable
> types that are not over-isolated and only sleeps when all of them
> are, waiting once for concurrent reclaimers exactly like the legacy
> path. The mask of the remaining types is passed to isolate_folios(),
> which restricts both its initial choice and its fallback to it, so
> isolation never lands on an over-isolated type and a type that is
> merely over-isolated never blocks the reclaim of the other one.
> This way the MGLRU eviction path backs off when isolated folios pile
> up instead of thrashing the shrinking LRU lists - the scenario the
> too_many_isolated() check exists for. A dying task fakes reclaim
> progress exactly like the legacy path so it exits reclaim quickly.
Hmm, could too_many_isolated gets over aggressive or over passive,
since the inactive number of MGLRU is just a compatiblity shim and
does not have the same meaning of classical LRU? Especially when you
run out of swap space or hit a memcg's swap limit, the anon inactive
number becomes a jumpy random number. Proactive aging also makes the
inactive number become huge for MGLRU. I still think even a "/
MIN_NR_GENS" is better than using inactive value, MGLRU used to use
that as the aging trigger like this:
if (young * MIN_NR_GENS > total)
return true;
if (old * (MIN_NR_GENS + 2) < total)
return true;
>
> Testing
> =======
>
> Test on 8G RAM qemu.
> The reproducer confines stress-ng workers in a 192M memcg and swaps
> through dm-delay (300ms write latency) so pageout is slow and isolated
> folios pile up; the workload is intentionally extreme. nr_isolated_*
> is sampled every 50ms against the per-type too_many_isolated
> threshold (inactive/8), and throttle events are counted via the
> mm_vmscan_throttled tracepoint.
> The test scripts and test log are in [1].
>
> Test 1, reclaim throttling, parallel direct reclaim in the memcg:
>
> before after
> throttle events (ISOLATED) 0 0
> - from kswapd 0 0
> nr_isolated_anon peak 0 3166
> nr_isolated_file peak 0 174
> - samples above the
> too_many_isolated threshold 0/1088 89/1077
> pgscan_direct 1540044096 858332151
> pswpout 6299497 725819
The throttling avoids premature OOM and over scan when there are too
many reclaimers, it is not for IO balancing. Because writeback folios
are not isolated, they are putback to LRU waiting to be rotated, not
isolated. So reclaimers never wait on the device however slow the
device is (except cgroup V1 have some special quirks I think we should
ignore), it's mostly CPU bound. I think the mechanism might be a bit
outdated, even for classical LRU.
The result just shows that swap became less aggressive when under
pressure, which is somewhat counter-intuitive. Overly aggressive
throttling will unnecessarily slows performance if anon pages are
actually cold or should be evicted. These results reflect a reclaim
behavior change, not a performance improvement. And I think we
shouldn't rely on throttling for balancing, those are two different
things. Heat or I/O cost-based methods are better.
For example the test you posted, it just stress-ng which only spawns
anon memory IIUC, so when you reclaim file, your code segments got
reclaimed and hence the process stall and has to wait for code
segments IO, which I think it will only slowdown the actual
performance?
^ permalink raw reply [flat|nested] 4+ messages in thread