* [PATCH resend 0/2] mm: vmscan: fix scan overshoot and ineligible folio scanning
@ 2026-09-01 8:47 john
2026-09-01 8:47 ` [PATCH resend 1/2] mm: vmscan: charge isolate overshoot against scan quota john
` (3 more replies)
0 siblings, 4 replies; 7+ messages in thread
From: john @ 2026-09-01 8:47 UTC (permalink / raw)
To: akpm, liuye, hannes, mhocko, david, ljs, hughd, mgorman, yang
Cc: zhangqiuhao, wangkefeng.wang, mawupeng1, love_goo, linux-mm,
linux-kernel, Wupeng Ma
From: Wupeng Ma <mawupeng1@huawei.com>
Rebase to the latest v7.3-rc-1.
These problems only surface when reclaim targets a lower zone
while the LRU holds folios from a higher zone. Normal userspace
allocations go to the highest zone. The zone-skip branch stays
dead under typical loads. Lower-zone-pressured configs (DMA32
module allocations, memory-constrained devices) hit the issues.
They inflate scan cost and delay the OOM.
shrink_lruvec() drives reclaim in SWAP_CLUSTER_MAX (32) chunks, but
isolate_lru_folios() may scan far more than that per call on a single
LRU. The excess is never charged back, so shrink_lruvec() keeps
rescanning the same folios round after round. When reclaim targets a
lower zone, the same scanner also keeps walking zone-ineligible folios
that can never satisfy the allocation, inflating nr_reclaimed into a
false progress that delays the OOM.
This series fixes both:
[1/2] Charge the isolate overshoot against the scan quota so the
next round skips already-scanned folios.
[2/2] Stop scanning once too many zone-ineligible folios have been
skipped, instead of force-isolating them.
Background
==========
We observed slow, unexpected OOM behavior during extreme stress testing,
and while digging into the reclaim code during that analysis we spotted
these latent risks in isolate_lru_folios() -- the overshoot never
being charged back, and the force-isolate path on lower-zone reclaim.
The concerns below are the ones surfaced from reading the code, then
confirmed by constructing the situation deliberately.
The problem only appears when reclaim targets a lower zone while the
LRU holds folios from a higher zone:
- reclaim_idx points at DMA32/DMA (the triggering allocation asked
for a lower zone, e.g. __GFP_DMA32), and
- the LRU holds folios from a higher zone (Normal/Movable).
Normal userspace allocations go to the highest zone, so reclaim_idx
never points below it and the zone-skip branch is never taken. It
needs a real lower-zone allocator to drain that zone below watermark;
that is also why it went unnoticed upstream -- the 1c7b17cf hard-lockup
fix that introduced the force-isolate path was found only on a ~1 TB
box running DMA32 module allocations.
Impact
======
Reproduced on x86 QEMU, 7.2-rc6, with a kernel module doing
__GFP_DMA32 allocations to drain DMA32 below watermark (LRU folios
sit in Movable):
- A single isolate_lru_folios() call scanned 32794 pages and took
25 (32769 skipped): 99.9% wasted on ineligible folios.
- Across one run, isolate was called 339 times, 140-165 of which
isolated nothing (taken=0, pure empty scans).
- shrink_lruvec() charges only the 32-page quota per round and
never refunds the overshoot, so the same skipped folios are
rescanned round after round. vmstat on a memcg OOM path:
pgscan_direct / pgsteal_direct = 2836624 / 112719 = 25.2x
(isolate -> shrink_folio_list returns the folio -> isolate again).
Once max_nr_skipped hits SWAP_CLUSTER_MAX_SKIPPED, the current code
force-isolates the remaining ineligible folios. On the inactive LRU
they reach shrink_folio_list() and get reclaimed though they can
never satisfy the allocation, inflating nr_reclaimed and resetting
no_progress_loops in should_reclaim_retry(), delaying the OOM.
A/B results (same .config, md5-identical):
baseline patched
empty scans (taken=0) 140-165 0
isolate calls 339 7
total pages scanned 343711 3140
single-call max scan 32794 3104
Empty-scan 0 is the stable evidence (holds every run). The
pgscan/pgsteal ratio is volatile (baseline 56-26675x, patched
65-1324x, ranges overlap) and is not relied on alone.
Caveats and reproduction
========================
- Triggering needs a lower-zone-pressured box. To confirm the code
analysis, the situation was constructed on a small x86 QEMU VM
(1500M, CONFIG_LRU_GEN=n) with kernel cmdline
`movable_zone=DMA32 kernelcore=256M` (DMA32 small, Normal empty,
Movable large), then a kernel module doing
`alloc_page(GFP_DMA32)` drains DMA32 below watermark. LRU folios
sit in Movable and are zone-skipped while reclaiming for DMA32.
Observed via the `mm_vmscan_lru_isolate` tracepoint and
/proc/vmstat (pgscan_direct, pgsteal_direct).
Wupeng Ma (2):
mm: vmscan: charge isolate overshoot against scan quota
mm: vmscan: stop scanning ineligible folios after max_nr_skipped
mm/vmscan.c | 48 +++++++++++++++++++++++++++++++++---------------
1 file changed, 33 insertions(+), 15 deletions(-)
--
2.53.0
^ permalink raw reply [flat|nested] 7+ messages in thread* [PATCH resend 1/2] mm: vmscan: charge isolate overshoot against scan quota 2026-09-01 8:47 [PATCH resend 0/2] mm: vmscan: fix scan overshoot and ineligible folio scanning john @ 2026-09-01 8:47 ` john 2026-09-09 9:52 ` Kunwu Chan 2026-09-01 8:47 ` [PATCH resend 2/2] mm: vmscan: stop scanning ineligible folios after max_nr_skipped john ` (2 subsequent siblings) 3 siblings, 1 reply; 7+ messages in thread From: john @ 2026-09-01 8:47 UTC (permalink / raw) To: akpm, liuye, hannes, mhocko, david, ljs, hughd, mgorman, yang Cc: zhangqiuhao, wangkefeng.wang, mawupeng1, love_goo, linux-mm, linux-kernel, Wupeng Ma From: Wupeng Ma <mawupeng1@huawei.com> shrink_lruvec() charges the per-LRU budget nr[lru] in SWAP_CLUSTER_MAX (32) chunks, but isolate_lru_folios() may scan far more per call: a large folio can jump scan by many pages at once (a PMD-sized folio counts 512), and a zone-ineligible LRU walks the whole list without feeding scan back. The overshoot is never refunded, so shrink_lruvec() keeps charging only 32 per round and rescans the same folios. Have isolate_lru_folios() record its scanned count in sc->nr_isolate_scanned and let shrink_lruvec() subtract the overshoot from the remaining quota so the next round skips already-scanned folios. The field is reset to 0 before each shrink_list() call, as shrink_list() only reaches isolate_lru_folios() on some paths (active + skipped_deactivate, too_many_isolated stall bail out early); a stale value would otherwise be charged. The budget floor stays nr_to_scan via max() so an empty LRU (sc->nr_isolate_scanned = 0) still advances and cannot deadlock. Co-developed-by: Qiuhao Zhang <zhangqiuhao@huawei.com> Signed-off-by: Qiuhao Zhang <zhangqiuhao@huawei.com> Signed-off-by: Wupeng Ma <mawupeng1@huawei.com> --- mm/vmscan.c | 32 +++++++++++++++++++++----------- 1 file changed, 21 insertions(+), 11 deletions(-) diff --git a/mm/vmscan.c b/mm/vmscan.c index f11491ee9ed5c..823af9e86efd3 100644 --- a/mm/vmscan.c +++ b/mm/vmscan.c @@ -166,6 +166,9 @@ struct scan_control { /* Incremented by the number of inactive pages that were scanned */ unsigned long nr_scanned; + /* Number of pages that were scanned from isolate_lru_folios() */ + unsigned long nr_isolate_scanned; + /* Number of pages freed so far during a call to shrink_zones() */ unsigned long nr_reclaimed; @@ -1670,7 +1673,6 @@ static __always_inline void update_lru_sizes(struct lruvec *lruvec, * @nr_to_scan: The number of eligible pages to look through on the list. * @lruvec: The LRU vector to pull pages from. * @dst: The temp list to put pages on to. - * @nr_scanned: The number of pages that were scanned. * @sc: The scan_control struct for this reclaim session * @lru: LRU list id for isolating * @@ -1678,8 +1680,7 @@ static __always_inline void update_lru_sizes(struct lruvec *lruvec, */ static unsigned long isolate_lru_folios(unsigned long nr_to_scan, struct lruvec *lruvec, struct list_head *dst, - unsigned long *nr_scanned, struct scan_control *sc, - enum lru_list lru) + struct scan_control *sc, enum lru_list lru) { struct list_head *src = &lruvec->lists[lru]; unsigned long nr_taken = 0; @@ -1763,7 +1764,7 @@ static unsigned long isolate_lru_folios(unsigned long nr_to_scan, skipped += nr_skipped[zid]; } } - *nr_scanned = total_scan; + sc->nr_isolate_scanned = total_scan; trace_mm_vmscan_lru_isolate(sc->reclaim_idx, sc->order, nr_to_scan, total_scan, skipped, nr_taken, lru); update_lru_sizes(lruvec, lru, nr_zone_taken); @@ -2011,8 +2012,8 @@ static unsigned long shrink_inactive_list(unsigned long nr_to_scan, lruvec_lock_irq(lruvec); - nr_taken = isolate_lru_folios(nr_to_scan, lruvec, &folio_list, - &nr_scanned, sc, lru); + nr_taken = isolate_lru_folios(nr_to_scan, lruvec, &folio_list, sc, lru); + nr_scanned = sc->nr_isolate_scanned; __mod_node_page_state(pgdat, NR_ISOLATED_ANON + file, nr_taken); item = PGSCAN_KSWAPD + reclaimer_offset(sc); @@ -2068,7 +2069,6 @@ static void shrink_active_list(unsigned long nr_to_scan, enum lru_list lru) { unsigned long nr_taken; - unsigned long nr_scanned; vma_flags_t vma_flags; LIST_HEAD(l_hold); /* The folios which were snipped off */ LIST_HEAD(l_active); @@ -2082,12 +2082,11 @@ static void shrink_active_list(unsigned long nr_to_scan, lruvec_lock_irq(lruvec); - nr_taken = isolate_lru_folios(nr_to_scan, lruvec, &l_hold, - &nr_scanned, sc, lru); + nr_taken = isolate_lru_folios(nr_to_scan, lruvec, &l_hold, sc, lru); __mod_node_page_state(pgdat, NR_ISOLATED_ANON + file, nr_taken); - mod_lruvec_state(lruvec, PGREFILL, nr_scanned); + mod_lruvec_state(lruvec, PGREFILL, sc->nr_isolate_scanned); lruvec_unlock_irq(lruvec); @@ -6013,10 +6012,21 @@ static void shrink_lruvec(struct lruvec *lruvec, struct scan_control *sc) for_each_evictable_lru(lru) { if (nr[lru]) { nr_to_scan = min(nr[lru], SWAP_CLUSTER_MAX); - nr[lru] -= nr_to_scan; + sc->nr_isolate_scanned = 0; nr_reclaimed += shrink_list(lru, nr_to_scan, lruvec, sc); + /* + * isolate_lru_folios() may scan far more + * than nr_to_scan when the LRU holds + * ineligible folios (zone-skip) or large + * folios. Charge that overshoot against the + * remaining quota (clamped by min() so it + * cannot go negative) so the next iteration + * does not rescan the same skipped folios. + */ + nr[lru] -= min(nr[lru], + max(nr_to_scan, sc->nr_isolate_scanned)); } } -- 2.53.0 ^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: [PATCH resend 1/2] mm: vmscan: charge isolate overshoot against scan quota 2026-09-01 8:47 ` [PATCH resend 1/2] mm: vmscan: charge isolate overshoot against scan quota john @ 2026-09-09 9:52 ` Kunwu Chan 0 siblings, 0 replies; 7+ messages in thread From: Kunwu Chan @ 2026-09-09 9:52 UTC (permalink / raw) To: john Cc: Kunwu Chan, akpm, liuye, hannes, mhocko, david, ljs, hughd, mgorman, yang, zhangqiuhao, wangkefeng.wang, mawupeng1, linux-mm, linux-kernel, Wupeng Ma On Tue, 1 Sep 2026 16:47:05 +0800 john <love_goo@163.com> wrote: > From: Wupeng Ma <mawupeng1@huawei.com> > > shrink_lruvec() charges the per-LRU budget nr[lru] in SWAP_CLUSTER_MAX > (32) chunks, but isolate_lru_folios() may scan far more per call: a > large folio can jump scan by many pages at once (a PMD-sized folio > counts 512), and a zone-ineligible LRU walks the whole list without > feeding scan back. The overshoot is never refunded, so shrink_lruvec() > keeps charging only 32 per round and rescans the same folios. > > Have isolate_lru_folios() record its scanned count in sc->nr_isolate_scanned > and let shrink_lruvec() subtract the overshoot from the remaining quota so > the next round skips already-scanned folios. The field is reset to 0 I'm not sure this is true across shrink_lruvec() invocations. isolate_lru_folios() splices folios_skipped back to the head of the LRU, while get_scan_count() provides a fresh nr[lru] each time shrink_lruvec() is entered. Thus, charging sc->nr_isolate_scanned against nr[lru] appears to prevent repeated 32-page scans within the same shrink_lruvec() invocation, but the same ineligible folios can still be encountered again by a subsequent invocation. Is the intended fix specifically to avoid repeated scanning within one shrink_lruvec() invocation, or is there another mechanism that prevents these skipped folios from being rescanned by a subsequent shrink_lruvec() invocation? Thanks, KunWu > before each shrink_list() call, as shrink_list() only reaches > isolate_lru_folios() on some paths (active + skipped_deactivate, > too_many_isolated stall bail out early); a stale value would otherwise > be charged. The budget floor stays nr_to_scan via max() so an empty > LRU (sc->nr_isolate_scanned = 0) still advances and cannot deadlock. > > Co-developed-by: Qiuhao Zhang <zhangqiuhao@huawei.com> > Signed-off-by: Qiuhao Zhang <zhangqiuhao@huawei.com> > Signed-off-by: Wupeng Ma <mawupeng1@huawei.com> > --- > mm/vmscan.c | 32 +++++++++++++++++++++----------- > 1 file changed, 21 insertions(+), 11 deletions(-) > > diff --git a/mm/vmscan.c b/mm/vmscan.c > index f11491ee9ed5c..823af9e86efd3 100644 > --- a/mm/vmscan.c > +++ b/mm/vmscan.c > @@ -166,6 +166,9 @@ struct scan_control { > /* Incremented by the number of inactive pages that were scanned */ > unsigned long nr_scanned; > > + /* Number of pages that were scanned from isolate_lru_folios() */ > + unsigned long nr_isolate_scanned; > + > /* Number of pages freed so far during a call to shrink_zones() */ > unsigned long nr_reclaimed; > > @@ -1670,7 +1673,6 @@ static __always_inline void update_lru_sizes(struct lruvec *lruvec, > * @nr_to_scan: The number of eligible pages to look through on the list. > * @lruvec: The LRU vector to pull pages from. > * @dst: The temp list to put pages on to. > - * @nr_scanned: The number of pages that were scanned. > * @sc: The scan_control struct for this reclaim session > * @lru: LRU list id for isolating > * > @@ -1678,8 +1680,7 @@ static __always_inline void update_lru_sizes(struct lruvec *lruvec, > */ > static unsigned long isolate_lru_folios(unsigned long nr_to_scan, > struct lruvec *lruvec, struct list_head *dst, > - unsigned long *nr_scanned, struct scan_control *sc, > - enum lru_list lru) > + struct scan_control *sc, enum lru_list lru) > { > struct list_head *src = &lruvec->lists[lru]; > unsigned long nr_taken = 0; > @@ -1763,7 +1764,7 @@ static unsigned long isolate_lru_folios(unsigned long nr_to_scan, > skipped += nr_skipped[zid]; > } > } > - *nr_scanned = total_scan; > + sc->nr_isolate_scanned = total_scan; > trace_mm_vmscan_lru_isolate(sc->reclaim_idx, sc->order, nr_to_scan, > total_scan, skipped, nr_taken, lru); > update_lru_sizes(lruvec, lru, nr_zone_taken); > @@ -2011,8 +2012,8 @@ static unsigned long shrink_inactive_list(unsigned long nr_to_scan, > > lruvec_lock_irq(lruvec); > > - nr_taken = isolate_lru_folios(nr_to_scan, lruvec, &folio_list, > - &nr_scanned, sc, lru); > + nr_taken = isolate_lru_folios(nr_to_scan, lruvec, &folio_list, sc, lru); > + nr_scanned = sc->nr_isolate_scanned; > > __mod_node_page_state(pgdat, NR_ISOLATED_ANON + file, nr_taken); > item = PGSCAN_KSWAPD + reclaimer_offset(sc); > @@ -2068,7 +2069,6 @@ static void shrink_active_list(unsigned long nr_to_scan, > enum lru_list lru) > { > unsigned long nr_taken; > - unsigned long nr_scanned; > vma_flags_t vma_flags; > LIST_HEAD(l_hold); /* The folios which were snipped off */ > LIST_HEAD(l_active); > @@ -2082,12 +2082,11 @@ static void shrink_active_list(unsigned long nr_to_scan, > > lruvec_lock_irq(lruvec); > > - nr_taken = isolate_lru_folios(nr_to_scan, lruvec, &l_hold, > - &nr_scanned, sc, lru); > + nr_taken = isolate_lru_folios(nr_to_scan, lruvec, &l_hold, sc, lru); > > __mod_node_page_state(pgdat, NR_ISOLATED_ANON + file, nr_taken); > > - mod_lruvec_state(lruvec, PGREFILL, nr_scanned); > + mod_lruvec_state(lruvec, PGREFILL, sc->nr_isolate_scanned); > > lruvec_unlock_irq(lruvec); > > @@ -6013,10 +6012,21 @@ static void shrink_lruvec(struct lruvec *lruvec, struct scan_control *sc) > for_each_evictable_lru(lru) { > if (nr[lru]) { > nr_to_scan = min(nr[lru], SWAP_CLUSTER_MAX); > - nr[lru] -= nr_to_scan; > > + sc->nr_isolate_scanned = 0; > nr_reclaimed += shrink_list(lru, nr_to_scan, > lruvec, sc); > + /* > + * isolate_lru_folios() may scan far more > + * than nr_to_scan when the LRU holds > + * ineligible folios (zone-skip) or large > + * folios. Charge that overshoot against the > + * remaining quota (clamped by min() so it > + * cannot go negative) so the next iteration > + * does not rescan the same skipped folios. > + */ > + nr[lru] -= min(nr[lru], > + max(nr_to_scan, sc->nr_isolate_scanned)); > } > } > > -- > 2.53.0 > > ^ permalink raw reply [flat|nested] 7+ messages in thread
* [PATCH resend 2/2] mm: vmscan: stop scanning ineligible folios after max_nr_skipped 2026-09-01 8:47 [PATCH resend 0/2] mm: vmscan: fix scan overshoot and ineligible folio scanning john 2026-09-01 8:47 ` [PATCH resend 1/2] mm: vmscan: charge isolate overshoot against scan quota john @ 2026-09-01 8:47 ` john 2026-09-09 9:56 ` Kunwu Chan 2026-09-08 1:49 ` [PATCH resend 0/2] mm: vmscan: fix scan overshoot and ineligible folio scanning john 2026-09-09 2:11 ` Andrew Morton 3 siblings, 1 reply; 7+ messages in thread From: john @ 2026-09-01 8:47 UTC (permalink / raw) To: akpm, liuye, hannes, mhocko, david, ljs, hughd, mgorman, yang Cc: zhangqiuhao, wangkefeng.wang, mawupeng1, love_goo, linux-mm, linux-kernel, Wupeng Ma From: Wupeng Ma <mawupeng1@huawei.com> When reclaiming for a lower zone, isolate_lru_folios() accounts folios from higher zones as skipped and, once max_nr_skipped hits SWAP_CLUSTER_MAX_SKIPPED, force-isolates the remaining ineligible folios to keep the loop from spinning on the skipped ones. Those folios are reclaimed even though they can never satisfy the current allocation, so nr_reclaimed is inflated into a false progress that keeps resetting no_progress_loops in should_reclaim_retry() and delays the OOM. Stop scanning once max_nr_skipped is reached instead of force-isolating the ineligible folios. The skipped folios are already accounted in total_scan, so shrink_lruvec() can charge the overshoot against its scan budget (see the previous commit) and will not rescan them. Fixes: 1c7b17cf0594 ("mm/vmscan: fix hard LOCKUP in function isolate_lru_folios") Signed-off-by: Wupeng Ma <mawupeng1@huawei.com> --- mm/vmscan.c | 16 ++++++++++++---- 1 file changed, 12 insertions(+), 4 deletions(-) diff --git a/mm/vmscan.c b/mm/vmscan.c index 823af9e86efd3..375f7cd5aa441 100644 --- a/mm/vmscan.c +++ b/mm/vmscan.c @@ -1701,12 +1701,20 @@ static unsigned long isolate_lru_folios(unsigned long nr_to_scan, nr_pages = folio_nr_pages(folio); total_scan += nr_pages; - /* Using max_nr_skipped to prevent hard LOCKUP*/ - if (max_nr_skipped < SWAP_CLUSTER_MAX_SKIPPED && - (folio_zonenum(folio) > sc->reclaim_idx)) { + /* + * Using max_nr_skipped to prevent hard LOCKUP. + * Once the cap is hit, stop rather than force-isolating: + * reclaiming ineligible folios only inflates nr_reclaimed + * into a false progress. + */ + if (folio_zonenum(folio) > sc->reclaim_idx) { nr_skipped[folio_zonenum(folio)] += nr_pages; - move_to = &folios_skipped; max_nr_skipped++; + if (max_nr_skipped >= SWAP_CLUSTER_MAX_SKIPPED) { + list_move(&folio->lru, &folios_skipped); + break; + } + move_to = &folios_skipped; goto move; } -- 2.53.0 ^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: [PATCH resend 2/2] mm: vmscan: stop scanning ineligible folios after max_nr_skipped 2026-09-01 8:47 ` [PATCH resend 2/2] mm: vmscan: stop scanning ineligible folios after max_nr_skipped john @ 2026-09-09 9:56 ` Kunwu Chan 0 siblings, 0 replies; 7+ messages in thread From: Kunwu Chan @ 2026-09-09 9:56 UTC (permalink / raw) To: john Cc: Kunwu Chan, akpm, liuye, hannes, mhocko, david, ljs, hughd, mgorman, yang, zhangqiuhao, wangkefeng.wang, mawupeng1, linux-mm, linux-kernel, Wupeng Ma On Tue, 1 Sep 2026 16:47:06 +0800 john <love_goo@163.com> wrote: > From: Wupeng Ma <mawupeng1@huawei.com> > > When reclaiming for a lower zone, isolate_lru_folios() accounts folios > from higher zones as skipped and, once max_nr_skipped hits > SWAP_CLUSTER_MAX_SKIPPED, force-isolates the remaining ineligible folios > to keep the loop from spinning on the skipped ones. Those folios are > reclaimed even though they can never satisfy the current allocation, so > nr_reclaimed is inflated into a false progress that keeps resetting > no_progress_loops in should_reclaim_retry() and delays the OOM. > > Stop scanning once max_nr_skipped is reached instead of force-isolating > the ineligible folios. The skipped folios are already accounted in > total_scan, so shrink_lruvec() can charge the overshoot against its scan > budget (see the previous commit) and will not rescan them. > > Fixes: 1c7b17cf0594 ("mm/vmscan: fix hard LOCKUP in function isolate_lru_folios") > Signed-off-by: Wupeng Ma <mawupeng1@huawei.com> > --- > mm/vmscan.c | 16 ++++++++++++---- > 1 file changed, 12 insertions(+), 4 deletions(-) > > diff --git a/mm/vmscan.c b/mm/vmscan.c > index 823af9e86efd3..375f7cd5aa441 100644 > --- a/mm/vmscan.c > +++ b/mm/vmscan.c > @@ -1701,12 +1701,20 @@ static unsigned long isolate_lru_folios(unsigned long nr_to_scan, > nr_pages = folio_nr_pages(folio); > total_scan += nr_pages; > > - /* Using max_nr_skipped to prevent hard LOCKUP*/ > - if (max_nr_skipped < SWAP_CLUSTER_MAX_SKIPPED && > - (folio_zonenum(folio) > sc->reclaim_idx)) { > + /* > + * Using max_nr_skipped to prevent hard LOCKUP. > + * Once the cap is hit, stop rather than force-isolating: > + * reclaiming ineligible folios only inflates nr_reclaimed > + * into a false progress. > + */ > + if (folio_zonenum(folio) > sc->reclaim_idx) { > nr_skipped[folio_zonenum(folio)] += nr_pages; > - move_to = &folios_skipped; > max_nr_skipped++; > + if (max_nr_skipped >= SWAP_CLUSTER_MAX_SKIPPED) { > + list_move(&folio->lru, &folios_skipped); > + break; One question about stopping the scan once max_nr_skipped reaches SWAP_CLUSTER_MAX_SKIPPED. Previously, reaching this limit caused the remaining ineligible folios to be force-isolated so that the scanner would not keep looping on skipped folios. With this change, we break out of isolate_lru_folios() instead. Could this cause us to stop before reaching eligible folios later in the LRU? Is SWAP_CLUSTER_MAX_SKIPPED intended to be sufficient to determine that continuing the scan cannot make useful reclaim progress? Thanks, KunWu > + } > + move_to = &folios_skipped; > goto move; > } > > -- > 2.53.0 > > Sent using hkml (https://github.com/sjp38/hackermail) ^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: [PATCH resend 0/2] mm: vmscan: fix scan overshoot and ineligible folio scanning 2026-09-01 8:47 [PATCH resend 0/2] mm: vmscan: fix scan overshoot and ineligible folio scanning john 2026-09-01 8:47 ` [PATCH resend 1/2] mm: vmscan: charge isolate overshoot against scan quota john 2026-09-01 8:47 ` [PATCH resend 2/2] mm: vmscan: stop scanning ineligible folios after max_nr_skipped john @ 2026-09-08 1:49 ` john 2026-09-09 2:11 ` Andrew Morton 3 siblings, 0 replies; 7+ messages in thread From: john @ 2026-09-08 1:49 UTC (permalink / raw) To: akpm, liuye, hannes, mhocko, david, ljs, hughd, mgorman, yang Cc: zhangqiuhao, wangkefeng.wang, linux-mm, linux-kernel, Wupeng Ma, mawupeng1 Hi, Maintainers kindly ping. 在 2026/9/1 16:47, john 写道: > From: Wupeng Ma <mawupeng1@huawei.com> > > Rebase to the latest v7.3-rc-1. > > These problems only surface when reclaim targets a lower zone > while the LRU holds folios from a higher zone. Normal userspace > allocations go to the highest zone. The zone-skip branch stays > dead under typical loads. Lower-zone-pressured configs (DMA32 > module allocations, memory-constrained devices) hit the issues. > They inflate scan cost and delay the OOM. > > shrink_lruvec() drives reclaim in SWAP_CLUSTER_MAX (32) chunks, but > isolate_lru_folios() may scan far more than that per call on a single > LRU. The excess is never charged back, so shrink_lruvec() keeps > rescanning the same folios round after round. When reclaim targets a > lower zone, the same scanner also keeps walking zone-ineligible folios > that can never satisfy the allocation, inflating nr_reclaimed into a > false progress that delays the OOM. > > This series fixes both: > > [1/2] Charge the isolate overshoot against the scan quota so the > next round skips already-scanned folios. > [2/2] Stop scanning once too many zone-ineligible folios have been > skipped, instead of force-isolating them. > > Background > ========== > > We observed slow, unexpected OOM behavior during extreme stress testing, > and while digging into the reclaim code during that analysis we spotted > these latent risks in isolate_lru_folios() -- the overshoot never > being charged back, and the force-isolate path on lower-zone reclaim. > The concerns below are the ones surfaced from reading the code, then > confirmed by constructing the situation deliberately. > > The problem only appears when reclaim targets a lower zone while the > LRU holds folios from a higher zone: > > - reclaim_idx points at DMA32/DMA (the triggering allocation asked > for a lower zone, e.g. __GFP_DMA32), and > - the LRU holds folios from a higher zone (Normal/Movable). > > Normal userspace allocations go to the highest zone, so reclaim_idx > never points below it and the zone-skip branch is never taken. It > needs a real lower-zone allocator to drain that zone below watermark; > that is also why it went unnoticed upstream -- the 1c7b17cf hard-lockup > fix that introduced the force-isolate path was found only on a ~1 TB > box running DMA32 module allocations. > > Impact > ====== > > Reproduced on x86 QEMU, 7.2-rc6, with a kernel module doing > __GFP_DMA32 allocations to drain DMA32 below watermark (LRU folios > sit in Movable): > > - A single isolate_lru_folios() call scanned 32794 pages and took > 25 (32769 skipped): 99.9% wasted on ineligible folios. > - Across one run, isolate was called 339 times, 140-165 of which > isolated nothing (taken=0, pure empty scans). > - shrink_lruvec() charges only the 32-page quota per round and > never refunds the overshoot, so the same skipped folios are > rescanned round after round. vmstat on a memcg OOM path: > pgscan_direct / pgsteal_direct = 2836624 / 112719 = 25.2x > (isolate -> shrink_folio_list returns the folio -> isolate again). > > Once max_nr_skipped hits SWAP_CLUSTER_MAX_SKIPPED, the current code > force-isolates the remaining ineligible folios. On the inactive LRU > they reach shrink_folio_list() and get reclaimed though they can > never satisfy the allocation, inflating nr_reclaimed and resetting > no_progress_loops in should_reclaim_retry(), delaying the OOM. > > A/B results (same .config, md5-identical): > > baseline patched > empty scans (taken=0) 140-165 0 > isolate calls 339 7 > total pages scanned 343711 3140 > single-call max scan 32794 3104 > > Empty-scan 0 is the stable evidence (holds every run). The > pgscan/pgsteal ratio is volatile (baseline 56-26675x, patched > 65-1324x, ranges overlap) and is not relied on alone. > > Caveats and reproduction > ======================== > > - Triggering needs a lower-zone-pressured box. To confirm the code > analysis, the situation was constructed on a small x86 QEMU VM > (1500M, CONFIG_LRU_GEN=n) with kernel cmdline > `movable_zone=DMA32 kernelcore=256M` (DMA32 small, Normal empty, > Movable large), then a kernel module doing > `alloc_page(GFP_DMA32)` drains DMA32 below watermark. LRU folios > sit in Movable and are zone-skipped while reclaiming for DMA32. > Observed via the `mm_vmscan_lru_isolate` tracepoint and > /proc/vmstat (pgscan_direct, pgsteal_direct). > > > Wupeng Ma (2): > mm: vmscan: charge isolate overshoot against scan quota > mm: vmscan: stop scanning ineligible folios after max_nr_skipped > > mm/vmscan.c | 48 +++++++++++++++++++++++++++++++++--------------- > 1 file changed, 33 insertions(+), 15 deletions(-) > ^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: [PATCH resend 0/2] mm: vmscan: fix scan overshoot and ineligible folio scanning 2026-09-01 8:47 [PATCH resend 0/2] mm: vmscan: fix scan overshoot and ineligible folio scanning john ` (2 preceding siblings ...) 2026-09-08 1:49 ` [PATCH resend 0/2] mm: vmscan: fix scan overshoot and ineligible folio scanning john @ 2026-09-09 2:11 ` Andrew Morton 3 siblings, 0 replies; 7+ messages in thread From: Andrew Morton @ 2026-09-09 2:11 UTC (permalink / raw) To: john Cc: liuye, hannes, mhocko, david, ljs, hughd, mgorman, yang, zhangqiuhao, wangkefeng.wang, mawupeng1, linux-mm, linux-kernel, Wupeng Ma On Tue, 1 Sep 2026 16:47:04 +0800 john <love_goo@163.com> wrote: > From: Wupeng Ma <mawupeng1@huawei.com> > > Rebase to the latest v7.3-rc-1. > > These problems only surface when reclaim targets a lower zone > while the LRU holds folios from a higher zone. Normal userspace > allocations go to the highest zone. The zone-skip branch stays > dead under typical loads. Lower-zone-pressured configs (DMA32 > module allocations, memory-constrained devices) hit the issues. > They inflate scan cost and delay the OOM. > > shrink_lruvec() drives reclaim in SWAP_CLUSTER_MAX (32) chunks, but > isolate_lru_folios() may scan far more than that per call on a single > LRU. The excess is never charged back, so shrink_lruvec() keeps > rescanning the same folios round after round. When reclaim targets a > lower zone, the same scanner also keeps walking zone-ineligible folios > that can never satisfy the allocation, inflating nr_reclaimed into a > false progress that delays the OOM. > > This series fixes both: > > [1/2] Charge the isolate overshoot against the scan quota so the > next round skips already-scanned folios. > [2/2] Stop scanning once too many zone-ineligible folios have been > skipped, instead of force-isolating them. > > Background > ========== > > We observed slow, unexpected OOM behavior during extreme stress testing, OK. But why should we care? Don't do extreme stress testing on your revenue-generating customer-facing computers! IOW, as long as the stress tests don't crash the kernel or lock up the box, we can spend our time thinking about kernel behavior which really matters. Now, if these changes can be shown to translate into improvement in real-world workloads then they're useful. Am I wrong? ^ permalink raw reply [flat|nested] 7+ messages in thread
end of thread, other threads:[~2026-09-09 9:56 UTC | newest] Thread overview: 7+ messages (download: mbox.gz / follow: Atom feed) -- links below jump to the message on this page -- 2026-09-01 8:47 [PATCH resend 0/2] mm: vmscan: fix scan overshoot and ineligible folio scanning john 2026-09-01 8:47 ` [PATCH resend 1/2] mm: vmscan: charge isolate overshoot against scan quota john 2026-09-09 9:52 ` Kunwu Chan 2026-09-01 8:47 ` [PATCH resend 2/2] mm: vmscan: stop scanning ineligible folios after max_nr_skipped john 2026-09-09 9:56 ` Kunwu Chan 2026-09-08 1:49 ` [PATCH resend 0/2] mm: vmscan: fix scan overshoot and ineligible folio scanning john 2026-09-09 2:11 ` Andrew Morton
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®