* [PATCH v2 0/2] mm/mglru: fix ineffective memory protection for non-kswapd reclaim @ 2026-08-28 11:09 Ridong Chen 2026-08-28 11:09 ` [PATCH v2 1/2] mm/page_counter: avoid integer overflow in effective_protection() Ridong Chen 2026-08-28 11:09 ` [PATCH v2 2/2] mm/mglru: fix ineffective memory protection for non-kswapd reclaim Ridong Chen 0 siblings, 2 replies; 8+ messages in thread From: Ridong Chen @ 2026-08-28 11:09 UTC (permalink / raw) To: Johannes Weiner, Michal Hocko, Roman Gushchin, Shakeel Butt, Andrew Morton Cc: Muchun Song, Kairui Song, Qi Zheng, Barry Song, Axel Rasmussen, Yuanchu Xie, Wei Xu, David Hildenbrand, Lorenzo Stoakes, Chris Down, Tejun Heo, Yu Zhao, open list:CONTROL GROUP - MEMORY RESOURCE CONTROLLER (MEMCG), open list:CONTROL GROUP - MEMORY RESOURCE CONTROLLER (MEMCG), linux-kernel, Ridong Chen, Ridong Chen From: Ridong Chen <chenridong@xiaomi.com> This series fixes memory.min/low being silently bypassed for MGLRU during non-kswapd global reclaim (global direct reclaim and root-level memory.reclaim). Patch 2 is the actual fix. Patch 1 is a prerequisite: an integer overflow in effective_protection(), spotted by the sashiko review tool, which patch 2's new caller would also be exposed to. --- v2: - Fix an issue caused by non-atomic read races in patch 1 [1] [1] https://sashiko.dev/#/patchset/20260828092432.1257917-1-ridong.chen@linux.dev?part=1 Ridong Chen (2): mm/page_counter: avoid integer overflow in effective_protection() mm/mglru: fix ineffective memory protection for non-kswapd reclaim include/linux/memcontrol.h | 11 ++++++++++ mm/memcontrol.c | 45 ++++++++++++++++++++++++++++++++++++++ mm/page_counter.c | 19 +++++++++++----- mm/vmscan.c | 8 ++++++- 4 files changed, 76 insertions(+), 7 deletions(-) -- 2.34.1 ^ permalink raw reply [flat|nested] 8+ messages in thread
* [PATCH v2 1/2] mm/page_counter: avoid integer overflow in effective_protection() 2026-08-28 11:09 [PATCH v2 0/2] mm/mglru: fix ineffective memory protection for non-kswapd reclaim Ridong Chen @ 2026-08-28 11:09 ` Ridong Chen 2026-08-30 7:59 ` Barry Song 2026-08-28 11:09 ` [PATCH v2 2/2] mm/mglru: fix ineffective memory protection for non-kswapd reclaim Ridong Chen 1 sibling, 1 reply; 8+ messages in thread From: Ridong Chen @ 2026-08-28 11:09 UTC (permalink / raw) To: Johannes Weiner, Michal Hocko, Roman Gushchin, Shakeel Butt, Andrew Morton Cc: Muchun Song, Kairui Song, Qi Zheng, Barry Song, Axel Rasmussen, Yuanchu Xie, Wei Xu, David Hildenbrand, Lorenzo Stoakes, Chris Down, Tejun Heo, Yu Zhao, open list:CONTROL GROUP - MEMORY RESOURCE CONTROLLER (MEMCG), open list:CONTROL GROUP - MEMORY RESOURCE CONTROLLER (MEMCG), linux-kernel, Ridong Chen, Ridong Chen, stable From: Ridong Chen <chenridong@xiaomi.com> effective_protection() scales a parent's protection by a ratio of page counts, e.g. for recursive protection: (parent_effective - siblings_protected) * (usage - protected) / (parent_usage - siblings_protected) The multiply is done at unsigned long width before dividing. On systems with >= 16TB RAM the product can exceed 2^64 and wrap, giving a bogus protection value and silently breaking memory.min/low enforcement. Use mul_u64_u64_div_u64() to multiply in a 128-bit intermediate. Because usage and parent_usage are not read atomically (a child is charged before its parent), usage - protected can briefly exceed the divisor, making the quotient overflow 64 bits and trap (#DE on x86). Cap it so the ratio stays <= 1. Reported by the sashiko review tool [1]. [1] https://sashiko.dev/#/patchset/20260826133054.88529-1-ridong.chen@linux.dev?part=1 Fixes: bc50bcc6e00b ("mm: memcontrol: clean up and document effective low/min calculations") Fixes: 8a931f801340 ("mm: memcontrol: recursive memory.low protection") Cc: stable@vger.kernel.org Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Ridong Chen <chenridong@xiaomi.com> --- mm/page_counter.c | 19 +++++++++++++------ 1 file changed, 13 insertions(+), 6 deletions(-) diff --git a/mm/page_counter.c b/mm/page_counter.c index 661e0f2a5127..e8bd512069c5 100644 --- a/mm/page_counter.c +++ b/mm/page_counter.c @@ -8,6 +8,7 @@ #include <linux/page_counter.h> #include <linux/atomic.h> #include <linux/kernel.h> +#include <linux/math64.h> #include <linux/string.h> #include <linux/sched.h> #include <linux/bug.h> @@ -356,7 +357,8 @@ static unsigned long effective_protection(unsigned long usage, * otherwise get a smaller chunk than what they claimed. */ if (siblings_protected > parent_effective) - return protected * parent_effective / siblings_protected; + return mul_u64_u64_div_u64(protected, parent_effective, + siblings_protected); /* * Ok, utilized protection of all children is within what the @@ -397,13 +399,18 @@ static unsigned long effective_protection(unsigned long usage, if (parent_effective > siblings_protected && parent_usage > siblings_protected && usage > protected) { - unsigned long unclaimed; + unsigned long unclaimed = parent_effective - siblings_protected; + unsigned long unprotected = usage - protected; + unsigned long parent_unprotected = parent_usage - siblings_protected; - unclaimed = parent_effective - siblings_protected; - unclaimed *= usage - protected; - unclaimed /= parent_usage - siblings_protected; + /* + * The usages aren't read atomically, so a child can transiently + * appear to use more than its parent, making the ratio exceed 1 + * and the quotient overflow 64 bits (#DE on x86). Cap it. + */ + unprotected = min(unprotected, parent_unprotected); - ep += unclaimed; + ep += mul_u64_u64_div_u64(unclaimed, unprotected, parent_unprotected); } return ep; -- 2.34.1 ^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH v2 1/2] mm/page_counter: avoid integer overflow in effective_protection() 2026-08-28 11:09 ` [PATCH v2 1/2] mm/page_counter: avoid integer overflow in effective_protection() Ridong Chen @ 2026-08-30 7:59 ` Barry Song 0 siblings, 0 replies; 8+ messages in thread From: Barry Song @ 2026-08-30 7:59 UTC (permalink / raw) To: Ridong Chen Cc: Johannes Weiner, Michal Hocko, Roman Gushchin, Shakeel Butt, Andrew Morton, Muchun Song, Kairui Song, Qi Zheng, Axel Rasmussen, Yuanchu Xie, Wei Xu, David Hildenbrand, Lorenzo Stoakes, Chris Down, Tejun Heo, Yu Zhao, open list:CONTROL GROUP - MEMORY RESOURCE CONTROLLER (MEMCG), open list:CONTROL GROUP - MEMORY RESOURCE CONTROLLER (MEMCG), linux-kernel, Ridong Chen, stable On Fri, Aug 28, 2026 at 7:09 PM Ridong Chen <ridong.chen@linux.dev> wrote: > > From: Ridong Chen <chenridong@xiaomi.com> > > effective_protection() scales a parent's protection by a ratio of page > counts, e.g. for recursive protection: > > (parent_effective - siblings_protected) * (usage - protected) > / (parent_usage - siblings_protected) > > The multiply is done at unsigned long width before dividing. On systems > with >= 16TB RAM the product can exceed 2^64 and wrap, giving a bogus > protection value and silently breaking memory.min/low enforcement. > > Use mul_u64_u64_div_u64() to multiply in a 128-bit intermediate. Because > usage and parent_usage are not read atomically (a child is charged > before its parent), usage - protected can briefly exceed the divisor, > making the quotient overflow 64 bits and trap (#DE on x86). Cap it so > the ratio stays <= 1. > > Reported by the sashiko review tool [1]. > > [1] https://sashiko.dev/#/patchset/20260826133054.88529-1-ridong.chen@linux.dev?part=1 > > Fixes: bc50bcc6e00b ("mm: memcontrol: clean up and document effective low/min calculations") > Fixes: 8a931f801340 ("mm: memcontrol: recursive memory.low protection") > Cc: stable@vger.kernel.org > Assisted-by: Claude:claude-opus-4-8 > Signed-off-by: Ridong Chen <chenridong@xiaomi.com> > --- LGTM, Reviewed-by: Barry Song <baohua@kernel.org> ^ permalink raw reply [flat|nested] 8+ messages in thread
* [PATCH v2 2/2] mm/mglru: fix ineffective memory protection for non-kswapd reclaim 2026-08-28 11:09 [PATCH v2 0/2] mm/mglru: fix ineffective memory protection for non-kswapd reclaim Ridong Chen 2026-08-28 11:09 ` [PATCH v2 1/2] mm/page_counter: avoid integer overflow in effective_protection() Ridong Chen @ 2026-08-28 11:09 ` Ridong Chen 2026-08-30 7:53 ` Barry Song 1 sibling, 1 reply; 8+ messages in thread From: Ridong Chen @ 2026-08-28 11:09 UTC (permalink / raw) To: Johannes Weiner, Michal Hocko, Roman Gushchin, Shakeel Butt, Andrew Morton Cc: Muchun Song, Kairui Song, Qi Zheng, Barry Song, Axel Rasmussen, Yuanchu Xie, Wei Xu, David Hildenbrand, Lorenzo Stoakes, Chris Down, Tejun Heo, Yu Zhao, open list:CONTROL GROUP - MEMORY RESOURCE CONTROLLER (MEMCG), open list:CONTROL GROUP - MEMORY RESOURCE CONTROLLER (MEMCG), linux-kernel, Ridong Chen, Ridong Chen, stable From: Ridong Chen <chenridong@xiaomi.com> memory.min/low is silently bypassed for MGLRU during global proactive reclaim (writing to the root memory.reclaim) and global direct reclaim. It can be reproduced as follows: # echo 7 > /sys/kernel/mm/lru_gen/enabled # cd /sys/fs/cgroup # mkdir -p a/b # echo 100M > a/memory.min # echo +memory > a/cgroup.subtree_control # echo 100M > a/b/memory.min # echo $$ > a/b/cgroup.procs # dd if=/dev/zero of=/tmp/testfile bs=1M count=200 # cat a/b/memory.current 222650368 # echo 500M > memory.reclaim -bash: echo: write error: Resource temporarily unavailable # cat a/b/memory.current 6070272 memory.min is 100M, yet reclaim drops a/b down to 6M, breaking the protection. The traditional LRU path is not affected because shrink_node() calls mem_cgroup_calculate_protection() for each memcg it visits during a top-down tree walk. Commit 30d77b7eef01 ("mm/mglru: fix ineffective protection calculation") moved the protection computation into lru_gen_age_node(), which only runs for kswapd. Non-kswapd global reclaim reaches shrink_one() through lru_gen_shrink_node() -> shrink_many() without any protection computation, so emin/elow remain stale or zero. Introduce mem_cgroup_protection_path() which computes emin/elow along the root-to-target path only by iterating through the cgroup ancestors array top-down. This avoids the full tree traversal that would be needed with mem_cgroup_calculate_protection(), limiting the cost to O(depth) per memcg - typically 3-5 levels. Call it from shrink_one() for the non-kswapd path so that each memcg about to be shrunk has correct protection values. Fixes: e4dde56cd208 ("mm: multi-gen LRU: per-node lru_gen_folio lists") Cc: stable@vger.kernel.org Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Ridong Chen <chenridong@xiaomi.com> --- include/linux/memcontrol.h | 11 ++++++++++ mm/memcontrol.c | 45 ++++++++++++++++++++++++++++++++++++++ mm/vmscan.c | 8 ++++++- 3 files changed, 63 insertions(+), 1 deletion(-) diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h index 7d1c0ce189a8..8066b798a759 100644 --- a/include/linux/memcontrol.h +++ b/include/linux/memcontrol.h @@ -605,6 +605,10 @@ static inline void mem_cgroup_protection(struct mem_cgroup *root, void mem_cgroup_calculate_protection(struct mem_cgroup *root, struct mem_cgroup *memcg); +#ifdef CONFIG_LRU_GEN +void mem_cgroup_protection_path(struct mem_cgroup *root, + struct mem_cgroup *memcg); +#endif static inline bool mem_cgroup_unprotected(struct mem_cgroup *target, struct mem_cgroup *memcg) @@ -1133,6 +1137,13 @@ static inline void mem_cgroup_calculate_protection(struct mem_cgroup *root, { } +#ifdef CONFIG_LRU_GEN +static inline void mem_cgroup_protection_path(struct mem_cgroup *root, + struct mem_cgroup *memcg) +{ +} +#endif + static inline bool mem_cgroup_unprotected(struct mem_cgroup *target, struct mem_cgroup *memcg) { diff --git a/mm/memcontrol.c b/mm/memcontrol.c index 1271d390b617..095050d4296a 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -5198,6 +5198,51 @@ void mem_cgroup_calculate_protection(struct mem_cgroup *root, page_counter_calculate_protection(&root->memory, &memcg->memory, recursive_protection); } +#ifdef CONFIG_LRU_GEN +/** + * mem_cgroup_protection_path - compute protection along root->memcg path + * @root: the top ancestor of the sub-tree being checked (NULL for root_mem_cgroup) + * @memcg: the target memory cgroup + * + * Walk the ancestor path from @root down to @memcg and compute the effective + * protection at each level. This is safe for isolated queries because it + * ensures parents are computed before children. + */ +void mem_cgroup_protection_path(struct mem_cgroup *root, + struct mem_cgroup *memcg) +{ + bool recursive_protection = + cgrp_dfl_root.flags & CGRP_ROOT_MEMORY_RECURSIVE_PROT; + struct cgroup *cg; + int root_level, i; + + if (mem_cgroup_disabled()) + return; + + if (!root) + root = root_mem_cgroup; + + if (memcg == root) + return; + + root_level = root->css.cgroup->level; + cg = memcg->css.cgroup; + + rcu_read_lock(); + for (i = root_level + 1; i <= cg->level; i++) { + struct mem_cgroup *cur; + + cur = mem_cgroup_from_css(cgroup_css(cg->ancestors[i], + &memory_cgrp_subsys)); + if (cur) + page_counter_calculate_protection(&root->memory, + &cur->memory, + recursive_protection); + } + rcu_read_unlock(); +} +#endif /* CONFIG_LRU_GEN */ + static int charge_memcg(struct folio *folio, struct mem_cgroup *memcg, gfp_t gfp) { diff --git a/mm/vmscan.c b/mm/vmscan.c index f11491ee9ed5..e0ba68ede745 100644 --- a/mm/vmscan.c +++ b/mm/vmscan.c @@ -5102,7 +5102,13 @@ static int shrink_one(struct lruvec *lruvec, struct scan_control *sc) struct mem_cgroup *memcg = lruvec_memcg(lruvec); struct pglist_data *pgdat = lruvec_pgdat(lruvec); - /* lru_gen_age_node() called mem_cgroup_calculate_protection() */ + /* + * For kswapd, lru_gen_age_node() has already called + * mem_cgroup_calculate_protection() + */ + if (!current_is_kswapd()) + mem_cgroup_protection_path(NULL, memcg); + if (mem_cgroup_below_min(NULL, memcg)) return MEMCG_LRU_YOUNG; -- 2.34.1 ^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH v2 2/2] mm/mglru: fix ineffective memory protection for non-kswapd reclaim 2026-08-28 11:09 ` [PATCH v2 2/2] mm/mglru: fix ineffective memory protection for non-kswapd reclaim Ridong Chen @ 2026-08-30 7:53 ` Barry Song 2026-08-30 10:13 ` Ridong Chen 0 siblings, 1 reply; 8+ messages in thread From: Barry Song @ 2026-08-30 7:53 UTC (permalink / raw) To: Ridong Chen Cc: Johannes Weiner, Michal Hocko, Roman Gushchin, Shakeel Butt, Andrew Morton, Muchun Song, Kairui Song, Qi Zheng, Axel Rasmussen, Yuanchu Xie, Wei Xu, David Hildenbrand, Lorenzo Stoakes, Chris Down, Tejun Heo, Yu Zhao, open list:CONTROL GROUP - MEMORY RESOURCE CONTROLLER (MEMCG), open list:CONTROL GROUP - MEMORY RESOURCE CONTROLLER (MEMCG), linux-kernel, Ridong Chen, stable On Fri, Aug 28, 2026 at 7:10 PM Ridong Chen <ridong.chen@linux.dev> wrote: > > From: Ridong Chen <chenridong@xiaomi.com> > > memory.min/low is silently bypassed for MGLRU during global proactive > reclaim (writing to the root memory.reclaim) and global direct reclaim. I guess nobody is silently bypassing anything. It's just that the effective min is stale data. If kswapd has run at least once, should the protection have been updated already? I guess we need to update the changelog a bit? > It can be reproduced as follows: > > # echo 7 > /sys/kernel/mm/lru_gen/enabled > # cd /sys/fs/cgroup > # mkdir -p a/b > # echo 100M > a/memory.min > # echo +memory > a/cgroup.subtree_control > # echo 100M > a/b/memory.min > # echo $$ > a/b/cgroup.procs > # dd if=/dev/zero of=/tmp/testfile bs=1M count=200 > # cat a/b/memory.current > 222650368 > # echo 500M > memory.reclaim > -bash: echo: write error: Resource temporarily unavailable > # cat a/b/memory.current > 6070272 > > memory.min is 100M, yet reclaim drops a/b down to 6M, breaking the > protection. The traditional LRU path is not affected because > shrink_node() calls mem_cgroup_calculate_protection() for each memcg it > visits during a top-down tree walk. > > Commit 30d77b7eef01 ("mm/mglru: fix ineffective protection calculation") > moved the protection computation into lru_gen_age_node(), which only > runs for kswapd. Non-kswapd global reclaim reaches shrink_one() through > lru_gen_shrink_node() -> shrink_many() without any protection > computation, so emin/elow remain stale or zero. > > Introduce mem_cgroup_protection_path() which computes emin/elow along > the root-to-target path only by iterating through the cgroup ancestors > array top-down. This avoids the full tree traversal that would be > needed with mem_cgroup_calculate_protection(), limiting the cost to > O(depth) per memcg - typically 3-5 levels. > > Call it from shrink_one() for the non-kswapd path so that each memcg > about to be shrunk has correct protection values. > > Fixes: e4dde56cd208 ("mm: multi-gen LRU: per-node lru_gen_folio lists") > Cc: stable@vger.kernel.org > Assisted-by: Claude:claude-opus-4-8 > Signed-off-by: Ridong Chen <chenridong@xiaomi.com> > --- > include/linux/memcontrol.h | 11 ++++++++++ > mm/memcontrol.c | 45 ++++++++++++++++++++++++++++++++++++++ > mm/vmscan.c | 8 ++++++- > 3 files changed, 63 insertions(+), 1 deletion(-) > > diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h > index 7d1c0ce189a8..8066b798a759 100644 > --- a/include/linux/memcontrol.h > +++ b/include/linux/memcontrol.h > @@ -605,6 +605,10 @@ static inline void mem_cgroup_protection(struct mem_cgroup *root, > > void mem_cgroup_calculate_protection(struct mem_cgroup *root, > struct mem_cgroup *memcg); > +#ifdef CONFIG_LRU_GEN > +void mem_cgroup_protection_path(struct mem_cgroup *root, > + struct mem_cgroup *memcg); > +#endif > > static inline bool mem_cgroup_unprotected(struct mem_cgroup *target, > struct mem_cgroup *memcg) > @@ -1133,6 +1137,13 @@ static inline void mem_cgroup_calculate_protection(struct mem_cgroup *root, > { > } > > +#ifdef CONFIG_LRU_GEN > +static inline void mem_cgroup_protection_path(struct mem_cgroup *root, > + struct mem_cgroup *memcg) > +{ > +} > +#endif > + I wonder if we could follow the zswap pattern? #if defined(CONFIG_MEMCG) && defined(CONFIG_ZSWAP) bool obj_cgroup_may_zswap(struct obj_cgroup *objcg); void obj_cgroup_charge_zswap(struct obj_cgroup *objcg, size_t size); void obj_cgroup_uncharge_zswap(struct obj_cgroup *objcg, size_t size); bool mem_cgroup_zswap_writeback_enabled(struct mem_cgroup *memcg); #else static inline bool obj_cgroup_may_zswap(struct obj_cgroup *objcg) { return true; } ... #endif > static inline bool mem_cgroup_unprotected(struct mem_cgroup *target, > struct mem_cgroup *memcg) > { > diff --git a/mm/memcontrol.c b/mm/memcontrol.c > index 1271d390b617..095050d4296a 100644 > --- a/mm/memcontrol.c > +++ b/mm/memcontrol.c > @@ -5198,6 +5198,51 @@ void mem_cgroup_calculate_protection(struct mem_cgroup *root, > page_counter_calculate_protection(&root->memory, &memcg->memory, recursive_protection); > } > > +#ifdef CONFIG_LRU_GEN > +/** > + * mem_cgroup_protection_path - compute protection along root->memcg path > + * @root: the top ancestor of the sub-tree being checked (NULL for root_mem_cgroup) > + * @memcg: the target memory cgroup > + * > + * Walk the ancestor path from @root down to @memcg and compute the effective > + * protection at each level. This is safe for isolated queries because it > + * ensures parents are computed before children. > + */ > +void mem_cgroup_protection_path(struct mem_cgroup *root, > + struct mem_cgroup *memcg) Can we rename it to `mem_cgroup_calculate_protection_path()`? BTW, I see that the only caller is in vmscan and it passes NULL as `root`. Do we need to keep the `root` argument if the new helper is only used for global reclaim? Best Regards Barry ^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH v2 2/2] mm/mglru: fix ineffective memory protection for non-kswapd reclaim 2026-08-30 7:53 ` Barry Song @ 2026-08-30 10:13 ` Ridong Chen 2026-08-30 10:40 ` Barry Song 0 siblings, 1 reply; 8+ messages in thread From: Ridong Chen @ 2026-08-30 10:13 UTC (permalink / raw) To: Barry Song Cc: Johannes Weiner, Michal Hocko, Roman Gushchin, Shakeel Butt, Andrew Morton, Muchun Song, Kairui Song, Qi Zheng, Axel Rasmussen, Yuanchu Xie, Wei Xu, David Hildenbrand, Lorenzo Stoakes, Chris Down, Tejun Heo, Yu Zhao, open list:CONTROL GROUP - MEMORY RESOURCE CONTROLLER (MEMCG), open list:CONTROL GROUP - MEMORY RESOURCE CONTROLLER (MEMCG), linux-kernel, Ridong Chen, stable On 8/30/2026 3:53 PM, Barry Song wrote: > On Fri, Aug 28, 2026 at 7:10 PM Ridong Chen <ridong.chen@linux.dev> wrote: >> >> From: Ridong Chen <chenridong@xiaomi.com> >> >> memory.min/low is silently bypassed for MGLRU during global proactive >> reclaim (writing to the root memory.reclaim) and global direct reclaim. > > I guess nobody is silently bypassing anything. It's just that the > effective min is stale data. If kswapd has run at least once, should > the protection have been updated already? > I guess we need to update the changelog a bit? > The children's emin/elow are derived from the parent's min/low settings and children_min_usage, both of which can change over time. As a result, emin/elow may become stale, even if kswapd has already run once. >> It can be reproduced as follows: >> >> # echo 7 > /sys/kernel/mm/lru_gen/enabled >> # cd /sys/fs/cgroup >> # mkdir -p a/b >> # echo 100M > a/memory.min >> # echo +memory > a/cgroup.subtree_control >> # echo 100M > a/b/memory.min >> # echo $$ > a/b/cgroup.procs >> # dd if=/dev/zero of=/tmp/testfile bs=1M count=200 >> # cat a/b/memory.current >> 222650368 >> # echo 500M > memory.reclaim >> -bash: echo: write error: Resource temporarily unavailable >> # cat a/b/memory.current >> 6070272 >> >> memory.min is 100M, yet reclaim drops a/b down to 6M, breaking the >> protection. The traditional LRU path is not affected because >> shrink_node() calls mem_cgroup_calculate_protection() for each memcg it >> visits during a top-down tree walk. >> >> Commit 30d77b7eef01 ("mm/mglru: fix ineffective protection calculation") >> moved the protection computation into lru_gen_age_node(), which only >> runs for kswapd. Non-kswapd global reclaim reaches shrink_one() through >> lru_gen_shrink_node() -> shrink_many() without any protection >> computation, so emin/elow remain stale or zero. >> >> Introduce mem_cgroup_protection_path() which computes emin/elow along >> the root-to-target path only by iterating through the cgroup ancestors >> array top-down. This avoids the full tree traversal that would be >> needed with mem_cgroup_calculate_protection(), limiting the cost to >> O(depth) per memcg - typically 3-5 levels. >> >> Call it from shrink_one() for the non-kswapd path so that each memcg >> about to be shrunk has correct protection values. >> >> Fixes: e4dde56cd208 ("mm: multi-gen LRU: per-node lru_gen_folio lists") >> Cc: stable@vger.kernel.org >> Assisted-by: Claude:claude-opus-4-8 >> Signed-off-by: Ridong Chen <chenridong@xiaomi.com> >> --- >> include/linux/memcontrol.h | 11 ++++++++++ >> mm/memcontrol.c | 45 ++++++++++++++++++++++++++++++++++++++ >> mm/vmscan.c | 8 ++++++- >> 3 files changed, 63 insertions(+), 1 deletion(-) >> >> diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h >> index 7d1c0ce189a8..8066b798a759 100644 >> --- a/include/linux/memcontrol.h >> +++ b/include/linux/memcontrol.h >> @@ -605,6 +605,10 @@ static inline void mem_cgroup_protection(struct mem_cgroup *root, >> >> void mem_cgroup_calculate_protection(struct mem_cgroup *root, >> struct mem_cgroup *memcg); >> +#ifdef CONFIG_LRU_GEN >> +void mem_cgroup_protection_path(struct mem_cgroup *root, >> + struct mem_cgroup *memcg); >> +#endif >> >> static inline bool mem_cgroup_unprotected(struct mem_cgroup *target, >> struct mem_cgroup *memcg) >> @@ -1133,6 +1137,13 @@ static inline void mem_cgroup_calculate_protection(struct mem_cgroup *root, >> { >> } >> >> +#ifdef CONFIG_LRU_GEN >> +static inline void mem_cgroup_protection_path(struct mem_cgroup *root, >> + struct mem_cgroup *memcg) >> +{ >> +} >> +#endif >> + > > I wonder if we could follow the zswap pattern? > > #if defined(CONFIG_MEMCG) && defined(CONFIG_ZSWAP) > bool obj_cgroup_may_zswap(struct obj_cgroup *objcg); > void obj_cgroup_charge_zswap(struct obj_cgroup *objcg, size_t size); > void obj_cgroup_uncharge_zswap(struct obj_cgroup *objcg, size_t size); > bool mem_cgroup_zswap_writeback_enabled(struct mem_cgroup *memcg); > #else > static inline bool obj_cgroup_may_zswap(struct obj_cgroup *objcg) > { > return true; > } That is fine for me, I will update it. > ... > #endif > >> static inline bool mem_cgroup_unprotected(struct mem_cgroup *target, >> struct mem_cgroup *memcg) >> { >> diff --git a/mm/memcontrol.c b/mm/memcontrol.c >> index 1271d390b617..095050d4296a 100644 >> --- a/mm/memcontrol.c >> +++ b/mm/memcontrol.c >> @@ -5198,6 +5198,51 @@ void mem_cgroup_calculate_protection(struct mem_cgroup *root, >> page_counter_calculate_protection(&root->memory, &memcg->memory, recursive_protection); >> } >> >> +#ifdef CONFIG_LRU_GEN >> +/** >> + * mem_cgroup_protection_path - compute protection along root->memcg path >> + * @root: the top ancestor of the sub-tree being checked (NULL for root_mem_cgroup) >> + * @memcg: the target memory cgroup >> + * >> + * Walk the ancestor path from @root down to @memcg and compute the effective >> + * protection at each level. This is safe for isolated queries because it >> + * ensures parents are computed before children. >> + */ >> +void mem_cgroup_protection_path(struct mem_cgroup *root, >> + struct mem_cgroup *memcg) > > Can we rename it to `mem_cgroup_calculate_protection_path()`? > > BTW, I see that the only caller is in vmscan and it passes NULL as > `root`. Do we need to keep the `root` argument if the new helper is > only used for global reclaim? > I'd suggest keeping it as is. This function updates protection along the path from root to memcg, and could be reused later. Note that mem_cgroup_calculate_protection() assumes the caller has already performed the top-down walk, each level's calculation depends on its parent being updated first. For mem_cgroup_calculate_protection_path(), it can be called in any context without such a precondition. -- Best regards Ridong ^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH v2 2/2] mm/mglru: fix ineffective memory protection for non-kswapd reclaim 2026-08-30 10:13 ` Ridong Chen @ 2026-08-30 10:40 ` Barry Song 2026-08-30 10:56 ` Ridong Chen 0 siblings, 1 reply; 8+ messages in thread From: Barry Song @ 2026-08-30 10:40 UTC (permalink / raw) To: Ridong Chen Cc: Johannes Weiner, Michal Hocko, Roman Gushchin, Shakeel Butt, Andrew Morton, Muchun Song, Kairui Song, Qi Zheng, Axel Rasmussen, Yuanchu Xie, Wei Xu, David Hildenbrand, Lorenzo Stoakes, Chris Down, Tejun Heo, Yu Zhao, open list:CONTROL GROUP - MEMORY RESOURCE CONTROLLER (MEMCG), open list:CONTROL GROUP - MEMORY RESOURCE CONTROLLER (MEMCG), linux-kernel, Ridong Chen, stable On Sun, Aug 30, 2026 at 6:13 PM Ridong Chen <ridong.chen@linux.dev> wrote: > > > > On 8/30/2026 3:53 PM, Barry Song wrote: > > On Fri, Aug 28, 2026 at 7:10 PM Ridong Chen <ridong.chen@linux.dev> wrote: > >> > >> From: Ridong Chen <chenridong@xiaomi.com> > >> > >> memory.min/low is silently bypassed for MGLRU during global proactive > >> reclaim (writing to the root memory.reclaim) and global direct reclaim. > > > > I guess nobody is silently bypassing anything. It's just that the > > effective min is stale data. If kswapd has run at least once, should > > the protection have been updated already? > > I guess we need to update the changelog a bit? > > > > The children's emin/elow are derived from the parent's min/low settings and > children_min_usage, both of which can change over time. As a result, emin/elow > may become stale, even if kswapd has already run once. right, let's just say this in changelog, we are *not* bypassing we are just checking against stable data. The current changelog seems to be misleading. [...] > >> +void mem_cgroup_protection_path(struct mem_cgroup *root, > >> + struct mem_cgroup *memcg) > > > > Can we rename it to `mem_cgroup_calculate_protection_path()`? > > > > BTW, I see that the only caller is in vmscan and it passes NULL as > > `root`. Do we need to keep the `root` argument if the new helper is > > only used for global reclaim? > > > I'd suggest keeping it as is. This function updates protection along the path > from root to memcg, and could be reused later. Note that > mem_cgroup_calculate_protection() assumes the caller has already performed the > top-down walk, each level's calculation depends on its parent being updated first. > > For mem_cgroup_calculate_protection_path(), it can be called in any context > without such a precondition. I am fine with this - keeping the root there. but I guess rename is worth it. Best Regards Barry ^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH v2 2/2] mm/mglru: fix ineffective memory protection for non-kswapd reclaim 2026-08-30 10:40 ` Barry Song @ 2026-08-30 10:56 ` Ridong Chen 0 siblings, 0 replies; 8+ messages in thread From: Ridong Chen @ 2026-08-30 10:56 UTC (permalink / raw) To: Barry Song Cc: Johannes Weiner, Michal Hocko, Roman Gushchin, Shakeel Butt, Andrew Morton, Muchun Song, Kairui Song, Qi Zheng, Axel Rasmussen, Yuanchu Xie, Wei Xu, David Hildenbrand, Lorenzo Stoakes, Chris Down, Tejun Heo, Yu Zhao, open list:CONTROL GROUP - MEMORY RESOURCE CONTROLLER (MEMCG), open list:CONTROL GROUP - MEMORY RESOURCE CONTROLLER (MEMCG), linux-kernel, Ridong Chen, stable On 8/30/2026 6:40 PM, Barry Song wrote: > On Sun, Aug 30, 2026 at 6:13 PM Ridong Chen <ridong.chen@linux.dev> wrote: >> >> >> >> On 8/30/2026 3:53 PM, Barry Song wrote: >>> On Fri, Aug 28, 2026 at 7:10 PM Ridong Chen <ridong.chen@linux.dev> wrote: >>>> >>>> From: Ridong Chen <chenridong@xiaomi.com> >>>> >>>> memory.min/low is silently bypassed for MGLRU during global proactive >>>> reclaim (writing to the root memory.reclaim) and global direct reclaim. >>> >>> I guess nobody is silently bypassing anything. It's just that the >>> effective min is stale data. If kswapd has run at least once, should >>> the protection have been updated already? >>> I guess we need to update the changelog a bit? >>> >> >> The children's emin/elow are derived from the parent's min/low settings and >> children_min_usage, both of which can change over time. As a result, emin/elow >> may become stale, even if kswapd has already run once. > > right, let's just say this in changelog, we are *not* bypassing we are > just checking > against stable data. The current changelog seems to be misleading. > Thanks, Will update. > [...] >>>> +void mem_cgroup_protection_path(struct mem_cgroup *root, >>>> + struct mem_cgroup *memcg) >>> >>> Can we rename it to `mem_cgroup_calculate_protection_path()`? >>> >>> BTW, I see that the only caller is in vmscan and it passes NULL as >>> `root`. Do we need to keep the `root` argument if the new helper is >>> only used for global reclaim? >>> >> I'd suggest keeping it as is. This function updates protection along the path >> from root to memcg, and could be reused later. Note that >> mem_cgroup_calculate_protection() assumes the caller has already performed the >> top-down walk, each level's calculation depends on its parent being updated first. >> >> For mem_cgroup_calculate_protection_path(), it can be called in any context >> without such a precondition. > > I am fine with this - keeping the root there. but I guess rename is worth it. > Yeah, I will rename in the next version. -- Best regards Ridong ^ permalink raw reply [flat|nested] 8+ messages in thread
end of thread, other threads:[~2026-08-30 10:57 UTC | newest] Thread overview: 8+ messages (download: mbox.gz / follow: Atom feed) -- links below jump to the message on this page -- 2026-08-28 11:09 [PATCH v2 0/2] mm/mglru: fix ineffective memory protection for non-kswapd reclaim Ridong Chen 2026-08-28 11:09 ` [PATCH v2 1/2] mm/page_counter: avoid integer overflow in effective_protection() Ridong Chen 2026-08-30 7:59 ` Barry Song 2026-08-28 11:09 ` [PATCH v2 2/2] mm/mglru: fix ineffective memory protection for non-kswapd reclaim Ridong Chen 2026-08-30 7:53 ` Barry Song 2026-08-30 10:13 ` Ridong Chen 2026-08-30 10:40 ` Barry Song 2026-08-30 10:56 ` Ridong Chen
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®