* [PATCH] mm: vmscan: don't count per-node proactive reclaim as memory pressure
@ 2026-10-05 6:22 Vernon Yang
2026-10-05 6:30 ` Barry Song
2026-10-05 6:36 ` Andrew Morton
0 siblings, 2 replies; 6+ messages in thread
From: Vernon Yang @ 2026-10-05 6:22 UTC (permalink / raw)
To: akpm, david, kasong, qi.zheng, shakeel.butt, baohua,
axelrasmussen, yuanchu, weixugc, baoquan.he, hannes, mhocko, ljs,
roman.gushchin, dave
Cc: linux-mm, linux-kernel, Vernon Yang, stable
From: Vernon Yang <yanglincheng@kylinos.cn>
When the cgroup has no memory pressure at all, writing to
/sys/devices/system/nodeX/reclaim triggers proactive reclaim on
NUMA node, causing increase in the writer cgroup's memory PSI.
Due to this reclaim is performed in the context of the write(),
accounted as memory pressure on the writer, like
commit e22c6ed90aa9 ("mm: memcontrol: don't count limit-setting reclaim
as memory pressure"). This is unexpected, the phenomenon resembling
senpai will appear again.
The Documentation/ABI/stable/sysfs-devices-node documentation also
notes that "This interface is equivalent to the memcg variant."
This patch unifies the semantics of the memcg and node interfaces:
per-node proactive reclaim is no longer counted as memory pressure,
and the per-node proactive reclaim interface no longer produces
phantom pressure.
I ran demo[1] that performs per-node proactive reclaim 10000 times
in qemu, writer cgroup memory.pressure as follows:
without patch:
some avg10=31.53 avg60=13.42 avg300=3.31 total=10602985
full avg10=31.53 avg60=13.42 avg300=3.31 total=10602985
with patch:
some avg10=9.59 avg60=3.41 avg300=0.81 total=2686221
full avg10=9.59 avg60=3.41 avg300=0.81 total=2686221
[1] https://github.com/vernon2gh/app_and_module/tree/main/reclaim_node_psi
Fixes: b980077899ea ("mm: introduce per-node proactive reclaim interface")
Cc: stable@vger.kernel.org
Signed-off-by: Vernon Yang <yanglincheng@kylinos.cn>
---
mm/vmscan.c | 9 +++++----
1 file changed, 5 insertions(+), 4 deletions(-)
diff --git a/mm/vmscan.c b/mm/vmscan.c
index 91295070ca33..8e22b8a6038b 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -7988,13 +7988,10 @@ static unsigned long __node_reclaim(struct pglist_data *pgdat,
{
struct task_struct *p = current;
unsigned int noreclaim_flag;
- unsigned long pflags;
trace_mm_vmscan_node_reclaim_begin(pgdat->node_id, sc->order,
sc->gfp_mask);
- cond_resched();
- psi_memstall_enter(&pflags);
delayacct_freepages_start();
fs_reclaim_acquire(sc->gfp_mask);
/*
@@ -8011,7 +8008,6 @@ static unsigned long __node_reclaim(struct pglist_data *pgdat,
memalloc_noreclaim_restore(noreclaim_flag);
fs_reclaim_release(sc->gfp_mask);
delayacct_freepages_end();
- psi_memstall_leave(&pflags);
trace_mm_vmscan_node_reclaim_end(sc->nr_reclaimed, NULL);
@@ -8021,6 +8017,7 @@ static unsigned long __node_reclaim(struct pglist_data *pgdat,
unsigned long node_reclaim(struct pglist_data *pgdat, gfp_t gfp_mask, unsigned int order)
{
unsigned long ret;
+ unsigned long pflags;
/* Minimum pages needed in order to stay on node */
const unsigned long nr_pages = 1 << order;
struct scan_control sc = {
@@ -8067,7 +8064,10 @@ unsigned long node_reclaim(struct pglist_data *pgdat, gfp_t gfp_mask, unsigned i
if (test_and_set_bit_lock(PGDAT_RECLAIM_LOCKED, &pgdat->flags))
return 0;
+ cond_resched();
+ psi_memstall_enter(&pflags);
ret = __node_reclaim(pgdat, nr_pages, &sc);
+ psi_memstall_leave(&pflags);
clear_bit_unlock(PGDAT_RECLAIM_LOCKED, &pgdat->flags);
if (ret >= nr_pages)
@@ -8193,6 +8193,7 @@ int user_proactive_reclaim(char *buf,
&pgdat->flags))
return -EBUSY;
+ cond_resched();
reclaimed = __node_reclaim(pgdat, batch_size, &sc);
clear_bit_unlock(PGDAT_RECLAIM_LOCKED, &pgdat->flags);
}
--
2.53.0
^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [PATCH] mm: vmscan: don't count per-node proactive reclaim as memory pressure
2026-10-05 6:22 [PATCH] mm: vmscan: don't count per-node proactive reclaim as memory pressure Vernon Yang
@ 2026-10-05 6:30 ` Barry Song
2026-10-05 10:21 ` Vernon Yang
2026-10-05 6:36 ` Andrew Morton
1 sibling, 1 reply; 6+ messages in thread
From: Barry Song @ 2026-10-05 6:30 UTC (permalink / raw)
To: Vernon Yang
Cc: akpm, david, kasong, qi.zheng, shakeel.butt, axelrasmussen,
yuanchu, weixugc, baoquan.he, hannes, mhocko, ljs,
roman.gushchin, dave, linux-mm, linux-kernel, Vernon Yang,
stable
On Mon, Oct 5, 2026 at 2:23 PM Vernon Yang <vernon2gm@gmail.com> wrote:
>
> From: Vernon Yang <yanglincheng@kylinos.cn>
>
> When the cgroup has no memory pressure at all, writing to
> /sys/devices/system/nodeX/reclaim triggers proactive reclaim on
> NUMA node, causing increase in the writer cgroup's memory PSI.
>
> Due to this reclaim is performed in the context of the write(),
> accounted as memory pressure on the writer, like
> commit e22c6ed90aa9 ("mm: memcontrol: don't count limit-setting reclaim
> as memory pressure"). This is unexpected, the phenomenon resembling
> senpai will appear again.
>
> The Documentation/ABI/stable/sysfs-devices-node documentation also
> notes that "This interface is equivalent to the memcg variant."
>
> This patch unifies the semantics of the memcg and node interfaces:
> per-node proactive reclaim is no longer counted as memory pressure,
> and the per-node proactive reclaim interface no longer produces
> phantom pressure.
>
> I ran demo[1] that performs per-node proactive reclaim 10000 times
> in qemu, writer cgroup memory.pressure as follows:
>
> without patch:
>
> some avg10=31.53 avg60=13.42 avg300=3.31 total=10602985
> full avg10=31.53 avg60=13.42 avg300=3.31 total=10602985
>
> with patch:
>
> some avg10=9.59 avg60=3.41 avg300=0.81 total=2686221
> full avg10=9.59 avg60=3.41 avg300=0.81 total=2686221
>
> [1] https://github.com/vernon2gh/app_and_module/tree/main/reclaim_node_psi
>
> Fixes: b980077899ea ("mm: introduce per-node proactive reclaim interface")
> Cc: stable@vger.kernel.org
> Signed-off-by: Vernon Yang <yanglincheng@kylinos.cn>
> ---
This seems to be a valid concern. Personally, I don't like
having the code depend on whether `__node_reclaim()` and
`node_reclaim()` are called from proactive reclaim or page
allocation.
Can't we check whether `sc->proactive` is true? Am I missing
something?
Best Regards
Barry
^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [PATCH] mm: vmscan: don't count per-node proactive reclaim as memory pressure
2026-10-05 6:22 [PATCH] mm: vmscan: don't count per-node proactive reclaim as memory pressure Vernon Yang
2026-10-05 6:30 ` Barry Song
@ 2026-10-05 6:36 ` Andrew Morton
2026-10-05 10:10 ` Vernon Yang
1 sibling, 1 reply; 6+ messages in thread
From: Andrew Morton @ 2026-10-05 6:36 UTC (permalink / raw)
To: Vernon Yang
Cc: david, kasong, qi.zheng, shakeel.butt, baohua, axelrasmussen,
yuanchu, weixugc, baoquan.he, hannes, mhocko, ljs,
roman.gushchin, dave, linux-mm, linux-kernel, Vernon Yang,
stable
On Mon, 5 Oct 2026 14:22:36 +0800 Vernon Yang <vernon2gm@gmail.com> wrote:
> When the cgroup has no memory pressure at all, writing to
> /sys/devices/system/nodeX/reclaim triggers proactive reclaim on
> NUMA node, causing increase in the writer cgroup's memory PSI.
>
> Due to this reclaim is performed in the context of the write(),
> accounted as memory pressure on the writer, like
> commit e22c6ed90aa9 ("mm: memcontrol: don't count limit-setting reclaim
> as memory pressure"). This is unexpected, the phenomenon resembling
> senpai will appear again.
What is this?
> The Documentation/ABI/stable/sysfs-devices-node documentation also
> notes that "This interface is equivalent to the memcg variant."
>
> This patch unifies the semantics of the memcg and node interfaces:
> per-node proactive reclaim is no longer counted as memory pressure,
> and the per-node proactive reclaim interface no longer produces
> phantom pressure.
>
> I ran demo[1] that performs per-node proactive reclaim 10000 times
> in qemu, writer cgroup memory.pressure as follows:
>
> without patch:
>
> some avg10=31.53 avg60=13.42 avg300=3.31 total=10602985
> full avg10=31.53 avg60=13.42 avg300=3.31 total=10602985
>
> with patch:
>
> some avg10=9.59 avg60=3.41 avg300=0.81 total=2686221
> full avg10=9.59 avg60=3.41 avg300=0.81 total=2686221
^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [PATCH] mm: vmscan: don't count per-node proactive reclaim as memory pressure
2026-10-05 6:36 ` Andrew Morton
@ 2026-10-05 10:10 ` Vernon Yang
2026-10-05 12:54 ` Andrew Morton
0 siblings, 1 reply; 6+ messages in thread
From: Vernon Yang @ 2026-10-05 10:10 UTC (permalink / raw)
To: Andrew Morton
Cc: david, kasong, qi.zheng, shakeel.butt, baohua, axelrasmussen,
yuanchu, weixugc, baoquan.he, hannes, mhocko, ljs,
roman.gushchin, dave, linux-mm, linux-kernel, Vernon Yang,
stable
On Sun, Oct 04, 2026 at 11:36:15PM -0700, Andrew Morton wrote:
> On Mon, 5 Oct 2026 14:22:36 +0800 Vernon Yang <vernon2gm@gmail.com> wrote:
>
> > When the cgroup has no memory pressure at all, writing to
> > /sys/devices/system/nodeX/reclaim triggers proactive reclaim on
> > NUMA node, causing increase in the writer cgroup's memory PSI.
> >
> > Due to this reclaim is performed in the context of the write(),
> > accounted as memory pressure on the writer, like
> > commit e22c6ed90aa9 ("mm: memcontrol: don't count limit-setting reclaim
> > as memory pressure"). This is unexpected, the phenomenon resembling
> > senpai will appear again.
>
> What is this?
This is commit e22c6ed90aa9, which addresses a scenario mentioned in
memory limits of a cgroup, in detail as follows:
Currently, this reclaim activity is accounted as memory pressure in the
cgroup that the writer(!) belongs to. This is unexpected. It
specifically causes problems for senpai
(https://github.com/facebookincubator/senpai), which is an agent that
routinely adjusts the memory limits and performs associated reclaim work
in tens or even hundreds of cgroups running on the host. The cgroup that
senpai is running in itself will report elevated levels of memory
pressure, even though it itself is under no memory shortage or any sort of
distress.
This is just to explain that similar scenarios will continue to occur,
only this time it is for the /sys/devices/system/nodeX/reclaim knob.
--
Cheers,
Vernon
> > The Documentation/ABI/stable/sysfs-devices-node documentation also
> > notes that "This interface is equivalent to the memcg variant."
> >
> > This patch unifies the semantics of the memcg and node interfaces:
> > per-node proactive reclaim is no longer counted as memory pressure,
> > and the per-node proactive reclaim interface no longer produces
> > phantom pressure.
> >
> > I ran demo[1] that performs per-node proactive reclaim 10000 times
> > in qemu, writer cgroup memory.pressure as follows:
> >
> > without patch:
> >
> > some avg10=31.53 avg60=13.42 avg300=3.31 total=10602985
> > full avg10=31.53 avg60=13.42 avg300=3.31 total=10602985
> >
> > with patch:
> >
> > some avg10=9.59 avg60=3.41 avg300=0.81 total=2686221
> > full avg10=9.59 avg60=3.41 avg300=0.81 total=2686221
>
^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [PATCH] mm: vmscan: don't count per-node proactive reclaim as memory pressure
2026-10-05 6:30 ` Barry Song
@ 2026-10-05 10:21 ` Vernon Yang
0 siblings, 0 replies; 6+ messages in thread
From: Vernon Yang @ 2026-10-05 10:21 UTC (permalink / raw)
To: Barry Song
Cc: akpm, david, kasong, qi.zheng, shakeel.butt, axelrasmussen,
yuanchu, weixugc, baoquan.he, hannes, mhocko, ljs,
roman.gushchin, dave, linux-mm, linux-kernel, Vernon Yang,
stable
On Mon, Oct 05, 2026 at 02:30:36PM +0800, Barry Song wrote:
> On Mon, Oct 5, 2026 at 2:23 PM Vernon Yang <vernon2gm@gmail.com> wrote:
> >
> > From: Vernon Yang <yanglincheng@kylinos.cn>
> >
> > When the cgroup has no memory pressure at all, writing to
> > /sys/devices/system/nodeX/reclaim triggers proactive reclaim on
> > NUMA node, causing increase in the writer cgroup's memory PSI.
> >
> > Due to this reclaim is performed in the context of the write(),
> > accounted as memory pressure on the writer, like
> > commit e22c6ed90aa9 ("mm: memcontrol: don't count limit-setting reclaim
> > as memory pressure"). This is unexpected, the phenomenon resembling
> > senpai will appear again.
> >
> > The Documentation/ABI/stable/sysfs-devices-node documentation also
> > notes that "This interface is equivalent to the memcg variant."
> >
> > This patch unifies the semantics of the memcg and node interfaces:
> > per-node proactive reclaim is no longer counted as memory pressure,
> > and the per-node proactive reclaim interface no longer produces
> > phantom pressure.
> >
> > I ran demo[1] that performs per-node proactive reclaim 10000 times
> > in qemu, writer cgroup memory.pressure as follows:
> >
> > without patch:
> >
> > some avg10=31.53 avg60=13.42 avg300=3.31 total=10602985
> > full avg10=31.53 avg60=13.42 avg300=3.31 total=10602985
> >
> > with patch:
> >
> > some avg10=9.59 avg60=3.41 avg300=0.81 total=2686221
> > full avg10=9.59 avg60=3.41 avg300=0.81 total=2686221
> >
> > [1] https://github.com/vernon2gh/app_and_module/tree/main/reclaim_node_psi
> >
> > Fixes: b980077899ea ("mm: introduce per-node proactive reclaim interface")
> > Cc: stable@vger.kernel.org
> > Signed-off-by: Vernon Yang <yanglincheng@kylinos.cn>
> > ---
>
> This seems to be a valid concern. Personally, I don't like
> having the code depend on whether `__node_reclaim()` and
> `node_reclaim()` are called from proactive reclaim or page
> allocation.
>
> Can't we check whether `sc->proactive` is true? Am I missing
> something?
It is also fine to directly check `sc->proactive` in __node_reclaim().
I chose the current coding because a previous similar fix commit
e22c6ed90aa9 was written this way, and it is also very clear.
Of course, it depends on everyone's preference. If everyone prefers to
directly check `sc->proactive`, please let me know clearly. Thanks!
--
Cheers,
Vernon
^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [PATCH] mm: vmscan: don't count per-node proactive reclaim as memory pressure
2026-10-05 10:10 ` Vernon Yang
@ 2026-10-05 12:54 ` Andrew Morton
0 siblings, 0 replies; 6+ messages in thread
From: Andrew Morton @ 2026-10-05 12:54 UTC (permalink / raw)
To: Vernon Yang
Cc: david, kasong, qi.zheng, shakeel.butt, baohua, axelrasmussen,
yuanchu, weixugc, baoquan.he, hannes, mhocko, ljs,
roman.gushchin, dave, linux-mm, linux-kernel, Vernon Yang,
stable
On Mon, 5 Oct 2026 18:10:45 +0800 Vernon Yang <vernon2gm@gmail.com> wrote:
> the phenomenon resembling
> > > senpai will appear again.
> >
> > What is this?
>
> This is commit e22c6ed90aa9, which addresses a scenario mentioned in
> memory limits of a cgroup, in detail as follows:
>
> Currently, this reclaim activity is accounted as memory pressure in the
> cgroup that the writer(!) belongs to. This is unexpected. It
> specifically causes problems for senpai
> (https://github.com/facebookincubator/senpai), which is an agent that
> routinely adjusts the memory limits and performs associated reclaim work
> in tens or even hundreds of cgroups running on the host. The cgroup that
> senpai is running in itself will report elevated levels of memory
> pressure, even though it itself is under no memory shortage or any sort of
> distress.
ah, thanks. It's arguably appropriate that the procfs-writing process
gets blamed for the I/O which it caused, but clearly that's the wrong
place to account the memory pressure.
^ permalink raw reply [flat|nested] 6+ messages in thread
end of thread, other threads:[~2026-10-05 12:54 UTC | newest]
Thread overview: 6+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-05 6:22 [PATCH] mm: vmscan: don't count per-node proactive reclaim as memory pressure Vernon Yang
2026-10-05 6:30 ` Barry Song
2026-10-05 10:21 ` Vernon Yang
2026-10-05 6:36 ` Andrew Morton
2026-10-05 10:10 ` Vernon Yang
2026-10-05 12:54 ` Andrew Morton
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®