From: Vernon Yang <vernon2gm@gmail.com>
To: akpm@linux-foundation.org, david@kernel.org, kasong@tencent.com,
qi.zheng@linux.dev, shakeel.butt@linux.dev, baohua@kernel.org,
axelrasmussen@google.com, yuanchu@google.com, weixugc@google.com,
baoquan.he@linux.dev, hannes@cmpxchg.org, mhocko@kernel.org,
ljs@kernel.org, roman.gushchin@linux.dev, dave@stgolabs.net
Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org,
Vernon Yang <yanglincheng@kylinos.cn>,
stable@vger.kernel.org
Subject: [PATCH] mm: vmscan: don't count per-node proactive reclaim as memory pressure
Date: Mon, 5 Oct 2026 14:22:36 +0800 [thread overview]
Message-ID: <20261005062236.564210-1-vernon2gm@gmail.com> (raw)
From: Vernon Yang <yanglincheng@kylinos.cn>
When the cgroup has no memory pressure at all, writing to
/sys/devices/system/nodeX/reclaim triggers proactive reclaim on
NUMA node, causing increase in the writer cgroup's memory PSI.
Due to this reclaim is performed in the context of the write(),
accounted as memory pressure on the writer, like
commit e22c6ed90aa9 ("mm: memcontrol: don't count limit-setting reclaim
as memory pressure"). This is unexpected, the phenomenon resembling
senpai will appear again.
The Documentation/ABI/stable/sysfs-devices-node documentation also
notes that "This interface is equivalent to the memcg variant."
This patch unifies the semantics of the memcg and node interfaces:
per-node proactive reclaim is no longer counted as memory pressure,
and the per-node proactive reclaim interface no longer produces
phantom pressure.
I ran demo[1] that performs per-node proactive reclaim 10000 times
in qemu, writer cgroup memory.pressure as follows:
without patch:
some avg10=31.53 avg60=13.42 avg300=3.31 total=10602985
full avg10=31.53 avg60=13.42 avg300=3.31 total=10602985
with patch:
some avg10=9.59 avg60=3.41 avg300=0.81 total=2686221
full avg10=9.59 avg60=3.41 avg300=0.81 total=2686221
[1] https://github.com/vernon2gh/app_and_module/tree/main/reclaim_node_psi
Fixes: b980077899ea ("mm: introduce per-node proactive reclaim interface")
Cc: stable@vger.kernel.org
Signed-off-by: Vernon Yang <yanglincheng@kylinos.cn>
---
mm/vmscan.c | 9 +++++----
1 file changed, 5 insertions(+), 4 deletions(-)
diff --git a/mm/vmscan.c b/mm/vmscan.c
index 91295070ca33..8e22b8a6038b 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -7988,13 +7988,10 @@ static unsigned long __node_reclaim(struct pglist_data *pgdat,
{
struct task_struct *p = current;
unsigned int noreclaim_flag;
- unsigned long pflags;
trace_mm_vmscan_node_reclaim_begin(pgdat->node_id, sc->order,
sc->gfp_mask);
- cond_resched();
- psi_memstall_enter(&pflags);
delayacct_freepages_start();
fs_reclaim_acquire(sc->gfp_mask);
/*
@@ -8011,7 +8008,6 @@ static unsigned long __node_reclaim(struct pglist_data *pgdat,
memalloc_noreclaim_restore(noreclaim_flag);
fs_reclaim_release(sc->gfp_mask);
delayacct_freepages_end();
- psi_memstall_leave(&pflags);
trace_mm_vmscan_node_reclaim_end(sc->nr_reclaimed, NULL);
@@ -8021,6 +8017,7 @@ static unsigned long __node_reclaim(struct pglist_data *pgdat,
unsigned long node_reclaim(struct pglist_data *pgdat, gfp_t gfp_mask, unsigned int order)
{
unsigned long ret;
+ unsigned long pflags;
/* Minimum pages needed in order to stay on node */
const unsigned long nr_pages = 1 << order;
struct scan_control sc = {
@@ -8067,7 +8064,10 @@ unsigned long node_reclaim(struct pglist_data *pgdat, gfp_t gfp_mask, unsigned i
if (test_and_set_bit_lock(PGDAT_RECLAIM_LOCKED, &pgdat->flags))
return 0;
+ cond_resched();
+ psi_memstall_enter(&pflags);
ret = __node_reclaim(pgdat, nr_pages, &sc);
+ psi_memstall_leave(&pflags);
clear_bit_unlock(PGDAT_RECLAIM_LOCKED, &pgdat->flags);
if (ret >= nr_pages)
@@ -8193,6 +8193,7 @@ int user_proactive_reclaim(char *buf,
&pgdat->flags))
return -EBUSY;
+ cond_resched();
reclaimed = __node_reclaim(pgdat, batch_size, &sc);
clear_bit_unlock(PGDAT_RECLAIM_LOCKED, &pgdat->flags);
}
--
2.53.0
next reply other threads:[~2026-10-05 6:22 UTC|newest]
Thread overview: 6+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-05 6:22 Vernon Yang [this message]
2026-10-05 6:30 ` Barry Song
2026-10-05 10:21 ` Vernon Yang
2026-10-05 6:36 ` Andrew Morton
2026-10-05 10:10 ` Vernon Yang
2026-10-05 12:54 ` Andrew Morton
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261005062236.564210-1-vernon2gm@gmail.com \
--to=vernon2gm@gmail.com \
--cc=akpm@linux-foundation.org \
--cc=axelrasmussen@google.com \
--cc=baohua@kernel.org \
--cc=baoquan.he@linux.dev \
--cc=dave@stgolabs.net \
--cc=david@kernel.org \
--cc=hannes@cmpxchg.org \
--cc=kasong@tencent.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=mhocko@kernel.org \
--cc=qi.zheng@linux.dev \
--cc=roman.gushchin@linux.dev \
--cc=shakeel.butt@linux.dev \
--cc=stable@vger.kernel.org \
--cc=weixugc@google.com \
--cc=yanglincheng@kylinos.cn \
--cc=yuanchu@google.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®