mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v2 0/2] mm/zswap: Fixes and improves the zswap global shrinker
@ 2026-07-17  8:51 Hao Jia
  2026-07-17  8:51 ` [PATCH v2 1/2] mm/zswap: Fix global shrinker when memory cgroup is disabled Hao Jia
                   ` (2 more replies)
  0 siblings, 3 replies; 31+ messages in thread
From: Hao Jia @ 2026-07-17  8:51 UTC (permalink / raw)
  To: akpm, tj, hannes, shakeel.butt, mhocko, yosry, mkoutny, nphamcs,
	chengming.zhou, muchun.song, roman.gushchin
  Cc: linux-mm, linux-kernel, linux-doc, Hao Jia

From: Hao Jia <jiahao1@lixiang.com>

This series fixes and improves the zswap global shrinker (shrink_worker()):
Patch 1: Fix missing global shrinker when memory cgroup is disabled.
Patch 2: Extend shrink_memcg() to support batch writeback and update its
         return value semantics, thereby improving the writeback efficiency
         in the shrink_worker() path.

v1->v2:
    - Add a reschedule check to the -ENOENT return path in shrink_memcg() to
      handle the theoretical issue of prolonged heavy concurrent zswap stores.
    - Remove the shrink_memcg() return value changes part, and include a more
       detailed test report in the commit message.


[v1] https://lore.kernel.org/all/20260714081510.16895-1-jiahao.kernel@gmail.com

Hao Jia (2):
  mm/zswap: Fix global shrinker when memory cgroup is disabled
  mm/zswap: Support batch writeback in shrink_memcg()

 mm/zswap.c | 60 ++++++++++++++++++++++++++++++++++++++++++------------
 1 file changed, 47 insertions(+), 13 deletions(-)

-- 
2.34.1


^ permalink raw reply	[flat|nested] 31+ messages in thread

* [PATCH v2 1/2] mm/zswap: Fix global shrinker when memory cgroup is disabled
  2026-07-17  8:51 [PATCH v2 0/2] mm/zswap: Fixes and improves the zswap global shrinker Hao Jia
@ 2026-07-17  8:51 ` Hao Jia
  2026-07-23  2:13   ` Johannes Weiner
  2026-07-17  8:51 ` [PATCH v2 2/2] mm/zswap: Support batch writeback in shrink_memcg() Hao Jia
  2026-07-18  1:18 ` [PATCH v2 0/2] mm/zswap: Fixes and improves the zswap global shrinker Andrew Morton
  2 siblings, 1 reply; 31+ messages in thread
From: Hao Jia @ 2026-07-17  8:51 UTC (permalink / raw)
  To: akpm, tj, hannes, shakeel.butt, mhocko, yosry, mkoutny, nphamcs,
	chengming.zhou, muchun.song, roman.gushchin
  Cc: linux-mm, linux-kernel, linux-doc, Hao Jia, stable

From: Hao Jia <jiahao1@lixiang.com>

Zswap writeback on hitting the pool limit is broken when memory cgroup
is disabled, because mem_cgroup_iter() always returns NULL. Therefore,
the global shrinker shrink_worker() always takes the !memcg branch.
After MAX_RECLAIM_RETRIES empty walks, the worker simply gives up, so it
fails to write back anything.

Therefore, when memory cgroup is disabled, fall through with the !memcg
branch and shrink the root memcg directly.

With memcg disabled, shrink_memcg() only returns -ENOENT when the root
LRU is empty, which means the total pages are already below thr. In the
absence of heavy concurrent zswap stores, the loop then safely bails out
via the zswap_total_pages() <= thr check; otherwise, it will resume
shrinking the memcg after processing the reschedule check. For any other
return value from shrink_memcg(), the loop is guaranteed to terminate,
either after MAX_RECLAIM_RETRIES failures or once the threshold is met.

Fixes: a65b0e7607cc ("zswap: make shrinking memcg-aware")
Cc: stable@vger.kernel.org
Suggested-by: Nhat Pham <nphamcs@gmail.com>
Acked-by: Nhat Pham <nphamcs@gmail.com>
Acked-by: Yosry Ahmed <yosry@kernel.org>
Reported-by: Yosry Ahmed <yosry@kernel.org>
Closes: https://lore.kernel.org/all/CAO9r8zPVzMKFbCixxD-qgtRrkFxWVrHiZZeLc=eyTPKPVQgX4g@mail.gmail.com
Signed-off-by: Hao Jia <jiahao1@lixiang.com>
---
 mm/zswap.c | 13 +++++++------
 1 file changed, 7 insertions(+), 6 deletions(-)

diff --git a/mm/zswap.c b/mm/zswap.c
index b5a17ea20237..48fc7b575e24 100644
--- a/mm/zswap.c
+++ b/mm/zswap.c
@@ -1356,11 +1356,12 @@ static void shrink_worker(struct work_struct *w)
 		} while (memcg && !mem_cgroup_tryget_online(memcg));
 		spin_unlock(&zswap_shrink_lock);
 
-		if (!memcg) {
-			/*
-			 * Continue shrinking without incrementing failures if
-			 * we found candidate memcgs in the last tree walk.
-			 */
+		/*
+		 * A NULL memcg ends a full hierarchy pass (except when memcg is
+		 * disabled, where it is always NULL: fall through to the root LRU).
+		 * Count a failure only if the last pass found no candidates.
+		 */
+		if (!memcg && !mem_cgroup_disabled()) {
 			if (!attempts && ++failures == MAX_RECLAIM_RETRIES)
 				break;
 
@@ -1379,7 +1380,7 @@ static void shrink_worker(struct work_struct *w)
 		 * and failures.
 		 */
 		if (ret == -ENOENT)
-			continue;
+			goto resched;
 		++attempts;
 
 		if (ret && ++failures == MAX_RECLAIM_RETRIES)
-- 
2.34.1


^ permalink raw reply	[flat|nested] 31+ messages in thread

* [PATCH v2 2/2] mm/zswap: Support batch writeback in shrink_memcg()
  2026-07-17  8:51 [PATCH v2 0/2] mm/zswap: Fixes and improves the zswap global shrinker Hao Jia
  2026-07-17  8:51 ` [PATCH v2 1/2] mm/zswap: Fix global shrinker when memory cgroup is disabled Hao Jia
@ 2026-07-17  8:51 ` Hao Jia
  2026-07-17 16:45   ` Yosry Ahmed
                     ` (2 more replies)
  2026-07-18  1:18 ` [PATCH v2 0/2] mm/zswap: Fixes and improves the zswap global shrinker Andrew Morton
  2 siblings, 3 replies; 31+ messages in thread
From: Hao Jia @ 2026-07-17  8:51 UTC (permalink / raw)
  To: akpm, tj, hannes, shakeel.butt, mhocko, yosry, mkoutny, nphamcs,
	chengming.zhou, muchun.song, roman.gushchin
  Cc: linux-mm, linux-kernel, linux-doc, Hao Jia

From: Hao Jia <jiahao1@lixiang.com>

Currently, shrink_memcg() writes back at most one entry per-node during
its traversal. This makes shrink_worker() inefficient, as it must
repeatedly re-enter shrink_memcg() to make any substantial progress.
Under high memory pressure, this can cause the writeback speed to be
too slow to keep up with refaults, leading to zswap store failures and
forcing pages to skip zswap and go directly to disk, which results in
an LRU inversion.

To address this, extend shrink_memcg() and rewrite its LRU iteration logic
to support batch writeback. Introduce the nr_to_scan parameter to bound how
many pages are scanned per call. This enables batch writeback in the
shrink_worker() path, while maintaining a low scan budget in the
zswap_store() path.

Test Setup:
- Total memory: 32 GB.
- zswap settings: max_pool_percent=1, accept_threshold_percent=50,
  shrinker_enabled=N.

Test Case 1:
Allocate 512MB of anonymous pages and fill them with random data (to avoid
compression), then use cgroup memory.reclaim to force a large amount of
anonymous pages into zswap. At an interval of 2ms, allocate a 4K anonymous
page where the first 4 bytes are random numbers and the rest are zeros, and
then trigger a reclamation of this 4K anonymous page through cgroup
memory.reclaim. When the pool threshold is reached, shrink_memcg() will
be triggered.
The test data after running for 120s is as follows:
                                Baseline       Patched
shrink_worker wakeups              5,363            85
shrink_memcg calls            11,373,201       180,928
written_back pages                40,212        40,236
zswap_store calls                161,190       168,741
   store succeeded (ret=1)       102,743       127,644
   store rejected (ret=0)         58,447        41,097
   store reject rate                ~36%          ~24%
pool_limit_hit delta              55,826        14,062
pswpout                           98,659        81,333
pswpin                                 2             1

Test Case 2:
To consistently force zswap store failures and trigger shrink_worker(),
the following stress-ng command was run for 120 seconds within a cgroup
limited to a memory.max of 1G:
   bash -c 'echo $$ > /sys/fs/cgroup/zswaptest/cgroup.procs ; \
   exec stress-ng --vm 4 --vm-bytes 4G --vm-keep --vm-method rand-set -t \
120s -q'
The test data after running for 120s is as follows:
                                Baseline       Patched
shrink_worker wakeups              5,640           987
shrink_memcg calls             8,481,500     2,504,818
written_back pages                   260       768,576
zswap_store calls              2,742,756     2,301,414
   store succeeded (ret=1)       934,640     1,308,686
   store rejected (ret=0)      1,808,116       992,728
   store reject rate                ~66%          ~43%
pool_limit_hit delta           1,181,310       101,593
pswpout                        1,808,376     1,761,304
pswpin                         4,288,497     3,902,658

Under identical workloads and runtimes, batching the zswap shrinker
exhibits a significant reduction in both shrink_worker wakeups and
shrink_memcg calls. Furthermore, the sharp drop in both pool_limit_hit
and zswap_store rejections demonstrates that batching the zswap shrinker
effectively mitigates zswap_store failures caused by hitting the pool
limit. This significantly prevents pages from bypassing zswap and falling
back directly to disk, thereby reducing LRU inversion.

Suggested-by: Yosry Ahmed <yosry@kernel.org>
Signed-off-by: Hao Jia <jiahao1@lixiang.com>
---
 mm/zswap.c | 47 ++++++++++++++++++++++++++++++++++++++++-------
 1 file changed, 40 insertions(+), 7 deletions(-)

diff --git a/mm/zswap.c b/mm/zswap.c
index 48fc7b575e24..6a09b9ebfd25 100644
--- a/mm/zswap.c
+++ b/mm/zswap.c
@@ -1275,9 +1275,27 @@ static struct shrinker *zswap_alloc_shrinker(void)
 	return shrinker;
 }
 
-static int shrink_memcg(struct mem_cgroup *memcg)
+#define NR_ZSWAP_WB_BATCH	64UL
+
+/*
+ * Scan up to @nr_to_scan pages across the per-node zswap LRUs of @memcg
+ * and write back the reclaimable ones.
+ *
+ * Since the second-chance algorithm rotates referenced entries to the
+ * LRU tail, the per-node scan is capped at the current LRU length so
+ * each entry is scanned at most once per call. It is up to the caller
+ * to handle retries, deciding whether to scan another memcg to complete
+ * the full iteration, or to rescan the current memcg to drain its zswap
+ * entries.
+ *
+ * Return: 0 if at least one entry was written back, -EAGAIN if entries
+ * were scanned but none could be written back, or -ENOENT if @memcg has
+ * writeback disabled, is a zombie cgroup, or has empty zswap LRUs.
+ */
+static int shrink_memcg(struct mem_cgroup *memcg, unsigned long nr_to_scan)
 {
-	int nid, shrunk = 0, scanned = 0;
+	unsigned long nr_remaining = nr_to_scan;
+	int nid, shrunk = 0;
 
 	if (!mem_cgroup_zswap_writeback_enabled(memcg))
 		return -ENOENT;
@@ -1290,14 +1308,29 @@ static int shrink_memcg(struct mem_cgroup *memcg)
 		return -ENOENT;
 
 	for_each_node_state(nid, N_NORMAL_MEMORY) {
-		unsigned long nr_to_walk = 1;
+		unsigned long nr_to_walk;
 
+		/*
+		 * Cap the scan at per-node LRU length so each entry is scanned
+		 * at most once per call.
+		 */
+		nr_to_walk = min(nr_remaining,
+				 list_lru_count_one(&zswap_list_lru, nid, memcg));
+		if (!nr_to_walk)
+			continue;
+
+		nr_remaining -= nr_to_walk;
 		shrunk += list_lru_walk_one(&zswap_list_lru, nid, memcg,
 					    &shrink_memcg_cb, NULL, &nr_to_walk);
-		scanned += 1 - nr_to_walk;
+		/* Return the unused share of the budget to the pool. */
+		nr_remaining += nr_to_walk;
+
+		if (!nr_remaining)
+			break;
 	}
 
-	if (!scanned)
+	/* Nothing was scanned: every LRU under @memcg was empty. */
+	if (nr_remaining == nr_to_scan)
 		return -ENOENT;
 
 	return shrunk ? 0 : -EAGAIN;
@@ -1369,7 +1402,7 @@ static void shrink_worker(struct work_struct *w)
 			goto resched;
 		}
 
-		ret = shrink_memcg(memcg);
+		ret = shrink_memcg(memcg, NR_ZSWAP_WB_BATCH);
 		/* drop the extra reference */
 		mem_cgroup_put(memcg);
 
@@ -1493,7 +1526,7 @@ bool zswap_store(struct folio *folio)
 	objcg = get_obj_cgroup_from_folio(folio);
 	if (objcg && !obj_cgroup_may_zswap(objcg)) {
 		memcg = get_mem_cgroup_from_objcg(objcg);
-		if (shrink_memcg(memcg)) {
+		if (shrink_memcg(memcg, 1)) {
 			mem_cgroup_put(memcg);
 			goto put_objcg;
 		}
-- 
2.34.1


^ permalink raw reply	[flat|nested] 31+ messages in thread

* Re: [PATCH v2 2/2] mm/zswap: Support batch writeback in shrink_memcg()
  2026-07-17  8:51 ` [PATCH v2 2/2] mm/zswap: Support batch writeback in shrink_memcg() Hao Jia
@ 2026-07-17 16:45   ` Yosry Ahmed
  2026-07-17 16:46   ` Nhat Pham
  2026-07-23  2:27   ` Johannes Weiner
  2 siblings, 0 replies; 31+ messages in thread
From: Yosry Ahmed @ 2026-07-17 16:45 UTC (permalink / raw)
  To: Hao Jia
  Cc: akpm, tj, hannes, shakeel.butt, mhocko, mkoutny, nphamcs,
	chengming.zhou, muchun.song, roman.gushchin, linux-mm,
	linux-kernel, linux-doc, Hao Jia

On Fri, Jul 17, 2026 at 04:51:51PM +0800, Hao Jia wrote:
> From: Hao Jia <jiahao1@lixiang.com>
> 
> Currently, shrink_memcg() writes back at most one entry per-node during
> its traversal. This makes shrink_worker() inefficient, as it must
> repeatedly re-enter shrink_memcg() to make any substantial progress.
> Under high memory pressure, this can cause the writeback speed to be
> too slow to keep up with refaults, leading to zswap store failures and
> forcing pages to skip zswap and go directly to disk, which results in
> an LRU inversion.
> 
> To address this, extend shrink_memcg() and rewrite its LRU iteration logic
> to support batch writeback. Introduce the nr_to_scan parameter to bound how
> many pages are scanned per call. This enables batch writeback in the
> shrink_worker() path, while maintaining a low scan budget in the
> zswap_store() path.
> 
> Test Setup:
> - Total memory: 32 GB.
> - zswap settings: max_pool_percent=1, accept_threshold_percent=50,
>   shrinker_enabled=N.
> 
> Test Case 1:
> Allocate 512MB of anonymous pages and fill them with random data (to avoid
> compression), then use cgroup memory.reclaim to force a large amount of
> anonymous pages into zswap. At an interval of 2ms, allocate a 4K anonymous
> page where the first 4 bytes are random numbers and the rest are zeros, and
> then trigger a reclamation of this 4K anonymous page through cgroup
> memory.reclaim. When the pool threshold is reached, shrink_memcg() will
> be triggered.
> The test data after running for 120s is as follows:
>                                 Baseline       Patched
> shrink_worker wakeups              5,363            85
> shrink_memcg calls            11,373,201       180,928
> written_back pages                40,212        40,236
> zswap_store calls                161,190       168,741
>    store succeeded (ret=1)       102,743       127,644
>    store rejected (ret=0)         58,447        41,097
>    store reject rate                ~36%          ~24%
> pool_limit_hit delta              55,826        14,062
> pswpout                           98,659        81,333
> pswpin                                 2             1
> 
> Test Case 2:
> To consistently force zswap store failures and trigger shrink_worker(),
> the following stress-ng command was run for 120 seconds within a cgroup
> limited to a memory.max of 1G:
>    bash -c 'echo $$ > /sys/fs/cgroup/zswaptest/cgroup.procs ; \
>    exec stress-ng --vm 4 --vm-bytes 4G --vm-keep --vm-method rand-set -t \
> 120s -q'
> The test data after running for 120s is as follows:
>                                 Baseline       Patched
> shrink_worker wakeups              5,640           987
> shrink_memcg calls             8,481,500     2,504,818
> written_back pages                   260       768,576
> zswap_store calls              2,742,756     2,301,414
>    store succeeded (ret=1)       934,640     1,308,686
>    store rejected (ret=0)      1,808,116       992,728
>    store reject rate                ~66%          ~43%
> pool_limit_hit delta           1,181,310       101,593
> pswpout                        1,808,376     1,761,304
> pswpin                         4,288,497     3,902,658
> 
> Under identical workloads and runtimes, batching the zswap shrinker
> exhibits a significant reduction in both shrink_worker wakeups and
> shrink_memcg calls. Furthermore, the sharp drop in both pool_limit_hit
> and zswap_store rejections demonstrates that batching the zswap shrinker
> effectively mitigates zswap_store failures caused by hitting the pool
> limit. This significantly prevents pages from bypassing zswap and falling
> back directly to disk, thereby reducing LRU inversion.
> 
> Suggested-by: Yosry Ahmed <yosry@kernel.org>
> Signed-off-by: Hao Jia <jiahao1@lixiang.com>

Acked-by: Yosry Ahmed <yosry@kernel.org>

^ permalink raw reply	[flat|nested] 31+ messages in thread

* Re: [PATCH v2 2/2] mm/zswap: Support batch writeback in shrink_memcg()
  2026-07-17  8:51 ` [PATCH v2 2/2] mm/zswap: Support batch writeback in shrink_memcg() Hao Jia
  2026-07-17 16:45   ` Yosry Ahmed
@ 2026-07-17 16:46   ` Nhat Pham
  2026-07-23  2:27   ` Johannes Weiner
  2 siblings, 0 replies; 31+ messages in thread
From: Nhat Pham @ 2026-07-17 16:46 UTC (permalink / raw)
  To: Hao Jia
  Cc: akpm, tj, hannes, shakeel.butt, mhocko, yosry, mkoutny,
	chengming.zhou, muchun.song, roman.gushchin, linux-mm,
	linux-kernel, linux-doc, Hao Jia

On Fri, Jul 17, 2026 at 1:52 AM Hao Jia <jiahao.kernel@gmail.com> wrote:
>
> From: Hao Jia <jiahao1@lixiang.com>
>
> Currently, shrink_memcg() writes back at most one entry per-node during
> its traversal. This makes shrink_worker() inefficient, as it must
> repeatedly re-enter shrink_memcg() to make any substantial progress.
> Under high memory pressure, this can cause the writeback speed to be
> too slow to keep up with refaults, leading to zswap store failures and
> forcing pages to skip zswap and go directly to disk, which results in
> an LRU inversion.
>
> To address this, extend shrink_memcg() and rewrite its LRU iteration logic
> to support batch writeback. Introduce the nr_to_scan parameter to bound how
> many pages are scanned per call. This enables batch writeback in the
> shrink_worker() path, while maintaining a low scan budget in the
> zswap_store() path.
>
> Test Setup:
> - Total memory: 32 GB.
> - zswap settings: max_pool_percent=1, accept_threshold_percent=50,
>   shrinker_enabled=N.
>
> Test Case 1:
> Allocate 512MB of anonymous pages and fill them with random data (to avoid
> compression), then use cgroup memory.reclaim to force a large amount of
> anonymous pages into zswap. At an interval of 2ms, allocate a 4K anonymous
> page where the first 4 bytes are random numbers and the rest are zeros, and
> then trigger a reclamation of this 4K anonymous page through cgroup
> memory.reclaim. When the pool threshold is reached, shrink_memcg() will
> be triggered.
> The test data after running for 120s is as follows:
>                                 Baseline       Patched
> shrink_worker wakeups              5,363            85
> shrink_memcg calls            11,373,201       180,928
> written_back pages                40,212        40,236
> zswap_store calls                161,190       168,741
>    store succeeded (ret=1)       102,743       127,644
>    store rejected (ret=0)         58,447        41,097
>    store reject rate                ~36%          ~24%
> pool_limit_hit delta              55,826        14,062
> pswpout                           98,659        81,333
> pswpin                                 2             1
>
> Test Case 2:
> To consistently force zswap store failures and trigger shrink_worker(),
> the following stress-ng command was run for 120 seconds within a cgroup
> limited to a memory.max of 1G:
>    bash -c 'echo $$ > /sys/fs/cgroup/zswaptest/cgroup.procs ; \
>    exec stress-ng --vm 4 --vm-bytes 4G --vm-keep --vm-method rand-set -t \
> 120s -q'
> The test data after running for 120s is as follows:
>                                 Baseline       Patched
> shrink_worker wakeups              5,640           987
> shrink_memcg calls             8,481,500     2,504,818
> written_back pages                   260       768,576
> zswap_store calls              2,742,756     2,301,414
>    store succeeded (ret=1)       934,640     1,308,686
>    store rejected (ret=0)      1,808,116       992,728
>    store reject rate                ~66%          ~43%
> pool_limit_hit delta           1,181,310       101,593
> pswpout                        1,808,376     1,761,304
> pswpin                         4,288,497     3,902,658
>

Acked-by: Nhat Pham <nphamcs@gmail.com>

^ permalink raw reply	[flat|nested] 31+ messages in thread

* Re: [PATCH v2 0/2] mm/zswap: Fixes and improves the zswap global shrinker
  2026-07-17  8:51 [PATCH v2 0/2] mm/zswap: Fixes and improves the zswap global shrinker Hao Jia
  2026-07-17  8:51 ` [PATCH v2 1/2] mm/zswap: Fix global shrinker when memory cgroup is disabled Hao Jia
  2026-07-17  8:51 ` [PATCH v2 2/2] mm/zswap: Support batch writeback in shrink_memcg() Hao Jia
@ 2026-07-18  1:18 ` Andrew Morton
  2026-07-18  1:22   ` Yosry Ahmed
  2026-07-18  1:28   ` Yosry Ahmed
  2 siblings, 2 replies; 31+ messages in thread
From: Andrew Morton @ 2026-07-18  1:18 UTC (permalink / raw)
  To: Hao Jia
  Cc: tj, hannes, shakeel.butt, mhocko, yosry, mkoutny, nphamcs,
	chengming.zhou, muchun.song, roman.gushchin, linux-mm,
	linux-kernel, linux-doc, Hao Jia

On Fri, 17 Jul 2026 16:51:49 +0800 Hao Jia <jiahao.kernel@gmail.com> wrote:

> This series fixes and improves the zswap global shrinker (shrink_worker()):
> Patch 1: Fix missing global shrinker when memory cgroup is disabled.
> Patch 2: Extend shrink_memcg() to support batch writeback and update its
>          return value semantics, thereby improving the writeback efficiency
>          in the shrink_worker() path.

Thanks.

[1/2] is a cc:stable fix so it isn't really appropriate to combine this
with [2/2] which doesn't fix any bugs.  Because the two patches may
take different paths into mainline, with different timings.  But that's
OK, I can deal with the splitup if needed.

The [1/2] changelog lacks a description of how the flaw impacts users. 
Please describe this fully and maintain that info within the
changelogging.  This info helps -stable maintainers and others
understand why we're proposing a backport and helps myself and others
with timing decisions.

Finally, AI review might have found an issue:
	https://sashiko.dev/#/patchset/20260717085151.22822-1-jiahao.kernel@gmail.com



^ permalink raw reply	[flat|nested] 31+ messages in thread

* Re: [PATCH v2 0/2] mm/zswap: Fixes and improves the zswap global shrinker
  2026-07-18  1:18 ` [PATCH v2 0/2] mm/zswap: Fixes and improves the zswap global shrinker Andrew Morton
@ 2026-07-18  1:22   ` Yosry Ahmed
  2026-07-18  1:28   ` Yosry Ahmed
  1 sibling, 0 replies; 31+ messages in thread
From: Yosry Ahmed @ 2026-07-18  1:22 UTC (permalink / raw)
  To: Andrew Morton
  Cc: Hao Jia, tj, hannes, shakeel.butt, mhocko, mkoutny, nphamcs,
	chengming.zhou, muchun.song, roman.gushchin, linux-mm,
	linux-kernel, linux-doc, Hao Jia

On Fri, Jul 17, 2026 at 6:18 PM Andrew Morton <akpm@linux-foundation.org> wrote:
>
> On Fri, 17 Jul 2026 16:51:49 +0800 Hao Jia <jiahao.kernel@gmail.com> wrote:
>
> > This series fixes and improves the zswap global shrinker (shrink_worker()):
> > Patch 1: Fix missing global shrinker when memory cgroup is disabled.
> > Patch 2: Extend shrink_memcg() to support batch writeback and update its
> >          return value semantics, thereby improving the writeback efficiency
> >          in the shrink_worker() path.
>
> Thanks.
>
> [1/2] is a cc:stable fix so it isn't really appropriate to combine this
> with [2/2] which doesn't fix any bugs.  Because the two patches may
> take different paths into mainline, with different timings.  But that's
> OK, I can deal with the splitup if needed.
>
> The [1/2] changelog lacks a description of how the flaw impacts users.
> Please describe this fully and maintain that info within the
> changelogging.  This info helps -stable maintainers and others
> understand why we're proposing a backport and helps myself and others
> with timing decisions.
>
> Finally, AI review might have found an issue:
>         https://sashiko.dev/#/patchset/20260717085151.22822-1-jiahao.kernel@gmail.com

We discussed this one in the previous version, it's a theoretical
scenario that can already happen today.

^ permalink raw reply	[flat|nested] 31+ messages in thread

* Re: [PATCH v2 0/2] mm/zswap: Fixes and improves the zswap global shrinker
  2026-07-18  1:18 ` [PATCH v2 0/2] mm/zswap: Fixes and improves the zswap global shrinker Andrew Morton
  2026-07-18  1:22   ` Yosry Ahmed
@ 2026-07-18  1:28   ` Yosry Ahmed
  2026-07-18  4:40     ` Andrew Morton
  1 sibling, 1 reply; 31+ messages in thread
From: Yosry Ahmed @ 2026-07-18  1:28 UTC (permalink / raw)
  To: Andrew Morton
  Cc: Hao Jia, tj, hannes, shakeel.butt, mhocko, mkoutny, nphamcs,
	chengming.zhou, muchun.song, roman.gushchin, linux-mm,
	linux-kernel, linux-doc, Hao Jia

On Fri, Jul 17, 2026 at 6:18 PM Andrew Morton <akpm@linux-foundation.org> wrote:
>
> On Fri, 17 Jul 2026 16:51:49 +0800 Hao Jia <jiahao.kernel@gmail.com> wrote:
>
> > This series fixes and improves the zswap global shrinker (shrink_worker()):
> > Patch 1: Fix missing global shrinker when memory cgroup is disabled.
> > Patch 2: Extend shrink_memcg() to support batch writeback and update its
> >          return value semantics, thereby improving the writeback efficiency
> >          in the shrink_worker() path.
>
> Thanks.
>
> [1/2] is a cc:stable fix so it isn't really appropriate to combine this
> with [2/2] which doesn't fix any bugs.  Because the two patches may
> take different paths into mainline, with different timings.  But that's
> OK, I can deal with the splitup if needed.

Thank you!

>
> The [1/2] changelog lacks a description of how the flaw impacts users.
> Please describe this fully and maintain that info within the
> changelogging.  This info helps -stable maintainers and others
> understand why we're proposing a backport and helps myself and others
> with timing decisions.

The first line in the changelog should be sufficient imo: "Zswap
writeback on hitting the pool limit is broken when memory cgroup is
disabled"

^ permalink raw reply	[flat|nested] 31+ messages in thread

* Re: [PATCH v2 0/2] mm/zswap: Fixes and improves the zswap global shrinker
  2026-07-18  1:28   ` Yosry Ahmed
@ 2026-07-18  4:40     ` Andrew Morton
  2026-07-20  1:26       ` Hao Jia
  0 siblings, 1 reply; 31+ messages in thread
From: Andrew Morton @ 2026-07-18  4:40 UTC (permalink / raw)
  To: Yosry Ahmed
  Cc: Hao Jia, tj, hannes, shakeel.butt, mhocko, mkoutny, nphamcs,
	chengming.zhou, muchun.song, roman.gushchin, linux-mm,
	linux-kernel, linux-doc, Hao Jia

On Fri, 17 Jul 2026 18:28:04 -0700 Yosry Ahmed <yosry@kernel.org> wrote:

> >
> > The [1/2] changelog lacks a description of how the flaw impacts users.
> > Please describe this fully and maintain that info within the
> > changelogging.  This info helps -stable maintainers and others
> > understand why we're proposing a backport and helps myself and others
> > with timing decisions.
> 
> The first line in the changelog should be sufficient imo: "Zswap
> writeback on hitting the pool limit is broken when memory cgroup is
> disabled"

"broken"?  Perhaps this means "fails to occur".

But what is the userspace-visible impact?  IOW, why are we proposing a
backport?


^ permalink raw reply	[flat|nested] 31+ messages in thread

* Re: [PATCH v2 0/2] mm/zswap: Fixes and improves the zswap global shrinker
  2026-07-18  4:40     ` Andrew Morton
@ 2026-07-20  1:26       ` Hao Jia
  2026-07-23  1:21         ` Hao Jia
  0 siblings, 1 reply; 31+ messages in thread
From: Hao Jia @ 2026-07-20  1:26 UTC (permalink / raw)
  To: Andrew Morton, Yosry Ahmed
  Cc: tj, hannes, shakeel.butt, mhocko, mkoutny, nphamcs,
	chengming.zhou, muchun.song, roman.gushchin, linux-mm,
	linux-kernel, linux-doc, Hao Jia



On 2026/7/18 12:40, Andrew Morton wrote:
> On Fri, 17 Jul 2026 18:28:04 -0700 Yosry Ahmed <yosry@kernel.org> wrote:
> 
>>>
>>> The [1/2] changelog lacks a description of how the flaw impacts users.
>>> Please describe this fully and maintain that info within the
>>> changelogging.  This info helps -stable maintainers and others
>>> understand why we're proposing a backport and helps myself and others
>>> with timing decisions.
>>
>> The first line in the changelog should be sufficient imo: "Zswap
>> writeback on hitting the pool limit is broken when memory cgroup is
>> disabled"
> 
> "broken"?  Perhaps this means "fails to occur".
> 
> But what is the userspace-visible impact?  IOW, why are we proposing a
> backport?
> 

Perhaps the first paragraph of the commit1 message could be modified as 
follows? I have added a description of the issues that occur without 
this patch.

Zswap writeback on hitting the pool limit fails to occur when memory 
cgroup is disabled, because mem_cgroup_iter() always returns NULL. 
Therefore, the global shrinker shrink_worker() always takes the !memcg 
branch. After MAX_RECLAIM_RETRIES empty walks, the worker simply gives 
up, so it fails to write back anything. As a result, once the pool 
reaches the zswap limit, every subsequent zswap shrink work run is a 
no-op. This leads to zswap store failures, forcing pages to bypass zswap 
and be written directly to the backing swap device, which can trigger 
issues such as LRU inversion.


Thanks,
Hao

^ permalink raw reply	[flat|nested] 31+ messages in thread

* Re: [PATCH v2 0/2] mm/zswap: Fixes and improves the zswap global shrinker
  2026-07-20  1:26       ` Hao Jia
@ 2026-07-23  1:21         ` Hao Jia
  2026-07-23  4:49           ` Yosry Ahmed
  0 siblings, 1 reply; 31+ messages in thread
From: Hao Jia @ 2026-07-23  1:21 UTC (permalink / raw)
  To: Andrew Morton, Yosry Ahmed, nphamcs
  Cc: tj, hannes, shakeel.butt, mhocko, mkoutny, chengming.zhou,
	muchun.song, roman.gushchin, linux-mm, linux-kernel, linux-doc,
	Hao Jia



On 2026/7/20 09:26, Hao Jia wrote:
> 
> 
> On 2026/7/18 12:40, Andrew Morton wrote:
>> On Fri, 17 Jul 2026 18:28:04 -0700 Yosry Ahmed <yosry@kernel.org> wrote:
>>
>>>>
>>>> The [1/2] changelog lacks a description of how the flaw impacts users.
>>>> Please describe this fully and maintain that info within the
>>>> changelogging.  This info helps -stable maintainers and others
>>>> understand why we're proposing a backport and helps myself and others
>>>> with timing decisions.
>>>
>>> The first line in the changelog should be sufficient imo: "Zswap
>>> writeback on hitting the pool limit is broken when memory cgroup is
>>> disabled"
>>
>> "broken"?  Perhaps this means "fails to occur".
>>
>> But what is the userspace-visible impact?  IOW, why are we proposing a
>> backport?
>>
> 
> Perhaps the first paragraph of the commit1 message could be modified as 
> follows? I have added a description of the issues that occur without 
> this patch.
> 

Hi Andrew, Yosry, and Nhat,
Any thoughts on this change?

Thanks,
Hao
> Zswap writeback on hitting the pool limit fails to occur when memory 
> cgroup is disabled, because mem_cgroup_iter() always returns NULL. 
> Therefore, the global shrinker shrink_worker() always takes the !memcg 
> branch. After MAX_RECLAIM_RETRIES empty walks, the worker simply gives 
> up, so it fails to write back anything. As a result, once the pool 
> reaches the zswap limit, every subsequent zswap shrink work run is a 
> no-op. This leads to zswap store failures, forcing pages to bypass zswap 
> and be written directly to the backing swap device, which can trigger 
> issues such as LRU inversion.
> 
> 
> Thanks,
> Hao

^ permalink raw reply	[flat|nested] 31+ messages in thread

* Re: [PATCH v2 1/2] mm/zswap: Fix global shrinker when memory cgroup is disabled
  2026-07-17  8:51 ` [PATCH v2 1/2] mm/zswap: Fix global shrinker when memory cgroup is disabled Hao Jia
@ 2026-07-23  2:13   ` Johannes Weiner
  0 siblings, 0 replies; 31+ messages in thread
From: Johannes Weiner @ 2026-07-23  2:13 UTC (permalink / raw)
  To: Hao Jia
  Cc: akpm, tj, shakeel.butt, mhocko, yosry, mkoutny, nphamcs,
	chengming.zhou, muchun.song, roman.gushchin, linux-mm,
	linux-kernel, linux-doc, Hao Jia, stable

On Fri, Jul 17, 2026 at 04:51:50PM +0800, Hao Jia wrote:
> From: Hao Jia <jiahao1@lixiang.com>
> 
> Zswap writeback on hitting the pool limit is broken when memory cgroup
> is disabled, because mem_cgroup_iter() always returns NULL. Therefore,
> the global shrinker shrink_worker() always takes the !memcg branch.
> After MAX_RECLAIM_RETRIES empty walks, the worker simply gives up, so it
> fails to write back anything.
> 
> Therefore, when memory cgroup is disabled, fall through with the !memcg
> branch and shrink the root memcg directly.
> 
> With memcg disabled, shrink_memcg() only returns -ENOENT when the root
> LRU is empty, which means the total pages are already below thr. In the
> absence of heavy concurrent zswap stores, the loop then safely bails out
> via the zswap_total_pages() <= thr check; otherwise, it will resume
> shrinking the memcg after processing the reschedule check. For any other
> return value from shrink_memcg(), the loop is guaranteed to terminate,
> either after MAX_RECLAIM_RETRIES failures or once the threshold is met.
> 
> Fixes: a65b0e7607cc ("zswap: make shrinking memcg-aware")
> Cc: stable@vger.kernel.org
> Suggested-by: Nhat Pham <nphamcs@gmail.com>
> Acked-by: Nhat Pham <nphamcs@gmail.com>
> Acked-by: Yosry Ahmed <yosry@kernel.org>
> Reported-by: Yosry Ahmed <yosry@kernel.org>
> Closes: https://lore.kernel.org/all/CAO9r8zPVzMKFbCixxD-qgtRrkFxWVrHiZZeLc=eyTPKPVQgX4g@mail.gmail.com
> Signed-off-by: Hao Jia <jiahao1@lixiang.com>

Acked-by: Johannes Weiner <hannes@cmpxchg.org>

^ permalink raw reply	[flat|nested] 31+ messages in thread

* Re: [PATCH v2 2/2] mm/zswap: Support batch writeback in shrink_memcg()
  2026-07-17  8:51 ` [PATCH v2 2/2] mm/zswap: Support batch writeback in shrink_memcg() Hao Jia
  2026-07-17 16:45   ` Yosry Ahmed
  2026-07-17 16:46   ` Nhat Pham
@ 2026-07-23  2:27   ` Johannes Weiner
  2026-07-23  4:52     ` Yosry Ahmed
  2 siblings, 1 reply; 31+ messages in thread
From: Johannes Weiner @ 2026-07-23  2:27 UTC (permalink / raw)
  To: Hao Jia
  Cc: akpm, tj, shakeel.butt, mhocko, yosry, mkoutny, nphamcs,
	chengming.zhou, muchun.song, roman.gushchin, linux-mm,
	linux-kernel, linux-doc, Hao Jia

On Fri, Jul 17, 2026 at 04:51:51PM +0800, Hao Jia wrote:
> @@ -1369,7 +1402,7 @@ static void shrink_worker(struct work_struct *w)
>  			goto resched;
>  		}
>  
> -		ret = shrink_memcg(memcg);
> +		ret = shrink_memcg(memcg, NR_ZSWAP_WB_BATCH);
>  		/* drop the extra reference */
>  		mem_cgroup_put(memcg);
>  
> @@ -1493,7 +1526,7 @@ bool zswap_store(struct folio *folio)
>  	objcg = get_obj_cgroup_from_folio(folio);
>  	if (objcg && !obj_cgroup_may_zswap(objcg)) {
>  		memcg = get_mem_cgroup_from_objcg(objcg);
> -		if (shrink_memcg(memcg)) {
> +		if (shrink_memcg(memcg, 1)) {

Why 64 for the global limit but only 1 for the cgroup limit? That
seems arbitrary in multiple ways.

Direct reclaim, kswapd, proactive reclaim, cgroup limit reclaim use
SWAP_CLUSTER_MAX for the batch size near-universally. It's magic too
to be sure, but at least you wouldn't have to make up new magic?

Otherwise, the patch looks good to me.

^ permalink raw reply	[flat|nested] 31+ messages in thread

* Re: [PATCH v2 0/2] mm/zswap: Fixes and improves the zswap global shrinker
  2026-07-23  1:21         ` Hao Jia
@ 2026-07-23  4:49           ` Yosry Ahmed
  2026-07-24 10:22             ` Hao Jia
  0 siblings, 1 reply; 31+ messages in thread
From: Yosry Ahmed @ 2026-07-23  4:49 UTC (permalink / raw)
  To: Hao Jia
  Cc: Andrew Morton, nphamcs, tj, hannes, shakeel.butt, mhocko,
	mkoutny, chengming.zhou, muchun.song, roman.gushchin, linux-mm,
	linux-kernel, linux-doc, Hao Jia

On Wed, Jul 22, 2026 at 6:21 PM Hao Jia <jiahao.kernel@gmail.com> wrote:
>
>
>
> On 2026/7/20 09:26, Hao Jia wrote:
> >
> >
> > On 2026/7/18 12:40, Andrew Morton wrote:
> >> On Fri, 17 Jul 2026 18:28:04 -0700 Yosry Ahmed <yosry@kernel.org> wrote:
> >>
> >>>>
> >>>> The [1/2] changelog lacks a description of how the flaw impacts users.
> >>>> Please describe this fully and maintain that info within the
> >>>> changelogging.  This info helps -stable maintainers and others
> >>>> understand why we're proposing a backport and helps myself and others
> >>>> with timing decisions.
> >>>
> >>> The first line in the changelog should be sufficient imo: "Zswap
> >>> writeback on hitting the pool limit is broken when memory cgroup is
> >>> disabled"
> >>
> >> "broken"?  Perhaps this means "fails to occur".
> >>
> >> But what is the userspace-visible impact?  IOW, why are we proposing a
> >> backport?
> >>
> >
> > Perhaps the first paragraph of the commit1 message could be modified as
> > follows? I have added a description of the issues that occur without
> > this patch.
> >
>
> Hi Andrew, Yosry, and Nhat,
> Any thoughts on this change?

I would front load the user impact in the first paragraph:

Zswap writeback when the global pool limit is hit fails when memory
cgroups are disabled. The pool remains full until it is organically
drained by swapins or memory freeing, leading to zswap store failures
and pages bypassing getting written directly to the backing swap device,
causing LRU inversion (hotter pages with higher fault latency).

...

^ permalink raw reply	[flat|nested] 31+ messages in thread

* Re: [PATCH v2 2/2] mm/zswap: Support batch writeback in shrink_memcg()
  2026-07-23  2:27   ` Johannes Weiner
@ 2026-07-23  4:52     ` Yosry Ahmed
  2026-07-23 13:55       ` Johannes Weiner
  0 siblings, 1 reply; 31+ messages in thread
From: Yosry Ahmed @ 2026-07-23  4:52 UTC (permalink / raw)
  To: Johannes Weiner
  Cc: Hao Jia, akpm, tj, shakeel.butt, mhocko, mkoutny, nphamcs,
	chengming.zhou, muchun.song, roman.gushchin, linux-mm,
	linux-kernel, linux-doc, Hao Jia

On Wed, Jul 22, 2026 at 7:27 PM Johannes Weiner <hannes@cmpxchg.org> wrote:
>
> On Fri, Jul 17, 2026 at 04:51:51PM +0800, Hao Jia wrote:
> > @@ -1369,7 +1402,7 @@ static void shrink_worker(struct work_struct *w)
> >                       goto resched;
> >               }
> >
> > -             ret = shrink_memcg(memcg);
> > +             ret = shrink_memcg(memcg, NR_ZSWAP_WB_BATCH);
> >               /* drop the extra reference */
> >               mem_cgroup_put(memcg);
> >
> > @@ -1493,7 +1526,7 @@ bool zswap_store(struct folio *folio)
> >       objcg = get_obj_cgroup_from_folio(folio);
> >       if (objcg && !obj_cgroup_may_zswap(objcg)) {
> >               memcg = get_mem_cgroup_from_objcg(objcg);
> > -             if (shrink_memcg(memcg)) {
> > +             if (shrink_memcg(memcg, 1)) {
>
> Why 64 for the global limit but only 1 for the cgroup limit? That
> seems arbitrary in multiple ways.

I suggested that we keep the writeback here without batching and do
that change separately, mainly out of abundance of caution as
writeback is done synchronously here so the extra latency could be
problematic. I think we probably want to measure the performance
impact of that separately.

That being said, this path is potentially too expensive anyway due to
the flush, but I would rather we do some basic measurements before
batching here.

What do you think?

> Direct reclaim, kswapd, proactive reclaim, cgroup limit reclaim use
> SWAP_CLUSTER_MAX for the batch size near-universally. It's magic too
> to be sure, but at least you wouldn't have to make up new magic?
>
> Otherwise, the patch looks good to me.

^ permalink raw reply	[flat|nested] 31+ messages in thread

* Re: [PATCH v2 2/2] mm/zswap: Support batch writeback in shrink_memcg()
  2026-07-23  4:52     ` Yosry Ahmed
@ 2026-07-23 13:55       ` Johannes Weiner
  2026-07-23 16:39         ` Yosry Ahmed
  0 siblings, 1 reply; 31+ messages in thread
From: Johannes Weiner @ 2026-07-23 13:55 UTC (permalink / raw)
  To: Yosry Ahmed
  Cc: Hao Jia, akpm, tj, shakeel.butt, mhocko, mkoutny, nphamcs,
	chengming.zhou, muchun.song, roman.gushchin, linux-mm,
	linux-kernel, linux-doc, Hao Jia

On Wed, Jul 22, 2026 at 09:52:18PM -0700, Yosry Ahmed wrote:
> On Wed, Jul 22, 2026 at 7:27 PM Johannes Weiner <hannes@cmpxchg.org> wrote:
> >
> > On Fri, Jul 17, 2026 at 04:51:51PM +0800, Hao Jia wrote:
> > > @@ -1369,7 +1402,7 @@ static void shrink_worker(struct work_struct *w)
> > >                       goto resched;
> > >               }
> > >
> > > -             ret = shrink_memcg(memcg);
> > > +             ret = shrink_memcg(memcg, NR_ZSWAP_WB_BATCH);
> > >               /* drop the extra reference */
> > >               mem_cgroup_put(memcg);
> > >
> > > @@ -1493,7 +1526,7 @@ bool zswap_store(struct folio *folio)
> > >       objcg = get_obj_cgroup_from_folio(folio);
> > >       if (objcg && !obj_cgroup_may_zswap(objcg)) {
> > >               memcg = get_mem_cgroup_from_objcg(objcg);
> > > -             if (shrink_memcg(memcg)) {
> > > +             if (shrink_memcg(memcg, 1)) {
> >
> > Why 64 for the global limit but only 1 for the cgroup limit? That
> > seems arbitrary in multiple ways.
> 
> I suggested that we keep the writeback here without batching and do
> that change separately, mainly out of abundance of caution as
> writeback is done synchronously here so the extra latency could be
> problematic. I think we probably want to measure the performance
> impact of that separately.
> 
> That being said, this path is potentially too expensive anyway due to
> the flush, but I would rather we do some basic measurements before
> batching here.
> 
> What do you think?

It's not an unknown, right? We know this works for direct reclaimers,
cgroup limit reclaim e.g., and what the latency implications are.

Because of how reclaim works, we also know it'll call zswap_store() in
batches of SWAP_CLUSTER_MAX. If we don't batch here, they're likely to
each call shrink_memcg() once we're at the limit - while still risking
rejections due to compressibility differences.

My worry is that if we start with an inconsistency, we'll be stuck
with it for a long time.

I'd rather start with the clean, consistent version. Dial it back only
if we have data to justfiy the complication that we can put into a
comment and the changelog that outlines why exactly it's different.

^ permalink raw reply	[flat|nested] 31+ messages in thread

* Re: [PATCH v2 2/2] mm/zswap: Support batch writeback in shrink_memcg()
  2026-07-23 13:55       ` Johannes Weiner
@ 2026-07-23 16:39         ` Yosry Ahmed
  2026-07-23 17:11           ` Johannes Weiner
  2026-07-24 10:20           ` Hao Jia
  0 siblings, 2 replies; 31+ messages in thread
From: Yosry Ahmed @ 2026-07-23 16:39 UTC (permalink / raw)
  To: Johannes Weiner
  Cc: Hao Jia, akpm, tj, shakeel.butt, mhocko, mkoutny, nphamcs,
	chengming.zhou, muchun.song, roman.gushchin, linux-mm,
	linux-kernel, linux-doc, Hao Jia

On Thu, Jul 23, 2026 at 6:55 AM Johannes Weiner <hannes@cmpxchg.org> wrote:
>
> On Wed, Jul 22, 2026 at 09:52:18PM -0700, Yosry Ahmed wrote:
> > On Wed, Jul 22, 2026 at 7:27 PM Johannes Weiner <hannes@cmpxchg.org> wrote:
> > >
> > > On Fri, Jul 17, 2026 at 04:51:51PM +0800, Hao Jia wrote:
> > > > @@ -1369,7 +1402,7 @@ static void shrink_worker(struct work_struct *w)
> > > >                       goto resched;
> > > >               }
> > > >
> > > > -             ret = shrink_memcg(memcg);
> > > > +             ret = shrink_memcg(memcg, NR_ZSWAP_WB_BATCH);
> > > >               /* drop the extra reference */
> > > >               mem_cgroup_put(memcg);
> > > >
> > > > @@ -1493,7 +1526,7 @@ bool zswap_store(struct folio *folio)
> > > >       objcg = get_obj_cgroup_from_folio(folio);
> > > >       if (objcg && !obj_cgroup_may_zswap(objcg)) {
> > > >               memcg = get_mem_cgroup_from_objcg(objcg);
> > > > -             if (shrink_memcg(memcg)) {
> > > > +             if (shrink_memcg(memcg, 1)) {
> > >
> > > Why 64 for the global limit but only 1 for the cgroup limit? That
> > > seems arbitrary in multiple ways.
> >
> > I suggested that we keep the writeback here without batching and do
> > that change separately, mainly out of abundance of caution as
> > writeback is done synchronously here so the extra latency could be
> > problematic. I think we probably want to measure the performance
> > impact of that separately.
> >
> > That being said, this path is potentially too expensive anyway due to
> > the flush, but I would rather we do some basic measurements before
> > batching here.
> >
> > What do you think?
>
> It's not an unknown, right? We know this works for direct reclaimers,
> cgroup limit reclaim e.g., and what the latency implications are.
>
> Because of how reclaim works, we also know it'll call zswap_store() in
> batches of SWAP_CLUSTER_MAX. If we don't batch here, they're likely to
> each call shrink_memcg() once we're at the limit - while still risking
> rejections due to compressibility differences.
>
> My worry is that if we start with an inconsistency, we'll be stuck
> with it for a long time.
>
> I'd rather start with the clean, consistent version. Dial it back only
> if we have data to justfiy the complication that we can put into a
> comment and the changelog that outlines why exactly it's different.

I am fine with doing that and basically always using NR_ZSWAP_WB_BATCH
as the batch size in shrink_memcg(), but I would be more comfortable
if we did some sanity testing.

Hao, would you be able to do some smoke testing with NR_ZSWAP_WB_BATCH
used for all paths, and memory.zswap.max set in a way that induces
writeback? You can probably set memory.zswap.max to 1% of total memory
instead of the global pool limit and rerun the same test.

^ permalink raw reply	[flat|nested] 31+ messages in thread

* Re: [PATCH v2 2/2] mm/zswap: Support batch writeback in shrink_memcg()
  2026-07-23 16:39         ` Yosry Ahmed
@ 2026-07-23 17:11           ` Johannes Weiner
  2026-07-24 10:20           ` Hao Jia
  1 sibling, 0 replies; 31+ messages in thread
From: Johannes Weiner @ 2026-07-23 17:11 UTC (permalink / raw)
  To: Yosry Ahmed
  Cc: Hao Jia, akpm, tj, shakeel.butt, mhocko, mkoutny, nphamcs,
	chengming.zhou, muchun.song, roman.gushchin, linux-mm,
	linux-kernel, linux-doc, Hao Jia

On Thu, Jul 23, 2026 at 09:39:35AM -0700, Yosry Ahmed wrote:
> On Thu, Jul 23, 2026 at 6:55 AM Johannes Weiner <hannes@cmpxchg.org> wrote:
> >
> > On Wed, Jul 22, 2026 at 09:52:18PM -0700, Yosry Ahmed wrote:
> > > On Wed, Jul 22, 2026 at 7:27 PM Johannes Weiner <hannes@cmpxchg.org> wrote:
> > > >
> > > > On Fri, Jul 17, 2026 at 04:51:51PM +0800, Hao Jia wrote:
> > > > > @@ -1369,7 +1402,7 @@ static void shrink_worker(struct work_struct *w)
> > > > >                       goto resched;
> > > > >               }
> > > > >
> > > > > -             ret = shrink_memcg(memcg);
> > > > > +             ret = shrink_memcg(memcg, NR_ZSWAP_WB_BATCH);
> > > > >               /* drop the extra reference */
> > > > >               mem_cgroup_put(memcg);
> > > > >
> > > > > @@ -1493,7 +1526,7 @@ bool zswap_store(struct folio *folio)
> > > > >       objcg = get_obj_cgroup_from_folio(folio);
> > > > >       if (objcg && !obj_cgroup_may_zswap(objcg)) {
> > > > >               memcg = get_mem_cgroup_from_objcg(objcg);
> > > > > -             if (shrink_memcg(memcg)) {
> > > > > +             if (shrink_memcg(memcg, 1)) {
> > > >
> > > > Why 64 for the global limit but only 1 for the cgroup limit? That
> > > > seems arbitrary in multiple ways.
> > >
> > > I suggested that we keep the writeback here without batching and do
> > > that change separately, mainly out of abundance of caution as
> > > writeback is done synchronously here so the extra latency could be
> > > problematic. I think we probably want to measure the performance
> > > impact of that separately.
> > >
> > > That being said, this path is potentially too expensive anyway due to
> > > the flush, but I would rather we do some basic measurements before
> > > batching here.
> > >
> > > What do you think?
> >
> > It's not an unknown, right? We know this works for direct reclaimers,
> > cgroup limit reclaim e.g., and what the latency implications are.
> >
> > Because of how reclaim works, we also know it'll call zswap_store() in
> > batches of SWAP_CLUSTER_MAX. If we don't batch here, they're likely to
> > each call shrink_memcg() once we're at the limit - while still risking
> > rejections due to compressibility differences.
> >
> > My worry is that if we start with an inconsistency, we'll be stuck
> > with it for a long time.
> >
> > I'd rather start with the clean, consistent version. Dial it back only
> > if we have data to justfiy the complication that we can put into a
> > comment and the changelog that outlines why exactly it's different.
> 
> I am fine with doing that and basically always using NR_ZSWAP_WB_BATCH
> as the batch size in shrink_memcg(), but I would be more comfortable
> if we did some sanity testing.
> 
> Hao, would you be able to do some smoke testing with NR_ZSWAP_WB_BATCH
> used for all paths, and memory.zswap.max set in a way that induces
> writeback? You can probably set memory.zswap.max to 1% of total memory
> instead of the global pool limit and rerun the same test.

Please just use SWAP_CLUSTER_MAX. That's what the reclaim side which
issues the stores is also using for batching.

^ permalink raw reply	[flat|nested] 31+ messages in thread

* Re: [PATCH v2 2/2] mm/zswap: Support batch writeback in shrink_memcg()
  2026-07-23 16:39         ` Yosry Ahmed
  2026-07-23 17:11           ` Johannes Weiner
@ 2026-07-24 10:20           ` Hao Jia
  2026-07-24 17:56             ` Yosry Ahmed
  2026-07-24 18:40             ` Johannes Weiner
  1 sibling, 2 replies; 31+ messages in thread
From: Hao Jia @ 2026-07-24 10:20 UTC (permalink / raw)
  To: Yosry Ahmed, Johannes Weiner
  Cc: akpm, tj, shakeel.butt, mhocko, mkoutny, nphamcs, chengming.zhou,
	muchun.song, roman.gushchin, linux-mm, linux-kernel, linux-doc,
	Hao Jia



On 2026/7/24 00:39, Yosry Ahmed wrote:
> On Thu, Jul 23, 2026 at 6:55 AM Johannes Weiner <hannes@cmpxchg.org> wrote:
>>
>> On Wed, Jul 22, 2026 at 09:52:18PM -0700, Yosry Ahmed wrote:
>>> On Wed, Jul 22, 2026 at 7:27 PM Johannes Weiner <hannes@cmpxchg.org> wrote:
>>>>
>>>> On Fri, Jul 17, 2026 at 04:51:51PM +0800, Hao Jia wrote:
>>>>> @@ -1369,7 +1402,7 @@ static void shrink_worker(struct work_struct *w)
>>>>>                        goto resched;
>>>>>                }
>>>>>
>>>>> -             ret = shrink_memcg(memcg);
>>>>> +             ret = shrink_memcg(memcg, NR_ZSWAP_WB_BATCH);
>>>>>                /* drop the extra reference */
>>>>>                mem_cgroup_put(memcg);
>>>>>
>>>>> @@ -1493,7 +1526,7 @@ bool zswap_store(struct folio *folio)
>>>>>        objcg = get_obj_cgroup_from_folio(folio);
>>>>>        if (objcg && !obj_cgroup_may_zswap(objcg)) {
>>>>>                memcg = get_mem_cgroup_from_objcg(objcg);
>>>>> -             if (shrink_memcg(memcg)) {
>>>>> +             if (shrink_memcg(memcg, 1)) {
>>>>
>>>> Why 64 for the global limit but only 1 for the cgroup limit? That
>>>> seems arbitrary in multiple ways.
>>>
>>> I suggested that we keep the writeback here without batching and do
>>> that change separately, mainly out of abundance of caution as
>>> writeback is done synchronously here so the extra latency could be
>>> problematic. I think we probably want to measure the performance
>>> impact of that separately.
>>>
>>> That being said, this path is potentially too expensive anyway due to
>>> the flush, but I would rather we do some basic measurements before
>>> batching here.
>>>
>>> What do you think?
>>
>> It's not an unknown, right? We know this works for direct reclaimers,
>> cgroup limit reclaim e.g., and what the latency implications are.
>>
>> Because of how reclaim works, we also know it'll call zswap_store() in
>> batches of SWAP_CLUSTER_MAX. If we don't batch here, they're likely to
>> each call shrink_memcg() once we're at the limit - while still risking
>> rejections due to compressibility differences.
>>
>> My worry is that if we start with an inconsistency, we'll be stuck
>> with it for a long time.
>>
>> I'd rather start with the clean, consistent version. Dial it back only
>> if we have data to justfiy the complication that we can put into a
>> comment and the changelog that outlines why exactly it's different.
> 
> I am fine with doing that and basically always using NR_ZSWAP_WB_BATCH
> as the batch size in shrink_memcg(), but I would be more comfortable
> if we did some sanity testing.
> 
> Hao, would you be able to do some smoke testing with NR_ZSWAP_WB_BATCH
> used for all paths, and memory.zswap.max set in a way that induces
> writeback? You can probably set memory.zswap.max to 1% of total memory
> instead of the global pool limit and rerun the same test.

Building on Test Case 2, I set zswap.max=320M (~1% of total system 
memory) and updated both invocation paths of shrink_memcg() to process 
batches of 32 or 64. The resulting benchmark data is shown below.
(Note: Test Case 2 also sets max_pool_percent=1.)

                       baseline-cgroup   batch-all-32-cgroup 
batch-all-64-cgroup
shrink_worker wakeups        7,238             766               367
shrink_memcg calls      12,059,142       1,961,194           983,878
written_back                28,277         301,157           327,997
zswap_store calls         1,349,572       1,168,190         1,114,549
   store succeeded            492,861         521,315           459,246
   store rejected             856,712         646,875           655,303
   store reject rate            ~63%           ~55%              ~58%
pool_limit_hit              510,130          50,096            57,715
pswpout                     884,989         948,032           983,300
pswpin                    1,251,268       1,638,668         1,878,453

Thanks,
Hao

^ permalink raw reply	[flat|nested] 31+ messages in thread

* Re: [PATCH v2 0/2] mm/zswap: Fixes and improves the zswap global shrinker
  2026-07-23  4:49           ` Yosry Ahmed
@ 2026-07-24 10:22             ` Hao Jia
  0 siblings, 0 replies; 31+ messages in thread
From: Hao Jia @ 2026-07-24 10:22 UTC (permalink / raw)
  To: Yosry Ahmed
  Cc: Andrew Morton, nphamcs, tj, hannes, shakeel.butt, mhocko,
	mkoutny, chengming.zhou, muchun.song, roman.gushchin, linux-mm,
	linux-kernel, linux-doc, Hao Jia



On 2026/7/23 12:49, Yosry Ahmed wrote:
> On Wed, Jul 22, 2026 at 6:21 PM Hao Jia <jiahao.kernel@gmail.com> wrote:
>>
>>
>>
>> On 2026/7/20 09:26, Hao Jia wrote:
>>>
>>>
>>> On 2026/7/18 12:40, Andrew Morton wrote:
>>>> On Fri, 17 Jul 2026 18:28:04 -0700 Yosry Ahmed <yosry@kernel.org> wrote:
>>>>
>>>>>>
>>>>>> The [1/2] changelog lacks a description of how the flaw impacts users.
>>>>>> Please describe this fully and maintain that info within the
>>>>>> changelogging.  This info helps -stable maintainers and others
>>>>>> understand why we're proposing a backport and helps myself and others
>>>>>> with timing decisions.
>>>>>
>>>>> The first line in the changelog should be sufficient imo: "Zswap
>>>>> writeback on hitting the pool limit is broken when memory cgroup is
>>>>> disabled"
>>>>
>>>> "broken"?  Perhaps this means "fails to occur".
>>>>
>>>> But what is the userspace-visible impact?  IOW, why are we proposing a
>>>> backport?
>>>>
>>>
>>> Perhaps the first paragraph of the commit1 message could be modified as
>>> follows? I have added a description of the issues that occur without
>>> this patch.
>>>
>>
>> Hi Andrew, Yosry, and Nhat,
>> Any thoughts on this change?
> 
> I would front load the user impact in the first paragraph:
> 
> Zswap writeback when the global pool limit is hit fails when memory
> cgroups are disabled. The pool remains full until it is organically
> drained by swapins or memory freeing, leading to zswap store failures
> and pages bypassing getting written directly to the backing swap device,
> causing LRU inversion (hotter pages with higher fault latency).
> 
> ...

Will done.

Thanks,
Hao

^ permalink raw reply	[flat|nested] 31+ messages in thread

* Re: [PATCH v2 2/2] mm/zswap: Support batch writeback in shrink_memcg()
  2026-07-24 10:20           ` Hao Jia
@ 2026-07-24 17:56             ` Yosry Ahmed
  2026-07-24 18:37               ` Nhat Pham
  2026-07-24 18:40             ` Johannes Weiner
  1 sibling, 1 reply; 31+ messages in thread
From: Yosry Ahmed @ 2026-07-24 17:56 UTC (permalink / raw)
  To: Hao Jia
  Cc: Johannes Weiner, akpm, tj, shakeel.butt, mhocko, mkoutny,
	nphamcs, chengming.zhou, muchun.song, roman.gushchin, linux-mm,
	linux-kernel, linux-doc, Hao Jia

On Fri, Jul 24, 2026 at 3:21 AM Hao Jia <jiahao.kernel@gmail.com> wrote:
>
>
>
> On 2026/7/24 00:39, Yosry Ahmed wrote:
> > On Thu, Jul 23, 2026 at 6:55 AM Johannes Weiner <hannes@cmpxchg.org> wrote:
> >>
> >> On Wed, Jul 22, 2026 at 09:52:18PM -0700, Yosry Ahmed wrote:
> >>> On Wed, Jul 22, 2026 at 7:27 PM Johannes Weiner <hannes@cmpxchg.org> wrote:
> >>>>
> >>>> On Fri, Jul 17, 2026 at 04:51:51PM +0800, Hao Jia wrote:
> >>>>> @@ -1369,7 +1402,7 @@ static void shrink_worker(struct work_struct *w)
> >>>>>                        goto resched;
> >>>>>                }
> >>>>>
> >>>>> -             ret = shrink_memcg(memcg);
> >>>>> +             ret = shrink_memcg(memcg, NR_ZSWAP_WB_BATCH);
> >>>>>                /* drop the extra reference */
> >>>>>                mem_cgroup_put(memcg);
> >>>>>
> >>>>> @@ -1493,7 +1526,7 @@ bool zswap_store(struct folio *folio)
> >>>>>        objcg = get_obj_cgroup_from_folio(folio);
> >>>>>        if (objcg && !obj_cgroup_may_zswap(objcg)) {
> >>>>>                memcg = get_mem_cgroup_from_objcg(objcg);
> >>>>> -             if (shrink_memcg(memcg)) {
> >>>>> +             if (shrink_memcg(memcg, 1)) {
> >>>>
> >>>> Why 64 for the global limit but only 1 for the cgroup limit? That
> >>>> seems arbitrary in multiple ways.
> >>>
> >>> I suggested that we keep the writeback here without batching and do
> >>> that change separately, mainly out of abundance of caution as
> >>> writeback is done synchronously here so the extra latency could be
> >>> problematic. I think we probably want to measure the performance
> >>> impact of that separately.
> >>>
> >>> That being said, this path is potentially too expensive anyway due to
> >>> the flush, but I would rather we do some basic measurements before
> >>> batching here.
> >>>
> >>> What do you think?
> >>
> >> It's not an unknown, right? We know this works for direct reclaimers,
> >> cgroup limit reclaim e.g., and what the latency implications are.
> >>
> >> Because of how reclaim works, we also know it'll call zswap_store() in
> >> batches of SWAP_CLUSTER_MAX. If we don't batch here, they're likely to
> >> each call shrink_memcg() once we're at the limit - while still risking
> >> rejections due to compressibility differences.
> >>
> >> My worry is that if we start with an inconsistency, we'll be stuck
> >> with it for a long time.
> >>
> >> I'd rather start with the clean, consistent version. Dial it back only
> >> if we have data to justfiy the complication that we can put into a
> >> comment and the changelog that outlines why exactly it's different.
> >
> > I am fine with doing that and basically always using NR_ZSWAP_WB_BATCH
> > as the batch size in shrink_memcg(), but I would be more comfortable
> > if we did some sanity testing.
> >
> > Hao, would you be able to do some smoke testing with NR_ZSWAP_WB_BATCH
> > used for all paths, and memory.zswap.max set in a way that induces
> > writeback? You can probably set memory.zswap.max to 1% of total memory
> > instead of the global pool limit and rerun the same test.
>
> Building on Test Case 2, I set zswap.max=320M (~1% of total system
> memory) and updated both invocation paths of shrink_memcg() to process
> batches of 32 or 64. The resulting benchmark data is shown below.
> (Note: Test Case 2 also sets max_pool_percent=1.)
>
>                        baseline-cgroup   batch-all-32-cgroup
> batch-all-64-cgroup
> shrink_worker wakeups        7,238             766               367
> shrink_memcg calls      12,059,142       1,961,194           983,878
> written_back                28,277         301,157           327,997
> zswap_store calls         1,349,572       1,168,190         1,114,549
>    store succeeded            492,861         521,315           459,246
>    store rejected             856,712         646,875           655,303
>    store reject rate            ~63%           ~55%              ~58%
> pool_limit_hit              510,130          50,096            57,715
> pswpout                     884,989         948,032           983,300
> pswpin                    1,251,268       1,638,668         1,878,453

The numbers generally look good, the store rejection rate is
decreasing and we are naturally doing more writeback which makes
sense. The store rejection rate actually increases when the batch size
increases to 64, maybe aggressive writeback is more likely to fail and
cause a store rejection, so that also makes sense.

The only part I can't immediately reason about is pswpin. How are we
reading from physical swap more than we ever wrote to it?

Also why does it significantly increase with batching? I would hope
that it actually decreases (like the commit message), because colder
pages end up on disk. But even if we eventually swapin everything on
disk, I would only expect a slight increase to match the slight
increase in pswpout.

^ permalink raw reply	[flat|nested] 31+ messages in thread

* Re: [PATCH v2 2/2] mm/zswap: Support batch writeback in shrink_memcg()
  2026-07-24 17:56             ` Yosry Ahmed
@ 2026-07-24 18:37               ` Nhat Pham
  2026-07-24 18:39                 ` Yosry Ahmed
  0 siblings, 1 reply; 31+ messages in thread
From: Nhat Pham @ 2026-07-24 18:37 UTC (permalink / raw)
  To: Yosry Ahmed
  Cc: Hao Jia, Johannes Weiner, akpm, tj, shakeel.butt, mhocko,
	mkoutny, chengming.zhou, muchun.song, roman.gushchin, linux-mm,
	linux-kernel, linux-doc, Hao Jia

On Fri, Jul 24, 2026 at 10:57 AM Yosry Ahmed <yosry@kernel.org> wrote:
>
> The numbers generally look good, the store rejection rate is
> decreasing and we are naturally doing more writeback which makes
> sense. The store rejection rate actually increases when the batch size
> increases to 64, maybe aggressive writeback is more likely to fail and
> cause a store rejection, so that also makes sense.
>
> The only part I can't immediately reason about is pswpin. How are we
> reading from physical swap more than we ever wrote to it?

The same page can be swapped out once, and swap in multiple times :)

Unlike zswap, physical swap in is not necessarily exclusive. You load
the page in memory, but if the swapfile is not full, etc. etc. you
don't invalidate the copy of the data on the swapfile. At reclaim
time, we notice the page is not dirty, so we just skip (z)swapout.

^ permalink raw reply	[flat|nested] 31+ messages in thread

* Re: [PATCH v2 2/2] mm/zswap: Support batch writeback in shrink_memcg()
  2026-07-24 18:37               ` Nhat Pham
@ 2026-07-24 18:39                 ` Yosry Ahmed
  2026-07-24 19:35                   ` Johannes Weiner
  0 siblings, 1 reply; 31+ messages in thread
From: Yosry Ahmed @ 2026-07-24 18:39 UTC (permalink / raw)
  To: Nhat Pham
  Cc: Hao Jia, Johannes Weiner, akpm, tj, shakeel.butt, mhocko,
	mkoutny, chengming.zhou, muchun.song, roman.gushchin, linux-mm,
	linux-kernel, linux-doc, Hao Jia

On Fri, Jul 24, 2026 at 11:37 AM Nhat Pham <nphamcs@gmail.com> wrote:
>
> On Fri, Jul 24, 2026 at 10:57 AM Yosry Ahmed <yosry@kernel.org> wrote:
> >
> > The numbers generally look good, the store rejection rate is
> > decreasing and we are naturally doing more writeback which makes
> > sense. The store rejection rate actually increases when the batch size
> > increases to 64, maybe aggressive writeback is more likely to fail and
> > cause a store rejection, so that also makes sense.
> >
> > The only part I can't immediately reason about is pswpin. How are we
> > reading from physical swap more than we ever wrote to it?
>
> The same page can be swapped out once, and swap in multiple times :)
>
> Unlike zswap, physical swap in is not necessarily exclusive. You load
> the page in memory, but if the swapfile is not full, etc. etc. you
> don't invalidate the copy of the data on the swapfile. At reclaim
> time, we notice the page is not dirty, so we just skip (z)swapout.

Oh right, I forgot about that, thanks.

I still don't fully understand why pswpin is significantly increased
with batching here. I would hope that with less LRU inversion we end
up with less disk swapin. Probably the test access patterns are just
too random compared to real workloads?

^ permalink raw reply	[flat|nested] 31+ messages in thread

* Re: [PATCH v2 2/2] mm/zswap: Support batch writeback in shrink_memcg()
  2026-07-24 10:20           ` Hao Jia
  2026-07-24 17:56             ` Yosry Ahmed
@ 2026-07-24 18:40             ` Johannes Weiner
  1 sibling, 0 replies; 31+ messages in thread
From: Johannes Weiner @ 2026-07-24 18:40 UTC (permalink / raw)
  To: Hao Jia
  Cc: Yosry Ahmed, akpm, tj, shakeel.butt, mhocko, mkoutny, nphamcs,
	chengming.zhou, muchun.song, roman.gushchin, linux-mm,
	linux-kernel, linux-doc, Hao Jia

On Fri, Jul 24, 2026 at 06:20:50PM +0800, Hao Jia wrote:
> 
> 
> On 2026/7/24 00:39, Yosry Ahmed wrote:
> > On Thu, Jul 23, 2026 at 6:55 AM Johannes Weiner <hannes@cmpxchg.org> wrote:
> >>
> >> On Wed, Jul 22, 2026 at 09:52:18PM -0700, Yosry Ahmed wrote:
> >>> On Wed, Jul 22, 2026 at 7:27 PM Johannes Weiner <hannes@cmpxchg.org> wrote:
> >>>>
> >>>> On Fri, Jul 17, 2026 at 04:51:51PM +0800, Hao Jia wrote:
> >>>>> @@ -1369,7 +1402,7 @@ static void shrink_worker(struct work_struct *w)
> >>>>>                        goto resched;
> >>>>>                }
> >>>>>
> >>>>> -             ret = shrink_memcg(memcg);
> >>>>> +             ret = shrink_memcg(memcg, NR_ZSWAP_WB_BATCH);
> >>>>>                /* drop the extra reference */
> >>>>>                mem_cgroup_put(memcg);
> >>>>>
> >>>>> @@ -1493,7 +1526,7 @@ bool zswap_store(struct folio *folio)
> >>>>>        objcg = get_obj_cgroup_from_folio(folio);
> >>>>>        if (objcg && !obj_cgroup_may_zswap(objcg)) {
> >>>>>                memcg = get_mem_cgroup_from_objcg(objcg);
> >>>>> -             if (shrink_memcg(memcg)) {
> >>>>> +             if (shrink_memcg(memcg, 1)) {
> >>>>
> >>>> Why 64 for the global limit but only 1 for the cgroup limit? That
> >>>> seems arbitrary in multiple ways.
> >>>
> >>> I suggested that we keep the writeback here without batching and do
> >>> that change separately, mainly out of abundance of caution as
> >>> writeback is done synchronously here so the extra latency could be
> >>> problematic. I think we probably want to measure the performance
> >>> impact of that separately.
> >>>
> >>> That being said, this path is potentially too expensive anyway due to
> >>> the flush, but I would rather we do some basic measurements before
> >>> batching here.
> >>>
> >>> What do you think?
> >>
> >> It's not an unknown, right? We know this works for direct reclaimers,
> >> cgroup limit reclaim e.g., and what the latency implications are.
> >>
> >> Because of how reclaim works, we also know it'll call zswap_store() in
> >> batches of SWAP_CLUSTER_MAX. If we don't batch here, they're likely to
> >> each call shrink_memcg() once we're at the limit - while still risking
> >> rejections due to compressibility differences.
> >>
> >> My worry is that if we start with an inconsistency, we'll be stuck
> >> with it for a long time.
> >>
> >> I'd rather start with the clean, consistent version. Dial it back only
> >> if we have data to justfiy the complication that we can put into a
> >> comment and the changelog that outlines why exactly it's different.
> > 
> > I am fine with doing that and basically always using NR_ZSWAP_WB_BATCH
> > as the batch size in shrink_memcg(), but I would be more comfortable
> > if we did some sanity testing.
> > 
> > Hao, would you be able to do some smoke testing with NR_ZSWAP_WB_BATCH
> > used for all paths, and memory.zswap.max set in a way that induces
> > writeback? You can probably set memory.zswap.max to 1% of total memory
> > instead of the global pool limit and rerun the same test.
> 
> Building on Test Case 2, I set zswap.max=320M (~1% of total system 
> memory) and updated both invocation paths of shrink_memcg() to process 
> batches of 32 or 64. The resulting benchmark data is shown below.
> (Note: Test Case 2 also sets max_pool_percent=1.)
> 
>                        baseline-cgroup   batch-all-32-cgroup 
> batch-all-64-cgroup
> shrink_worker wakeups        7,238             766               367
> shrink_memcg calls      12,059,142       1,961,194           983,878
> written_back                28,277         301,157           327,997
> zswap_store calls         1,349,572       1,168,190         1,114,549
>    store succeeded            492,861         521,315           459,246
>    store rejected             856,712         646,875           655,303
>    store reject rate            ~63%           ~55%              ~58%
> pool_limit_hit              510,130          50,096            57,715
> pswpout                     884,989         948,032           983,300
> pswpin                    1,251,268       1,638,668         1,878,453

Thanks for testing both!

Looks like 32 shows the better matching with the reclaim batches than
64: it writes back less and swaps in less, while still having the
improved rejection rate. It even rejects slightly less than 64, but
that might be noise? Absolute stores win handily in any case - not
sure if that's meaningful in your test design.

^ permalink raw reply	[flat|nested] 31+ messages in thread

* Re: [PATCH v2 2/2] mm/zswap: Support batch writeback in shrink_memcg()
  2026-07-24 18:39                 ` Yosry Ahmed
@ 2026-07-24 19:35                   ` Johannes Weiner
  2026-07-24 22:22                     ` Yosry Ahmed
  0 siblings, 1 reply; 31+ messages in thread
From: Johannes Weiner @ 2026-07-24 19:35 UTC (permalink / raw)
  To: Yosry Ahmed
  Cc: Nhat Pham, Hao Jia, akpm, tj, shakeel.butt, mhocko, mkoutny,
	chengming.zhou, muchun.song, roman.gushchin, linux-mm,
	linux-kernel, linux-doc, Hao Jia

On Fri, Jul 24, 2026 at 11:39:41AM -0700, Yosry Ahmed wrote:
> On Fri, Jul 24, 2026 at 11:37 AM Nhat Pham <nphamcs@gmail.com> wrote:
> >
> > On Fri, Jul 24, 2026 at 10:57 AM Yosry Ahmed <yosry@kernel.org> wrote:
> > >
> > > The numbers generally look good, the store rejection rate is
> > > decreasing and we are naturally doing more writeback which makes
> > > sense. The store rejection rate actually increases when the batch size
> > > increases to 64, maybe aggressive writeback is more likely to fail and
> > > cause a store rejection, so that also makes sense.
> > >
> > > The only part I can't immediately reason about is pswpin. How are we
> > > reading from physical swap more than we ever wrote to it?
> >
> > The same page can be swapped out once, and swap in multiple times :)
> >
> > Unlike zswap, physical swap in is not necessarily exclusive. You load
> > the page in memory, but if the swapfile is not full, etc. etc. you
> > don't invalidate the copy of the data on the swapfile. At reclaim
> > time, we notice the page is not dirty, so we just skip (z)swapout.
> 
> Oh right, I forgot about that, thanks.
> 
> I still don't fully understand why pswpin is significantly increased
> with batching here. I would hope that with less LRU inversion we end
> up with less disk swapin. Probably the test access patterns are just
> too random compared to real workloads?

Maybe more time in zswap for those tail entries left behind with the
smaller batch?

That's a 32 entry window that could zswpin with the smaller batch but
pswpin with the larger one.

^ permalink raw reply	[flat|nested] 31+ messages in thread

* Re: [PATCH v2 2/2] mm/zswap: Support batch writeback in shrink_memcg()
  2026-07-24 19:35                   ` Johannes Weiner
@ 2026-07-24 22:22                     ` Yosry Ahmed
  2026-07-26  2:37                       ` Hao Jia
  0 siblings, 1 reply; 31+ messages in thread
From: Yosry Ahmed @ 2026-07-24 22:22 UTC (permalink / raw)
  To: Johannes Weiner
  Cc: Nhat Pham, Hao Jia, akpm, tj, shakeel.butt, mhocko, mkoutny,
	chengming.zhou, muchun.song, roman.gushchin, linux-mm,
	linux-kernel, linux-doc, Hao Jia

On Fri, Jul 24, 2026 at 12:35 PM Johannes Weiner <hannes@cmpxchg.org> wrote:
>
> On Fri, Jul 24, 2026 at 11:39:41AM -0700, Yosry Ahmed wrote:
> > On Fri, Jul 24, 2026 at 11:37 AM Nhat Pham <nphamcs@gmail.com> wrote:
> > >
> > > On Fri, Jul 24, 2026 at 10:57 AM Yosry Ahmed <yosry@kernel.org> wrote:
> > > >
> > > > The numbers generally look good, the store rejection rate is
> > > > decreasing and we are naturally doing more writeback which makes
> > > > sense. The store rejection rate actually increases when the batch size
> > > > increases to 64, maybe aggressive writeback is more likely to fail and
> > > > cause a store rejection, so that also makes sense.
> > > >
> > > > The only part I can't immediately reason about is pswpin. How are we
> > > > reading from physical swap more than we ever wrote to it?
> > >
> > > The same page can be swapped out once, and swap in multiple times :)
> > >
> > > Unlike zswap, physical swap in is not necessarily exclusive. You load
> > > the page in memory, but if the swapfile is not full, etc. etc. you
> > > don't invalidate the copy of the data on the swapfile. At reclaim
> > > time, we notice the page is not dirty, so we just skip (z)swapout.
> >
> > Oh right, I forgot about that, thanks.
> >
> > I still don't fully understand why pswpin is significantly increased
> > with batching here. I would hope that with less LRU inversion we end
> > up with less disk swapin. Probably the test access patterns are just
> > too random compared to real workloads?
>
> Maybe more time in zswap for those tail entries left behind with the
> smaller batch?
>
> That's a 32 entry window that could zswpin with the smaller batch but
> pswpin with the larger one.

Oh I was talking about going from no batching to batch=32. pspwin
increased by 387,400 (~30%). We are writing back 272,880 more pages
but we also have 209,837 less store rejections. So overall ~ 63,043
more pages should end up on disk, which doesn't explain the increase
in pspwin.

The only explanation I can think of is that the test is just randomly
accessing memory, so it ends up accessing colder memory that we moved
to disk more than hotter memory that we kept in zswap. IOW, the access
patterns do not conform to a "normal" workload that benefits from the
LRU.

^ permalink raw reply	[flat|nested] 31+ messages in thread

* Re: [PATCH v2 2/2] mm/zswap: Support batch writeback in shrink_memcg()
  2026-07-24 22:22                     ` Yosry Ahmed
@ 2026-07-26  2:37                       ` Hao Jia
  2026-07-27 16:18                         ` Yosry Ahmed
  0 siblings, 1 reply; 31+ messages in thread
From: Hao Jia @ 2026-07-26  2:37 UTC (permalink / raw)
  To: Yosry Ahmed, Johannes Weiner
  Cc: Nhat Pham, akpm, tj, shakeel.butt, mhocko, mkoutny,
	chengming.zhou, muchun.song, roman.gushchin, linux-mm,
	linux-kernel, linux-doc, Hao Jia



On 2026/7/25 06:22, Yosry Ahmed wrote:
> On Fri, Jul 24, 2026 at 12:35 PM Johannes Weiner <hannes@cmpxchg.org> wrote:
>>
>> On Fri, Jul 24, 2026 at 11:39:41AM -0700, Yosry Ahmed wrote:
>>> On Fri, Jul 24, 2026 at 11:37 AM Nhat Pham <nphamcs@gmail.com> wrote:
>>>>
>>>> On Fri, Jul 24, 2026 at 10:57 AM Yosry Ahmed <yosry@kernel.org> wrote:
>>>>>
>>>>> The numbers generally look good, the store rejection rate is
>>>>> decreasing and we are naturally doing more writeback which makes
>>>>> sense. The store rejection rate actually increases when the batch size
>>>>> increases to 64, maybe aggressive writeback is more likely to fail and
>>>>> cause a store rejection, so that also makes sense.
>>>>>
>>>>> The only part I can't immediately reason about is pswpin. How are we
>>>>> reading from physical swap more than we ever wrote to it?
>>>>
>>>> The same page can be swapped out once, and swap in multiple times :)
>>>>
>>>> Unlike zswap, physical swap in is not necessarily exclusive. You load
>>>> the page in memory, but if the swapfile is not full, etc. etc. you
>>>> don't invalidate the copy of the data on the swapfile. At reclaim
>>>> time, we notice the page is not dirty, so we just skip (z)swapout.
>>>
>>> Oh right, I forgot about that, thanks.
>>>
>>> I still don't fully understand why pswpin is significantly increased
>>> with batching here. I would hope that with less LRU inversion we end
>>> up with less disk swapin. Probably the test access patterns are just
>>> too random compared to real workloads?
>>
>> Maybe more time in zswap for those tail entries left behind with the
>> smaller batch?
>>
>> That's a 32 entry window that could zswpin with the smaller batch but
>> pswpin with the larger one.
> 
> Oh I was talking about going from no batching to batch=32. pspwin
> increased by 387,400 (~30%). We are writing back 272,880 more pages
> but we also have 209,837 less store rejections. So overall ~ 63,043
> more pages should end up on disk, which doesn't explain the increase
> in pspwin.
> 

In my previous runs, I also collected zswpin. could the following 
explanation help account for what we are seeing?

Assuming stress-ng maintains a roughly similar memory access rate, 
batching flushes significantly more pages back to disk (meaning pages in 
the zswap pool are evicted faster, keeping zswap pool residency lower). 
As a result, when stress-ng accesses swapped memory, it is more likely 
to fault pages in from disk (pswpin) rather than hit zswap (zswpin).

In fact, zswpin in batch-32 dropped by 143,425 compared to 
baseline—which aligns with batching having a lower zswpin and a higher 
pswpin.

To be frank, under the Test Case 2 workload, we cannot strictly 
guarantee that the total volume of page reads and writes remains 
consistent across all three runs within the same time window.

                      baseline-cgroup   batch-all-32-cgroup 
batch-all-64-cgroup
shrink_worker wakeups       7,238             766                 367
shrink_memcg calls     12,059,142       1,961,194             983,878
written_back               28,277         301,157             327,997
zswap_store calls        1,349,572       1,168,190          1,114,549
    store succeeded           492,861       521,315           459,246
    store rejected            856,712       646,875           655,303
    store reject rate           ~63%           ~55%               ~58%
pool_limit_hit             510,130          50,096             57,715
pswpout                    884,989         948,032            983,300
pswpin                   1,251,268       1,638,668          1,878,453
zswpout                    492,861         521,314            459,245
zswpin                     309,631         166,206            84,977 <-

Thanks,
Hao
> The only explanation I can think of is that the test is just randomly
> accessing memory, so it ends up accessing colder memory that we moved
> to disk more than hotter memory that we kept in zswap. IOW, the access
> patterns do not conform to a "normal" workload that benefits from the
> LRU.

^ permalink raw reply	[flat|nested] 31+ messages in thread

* Re: [PATCH v2 2/2] mm/zswap: Support batch writeback in shrink_memcg()
  2026-07-26  2:37                       ` Hao Jia
@ 2026-07-27 16:18                         ` Yosry Ahmed
  2026-07-28 11:22                           ` Hao Jia
  0 siblings, 1 reply; 31+ messages in thread
From: Yosry Ahmed @ 2026-07-27 16:18 UTC (permalink / raw)
  To: Hao Jia
  Cc: Johannes Weiner, Nhat Pham, akpm, tj, shakeel.butt, mhocko,
	mkoutny, chengming.zhou, muchun.song, roman.gushchin, linux-mm,
	linux-kernel, linux-doc, Hao Jia

On Sat, Jul 25, 2026 at 7:39 PM Hao Jia <jiahao.kernel@gmail.com> wrote:
>
>
>
> On 2026/7/25 06:22, Yosry Ahmed wrote:
> > On Fri, Jul 24, 2026 at 12:35 PM Johannes Weiner <hannes@cmpxchg.org> wrote:
> >>
> >> On Fri, Jul 24, 2026 at 11:39:41AM -0700, Yosry Ahmed wrote:
> >>> On Fri, Jul 24, 2026 at 11:37 AM Nhat Pham <nphamcs@gmail.com> wrote:
> >>>>
> >>>> On Fri, Jul 24, 2026 at 10:57 AM Yosry Ahmed <yosry@kernel.org> wrote:
> >>>>>
> >>>>> The numbers generally look good, the store rejection rate is
> >>>>> decreasing and we are naturally doing more writeback which makes
> >>>>> sense. The store rejection rate actually increases when the batch size
> >>>>> increases to 64, maybe aggressive writeback is more likely to fail and
> >>>>> cause a store rejection, so that also makes sense.
> >>>>>
> >>>>> The only part I can't immediately reason about is pswpin. How are we
> >>>>> reading from physical swap more than we ever wrote to it?
> >>>>
> >>>> The same page can be swapped out once, and swap in multiple times :)
> >>>>
> >>>> Unlike zswap, physical swap in is not necessarily exclusive. You load
> >>>> the page in memory, but if the swapfile is not full, etc. etc. you
> >>>> don't invalidate the copy of the data on the swapfile. At reclaim
> >>>> time, we notice the page is not dirty, so we just skip (z)swapout.
> >>>
> >>> Oh right, I forgot about that, thanks.
> >>>
> >>> I still don't fully understand why pswpin is significantly increased
> >>> with batching here. I would hope that with less LRU inversion we end
> >>> up with less disk swapin. Probably the test access patterns are just
> >>> too random compared to real workloads?
> >>
> >> Maybe more time in zswap for those tail entries left behind with the
> >> smaller batch?
> >>
> >> That's a 32 entry window that could zswpin with the smaller batch but
> >> pswpin with the larger one.
> >
> > Oh I was talking about going from no batching to batch=32. pspwin
> > increased by 387,400 (~30%). We are writing back 272,880 more pages
> > but we also have 209,837 less store rejections. So overall ~ 63,043
> > more pages should end up on disk, which doesn't explain the increase
> > in pspwin.
> >
>
> In my previous runs, I also collected zswpin. could the following
> explanation help account for what we are seeing?
>
> Assuming stress-ng maintains a roughly similar memory access rate,
> batching flushes significantly more pages back to disk (meaning pages in
> the zswap pool are evicted faster, keeping zswap pool residency lower).
> As a result, when stress-ng accesses swapped memory, it is more likely
> to fault pages in from disk (pswpin) rather than hit zswap (zswpin).

We are not just writing back more pages to disk, we are also accepting
more pages into zswap. The whole point of writeback here is to evict
colder pages to disk and accept (relatively) hotter pages into zswap,
as they are more likely to be swapped in again.

>
> In fact, zswpin in batch-32 dropped by 143,425 compared to
> baseline—which aligns with batching having a lower zswpin and a higher
> pswpin.
>
> To be frank, under the Test Case 2 workload, we cannot strictly
> guarantee that the total volume of page reads and writes remains
> consistent across all three runs within the same time window.
>
>                       baseline-cgroup   batch-all-32-cgroup
> batch-all-64-cgroup
> shrink_worker wakeups       7,238             766                 367
> shrink_memcg calls     12,059,142       1,961,194             983,878
> written_back               28,277         301,157             327,997
> zswap_store calls        1,349,572       1,168,190          1,114,549
>     store succeeded           492,861       521,315           459,246
>     store rejected            856,712       646,875           655,303
>     store reject rate           ~63%           ~55%               ~58%
> pool_limit_hit             510,130          50,096             57,715
> pswpout                    884,989         948,032            983,300
> pswpin                   1,251,268       1,638,668          1,878,453
> zswpout                    492,861         521,314            459,245
> zswpin                     309,631         166,206            84,977 <-

One possible explanation is that we are shrinking too aggressively due
to concurrent shrinkers. That's one difference between hitting the
global limit and the memcg limit. There's a single global shrinker
that all store failures will just wake up, while all concurrent stores
hitting the memcg limit will start shrinking.

The other explanation is that, as I mentioned earlier, the access
pattern is just too random and we are consistently refaulting cold
memory. I am honestly a bit concerned about these results.

Hao, would you be able to try re-running the test with a single
stressor? It would help us figure out if concurrent shrinking is a
problem. It would also be useful to increase the global limit, to
isolate writeback from hitting the memcg limit. I don't think both
limits will often be used in conjunction, at least not in a way where
both limits are hit simultaneously.

It may also be useful to run an actual workload if you are able to,
like a kernel build or something, to see if the access patterns change
how much we end up swapping in from zswap vs. disk.

^ permalink raw reply	[flat|nested] 31+ messages in thread

* Re: [PATCH v2 2/2] mm/zswap: Support batch writeback in shrink_memcg()
  2026-07-27 16:18                         ` Yosry Ahmed
@ 2026-07-28 11:22                           ` Hao Jia
  2026-07-28 15:52                             ` Nhat Pham
  0 siblings, 1 reply; 31+ messages in thread
From: Hao Jia @ 2026-07-28 11:22 UTC (permalink / raw)
  To: Yosry Ahmed
  Cc: Johannes Weiner, Nhat Pham, akpm, tj, shakeel.butt, mhocko,
	mkoutny, chengming.zhou, muchun.song, roman.gushchin, linux-mm,
	linux-kernel, linux-doc, Hao Jia



On 2026/7/28 00:18, Yosry Ahmed wrote:
> On Sat, Jul 25, 2026 at 7:39 PM Hao Jia <jiahao.kernel@gmail.com> wrote:
>>
>>
>>
>> On 2026/7/25 06:22, Yosry Ahmed wrote:
>>> On Fri, Jul 24, 2026 at 12:35 PM Johannes Weiner <hannes@cmpxchg.org> wrote:
>>>>
>>>> On Fri, Jul 24, 2026 at 11:39:41AM -0700, Yosry Ahmed wrote:
>>>>> On Fri, Jul 24, 2026 at 11:37 AM Nhat Pham <nphamcs@gmail.com> wrote:
>>>>>>
>>>>>> On Fri, Jul 24, 2026 at 10:57 AM Yosry Ahmed <yosry@kernel.org> wrote:
>>>>>>>
>>>>>>> The numbers generally look good, the store rejection rate is
>>>>>>> decreasing and we are naturally doing more writeback which makes
>>>>>>> sense. The store rejection rate actually increases when the batch size
>>>>>>> increases to 64, maybe aggressive writeback is more likely to fail and
>>>>>>> cause a store rejection, so that also makes sense.
>>>>>>>
>>>>>>> The only part I can't immediately reason about is pswpin. How are we
>>>>>>> reading from physical swap more than we ever wrote to it?
>>>>>>
>>>>>> The same page can be swapped out once, and swap in multiple times :)
>>>>>>
>>>>>> Unlike zswap, physical swap in is not necessarily exclusive. You load
>>>>>> the page in memory, but if the swapfile is not full, etc. etc. you
>>>>>> don't invalidate the copy of the data on the swapfile. At reclaim
>>>>>> time, we notice the page is not dirty, so we just skip (z)swapout.
>>>>>
>>>>> Oh right, I forgot about that, thanks.
>>>>>
>>>>> I still don't fully understand why pswpin is significantly increased
>>>>> with batching here. I would hope that with less LRU inversion we end
>>>>> up with less disk swapin. Probably the test access patterns are just
>>>>> too random compared to real workloads?
>>>>
>>>> Maybe more time in zswap for those tail entries left behind with the
>>>> smaller batch?
>>>>
>>>> That's a 32 entry window that could zswpin with the smaller batch but
>>>> pswpin with the larger one.
>>>
>>> Oh I was talking about going from no batching to batch=32. pspwin
>>> increased by 387,400 (~30%). We are writing back 272,880 more pages
>>> but we also have 209,837 less store rejections. So overall ~ 63,043
>>> more pages should end up on disk, which doesn't explain the increase
>>> in pspwin.
>>>
>>
>> In my previous runs, I also collected zswpin. could the following
>> explanation help account for what we are seeing?
>>
>> Assuming stress-ng maintains a roughly similar memory access rate,
>> batching flushes significantly more pages back to disk (meaning pages in
>> the zswap pool are evicted faster, keeping zswap pool residency lower).
>> As a result, when stress-ng accesses swapped memory, it is more likely
>> to fault pages in from disk (pswpin) rather than hit zswap (zswpin).
> 
> We are not just writing back more pages to disk, we are also accepting
> more pages into zswap. The whole point of writeback here is to evict
> colder pages to disk and accept (relatively) hotter pages into zswap,
> as they are more likely to be swapped in again.
> 
>>
>> In fact, zswpin in batch-32 dropped by 143,425 compared to
>> baseline—which aligns with batching having a lower zswpin and a higher
>> pswpin.
>>
>> To be frank, under the Test Case 2 workload, we cannot strictly
>> guarantee that the total volume of page reads and writes remains
>> consistent across all three runs within the same time window.
>>
>>                        baseline-cgroup   batch-all-32-cgroup
>> batch-all-64-cgroup
>> shrink_worker wakeups       7,238             766                 367
>> shrink_memcg calls     12,059,142       1,961,194             983,878
>> written_back               28,277         301,157             327,997
>> zswap_store calls        1,349,572       1,168,190          1,114,549
>>      store succeeded           492,861       521,315           459,246
>>      store rejected            856,712       646,875           655,303
>>      store reject rate           ~63%           ~55%               ~58%
>> pool_limit_hit             510,130          50,096             57,715
>> pswpout                    884,989         948,032            983,300
>> pswpin                   1,251,268       1,638,668          1,878,453
>> zswpout                    492,861         521,314            459,245
>> zswpin                     309,631         166,206            84,977 <-
> 
> One possible explanation is that we are shrinking too aggressively due
> to concurrent shrinkers. That's one difference between hitting the
> global limit and the memcg limit. There's a single global shrinker
> that all store failures will just wake up, while all concurrent stores
> hitting the memcg limit will start shrinking.
> 
> The other explanation is that, as I mentioned earlier, the access
> pattern is just too random and we are consistently refaulting cold
> memory. I am honestly a bit concerned about these results.
> 
> Hao, would you be able to try re-running the test with a single
> stressor? It would help us figure out if concurrent shrinking is a
> problem. It would also be useful to increase the global limit, to
> isolate writeback from hitting the memcg limit. I don't think both
> limits will often be used in conjunction, at least not in a way where
> both limits are hit simultaneously.
> 
> It may also be useful to run an actual workload if you are able to,
> like a kernel build or something, to see if the access patterns change
> how much we end up swapping in from zswap vs. disk.

Apologies—there was a flaw in my previous test setup. I had configured 
zswap.max=320M alongside max_pool_percent=1. Since the global zswap pool 
capacity was roughly 314M (< 320M), the global pool limit was actually 
being hit before the cgroup zswap.max limit. Both limits were likely 
interleaving throughout the run, making the results difficult to analyze.

Re-running with **zswap.max=280M** yields results that align with 
expectations:

Test Case: zmax280 stress-4 pool1%
                             baseline        batch-all32      batch-all64
   shrink_worker wakeups            0                0                0
   shrink_memcg calls         698,990           57,187           29,649
   written_back               655,570          865,131          883,508
   store succeeded          1,054,928        1,124,990        1,190,608
   store rejected             229,722           22,642           11,659
   store reject rate             ~17%              ~1%              ~0%
   reject_reclaim_fail            110           15,454           12,985
   pswpout                    885,437          888,027          895,478
   pswpin                   2,035,415        1,305,341        1,384,751
   zswpout                  1,054,928        1,124,990        1,190,608
   zswpin                     280,675          162,684          158,556

To isolate the cgroup zswap limit, I also tested with 
max_pool_percent=50 (which should behave similarly to the zswap.max=280M 
case above, bounded strictly by cgroup zswap.max):

Test Case: zmax320 stress-4 pool50%
                            baseline        batch-all32      batch-all64
   shrink_worker wakeups            0                0                0
   shrink_memcg calls         687,608           54,002           28,116
   written_back               639,176          846,663          862,347
   store succeeded            992,816        1,208,123        1,213,286
   store rejected             231,431           20,425           10,915
   store reject rate             ~18%              ~1%              ~0%
   pswpout                    870,745          867,360          873,553
   pswpin                   1,707,823        1,216,814        1,378,381
   zswpout                    992,816        1,208,123        1,213,286
   zswpin                     250,220          210,791          185,716


Finally, here is the dataset for zswap.max=320M and max_pool_percent=1 
with a single stress-ng worker (stress-1), which also looks consistent 
with your expectations:

Test Case: zmax320 stress-1 pool1%
                           baseline        batch-all32      batch-all64
   shrink_worker wakeups        9,279              980              557
   shrink_memcg calls      14,247,647        2,638,751        1,465,915
   written_back                 7,702          439,809          482,809
   store succeeded            329,603          531,224          566,150
   store rejected             883,945          517,578          482,644
   store reject rate             ~73%             ~49%             ~46%
   pswpout                    891,647          957,388          965,453
   pswpin                     692,255          628,246          582,781
   zswpout                    329,603          531,224          566,150
   zswpin                     248,871                3            7,872

Thanks,
Hao



^ permalink raw reply	[flat|nested] 31+ messages in thread

* Re: [PATCH v2 2/2] mm/zswap: Support batch writeback in shrink_memcg()
  2026-07-28 11:22                           ` Hao Jia
@ 2026-07-28 15:52                             ` Nhat Pham
  2026-07-28 15:53                               ` Yosry Ahmed
  0 siblings, 1 reply; 31+ messages in thread
From: Nhat Pham @ 2026-07-28 15:52 UTC (permalink / raw)
  To: Hao Jia
  Cc: Yosry Ahmed, Johannes Weiner, akpm, tj, shakeel.butt, mhocko,
	mkoutny, chengming.zhou, muchun.song, roman.gushchin, linux-mm,
	linux-kernel, linux-doc, Hao Jia

On Tue, Jul 28, 2026 at 4:22 AM Hao Jia <jiahao.kernel@gmail.com> wrote:
>
>
> On 2026/7/28 00:18, Yosry Ahmed wrote:
> Apologies—there was a flaw in my previous test setup. I had configured
> zswap.max=320M alongside max_pool_percent=1. Since the global zswap pool
> capacity was roughly 314M (< 320M), the global pool limit was actually
> being hit before the cgroup zswap.max limit. Both limits were likely
> interleaving throughout the run, making the results difficult to analyze.
>
> Re-running with **zswap.max=280M** yields results that align with
> expectations:
>
> Test Case: zmax280 stress-4 pool1%
>                              baseline        batch-all32      batch-all64
>    shrink_worker wakeups            0                0                0
>    shrink_memcg calls         698,990           57,187           29,649
>    written_back               655,570          865,131          883,508
>    store succeeded          1,054,928        1,124,990        1,190,608
>    store rejected             229,722           22,642           11,659
>    store reject rate             ~17%              ~1%              ~0%
>    reject_reclaim_fail            110           15,454           12,985
>    pswpout                    885,437          888,027          895,478
>    pswpin                   2,035,415        1,305,341        1,384,751
>    zswpout                  1,054,928        1,124,990        1,190,608
>    zswpin                     280,675          162,684          158,556
>
> To isolate the cgroup zswap limit, I also tested with
> max_pool_percent=50 (which should behave similarly to the zswap.max=280M
> case above, bounded strictly by cgroup zswap.max):
>
> Test Case: zmax320 stress-4 pool50%
>                             baseline        batch-all32      batch-all64
>    shrink_worker wakeups            0                0                0
>    shrink_memcg calls         687,608           54,002           28,116
>    written_back               639,176          846,663          862,347
>    store succeeded            992,816        1,208,123        1,213,286
>    store rejected             231,431           20,425           10,915
>    store reject rate             ~18%              ~1%              ~0%
>    pswpout                    870,745          867,360          873,553
>    pswpin                   1,707,823        1,216,814        1,378,381
>    zswpout                    992,816        1,208,123        1,213,286
>    zswpin                     250,220          210,791          185,716
>
>
> Finally, here is the dataset for zswap.max=320M and max_pool_percent=1
> with a single stress-ng worker (stress-1), which also looks consistent
> with your expectations:
>
> Test Case: zmax320 stress-1 pool1%
>                            baseline        batch-all32      batch-all64
>    shrink_worker wakeups        9,279              980              557
>    shrink_memcg calls      14,247,647        2,638,751        1,465,915
>    written_back                 7,702          439,809          482,809
>    store succeeded            329,603          531,224          566,150
>    store rejected             883,945          517,578          482,644
>    store reject rate             ~73%             ~49%             ~46%
>    pswpout                    891,647          957,388          965,453
>    pswpin                     692,255          628,246          582,781
>    zswpout                    329,603          531,224          566,150
>    zswpin                     248,871                3            7,872

Shall we just go with SWAP_CLUSTER_MAX per Johannes' suggestion then?
32 works the best here FWIW, but seems like it's all relatively good
(as long as it's not 1).

(I wonder if this works well with your proactive reclaimer too but we
shall see).

^ permalink raw reply	[flat|nested] 31+ messages in thread

* Re: [PATCH v2 2/2] mm/zswap: Support batch writeback in shrink_memcg()
  2026-07-28 15:52                             ` Nhat Pham
@ 2026-07-28 15:53                               ` Yosry Ahmed
  0 siblings, 0 replies; 31+ messages in thread
From: Yosry Ahmed @ 2026-07-28 15:53 UTC (permalink / raw)
  To: Nhat Pham
  Cc: Hao Jia, Johannes Weiner, akpm, tj, shakeel.butt, mhocko,
	mkoutny, chengming.zhou, muchun.song, roman.gushchin, linux-mm,
	linux-kernel, linux-doc, Hao Jia

On Tue, Jul 28, 2026 at 8:52 AM Nhat Pham <nphamcs@gmail.com> wrote:
>
> On Tue, Jul 28, 2026 at 4:22 AM Hao Jia <jiahao.kernel@gmail.com> wrote:
> >
> >
> > On 2026/7/28 00:18, Yosry Ahmed wrote:
> > Apologies—there was a flaw in my previous test setup. I had configured
> > zswap.max=320M alongside max_pool_percent=1. Since the global zswap pool
> > capacity was roughly 314M (< 320M), the global pool limit was actually
> > being hit before the cgroup zswap.max limit. Both limits were likely
> > interleaving throughout the run, making the results difficult to analyze.

Thanks for running the numbers again, these do look better!

> >
> > Re-running with **zswap.max=280M** yields results that align with
> > expectations:
> >
> > Test Case: zmax280 stress-4 pool1%
> >                              baseline        batch-all32      batch-all64
> >    shrink_worker wakeups            0                0                0
> >    shrink_memcg calls         698,990           57,187           29,649
> >    written_back               655,570          865,131          883,508
> >    store succeeded          1,054,928        1,124,990        1,190,608
> >    store rejected             229,722           22,642           11,659
> >    store reject rate             ~17%              ~1%              ~0%
> >    reject_reclaim_fail            110           15,454           12,985
> >    pswpout                    885,437          888,027          895,478
> >    pswpin                   2,035,415        1,305,341        1,384,751
> >    zswpout                  1,054,928        1,124,990        1,190,608
> >    zswpin                     280,675          162,684          158,556
> >
> > To isolate the cgroup zswap limit, I also tested with
> > max_pool_percent=50 (which should behave similarly to the zswap.max=280M
> > case above, bounded strictly by cgroup zswap.max):
> >
> > Test Case: zmax320 stress-4 pool50%
> >                             baseline        batch-all32      batch-all64
> >    shrink_worker wakeups            0                0                0
> >    shrink_memcg calls         687,608           54,002           28,116
> >    written_back               639,176          846,663          862,347
> >    store succeeded            992,816        1,208,123        1,213,286
> >    store rejected             231,431           20,425           10,915
> >    store reject rate             ~18%              ~1%              ~0%
> >    pswpout                    870,745          867,360          873,553
> >    pswpin                   1,707,823        1,216,814        1,378,381
> >    zswpout                    992,816        1,208,123        1,213,286
> >    zswpin                     250,220          210,791          185,716
> >
> >
> > Finally, here is the dataset for zswap.max=320M and max_pool_percent=1
> > with a single stress-ng worker (stress-1), which also looks consistent
> > with your expectations:
> >
> > Test Case: zmax320 stress-1 pool1%
> >                            baseline        batch-all32      batch-all64
> >    shrink_worker wakeups        9,279              980              557
> >    shrink_memcg calls      14,247,647        2,638,751        1,465,915
> >    written_back                 7,702          439,809          482,809
> >    store succeeded            329,603          531,224          566,150
> >    store rejected             883,945          517,578          482,644
> >    store reject rate             ~73%             ~49%             ~46%
> >    pswpout                    891,647          957,388          965,453
> >    pswpin                     692,255          628,246          582,781
> >    zswpout                    329,603          531,224          566,150
> >    zswpin                     248,871                3            7,872
>
> Shall we just go with SWAP_CLUSTER_MAX per Johannes' suggestion then?
> 32 works the best here FWIW, but seems like it's all relatively good
> (as long as it's not 1).

Yeah I think so, the less magic macros the better.

>
> (I wonder if this works well with your proactive reclaimer too but we
> shall see).

^ permalink raw reply	[flat|nested] 31+ messages in thread

end of thread, other threads:[~2026-07-28 15:53 UTC | newest]

Thread overview: 31+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-07-17  8:51 [PATCH v2 0/2] mm/zswap: Fixes and improves the zswap global shrinker Hao Jia
2026-07-17  8:51 ` [PATCH v2 1/2] mm/zswap: Fix global shrinker when memory cgroup is disabled Hao Jia
2026-07-23  2:13   ` Johannes Weiner
2026-07-17  8:51 ` [PATCH v2 2/2] mm/zswap: Support batch writeback in shrink_memcg() Hao Jia
2026-07-17 16:45   ` Yosry Ahmed
2026-07-17 16:46   ` Nhat Pham
2026-07-23  2:27   ` Johannes Weiner
2026-07-23  4:52     ` Yosry Ahmed
2026-07-23 13:55       ` Johannes Weiner
2026-07-23 16:39         ` Yosry Ahmed
2026-07-23 17:11           ` Johannes Weiner
2026-07-24 10:20           ` Hao Jia
2026-07-24 17:56             ` Yosry Ahmed
2026-07-24 18:37               ` Nhat Pham
2026-07-24 18:39                 ` Yosry Ahmed
2026-07-24 19:35                   ` Johannes Weiner
2026-07-24 22:22                     ` Yosry Ahmed
2026-07-26  2:37                       ` Hao Jia
2026-07-27 16:18                         ` Yosry Ahmed
2026-07-28 11:22                           ` Hao Jia
2026-07-28 15:52                             ` Nhat Pham
2026-07-28 15:53                               ` Yosry Ahmed
2026-07-24 18:40             ` Johannes Weiner
2026-07-18  1:18 ` [PATCH v2 0/2] mm/zswap: Fixes and improves the zswap global shrinker Andrew Morton
2026-07-18  1:22   ` Yosry Ahmed
2026-07-18  1:28   ` Yosry Ahmed
2026-07-18  4:40     ` Andrew Morton
2026-07-20  1:26       ` Hao Jia
2026-07-23  1:21         ` Hao Jia
2026-07-23  4:49           ` Yosry Ahmed
2026-07-24 10:22             ` Hao Jia

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®