mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [RFC PATCH 0/4] mm/zswap: make shrinker writeback work with iocost
@ 2026-09-28  8:18 Alexandre Ghiti
  2026-09-28  8:18 ` [RFC PATCH 1/4] block: do not issue background swap bios as root Alexandre Ghiti
                   ` (3 more replies)
  0 siblings, 4 replies; 6+ messages in thread
From: Alexandre Ghiti @ 2026-09-28  8:18 UTC (permalink / raw)
  To: linux-mm
  Cc: linux-kernel, linux-block, cgroups, Andrew Morton, david,
	Johannes Weiner, Yosry Ahmed, Nhat Pham, Chengming Zhou,
	Jens Axboe, Tejun Heo, Josef Bacik, Chris Li, Kairui Song,
	Kemeng Shi, Baoquan He, Barry Song, Youngjun Park,
	Alexandre Ghiti

We want to enable the zswap shrinker on our fleet, which runs iocost,
but with it on, a large production workload is slower than with it
off. Two reasons:

1. zswap writeback is REQ_SWAP, so the IO controllers issue it as root
   and charge it to the cgroup afterwards, as debt. The workload pays it
   back: its swapin and page cache reads wait, and iocost eventually
   stalls its allocations, which in turn stalls the application.
   Patches 1-3 let zswap writeback be throttled instead, by marking it
   REQ_BACKGROUND.

2. A throttled write sleeps in its submitter, which for the shrinker is
   whoever entered reclaim, often an application thread in a page fault.
   Patch 4 moves that writeback to a per-lruvec kworker.

Large production workload, 10 hosts per arm, CPU matched, one 24h run
per comparison. Change with this series:

                     vs shrinker off    vs stock shrinker
  p50                    -0.7%               -5.8%
  p99                    -0.2%               -6.5%
  memory PSI full        -52%                -94%
  zswap pool             -97%                +12%

The benchmarks below run with iocost enabled, in a cgroup whose
memory.max (the "cap") is smaller than their working set, so they are
under constant memcg reclaim.

UCacheBench, 12G cap, 5 runs:

                     shrinker off    stock shrinker    this series
  ops/s                  469k             280k             634k
  errors                  0%               44%              0%
  p99.9                 5.8 ms           16.6 ms          4.3 ms

MySQL, sysbench OLTP, 256M cap, 5 runs per host:

                     shrinker off    stock shrinker    this series
  TPS, host 1            OOM               413              609
  TPS, host 2            OOM               399              666
  p99, host 1            OOM             334 ms            90 ms
  p99, host 2            OOM             273 ms            44 ms

FeedSim, 16G cap:

                     shrinker off    stock shrinker    this series
  QPS                    4.7              30.2             32.4
  p95                  1306 ms           495 ms           494 ms

kbuild, defconfig -j4, 600M cap, 3 runs per host:

                     shrinker off    stock shrinker    this series
  build time, host 1    784 s             788 s            778 s
  build time, host 2    691 s             696 s            680 s

Alexandre Ghiti (4):
  block: do not issue background swap bios as root
  mm/zswap: turn shrink_memcg_cb()'s argument into a flags word
  mm/zswap: allow writeback to be throttled by the cgroup IO controllers
  mm/zswap: move reclaim-driven writeback to a kworker

 block/blk-cgroup.h       |   8 ++-
 include/linux/swap_ops.h |   1 +
 include/linux/zswap.h    |   4 ++
 mm/page_io.c             |   7 +-
 mm/zswap.c               | 148 ++++++++++++++++++++++++++++++++++-----
 5 files changed, 146 insertions(+), 22 deletions(-)


base-commit: fe2ec83746e501645709761605c2464a44fd2929
-- 
2.53.0-Meta


^ permalink raw reply	[flat|nested] 6+ messages in thread

* [RFC PATCH 1/4] block: do not issue background swap bios as root
  2026-09-28  8:18 [RFC PATCH 0/4] mm/zswap: make shrinker writeback work with iocost Alexandre Ghiti
@ 2026-09-28  8:18 ` Alexandre Ghiti
  2026-09-28  9:44   ` Christoph Hellwig
  2026-09-28  8:18 ` [RFC PATCH 2/4] mm/zswap: turn shrink_memcg_cb()'s argument into a flags word Alexandre Ghiti
                   ` (2 subsequent siblings)
  3 siblings, 1 reply; 6+ messages in thread
From: Alexandre Ghiti @ 2026-09-28  8:18 UTC (permalink / raw)
  To: linux-mm
  Cc: linux-kernel, linux-block, cgroups, Andrew Morton, david,
	Johannes Weiner, Yosry Ahmed, Nhat Pham, Chengming Zhou,
	Jens Axboe, Tejun Heo, Josef Bacik, Chris Li, Kairui Song,
	Kemeng Shi, Baoquan He, Barry Song, Youngjun Park,
	Alexandre Ghiti

bio_issue_as_root_blkg() issues every REQ_SWAP bio as root, bypassing
the cgroup IO controllers, because reclaim cannot stall indefinitely on
swap writes. Some swap writes are not urgent though: zswap writeback is
best effort.

Such a bio cannot simply drop REQ_SWAP: dm-crypt and dm-thin rely on it
for limit_swap_bios, which caps in-flight swap bios to avoid a memory
deadlock, see commit a666e5c05e7c ("dm: fix deadlock when swapping to
encrypted device").

So let REQ_BACKGROUND, which already marks IO that is not urgent, opt a
swap bio out of being issued as root. No swap bio sets it today, so
nothing changes until a caller does.

Signed-off-by: Alexandre Ghiti <alex@ghiti.fr>
---
 block/blk-cgroup.h | 8 +++++++-
 1 file changed, 7 insertions(+), 1 deletion(-)

diff --git a/block/blk-cgroup.h b/block/blk-cgroup.h
index e67c69839129..02009cb45495 100644
--- a/block/blk-cgroup.h
+++ b/block/blk-cgroup.h
@@ -242,10 +242,16 @@ void blkg_conf_close_bdev(struct blkg_conf_ctx *ctx)
  * the bio and attach the appropriate blkg to the bio.  Then we call this helper
  * and if it is true run with the root blkg for that queue and then do any
  * backcharging to the originating cgroup once the io is complete.
+ *
+ * A REQ_SWAP bio is issued as root because reclaim waits on it, unless it is
+ * also REQ_BACKGROUND, which marks it as not urgent.
  */
 static inline bool bio_issue_as_root_blkg(struct bio *bio)
 {
-	return (bio->bi_opf & (REQ_META | REQ_SWAP)) != 0;
+	blk_opf_t opf = bio->bi_opf;
+
+	return (opf & REQ_META) ||
+	       (opf & (REQ_SWAP | REQ_BACKGROUND)) == REQ_SWAP;
 }
 
 /**
-- 
2.53.0-Meta


^ permalink raw reply	[flat|nested] 6+ messages in thread

* [RFC PATCH 2/4] mm/zswap: turn shrink_memcg_cb()'s argument into a flags word
  2026-09-28  8:18 [RFC PATCH 0/4] mm/zswap: make shrinker writeback work with iocost Alexandre Ghiti
  2026-09-28  8:18 ` [RFC PATCH 1/4] block: do not issue background swap bios as root Alexandre Ghiti
@ 2026-09-28  8:18 ` Alexandre Ghiti
  2026-09-28  8:18 ` [RFC PATCH 3/4] mm/zswap: allow writeback to be throttled by the cgroup IO controllers Alexandre Ghiti
  2026-09-28  8:18 ` [RFC PATCH 4/4] mm/zswap: move reclaim-driven writeback to a kworker Alexandre Ghiti
  3 siblings, 0 replies; 6+ messages in thread
From: Alexandre Ghiti @ 2026-09-28  8:18 UTC (permalink / raw)
  To: linux-mm
  Cc: linux-kernel, linux-block, cgroups, Andrew Morton, david,
	Johannes Weiner, Yosry Ahmed, Nhat Pham, Chengming Zhou,
	Jens Axboe, Tejun Heo, Josef Bacik, Chris Li, Kairui Song,
	Kemeng Shi, Baoquan He, Barry Song, Youngjun Park,
	Alexandre Ghiti

shrink_memcg_cb() takes a bool pointer as the list_lru walk's argument,
which it sets when the walk stopped at an entry already in the swap
cache. Make it a pointer to a flags word instead, so that callers can
also pass flags in.

No functional change.

Signed-off-by: Alexandre Ghiti <alex@ghiti.fr>
---
 mm/zswap.c | 20 ++++++++++++++------
 1 file changed, 14 insertions(+), 6 deletions(-)

diff --git a/mm/zswap.c b/mm/zswap.c
index 37f34e406c8e..8925b6100402 100644
--- a/mm/zswap.c
+++ b/mm/zswap.c
@@ -1065,6 +1065,14 @@ static int zswap_writeback_entry(struct zswap_entry *entry,
 /*********************************
 * shrinker functions
 **********************************/
+/*
+ * shrink_memcg_cb() flags, passed by pointer as the list_lru walk's @arg.
+ *
+ * ZSWAP_SHRINK_SWAPCACHE (out) is set when the walk stopped at an entry whose
+ * folio is already in the swap cache.
+ */
+#define ZSWAP_SHRINK_SWAPCACHE	BIT(0)
+
 /*
  * The dynamic shrinker is modulated by the following factors:
  *
@@ -1091,7 +1099,7 @@ static enum lru_status shrink_memcg_cb(struct list_head *item, struct list_lru_o
 				       void *arg)
 {
 	struct zswap_entry *entry = container_of(item, struct zswap_entry, lru);
-	bool *encountered_page_in_swapcache = (bool *)arg;
+	unsigned int *flags = arg;
 	swp_entry_t swpentry;
 	enum lru_status ret = LRU_REMOVED_RETRY;
 	int writeback_result;
@@ -1157,9 +1165,9 @@ static enum lru_status shrink_memcg_cb(struct list_head *item, struct list_lru_o
 		 * into the warmer region. We should terminate shrinking (if we're in the dynamic
 		 * shrinker context).
 		 */
-		if (writeback_result == -EEXIST && encountered_page_in_swapcache) {
+		if (writeback_result == -EEXIST && flags) {
 			ret = LRU_STOP;
-			*encountered_page_in_swapcache = true;
+			*flags |= ZSWAP_SHRINK_SWAPCACHE;
 		}
 	} else {
 		zswap_written_back_pages++;
@@ -1172,7 +1180,7 @@ static unsigned long zswap_shrinker_scan(struct shrinker *shrinker,
 		struct shrink_control *sc)
 {
 	unsigned long shrink_ret;
-	bool encountered_page_in_swapcache = false;
+	unsigned int flags = 0;
 
 	if (!zswap_shrinker_enabled ||
 			!mem_cgroup_zswap_writeback_enabled(sc->memcg)) {
@@ -1181,9 +1189,9 @@ static unsigned long zswap_shrinker_scan(struct shrinker *shrinker,
 	}
 
 	shrink_ret = list_lru_shrink_walk(&zswap_list_lru, sc, &shrink_memcg_cb,
-		&encountered_page_in_swapcache);
+		&flags);
 
-	if (encountered_page_in_swapcache)
+	if (flags & ZSWAP_SHRINK_SWAPCACHE)
 		return SHRINK_STOP;
 
 	return shrink_ret ? shrink_ret : SHRINK_STOP;
-- 
2.53.0-Meta


^ permalink raw reply	[flat|nested] 6+ messages in thread

* [RFC PATCH 3/4] mm/zswap: allow writeback to be throttled by the cgroup IO controllers
  2026-09-28  8:18 [RFC PATCH 0/4] mm/zswap: make shrinker writeback work with iocost Alexandre Ghiti
  2026-09-28  8:18 ` [RFC PATCH 1/4] block: do not issue background swap bios as root Alexandre Ghiti
  2026-09-28  8:18 ` [RFC PATCH 2/4] mm/zswap: turn shrink_memcg_cb()'s argument into a flags word Alexandre Ghiti
@ 2026-09-28  8:18 ` Alexandre Ghiti
  2026-09-28  8:18 ` [RFC PATCH 4/4] mm/zswap: move reclaim-driven writeback to a kworker Alexandre Ghiti
  3 siblings, 0 replies; 6+ messages in thread
From: Alexandre Ghiti @ 2026-09-28  8:18 UTC (permalink / raw)
  To: linux-mm
  Cc: linux-kernel, linux-block, cgroups, Andrew Morton, david,
	Johannes Weiner, Yosry Ahmed, Nhat Pham, Chengming Zhou,
	Jens Axboe, Tejun Heo, Josef Bacik, Chris Li, Kairui Song,
	Kemeng Shi, Baoquan He, Barry Song, Youngjun Park,
	Alexandre Ghiti

bio_issue_as_root_blkg() exempts REQ_SWAP bios from the cgroup IO
controllers. The exemption avoids priority inversion: reclaim issues
swap writes on behalf of the whole system, so making them wait on one
cgroup's IO budget stalls memory reclaim for everyone.

For zswap writeback the exemption does harm: the write is charged to the
cgroup's IO budget anyway, afterwards, but is issued ahead of the
synchronous swapin and page cache reads that the workload blocks on.

Let zswap writeback opt out of it with REQ_BACKGROUND.

Signed-off-by: Alexandre Ghiti <alex@ghiti.fr>
---
 include/linux/swap_ops.h |  1 +
 mm/page_io.c             |  7 +++++--
 mm/zswap.c               | 20 ++++++++++++++++----
 3 files changed, 22 insertions(+), 6 deletions(-)

diff --git a/include/linux/swap_ops.h b/include/linux/swap_ops.h
index 57ac6c703f68..152598ce6c4f 100644
--- a/include/linux/swap_ops.h
+++ b/include/linux/swap_ops.h
@@ -17,6 +17,7 @@ struct swap_iocb {
 struct swap_io_ctx {
 	struct swap_iocb	*sio;
 	struct swap_info_struct	*sis;
+	bool			throttled;
 };
 
 /*
diff --git a/mm/page_io.c b/mm/page_io.c
index 88962571cb93..14dd54ae6d65 100644
--- a/mm/page_io.c
+++ b/mm/page_io.c
@@ -591,9 +591,12 @@ static void swap_bdev_submit_write(struct swap_io_ctx *ctx)
 {
 	struct swap_iocb *sio = ctx->sio;
 	struct bio *bio = &sio->bio;
+	blk_opf_t opf = REQ_OP_WRITE | REQ_SWAP;
 
-	bio_init(bio, ctx->sis->bdev, sio->bvecs, ARRAY_SIZE(sio->bvecs),
-			REQ_OP_WRITE | REQ_SWAP);
+	if (ctx->throttled)
+		opf |= REQ_BACKGROUND;
+
+	bio_init(bio, ctx->sis->bdev, sio->bvecs, ARRAY_SIZE(sio->bvecs), opf);
 	bio->bi_iter.bi_size = sio->len;
 	bio->bi_iter.bi_sector = swap_folio_sector(bio_first_folio_all(bio));
 	bio_associate_blkg_from_page(bio, bio_first_folio_all(bio));
diff --git a/mm/zswap.c b/mm/zswap.c
index 8925b6100402..3b902f29ed4c 100644
--- a/mm/zswap.c
+++ b/mm/zswap.c
@@ -983,16 +983,21 @@ static bool zswap_decompress(struct zswap_entry *entry, struct folio *folio)
  * in the first place.  After the folio has been decompressed into
  * the swap cache, the compressed version stored by zswap can be
  * freed.
+ *
+ * @throttled lets the cgroup IO controllers throttle the write rather than
+ * issue it as root.
  */
 static int zswap_writeback_entry(struct zswap_entry *entry,
-				 swp_entry_t swpentry)
+				 swp_entry_t swpentry, bool throttled)
 {
 	struct xarray *tree;
 	pgoff_t offset = swp_offset(swpentry);
 	struct folio *folio;
 	struct mempolicy *mpol;
 	struct swap_info_struct *si;
-	struct swap_io_ctx ctx = {};
+	struct swap_io_ctx ctx = {
+		.throttled = throttled,
+	};
 	int ret = 0;
 
 	/* try to allocate swap cache folio */
@@ -1070,8 +1075,14 @@ static int zswap_writeback_entry(struct zswap_entry *entry,
  *
  * ZSWAP_SHRINK_SWAPCACHE (out) is set when the walk stopped at an entry whose
  * folio is already in the swap cache.
+ *
+ * ZSWAP_SHRINK_THROTTLED (in) lets the cgroup IO controllers throttle the
+ * writeback rather than issue it as root. Only the memory pressure shrinker
+ * sets it. The pool limit path must not: once the pool is full zswap_store()
+ * rejects everything until this writeback drains it.
  */
 #define ZSWAP_SHRINK_SWAPCACHE	BIT(0)
+#define ZSWAP_SHRINK_THROTTLED	BIT(1)
 
 /*
  * The dynamic shrinker is modulated by the following factors:
@@ -1154,7 +1165,8 @@ static enum lru_status shrink_memcg_cb(struct list_head *item, struct list_lru_o
 	 */
 	spin_unlock(&l->lock);
 
-	writeback_result = zswap_writeback_entry(entry, swpentry);
+	writeback_result = zswap_writeback_entry(entry, swpentry, flags &&
+						 (*flags & ZSWAP_SHRINK_THROTTLED));
 
 	if (writeback_result) {
 		zswap_reject_reclaim_fail++;
@@ -1180,7 +1192,7 @@ static unsigned long zswap_shrinker_scan(struct shrinker *shrinker,
 		struct shrink_control *sc)
 {
 	unsigned long shrink_ret;
-	unsigned int flags = 0;
+	unsigned int flags = ZSWAP_SHRINK_THROTTLED;
 
 	if (!zswap_shrinker_enabled ||
 			!mem_cgroup_zswap_writeback_enabled(sc->memcg)) {
-- 
2.53.0-Meta


^ permalink raw reply	[flat|nested] 6+ messages in thread

* [RFC PATCH 4/4] mm/zswap: move reclaim-driven writeback to a kworker
  2026-09-28  8:18 [RFC PATCH 0/4] mm/zswap: make shrinker writeback work with iocost Alexandre Ghiti
                   ` (2 preceding siblings ...)
  2026-09-28  8:18 ` [RFC PATCH 3/4] mm/zswap: allow writeback to be throttled by the cgroup IO controllers Alexandre Ghiti
@ 2026-09-28  8:18 ` Alexandre Ghiti
  3 siblings, 0 replies; 6+ messages in thread
From: Alexandre Ghiti @ 2026-09-28  8:18 UTC (permalink / raw)
  To: linux-mm
  Cc: linux-kernel, linux-block, cgroups, Andrew Morton, david,
	Johannes Weiner, Yosry Ahmed, Nhat Pham, Chengming Zhou,
	Jens Axboe, Tejun Heo, Josef Bacik, Chris Li, Kairui Song,
	Kemeng Shi, Baoquan He, Barry Song, Youngjun Park,
	Alexandre Ghiti

zswap_shrinker_scan() writes back in whatever context reclaim called it
from: an application thread inside a page fault, or kswapd reclaiming
for the whole node. Now that the cgroup IO controllers throttle zswap
writeback, that context sleeps uninterruptibly until the owning cgroup's
IO budget allows the write. Such stalls were observed on a large
production workload.

So defer the writeback to a kworker: each lruvec has its own work item
on shrink_wq, which writes back what reclaim asked for, SWAP_CLUSTER_MAX
entries at a time, while holding a reference on the memcg. The shrinker
does not queue more work while the worker has yet to claim the previous
one, and reports nothing freed, as nothing is until the worker runs.

Signed-off-by: Alexandre Ghiti <alex@ghiti.fr>
---
 include/linux/zswap.h |   4 ++
 mm/zswap.c            | 114 +++++++++++++++++++++++++++++++++++++-----
 2 files changed, 106 insertions(+), 12 deletions(-)

diff --git a/include/linux/zswap.h b/include/linux/zswap.h
index 30c193a1207e..80391c440e97 100644
--- a/include/linux/zswap.h
+++ b/include/linux/zswap.h
@@ -4,6 +4,7 @@
 
 #include <linux/types.h>
 #include <linux/mm_types.h>
+#include <linux/workqueue_types.h>
 
 struct lruvec;
 
@@ -22,6 +23,9 @@ struct zswap_lruvec_state {
 	 * swapped them in.
 	 */
 	atomic_long_t nr_disk_swapins;
+
+	atomic_long_t nr_deferred_writeback;
+	struct work_struct deferred_writeback_work;
 };
 
 unsigned long zswap_total_pages(void);
diff --git a/mm/zswap.c b/mm/zswap.c
index 3b902f29ed4c..cf2dd7a5aff0 100644
--- a/mm/zswap.c
+++ b/mm/zswap.c
@@ -237,6 +237,8 @@ static inline struct xarray *swap_zswap_tree(swp_entry_t swp)
 #define zswap_pool_debug(msg, p)			\
 	pr_debug("%s pool %s\n", msg, (p)->tfm_name)
 
+static void zswap_deferred_writeback_work(struct work_struct *w);
+
 /*********************************
 * pool functions
 **********************************/
@@ -702,7 +704,11 @@ static void zswap_lru_del(struct zswap_entry *entry)
 
 void zswap_lruvec_state_init(struct lruvec *lruvec)
 {
-	atomic_long_set(&lruvec->zswap_lruvec_state.nr_disk_swapins, 0);
+	struct zswap_lruvec_state *zls = &lruvec->zswap_lruvec_state;
+
+	atomic_long_set(&zls->nr_disk_swapins, 0);
+	atomic_long_set(&zls->nr_deferred_writeback, 0);
+	INIT_WORK(&zls->deferred_writeback_work, zswap_deferred_writeback_work);
 }
 
 void zswap_folio_swapin(struct folio *folio)
@@ -726,9 +732,15 @@ void zswap_folio_swapin(struct folio *folio)
  *
  * shrink_worker() must handle the case where this function releases
  * the reference of memcg being shrunk.
+ *
+ * The deferred writeback workers hold a reference of the memcg too. Stop
+ * queued ones here so they never start; running ones drop theirs when they
+ * finish.
  */
 void zswap_memcg_offline_cleanup(struct mem_cgroup *memcg)
 {
+	int nid;
+
 	/* lock out zswap shrinker walking memcg tree */
 	spin_lock(&zswap_shrink_lock);
 	if (zswap_next_shrink == memcg) {
@@ -737,6 +749,20 @@ void zswap_memcg_offline_cleanup(struct mem_cgroup *memcg)
 		} while (zswap_next_shrink && !mem_cgroup_online(zswap_next_shrink));
 	}
 	spin_unlock(&zswap_shrink_lock);
+
+	for_each_node_state(nid, N_MEMORY) {
+		struct lruvec *lruvec = mem_cgroup_lruvec(memcg, NODE_DATA(nid));
+		struct zswap_lruvec_state *zls = &lruvec->zswap_lruvec_state;
+
+		/*
+		 * Returning true means the work was pending: it will not run,
+		 * so the reference zswap_defer_writeback() took for it has to
+		 * be dropped here. A work item that is already executing is
+		 * not cancelled and drops its own reference when it finishes.
+		 */
+		if (cancel_work(&zls->deferred_writeback_work))
+			mem_cgroup_put(memcg);
+	}
 }
 
 /*********************************
@@ -1188,25 +1214,89 @@ static enum lru_status shrink_memcg_cb(struct list_head *item, struct list_lru_o
 	return ret;
 }
 
-static unsigned long zswap_shrinker_scan(struct shrinker *shrinker,
-		struct shrink_control *sc)
+static void zswap_deferred_writeback_work(struct work_struct *w)
 {
-	unsigned long shrink_ret;
 	unsigned int flags = ZSWAP_SHRINK_THROTTLED;
+	struct mem_cgroup *memcg, *old_memcg;
+	struct zswap_lruvec_state *zls;
+	unsigned int noreclaim_flag;
+	struct lruvec *lruvec;
+	unsigned long nr;
+	int nid;
+
+	zls = container_of(w, struct zswap_lruvec_state,
+			   deferred_writeback_work);
+	lruvec = container_of(zls, struct lruvec, zswap_lruvec_state);
+	memcg = lruvec_memcg(lruvec);
+	nid = lruvec_pgdat(lruvec)->node_id;
+
+	nr = atomic_long_xchg(&zls->nr_deferred_writeback, 0);
 
 	if (!zswap_shrinker_enabled ||
-			!mem_cgroup_zswap_writeback_enabled(sc->memcg)) {
-		sc->nr_scanned = 0;
-		return SHRINK_STOP;
+	    !mem_cgroup_zswap_writeback_enabled(memcg))
+		nr = 0;
+
+	noreclaim_flag = memalloc_noreclaim_save();
+	/*
+	 * Once the memcg is offline, the IO falls back to current's memcg,
+	 * which is wrong for a kworker.
+	 */
+	old_memcg = set_active_memcg(memcg);
+	while (nr) {
+		unsigned long nr_to_walk = min(nr, SWAP_CLUSTER_MAX);
+		unsigned long budget = nr_to_walk;
+
+		list_lru_walk_one(&zswap_list_lru, nid, memcg, &shrink_memcg_cb,
+				  &flags, &nr_to_walk);
+
+		if (nr_to_walk == budget)
+			break;
+
+		nr -= budget - nr_to_walk;
+
+		if (flags & ZSWAP_SHRINK_SWAPCACHE)
+			break;
+
+		cond_resched();
 	}
+	set_active_memcg(old_memcg);
+	memalloc_noreclaim_restore(noreclaim_flag);
+
+	/* Paired with mem_cgroup_tryget_online() in zswap_defer_writeback(). */
+	mem_cgroup_put(memcg);
+}
+
+static bool zswap_defer_writeback(struct shrink_control *sc)
+{
+	struct lruvec *lruvec = mem_cgroup_lruvec(sc->memcg, NODE_DATA(sc->nid));
+	struct zswap_lruvec_state *zls = &lruvec->zswap_lruvec_state;
+	long old = 0;
+
+	/* Do not accumulate work: back off if the worker is lagging behind. */
+	if (!atomic_long_try_cmpxchg(&zls->nr_deferred_writeback, &old,
+				     sc->nr_to_scan))
+		return false;
+
+	if (!mem_cgroup_tryget_online(sc->memcg))
+		return false;
 
-	shrink_ret = list_lru_shrink_walk(&zswap_list_lru, sc, &shrink_memcg_cb,
-		&flags);
+	if (!queue_work(shrink_wq, &zls->deferred_writeback_work))
+		mem_cgroup_put(sc->memcg);
+
+	return true;
+}
 
-	if (flags & ZSWAP_SHRINK_SWAPCACHE)
+static unsigned long zswap_shrinker_scan(struct shrinker *shrinker,
+					 struct shrink_control *sc)
+{
+	if (!zswap_shrinker_enabled ||
+	    !mem_cgroup_zswap_writeback_enabled(sc->memcg)) {
+		sc->nr_scanned = 0;
 		return SHRINK_STOP;
+	}
 
-	return shrink_ret ? shrink_ret : SHRINK_STOP;
+	/* Nothing is freed until the worker runs, so report nothing. */
+	return zswap_defer_writeback(sc) ? 0 : SHRINK_STOP;
 }
 
 static unsigned long zswap_shrinker_count(struct shrinker *shrinker,
@@ -1807,7 +1897,7 @@ static int zswap_setup(void)
 		goto hp_fail;
 
 	shrink_wq = alloc_workqueue("zswap-shrink",
-			WQ_UNBOUND|WQ_MEM_RECLAIM, 1);
+			WQ_UNBOUND | WQ_MEM_RECLAIM, 0);
 	if (!shrink_wq)
 		goto shrink_wq_fail;
 
-- 
2.53.0-Meta


^ permalink raw reply	[flat|nested] 6+ messages in thread

* Re: [RFC PATCH 1/4] block: do not issue background swap bios as root
  2026-09-28  8:18 ` [RFC PATCH 1/4] block: do not issue background swap bios as root Alexandre Ghiti
@ 2026-09-28  9:44   ` Christoph Hellwig
  0 siblings, 0 replies; 6+ messages in thread
From: Christoph Hellwig @ 2026-09-28  9:44 UTC (permalink / raw)
  To: Alexandre Ghiti
  Cc: linux-mm, linux-kernel, linux-block, cgroups, Andrew Morton,
	david, Johannes Weiner, Yosry Ahmed, Nhat Pham, Chengming Zhou,
	Jens Axboe, Tejun Heo, Josef Bacik, Chris Li, Kairui Song,
	Kemeng Shi, Baoquan He, Barry Song, Youngjun Park

On Mon, Sep 28, 2026 at 10:18:36AM +0200, Alexandre Ghiti wrote:
> bio_issue_as_root_blkg() issues every REQ_SWAP bio as root, bypassing
> the cgroup IO controllers, because reclaim cannot stall indefinitely on
> swap writes. Some swap writes are not urgent though: zswap writeback is
> best effort.
> 
> Such a bio cannot simply drop REQ_SWAP: dm-crypt and dm-thin rely on it
> for limit_swap_bios, which caps in-flight swap bios to avoid a memory
> deadlock, see commit a666e5c05e7c ("dm: fix deadlock when swapping to
> encrypted device").

Is that true?  REQ_BACKGROUND very much implies the I/O can the throttled
and thus better not be foreground paging, and thus the caller should be
able to tolerate a failure.

^ permalink raw reply	[flat|nested] 6+ messages in thread

end of thread, other threads:[~2026-09-28  9:44 UTC | newest]

Thread overview: 6+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-28  8:18 [RFC PATCH 0/4] mm/zswap: make shrinker writeback work with iocost Alexandre Ghiti
2026-09-28  8:18 ` [RFC PATCH 1/4] block: do not issue background swap bios as root Alexandre Ghiti
2026-09-28  9:44   ` Christoph Hellwig
2026-09-28  8:18 ` [RFC PATCH 2/4] mm/zswap: turn shrink_memcg_cb()'s argument into a flags word Alexandre Ghiti
2026-09-28  8:18 ` [RFC PATCH 3/4] mm/zswap: allow writeback to be throttled by the cgroup IO controllers Alexandre Ghiti
2026-09-28  8:18 ` [RFC PATCH 4/4] mm/zswap: move reclaim-driven writeback to a kworker Alexandre Ghiti

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®