* [PATCH] mm: zswap: return -ENOENT when the swap device is gone
@ 2026-09-13 6:30 Baoquan He
2026-09-13 7:48 ` Andrew Morton
2026-09-13 7:51 ` Andrew Morton
0 siblings, 2 replies; 5+ messages in thread
From: Baoquan He @ 2026-09-13 6:30 UTC (permalink / raw)
To: akpm
Cc: hannes, yosry, nphamcs, chengming.zhou, linux-mm, linux-kernel,
kasong, chrisl, Baoquan He
zswap_writeback_entry() returns -EEXIST when get_swap_device() finds no
device. -EEXIST is the shrinker's "page already in swap cache" signal,
which makes zswap_shrinker_scan() stop shrinking entirely. A NULL
get_swap_device() instead means the device is being swapped off, so the
entry is simply stale.
Return -ENOENT so the shrinker skips the stale entry and keeps scanning.
Independent of xswap; affects all swap devices.
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
Acked-by: Nhat Pham <nphamcs@gmail.com>
---
Note:
This is taken from xswap patchset. Nhat suggested this is a fix, should
be sent out independently.
mm/zswap.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/mm/zswap.c b/mm/zswap.c
index 37f34e406c8e..b9948d4657d2 100644
--- a/mm/zswap.c
+++ b/mm/zswap.c
@@ -998,7 +998,7 @@ static int zswap_writeback_entry(struct zswap_entry *entry,
/* try to allocate swap cache folio */
si = get_swap_device(swpentry);
if (!si)
- return -EEXIST;
+ return -ENOENT;
mpol = get_task_policy(current);
folio = swap_cache_alloc_folio(swpentry, GFP_KERNEL, BIT(0), NULL, mpol,
--
2.54.0
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH] mm: zswap: return -ENOENT when the swap device is gone
2026-09-13 6:30 [PATCH] mm: zswap: return -ENOENT when the swap device is gone Baoquan He
@ 2026-09-13 7:48 ` Andrew Morton
2026-09-14 6:27 ` Baoquan He
2026-09-13 7:51 ` Andrew Morton
1 sibling, 1 reply; 5+ messages in thread
From: Andrew Morton @ 2026-09-13 7:48 UTC (permalink / raw)
To: Baoquan He
Cc: hannes, yosry, nphamcs, chengming.zhou, linux-mm, linux-kernel,
kasong, chrisl
On Sun, 13 Sep 2026 14:30:31 +0800 Baoquan He <hebaoquan@kylinos.cn> wrote:
> zswap_writeback_entry() returns -EEXIST when get_swap_device() finds no
> device. -EEXIST is the shrinker's "page already in swap cache" signal,
> which makes zswap_shrinker_scan() stop shrinking entirely. A NULL
> get_swap_device() instead means the device is being swapped off, so the
> entry is simply stale.
>
> Return -ENOENT so the shrinker skips the stale entry and keeps scanning.
> Independent of xswap; affects all swap devices.
I'm struggling to understand the userspace-visible runtime effects of this.
I see that reclaim will prematurely abort, but is this a once-off thing
which will resolve on the next reclaim attempt, or will the reclaim
failure persist for a significant period?
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH] mm: zswap: return -ENOENT when the swap device is gone
2026-09-13 6:30 [PATCH] mm: zswap: return -ENOENT when the swap device is gone Baoquan He
2026-09-13 7:48 ` Andrew Morton
@ 2026-09-13 7:51 ` Andrew Morton
2026-09-14 6:31 ` Baoquan He
1 sibling, 1 reply; 5+ messages in thread
From: Andrew Morton @ 2026-09-13 7:51 UTC (permalink / raw)
To: Baoquan He
Cc: hannes, yosry, nphamcs, chengming.zhou, linux-mm, linux-kernel,
kasong, chrisl
On Sun, 13 Sep 2026 14:30:31 +0800 Baoquan He <hebaoquan@kylinos.cn> wrote:
> zswap_writeback_entry() returns -EEXIST when get_swap_device() finds no
> device. -EEXIST is the shrinker's "page already in swap cache" signal,
> which makes zswap_shrinker_scan() stop shrinking entirely. A NULL
> get_swap_device() instead means the device is being swapped off, so the
> entry is simply stale.
>
> Return -ENOENT so the shrinker skips the stale entry and keeps scanning.
> Independent of xswap; affects all swap devices.
>
> ...
>
> --- a/mm/zswap.c
> +++ b/mm/zswap.c
> @@ -998,7 +998,7 @@ static int zswap_writeback_entry(struct zswap_entry *entry,
> /* try to allocate swap cache folio */
Comment was always lame ("say why, not what"). It now seems flat out wrong?
> si = get_swap_device(swpentry);
> if (!si)
> - return -EEXIST;
> + return -ENOENT;
mm-new has changed. I made this
/* try to allocate swap cache folio */
si = get_swap_device(swpentry);
if (IS_ERR_OR_NULL(si))
return -ENOENT;
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH] mm: zswap: return -ENOENT when the swap device is gone
2026-09-13 7:48 ` Andrew Morton
@ 2026-09-14 6:27 ` Baoquan He
0 siblings, 0 replies; 5+ messages in thread
From: Baoquan He @ 2026-09-14 6:27 UTC (permalink / raw)
To: Andrew Morton
Cc: Baoquan He, hannes, yosry, nphamcs, chengming.zhou, linux-mm,
linux-kernel, kasong, chrisl
On 09/13/26 at 12:48am, Andrew Morton wrote:
> On Sun, 13 Sep 2026 14:30:31 +0800 Baoquan He <hebaoquan@kylinos.cn> wrote:
>
> > zswap_writeback_entry() returns -EEXIST when get_swap_device() finds no
> > device. -EEXIST is the shrinker's "page already in swap cache" signal,
> > which makes zswap_shrinker_scan() stop shrinking entirely. A NULL
> > get_swap_device() instead means the device is being swapped off, so the
> > entry is simply stale.
> >
> > Return -ENOENT so the shrinker skips the stale entry and keeps scanning.
> > Independent of xswap; affects all swap devices.
>
> I'm struggling to understand the userspace-visible runtime effects of this.
>
> I see that reclaim will prematurely abort, but is this a once-off thing
> which will resolve on the next reclaim attempt, or will the reclaim
> failure persist for a significant period?
Not a one-off, and not permanent either: it lasts the whole swapoff.
Assume I have two swap disks. zswap is enabled. By default 20% of RAM is
the zswap upper limit. So now if I swapoff /dev/vdb, at the same time
reclaimer call shrinker to writeback, -EEXIST makes shrink_memcg_cb()
return LRU_STOP, which ends the shrink pass right there. get_swap_device()
returns NULL for an entry whose device is gone, so every such entry still
on the zswap LRU stops a pass where it stands. The pool gets almost nothing
written back for the length of the swapoff.
# swapon
NAME TYPE SIZE USED PRIO
/dev/vdb partition 4G 0B -1
/dev/vdc partition 2G 0B -1
The entry is rotated before writeback, so later passes get past it. It is
a throughput collapse, not a deadlock.
static enum lru_status shrink_memcg_cb(struct list_head *item, struct list_lru_one *l,
void *arg)
{
......
list_move_tail(item, &l->list);
......
writeback_result = zswap_writeback_entry(entry, swpentry);
......
}
I can't reproduce it now. And I forget how I met this, just did too many
tiems of testing and code change. this probably comes from reading the
code. So this may be a logic bug that rarely happens rather than a easily
seen regression.
Thanks
Baoquan
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH] mm: zswap: return -ENOENT when the swap device is gone
2026-09-13 7:51 ` Andrew Morton
@ 2026-09-14 6:31 ` Baoquan He
0 siblings, 0 replies; 5+ messages in thread
From: Baoquan He @ 2026-09-14 6:31 UTC (permalink / raw)
To: Andrew Morton
Cc: Baoquan He, hannes, yosry, nphamcs, chengming.zhou, linux-mm,
linux-kernel, kasong, chrisl
On 09/13/26 at 12:51am, Andrew Morton wrote:
> On Sun, 13 Sep 2026 14:30:31 +0800 Baoquan He <hebaoquan@kylinos.cn> wrote:
>
> > zswap_writeback_entry() returns -EEXIST when get_swap_device() finds no
> > device. -EEXIST is the shrinker's "page already in swap cache" signal,
> > which makes zswap_shrinker_scan() stop shrinking entirely. A NULL
> > get_swap_device() instead means the device is being swapped off, so the
> > entry is simply stale.
> >
> > Return -ENOENT so the shrinker skips the stale entry and keeps scanning.
> > Independent of xswap; affects all swap devices.
> >
> > ...
> >
> > --- a/mm/zswap.c
> > +++ b/mm/zswap.c
> > @@ -998,7 +998,7 @@ static int zswap_writeback_entry(struct zswap_entry *entry,
> > /* try to allocate swap cache folio */
>
> Comment was always lame ("say why, not what"). It now seems flat out wrong?
Agree, the comment should be moved down to be cloe to
"mpol = get_task_policy(current);" line.
>
> > si = get_swap_device(swpentry);
> > if (!si)
> > - return -EEXIST;
> > + return -ENOENT;
>
> mm-new has changed. I made this
Thanks. Does it need a v2? or just use ther version you tuned.
By the way, which mm branch is suggested to take as a base for mm
patches posting? I usually take mm-unstable branch, seems it's changed
to mm-new now?
>
>
> /* try to allocate swap cache folio */
> si = get_swap_device(swpentry);
> if (IS_ERR_OR_NULL(si))
> return -ENOENT;
>
>
^ permalink raw reply [flat|nested] 5+ messages in thread
end of thread, other threads:[~2026-09-14 6:31 UTC | newest]
Thread overview: 5+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-13 6:30 [PATCH] mm: zswap: return -ENOENT when the swap device is gone Baoquan He
2026-09-13 7:48 ` Andrew Morton
2026-09-14 6:27 ` Baoquan He
2026-09-13 7:51 ` Andrew Morton
2026-09-14 6:31 ` Baoquan He
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®