mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH 1/1] mm/zswap: enable static key after runtime pool recovery
@ 2026-09-05 12:50 Longlong Xia
  2026-09-05 23:09 ` Andrew Morton
  2026-09-06  9:09 ` Yosry Ahmed
  0 siblings, 2 replies; 11+ messages in thread
From: Longlong Xia @ 2026-09-05 12:50 UTC (permalink / raw)
  To: hannes, yosry, nphamcs, akpm
  Cc: chengming.zhou, linux-mm, linux-kernel, stable, Longlong Xia

From: Longlong Xia <xialonglong@kylinos.cn>

When CONFIG_ZSWAP_DEFAULT_ON is disabled, zswap_setup() can complete
without a pool after a failed initial pool creation. A later compressor
parameter update can create and publish a pool, but does not enable
zswap_ever_enabled.

If users then enable zswap, zswap_store() intercepts swapout while
zswap_load() still returns -ENOENT without consulting the xarray. The
swapin path therefore reads a stale backing swap slot because the store
skipped writing it.

Enable the static key after a successful compressor and pool update. Do
this outside zswap_pools_lock because static key updates may sleep.

Verified with fault injection on a stock kernel (compressor builtin,
CONFIG_ZSWAP_DEFAULT_ON=n):

  1. Boot with zswap.enabled=1; pool creation fails, init completes
     pool-less (static key off).
  2. Echo an available compressor name to zswap.compressor; a pool is
     recovered but the key stays off.
  3. Enable zswap.
  4. madvise(MADV_PAGEOUT) a pattern-verified 512 MiB region, then
     fault it back in and verify.

Step 4 reads back 131072/131072 zeroed pages (zswpin=0, zswpout=131072)
without this patch; all pages intact (zswpin=131072) with it.

Fixes: 2d4d2b1cfb85 ("mm: zswap: add zswap_never_enabled()")
Cc: stable@vger.kernel.org
Assisted-by: Zcode:GLM-5.3
Signed-off-by: Longlong Xia <xialonglong@kylinos.cn>
---
 mm/zswap.c | 3 +++
 1 file changed, 3 insertions(+)

diff --git a/mm/zswap.c b/mm/zswap.c
index 37f34e406c8e3..c48c4df63f188 100644
--- a/mm/zswap.c
+++ b/mm/zswap.c
@@ -586,6 +586,9 @@ static int zswap_compressor_param_set(const char *val, const struct kernel_param
 	else
 		ret = -EINVAL;
 
+	if (!ret)
+		static_branch_enable(&zswap_ever_enabled);
+
 	spin_lock_bh(&zswap_pools_lock);
 
 	if (!ret) {
-- 
2.43.0


^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH 1/1] mm/zswap: enable static key after runtime pool recovery
  2026-09-05 12:50 [PATCH 1/1] mm/zswap: enable static key after runtime pool recovery Longlong Xia
@ 2026-09-05 23:09 ` Andrew Morton
  2026-09-06  0:31   ` Longlong Xia
  2026-09-06  9:19   ` Yosry Ahmed
  2026-09-06  9:09 ` Yosry Ahmed
  1 sibling, 2 replies; 11+ messages in thread
From: Andrew Morton @ 2026-09-05 23:09 UTC (permalink / raw)
  To: Longlong Xia
  Cc: hannes, yosry, nphamcs, chengming.zhou, linux-mm, linux-kernel,
	stable, Longlong Xia

On Sat,  5 Sep 2026 20:50:28 +0800 Longlong Xia <xialonglong2025@163.com> wrote:

> From: Longlong Xia <xialonglong@kylinos.cn>
> 
> When CONFIG_ZSWAP_DEFAULT_ON is disabled, zswap_setup() can complete
> without a pool after a failed initial pool creation. A later compressor
> parameter update can create and publish a pool, but does not enable
> zswap_ever_enabled.
> 
> If users then enable zswap, zswap_store() intercepts swapout while
> zswap_load() still returns -ENOENT without consulting the xarray. The
> swapin path therefore reads a stale backing swap slot because the store
> skipped writing it.

That sounds bad.  I'll leave it to reviewers to suggest whether this is
a sufficient description of the runtime effects, and to decide whether
a backport is appropriate.  Please.

> Enable the static key after a successful compressor and pool update. Do
> this outside zswap_pools_lock because static key updates may sleep.
> 
> Verified with fault injection on a stock kernel (compressor builtin,
> CONFIG_ZSWAP_DEFAULT_ON=n):
> 
>   1. Boot with zswap.enabled=1; pool creation fails, init completes
>      pool-less (static key off).
>   2. Echo an available compressor name to zswap.compressor; a pool is
>      recovered but the key stays off.
>   3. Enable zswap.
>   4. madvise(MADV_PAGEOUT) a pattern-verified 512 MiB region, then
>      fault it back in and verify.
> 
> Step 4 reads back 131072/131072 zeroed pages (zswpin=0, zswpout=131072)
> without this patch; all pages intact (zswpin=131072) with it.

And thanks.  Sashiko might have found another issue in this zswap code:
	https://sashiko.dev/#/patchset/20260905125101.2970456-1-xialonglong2025@163.com


^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH 1/1] mm/zswap: enable static key after runtime pool recovery
  2026-09-05 23:09 ` Andrew Morton
@ 2026-09-06  0:31   ` Longlong Xia
  2026-09-06  9:19   ` Yosry Ahmed
  1 sibling, 0 replies; 11+ messages in thread
From: Longlong Xia @ 2026-09-06  0:31 UTC (permalink / raw)
  To: Andrew Morton
  Cc: hannes, yosry, nphamcs, chengming.zhou, linux-mm, linux-kernel,
	stable, Longlong Xia

Thanks for taking a look.

在 2026/9/6 7:09, Andrew Morton 写道:
> On Sat,  5 Sep 2026 20:50:28 +0800 Longlong Xia <xialonglong2025@163.com> wrote:
>
>> From: Longlong Xia <xialonglong@kylinos.cn>
>>
>> When CONFIG_ZSWAP_DEFAULT_ON is disabled, zswap_setup() can complete
>> without a pool after a failed initial pool creation. A later compressor
>> parameter update can create and publish a pool, but does not enable
>> zswap_ever_enabled.
>>
>> If users then enable zswap, zswap_store() intercepts swapout while
>> zswap_load() still returns -ENOENT without consulting the xarray. The
>> swapin path therefore reads a stale backing swap slot because the store
>> skipped writing it.
> That sounds bad.  I'll leave it to reviewers to suggest whether this is
> a sufficient description of the runtime effects, and to decide whether
> a backport is appropriate.  Please.
>
>> Enable the static key after a successful compressor and pool update. Do
>> this outside zswap_pools_lock because static key updates may sleep.
>>
>> Verified with fault injection on a stock kernel (compressor builtin,
>> CONFIG_ZSWAP_DEFAULT_ON=n):
>>
>>    1. Boot with zswap.enabled=1; pool creation fails, init completes
>>       pool-less (static key off).
>>    2. Echo an available compressor name to zswap.compressor; a pool is
>>       recovered but the key stays off.
>>    3. Enable zswap.
>>    4. madvise(MADV_PAGEOUT) a pattern-verified 512 MiB region, then
>>       fault it back in and verify.
>>
>> Step 4 reads back 131072/131072 zeroed pages (zswpin=0, zswpout=131072)
>> without this patch; all pages intact (zswpin=131072) with it.
> And thanks.  Sashiko might have found another issue in this zswap code:
> 	https://sashiko.dev/#/patchset/20260905125101.2970456-1-xialonglong2025@163.com

I'll send a separate fix patch.


Thanks,

Longlong



^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH 1/1] mm/zswap: enable static key after runtime pool recovery
  2026-09-05 12:50 [PATCH 1/1] mm/zswap: enable static key after runtime pool recovery Longlong Xia
  2026-09-05 23:09 ` Andrew Morton
@ 2026-09-06  9:09 ` Yosry Ahmed
  2026-09-06 13:36   ` [PATCH v2 1/1] mm/zswap: enable zswap_ever_enabled in zswap_pool_create() Longlong Xia
  1 sibling, 1 reply; 11+ messages in thread
From: Yosry Ahmed @ 2026-09-06  9:09 UTC (permalink / raw)
  To: Longlong Xia
  Cc: hannes, nphamcs, akpm, chengming.zhou, linux-mm, linux-kernel,
	stable, Longlong Xia

On Sat, Sep 5, 2026 at 5:51 AM Longlong Xia <xialonglong2025@163.com> wrote:
>
> From: Longlong Xia <xialonglong@kylinos.cn>
>
> When CONFIG_ZSWAP_DEFAULT_ON is disabled, zswap_setup() can complete
> without a pool after a failed initial pool creation. A later compressor
> parameter update can create and publish a pool, but does not enable
> zswap_ever_enabled.
>
> If users then enable zswap, zswap_store() intercepts swapout while
> zswap_load() still returns -ENOENT without consulting the xarray. The
> swapin path therefore reads a stale backing swap slot because the store
> skipped writing it.
>
> Enable the static key after a successful compressor and pool update. Do
> this outside zswap_pools_lock because static key updates may sleep.
>
> Verified with fault injection on a stock kernel (compressor builtin,
> CONFIG_ZSWAP_DEFAULT_ON=n):
>
>   1. Boot with zswap.enabled=1; pool creation fails, init completes
>      pool-less (static key off).
>   2. Echo an available compressor name to zswap.compressor; a pool is
>      recovered but the key stays off.
>   3. Enable zswap.
>   4. madvise(MADV_PAGEOUT) a pattern-verified 512 MiB region, then
>      fault it back in and verify.
>
> Step 4 reads back 131072/131072 zeroed pages (zswpin=0, zswpout=131072)
> without this patch; all pages intact (zswpin=131072) with it.
>
> Fixes: 2d4d2b1cfb85 ("mm: zswap: add zswap_never_enabled()")
> Cc: stable@vger.kernel.org
> Assisted-by: Zcode:GLM-5.3
> Signed-off-by: Longlong Xia <xialonglong@kylinos.cn>
> ---
>  mm/zswap.c | 3 +++
>  1 file changed, 3 insertions(+)
>
> diff --git a/mm/zswap.c b/mm/zswap.c
> index 37f34e406c8e3..c48c4df63f188 100644
> --- a/mm/zswap.c
> +++ b/mm/zswap.c
> @@ -586,6 +586,9 @@ static int zswap_compressor_param_set(const char *val, const struct kernel_param
>         else
>                 ret = -EINVAL;
>
> +       if (!ret)
> +               static_branch_enable(&zswap_ever_enabled);

What if we move it from zswap_setup() to zswap_pool_create()? IIUC
this would cover both cases?

> +
>         spin_lock_bh(&zswap_pools_lock);
>
>         if (!ret) {
> --
> 2.43.0
>

^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH 1/1] mm/zswap: enable static key after runtime pool recovery
  2026-09-05 23:09 ` Andrew Morton
  2026-09-06  0:31   ` Longlong Xia
@ 2026-09-06  9:19   ` Yosry Ahmed
  2026-09-07 11:00     ` Usama Arif
  1 sibling, 1 reply; 11+ messages in thread
From: Yosry Ahmed @ 2026-09-06  9:19 UTC (permalink / raw)
  To: Andrew Morton
  Cc: Longlong Xia, hannes, nphamcs, chengming.zhou, linux-mm,
	linux-kernel, stable, Longlong Xia, Alexandre Ghiti, Usama Arif

On Sat, Sep 5, 2026 at 4:09 PM Andrew Morton <akpm@linux-foundation.org> wrote:
>
> On Sat,  5 Sep 2026 20:50:28 +0800 Longlong Xia <xialonglong2025@163.com> wrote:
>
> > From: Longlong Xia <xialonglong@kylinos.cn>
> >
> > When CONFIG_ZSWAP_DEFAULT_ON is disabled, zswap_setup() can complete
> > without a pool after a failed initial pool creation. A later compressor
> > parameter update can create and publish a pool, but does not enable
> > zswap_ever_enabled.
> >
> > If users then enable zswap, zswap_store() intercepts swapout while
> > zswap_load() still returns -ENOENT without consulting the xarray. The
> > swapin path therefore reads a stale backing swap slot because the store
> > skipped writing it.
>
> That sounds bad.  I'll leave it to reviewers to suggest whether this is
> a sufficient description of the runtime effects, and to decide whether
> a backport is appropriate.  Please.

Yes this needs a stable backport AFAICT.

A more high-level description would be:

If zswap is enabled by default at boot and pool creation fails, then a
pool is later created by updating the compressor, data written to
zswap is corrupted on swapin.

>
> > Enable the static key after a successful compressor and pool update. Do
> > this outside zswap_pools_lock because static key updates may sleep.
> >
> > Verified with fault injection on a stock kernel (compressor builtin,
> > CONFIG_ZSWAP_DEFAULT_ON=n):
> >
> >   1. Boot with zswap.enabled=1; pool creation fails, init completes
> >      pool-less (static key off).
> >   2. Echo an available compressor name to zswap.compressor; a pool is
> >      recovered but the key stays off.
> >   3. Enable zswap.
> >   4. madvise(MADV_PAGEOUT) a pattern-verified 512 MiB region, then
> >      fault it back in and verify.
> >
> > Step 4 reads back 131072/131072 zeroed pages (zswpin=0, zswpout=131072)
> > without this patch; all pages intact (zswpin=131072) with it.
>
> And thanks.  Sashiko might have found another issue in this zswap code:
>         https://sashiko.dev/#/patchset/20260905125101.2970456-1-xialonglong2025@163.com

Hmm I think this might be fixed by Alexandre's patch (in Usama's
series): https://lore.kernel.org/linux-mm/20260818131202.494754-6-usama.arif@linux.dev/.

Instead of always returning -EINVAL for large folios we only do so if
they are actually in zswap. Usama/Alexandre, assuming I got this
right, can I interest you in sending the zswap bits of that patch as a
standalone fix? :)

^ permalink raw reply	[flat|nested] 11+ messages in thread

* [PATCH v2 1/1] mm/zswap: enable zswap_ever_enabled in zswap_pool_create()
  2026-09-06  9:09 ` Yosry Ahmed
@ 2026-09-06 13:36   ` Longlong Xia
  2026-09-06 13:43     ` Yosry Ahmed
  0 siblings, 1 reply; 11+ messages in thread
From: Longlong Xia @ 2026-09-06 13:36 UTC (permalink / raw)
  To: yosry
  Cc: akpm, chengming.zhou, hannes, linux-kernel, linux-mm, nphamcs,
	stable, xialonglong2025, xialonglong

From: Longlong Xia <xialonglong@kylinos.cn>

If zswap is enabled by default at boot and pool creation fails, then a
pool is later created by updating the compressor, data written to
zswap is corrupted on swapin.

Enable the static key in zswap_pool_create(), covering boot-time and
runtime pool creation with a single site.

Verified with fault injection on a stock kernel (compressor builtin,
CONFIG_ZSWAP_DEFAULT_ON=n):

  1. Boot with zswap.enabled=1; pool creation fails, init completes
     pool-less (static key off).
  2. Echo an available compressor name to zswap.compressor.
  3. Enable zswap.
  4. madvise(MADV_PAGEOUT) a pattern-verified 512 MiB region, then
     fault it back in and verify.

Step 4 reads back 131072/131072 zeroed pages (zswpin=0, zswpout=131072)
without this patch; all pages intact (zswpin=131072) with it.

Fixes: 2d4d2b1cfb85 ("mm: zswap: add zswap_never_enabled()")
Suggested-by: Yosry Ahmed <yosry@kernel.org>
Cc: stable@vger.kernel.org
Assisted-by: Zcode:GLM-5.3
Signed-off-by: Longlong Xia <xialonglong@kylinos.cn>
---
Changes in v2:
- Enable the static key in zswap_pool_create() instead of adding a
  second site next to the compressor parameter update, as suggested
  by Yosry; boot-time and runtime pool creation are now covered by a
  single site, and the zswap_setup() site is gone.
- Use the high-level description from Yosry's reply to Andrew for the
  corruption scenario, and rename the subject accordingly.

Link: https://lore.kernel.org/r/20260905125101.2970456-1-xialonglong2025@163.com

 mm/zswap.c | 4 +++-
 1 file changed, 3 insertions(+), 1 deletion(-)

diff --git a/mm/zswap.c b/mm/zswap.c
index 37f34e406c8e3..6dfb6ae709e97 100644
--- a/mm/zswap.c
+++ b/mm/zswap.c
@@ -324,6 +324,9 @@ static struct zswap_pool *zswap_pool_create(char *compressor)
 
 	zswap_pool_debug("created", pool);
 
+	/* Enable the key here so every pool creation path is covered. */
+	static_branch_enable(&zswap_ever_enabled);
+
 	return pool;
 
 ref_fail:
@@ -1805,7 +1808,6 @@ static int zswap_setup(void)
 		pr_info("loaded using pool %s\n", pool->tfm_name);
 		list_add(&pool->list, &zswap_pools);
 		zswap_has_pool = true;
-		static_branch_enable(&zswap_ever_enabled);
 	} else {
 		pr_err("pool creation failed\n");
 		zswap_enabled = false;
-- 
2.43.0


^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH v2 1/1] mm/zswap: enable zswap_ever_enabled in zswap_pool_create()
  2026-09-06 13:36   ` [PATCH v2 1/1] mm/zswap: enable zswap_ever_enabled in zswap_pool_create() Longlong Xia
@ 2026-09-06 13:43     ` Yosry Ahmed
  2026-09-06 13:59       ` [PATCH v3 " Longlong Xia
  0 siblings, 1 reply; 11+ messages in thread
From: Yosry Ahmed @ 2026-09-06 13:43 UTC (permalink / raw)
  To: Longlong Xia
  Cc: akpm, chengming.zhou, hannes, linux-kernel, linux-mm, nphamcs,
	stable, xialonglong

On Sun, Sep 6, 2026 at 6:37 AM Longlong Xia <xialonglong2025@163.com> wrote:
>
> From: Longlong Xia <xialonglong@kylinos.cn>
>
> If zswap is enabled by default at boot and pool creation fails, then a
> pool is later created by updating the compressor, data written to
> zswap is corrupted on swapin.
>
> Enable the static key in zswap_pool_create(), covering boot-time and
> runtime pool creation with a single site.
>
> Verified with fault injection on a stock kernel (compressor builtin,
> CONFIG_ZSWAP_DEFAULT_ON=n):
>
>   1. Boot with zswap.enabled=1; pool creation fails, init completes
>      pool-less (static key off).
>   2. Echo an available compressor name to zswap.compressor.
>   3. Enable zswap.
>   4. madvise(MADV_PAGEOUT) a pattern-verified 512 MiB region, then
>      fault it back in and verify.
>
> Step 4 reads back 131072/131072 zeroed pages (zswpin=0, zswpout=131072)
> without this patch; all pages intact (zswpin=131072) with it.
>
> Fixes: 2d4d2b1cfb85 ("mm: zswap: add zswap_never_enabled()")
> Suggested-by: Yosry Ahmed <yosry@kernel.org>
> Cc: stable@vger.kernel.org
> Assisted-by: Zcode:GLM-5.3
> Signed-off-by: Longlong Xia <xialonglong@kylinos.cn>
> ---
> Changes in v2:
> - Enable the static key in zswap_pool_create() instead of adding a
>   second site next to the compressor parameter update, as suggested
>   by Yosry; boot-time and runtime pool creation are now covered by a
>   single site, and the zswap_setup() site is gone.
> - Use the high-level description from Yosry's reply to Andrew for the
>   corruption scenario, and rename the subject accordingly.
>
> Link: https://lore.kernel.org/r/20260905125101.2970456-1-xialonglong2025@163.com
>
>  mm/zswap.c | 4 +++-
>  1 file changed, 3 insertions(+), 1 deletion(-)
>
> diff --git a/mm/zswap.c b/mm/zswap.c
> index 37f34e406c8e3..6dfb6ae709e97 100644
> --- a/mm/zswap.c
> +++ b/mm/zswap.c
> @@ -324,6 +324,9 @@ static struct zswap_pool *zswap_pool_create(char *compressor)
>
>         zswap_pool_debug("created", pool);
>
> +       /* Enable the key here so every pool creation path is covered. */
> +       static_branch_enable(&zswap_ever_enabled);

I would drop the comment, otherwise LGTM:

Acked-by: Yosry Ahmed <yosry@kernel.org>

> +
>         return pool;
>
>  ref_fail:
> @@ -1805,7 +1808,6 @@ static int zswap_setup(void)
>                 pr_info("loaded using pool %s\n", pool->tfm_name);
>                 list_add(&pool->list, &zswap_pools);
>                 zswap_has_pool = true;
> -               static_branch_enable(&zswap_ever_enabled);
>         } else {
>                 pr_err("pool creation failed\n");
>                 zswap_enabled = false;
> --
> 2.43.0
>

^ permalink raw reply	[flat|nested] 11+ messages in thread

* [PATCH v3 1/1] mm/zswap: enable zswap_ever_enabled in zswap_pool_create()
  2026-09-06 13:43     ` Yosry Ahmed
@ 2026-09-06 13:59       ` Longlong Xia
  0 siblings, 0 replies; 11+ messages in thread
From: Longlong Xia @ 2026-09-06 13:59 UTC (permalink / raw)
  To: yosry
  Cc: akpm, chengming.zhou, hannes, linux-kernel, linux-mm, nphamcs,
	stable, xialonglong2025, xialonglong

From: Longlong Xia <xialonglong@kylinos.cn>

If zswap is enabled by default at boot and pool creation fails, then a
pool is later created by updating the compressor, data written to
zswap is corrupted on swapin.

Enable the static key in zswap_pool_create(), covering boot-time and
runtime pool creation with a single site.

Verified with fault injection on a stock kernel (compressor builtin,
CONFIG_ZSWAP_DEFAULT_ON=n):

  1. Boot with zswap.enabled=1; pool creation fails, init completes
     pool-less (static key off).
  2. Echo an available compressor name to zswap.compressor.
  3. Enable zswap.
  4. madvise(MADV_PAGEOUT) a pattern-verified 512 MiB region, then
     fault it back in and verify.

Step 4 reads back 131072/131072 zeroed pages (zswpin=0, zswpout=131072)
without this patch; all pages intact (zswpin=131072) with it.

Fixes: 2d4d2b1cfb85 ("mm: zswap: add zswap_never_enabled()")
Suggested-by: Yosry Ahmed <yosry@kernel.org>
Acked-by: Yosry Ahmed <yosry@kernel.org>
Cc: stable@vger.kernel.org
Assisted-by: Zcode:GLM-5.3
Signed-off-by: Longlong Xia <xialonglong@kylinos.cn>
---
Changes in v3:
- Drop the code comment per Yosry, and add his Acked-by.

Changes in v2:
- Enable the static key in zswap_pool_create() instead of adding a
  second site next to the compressor parameter update, as suggested
  by Yosry; boot-time and runtime pool creation are now covered by a
  single site, and the zswap_setup() site is gone.
- Use the high-level description from Yosry's reply to Andrew for the
  corruption scenario, and rename the subject accordingly.

Link: https://lore.kernel.org/r/20260905125101.2970456-1-xialonglong2025@163.com

 mm/zswap.c | 3 ++-
 1 file changed, 2 insertions(+), 1 deletion(-)

diff --git a/mm/zswap.c b/mm/zswap.c
index 37f34e406c8e3..3abcd8433425b 100644
--- a/mm/zswap.c
+++ b/mm/zswap.c
@@ -324,6 +324,8 @@ static struct zswap_pool *zswap_pool_create(char *compressor)
 
 	zswap_pool_debug("created", pool);
 
+	static_branch_enable(&zswap_ever_enabled);
+
 	return pool;
 
 ref_fail:
@@ -1805,7 +1807,6 @@ static int zswap_setup(void)
 		pr_info("loaded using pool %s\n", pool->tfm_name);
 		list_add(&pool->list, &zswap_pools);
 		zswap_has_pool = true;
-		static_branch_enable(&zswap_ever_enabled);
 	} else {
 		pr_err("pool creation failed\n");
 		zswap_enabled = false;
-- 
2.43.0


^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH 1/1] mm/zswap: enable static key after runtime pool recovery
  2026-09-06  9:19   ` Yosry Ahmed
@ 2026-09-07 11:00     ` Usama Arif
  2026-09-07 11:34       ` Yosry Ahmed
  0 siblings, 1 reply; 11+ messages in thread
From: Usama Arif @ 2026-09-07 11:00 UTC (permalink / raw)
  To: Yosry Ahmed, Andrew Morton
  Cc: Longlong Xia, hannes, nphamcs, chengming.zhou, linux-mm,
	linux-kernel, stable, Longlong Xia, Alexandre Ghiti



On 06/09/2026 10:19, Yosry Ahmed wrote:
> On Sat, Sep 5, 2026 at 4:09 PM Andrew Morton <akpm@linux-foundation.org> wrote:
>>
>> On Sat,  5 Sep 2026 20:50:28 +0800 Longlong Xia <xialonglong2025@163.com> wrote:
>>
>>> From: Longlong Xia <xialonglong@kylinos.cn>
>>>
>>> When CONFIG_ZSWAP_DEFAULT_ON is disabled, zswap_setup() can complete
>>> without a pool after a failed initial pool creation. A later compressor
>>> parameter update can create and publish a pool, but does not enable
>>> zswap_ever_enabled.
>>>
>>> If users then enable zswap, zswap_store() intercepts swapout while
>>> zswap_load() still returns -ENOENT without consulting the xarray. The
>>> swapin path therefore reads a stale backing swap slot because the store
>>> skipped writing it.
>>
>> That sounds bad.  I'll leave it to reviewers to suggest whether this is
>> a sufficient description of the runtime effects, and to decide whether
>> a backport is appropriate.  Please.
> 
> Yes this needs a stable backport AFAICT.
> 
> A more high-level description would be:
> 
> If zswap is enabled by default at boot and pool creation fails, then a
> pool is later created by updating the compressor, data written to
> zswap is corrupted on swapin.
> 
>>
>>> Enable the static key after a successful compressor and pool update. Do
>>> this outside zswap_pools_lock because static key updates may sleep.
>>>
>>> Verified with fault injection on a stock kernel (compressor builtin,
>>> CONFIG_ZSWAP_DEFAULT_ON=n):
>>>
>>>   1. Boot with zswap.enabled=1; pool creation fails, init completes
>>>      pool-less (static key off).
>>>   2. Echo an available compressor name to zswap.compressor; a pool is
>>>      recovered but the key stays off.
>>>   3. Enable zswap.
>>>   4. madvise(MADV_PAGEOUT) a pattern-verified 512 MiB region, then
>>>      fault it back in and verify.
>>>
>>> Step 4 reads back 131072/131072 zeroed pages (zswpin=0, zswpout=131072)
>>> without this patch; all pages intact (zswpin=131072) with it.
>>
>> And thanks.  Sashiko might have found another issue in this zswap code:
>>         https://sashiko.dev/#/patchset/20260905125101.2970456-1-xialonglong2025@163.com
> 
> Hmm I think this might be fixed by Alexandre's patch (in Usama's
> series): https://lore.kernel.org/linux-mm/20260818131202.494754-6-usama.arif@linux.dev/.
> 
> Instead of always returning -EINVAL for large folios we only do so if
> they are actually in zswap. Usama/Alexandre, assuming I got this
> right, can I interest you in sending the zswap bits of that patch as a
> standalone fix? :)



Hello!

Below is what the patch looks like in my tree now. It can be sent independently of
the PMD swap series. Yosry if you are happy with it, will send it on the list

From a5b70b6d72eda15bd7a16b8fc51db1d271b08520 Mon Sep 17 00:00:00 2001
From: Alexandre Ghiti <alexghiti@fb.com>
Date: Wed, 22 Jul 2026 08:19:35 -0700
Subject: [PATCH 01/19] mm: zswap: add range lookup for large-folio swapin

A large folio reaches zswap_load() only when the caller expects
the whole range to be on disk. Zswap still stores large folios as
independent order-0 entries, so reconstructing a large folio from
zswap entries would risk returning partially initialized data.

Teach zswap_load() to scan the covered range. If no slot is in zswap,
return -ENOENT so swap_read_folio() reads the backing device. If any
slot is still in zswap, fail the large-folio read so the caller can
fall back to per-page swapin.

Return -EIO rather than -EINVAL for that conflict. Large-folio loads
are now valid requests; the error means zswap cannot safely satisfy
the request from partial per-page compressed state, not that the
request is unsupported. Existing callers only distinguish -ENOENT,
so this is a semantic clarification rather than a behavioral change.

Add zswap_is_present() so PMD swap-entry consumers can make the same
range decision before attempting PMD-order swapin. For high-order swap
cache allocations, check the zswap range after inserting the folio into
swap cache. The insertion stabilizes the range against zswap store and
writeback; if pre-existing per-page zswap entries are found, remove the
folio through the existing allocation rollback path and return -EBUSY.

Signed-off-by: Alexandre Ghiti <alexghiti@fb.com>
Signed-off-by: Usama Arif <usama.arif@linux.dev>
---
 include/linux/zswap.h |  6 ++++++
 mm/swap_state.c       | 29 ++++++++++++++++++++-------
 mm/zswap.c            | 46 +++++++++++++++++++++++++++++++------------
 3 files changed, 61 insertions(+), 20 deletions(-)

diff --git a/include/linux/zswap.h b/include/linux/zswap.h
index 30c193a1207e1..cd9efcf9dec94 100644
--- a/include/linux/zswap.h
+++ b/include/linux/zswap.h
@@ -35,6 +35,7 @@ void zswap_lruvec_state_init(struct lruvec *lruvec);
 void zswap_folio_swapin(struct folio *folio);
 bool zswap_is_enabled(void);
 bool zswap_never_enabled(void);
+bool zswap_is_present(swp_entry_t entry, unsigned int nr);
 #else
 
 struct zswap_lruvec_state {};
@@ -69,6 +70,11 @@ static inline bool zswap_never_enabled(void)
 	return true;
 }
 
+static inline bool zswap_is_present(swp_entry_t entry, unsigned int nr)
+{
+	return false;
+}
+
 #endif
 
 #endif /* _LINUX_ZSWAP_H */
diff --git a/mm/swap_state.c b/mm/swap_state.c
index b76eb3d876fd7..103ae7ae8a4a6 100644
--- a/mm/swap_state.c
+++ b/mm/swap_state.c
@@ -12,6 +12,7 @@
 #include <linux/kernel_stat.h>
 #include <linux/mempolicy.h>
 #include <linux/swap.h>
+#include <linux/zswap.h>
 #include <linux/leafops.h>
 #include <linux/init.h>
 #include <linux/pagemap.h>
@@ -459,16 +460,21 @@ static struct folio *__swap_cache_alloc(struct swap_cluster_info *ci,
 	__swap_cache_do_add_folio(ci, folio, entry);
 	spin_unlock(&ci->lock);
 
+	/*
+	 * Once the folio is in swap cache, zswap cannot start storing or
+	 * writing back any slot in the range. Reject high-order allocations
+	 * that raced with pre-existing per-page zswap entries.
+	 */
+	if (order && zswap_is_present(entry, nr_pages)) {
+		err = -EBUSY;
+		goto delete_folio;
+	}
+
 	if (mem_cgroup_swapin_charge_folio(folio, memcg_id,
 					   vmf ? vmf->vma->vm_mm : NULL, gfp)) {
-		spin_lock(&ci->lock);
-		__swap_cache_do_del_folio(ci, folio, entry, shadow);
-		spin_unlock(&ci->lock);
-		folio_unlock(folio);
-		/* nr_pages refs from swap cache, 1 from allocation */
-		folio_put_refs(folio, nr_pages + 1);
+		err = -ENOMEM;
 		count_mthp_stat(order, MTHP_STAT_SWPIN_FALLBACK_CHARGE);
-		return ERR_PTR(-ENOMEM);
+		goto delete_folio;
 	}
 
 	if (order > 1 && folio_memcg_alloc_deferred(folio)) {
@@ -492,6 +498,15 @@ static struct folio *__swap_cache_alloc(struct swap_cluster_info *ci,
 	/* Caller will initiate read into locked new_folio */
 	folio_add_lru(folio);
 	return folio;
+
+delete_folio:
+	spin_lock(&ci->lock);
+	__swap_cache_do_del_folio(ci, folio, entry, shadow);
+	spin_unlock(&ci->lock);
+	folio_unlock(folio);
+	/* nr_pages refs from swap cache, 1 from allocation */
+	folio_put_refs(folio, nr_pages + 1);
+	return ERR_PTR(err);
 }
 
 /**
diff --git a/mm/zswap.c b/mm/zswap.c
index 37f34e406c8e3..32671dc2bf84d 100644
--- a/mm/zswap.c
+++ b/mm/zswap.c
@@ -1571,6 +1571,23 @@ bool zswap_store(struct folio *folio)
 	return ret;
 }
 
+/**
+ * zswap_is_present() - is any slot in [entry, entry + nr) in zswap?
+ * @entry: base swap entry of the range
+ * @nr: number of contiguous slots to check (pass 1 for a single-slot query)
+ */
+bool zswap_is_present(swp_entry_t entry, unsigned int nr)
+{
+	pgoff_t offset = swp_offset(entry);
+	struct xarray *tree = swap_zswap_tree(entry);
+	unsigned long index = offset;
+
+	if (!nr || zswap_never_enabled())
+		return false;
+
+	return xa_find(tree, &index, offset + nr - 1, XA_PRESENT);
+}
+
 /**
  * zswap_load() - load a folio from zswap
  * @folio: folio to load
@@ -1578,13 +1595,9 @@ bool zswap_store(struct folio *folio)
  * Return: 0 on success, with the folio unlocked and marked up-to-date, or one
  * of the following error codes:
  *
- *  -EIO: if the swapped out content was in zswap, but could not be loaded
- *  into the page due to a decompression failure. The folio is unlocked, but
- *  NOT marked up-to-date, so that an IO error is emitted (e.g. do_swap_page()
- *  will SIGBUS).
- *
- *  -EINVAL: if the swapped out content was in zswap, but the page belongs
- *  to a large folio, which is not supported by zswap. The folio is unlocked,
+ *  -EIO: if the swapped out content was in zswap but could not be handed
+ *  back, either because decompression failed or because a slot in a
+ *  large-folio range is unexpectedly still in zswap. The folio is unlocked,
  *  but NOT marked up-to-date, so that an IO error is emitted (e.g.
  *  do_swap_page() will SIGBUS).
  *
@@ -1605,13 +1618,20 @@ int zswap_load(struct folio *folio)
 		return -ENOENT;
 
 	/*
-	 * Large folios should not be swapped in while zswap is being used, as
-	 * they are not properly handled. Zswap does not properly load large
-	 * folios, and a large folio may only be partially in zswap.
+	 * A large folio reaches zswap_load() only when its whole range is
+	 * expected to be on disk: PMD swap-entry consumers split before
+	 * calling into PMD-order swapin whenever any slot is still in zswap.
+	 * Confirm the range is entirely absent from zswap and return -ENOENT
+	 * so the caller reads it from disk; if a slot is unexpectedly still in
+	 * zswap, fail the read rather than return partially-initialized data.
 	 */
-	if (WARN_ON_ONCE(folio_test_large(folio))) {
-		folio_unlock(folio);
-		return -EINVAL;
+	if (folio_test_large(folio)) {
+		if (WARN_ON_ONCE(zswap_is_present(swp,
+						  folio_nr_pages(folio)))) {
+			folio_unlock(folio);
+			return -EIO;
+		}
+		return -ENOENT;
 	}
 
 	entry = xa_load(tree, offset);
-- 
2.53.0-Meta



^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH 1/1] mm/zswap: enable static key after runtime pool recovery
  2026-09-07 11:00     ` Usama Arif
@ 2026-09-07 11:34       ` Yosry Ahmed
  2026-09-07 16:22         ` Usama Arif
  0 siblings, 1 reply; 11+ messages in thread
From: Yosry Ahmed @ 2026-09-07 11:34 UTC (permalink / raw)
  To: Usama Arif
  Cc: Andrew Morton, Longlong Xia, hannes, nphamcs, chengming.zhou,
	linux-mm, linux-kernel, stable, Longlong Xia, Alexandre Ghiti

> >> And thanks.  Sashiko might have found another issue in this zswap code:
> >>         https://sashiko.dev/#/patchset/20260905125101.2970456-1-xialonglong2025@163.com
> >
> > Hmm I think this might be fixed by Alexandre's patch (in Usama's
> > series): https://lore.kernel.org/linux-mm/20260818131202.494754-6-usama.arif@linux.dev/.
> >
> > Instead of always returning -EINVAL for large folios we only do so if
> > they are actually in zswap. Usama/Alexandre, assuming I got this
> > right, can I interest you in sending the zswap bits of that patch as a
> > standalone fix? :)
>
>
>
> Hello!
>
> Below is what the patch looks like in my tree now. It can be sent independently of
> the PMD swap series. Yosry if you are happy with it, will send it on the list

Would you be able to send the zswap bits only as a standalone (and
hopefully backportable) fix to the issue Sashiko surfaced?

>
> From a5b70b6d72eda15bd7a16b8fc51db1d271b08520 Mon Sep 17 00:00:00 2001
> From: Alexandre Ghiti <alexghiti@fb.com>
> Date: Wed, 22 Jul 2026 08:19:35 -0700
> Subject: [PATCH 01/19] mm: zswap: add range lookup for large-folio swapin
>
> A large folio reaches zswap_load() only when the caller expects
> the whole range to be on disk. Zswap still stores large folios as
> independent order-0 entries, so reconstructing a large folio from
> zswap entries would risk returning partially initialized data.
>
> Teach zswap_load() to scan the covered range. If no slot is in zswap,
> return -ENOENT so swap_read_folio() reads the backing device. If any
> slot is still in zswap, fail the large-folio read so the caller can
> fall back to per-page swapin.
>
> Return -EIO rather than -EINVAL for that conflict. Large-folio loads
> are now valid requests; the error means zswap cannot safely satisfy
> the request from partial per-page compressed state, not that the
> request is unsupported. Existing callers only distinguish -ENOENT,
> so this is a semantic clarification rather than a behavioral change.
>
> Add zswap_is_present() so PMD swap-entry consumers can make the same
> range decision before attempting PMD-order swapin. For high-order swap
> cache allocations, check the zswap range after inserting the folio into
> swap cache. The insertion stabilizes the range against zswap store and
> writeback; if pre-existing per-page zswap entries are found, remove the
> folio through the existing allocation rollback path and return -EBUSY.
>
> Signed-off-by: Alexandre Ghiti <alexghiti@fb.com>
> Signed-off-by: Usama Arif <usama.arif@linux.dev>
> ---
>  include/linux/zswap.h |  6 ++++++
>  mm/swap_state.c       | 29 ++++++++++++++++++++-------
>  mm/zswap.c            | 46 +++++++++++++++++++++++++++++++------------
>  3 files changed, 61 insertions(+), 20 deletions(-)
>
> diff --git a/include/linux/zswap.h b/include/linux/zswap.h
> index 30c193a1207e1..cd9efcf9dec94 100644
> --- a/include/linux/zswap.h
> +++ b/include/linux/zswap.h
> @@ -35,6 +35,7 @@ void zswap_lruvec_state_init(struct lruvec *lruvec);
>  void zswap_folio_swapin(struct folio *folio);
>  bool zswap_is_enabled(void);
>  bool zswap_never_enabled(void);
> +bool zswap_is_present(swp_entry_t entry, unsigned int nr);
>  #else
>
>  struct zswap_lruvec_state {};
> @@ -69,6 +70,11 @@ static inline bool zswap_never_enabled(void)
>         return true;
>  }
>
> +static inline bool zswap_is_present(swp_entry_t entry, unsigned int nr)
> +{
> +       return false;
> +}
> +
>  #endif
>
>  #endif /* _LINUX_ZSWAP_H */
> diff --git a/mm/swap_state.c b/mm/swap_state.c
> index b76eb3d876fd7..103ae7ae8a4a6 100644
> --- a/mm/swap_state.c
> +++ b/mm/swap_state.c
> @@ -12,6 +12,7 @@
>  #include <linux/kernel_stat.h>
>  #include <linux/mempolicy.h>
>  #include <linux/swap.h>
> +#include <linux/zswap.h>
>  #include <linux/leafops.h>
>  #include <linux/init.h>
>  #include <linux/pagemap.h>
> @@ -459,16 +460,21 @@ static struct folio *__swap_cache_alloc(struct swap_cluster_info *ci,
>         __swap_cache_do_add_folio(ci, folio, entry);
>         spin_unlock(&ci->lock);
>
> +       /*
> +        * Once the folio is in swap cache, zswap cannot start storing or
> +        * writing back any slot in the range. Reject high-order allocations
> +        * that raced with pre-existing per-page zswap entries.
> +        */
> +       if (order && zswap_is_present(entry, nr_pages)) {
> +               err = -EBUSY;
> +               goto delete_folio;
> +       }
> +
>         if (mem_cgroup_swapin_charge_folio(folio, memcg_id,
>                                            vmf ? vmf->vma->vm_mm : NULL, gfp)) {
> -               spin_lock(&ci->lock);
> -               __swap_cache_do_del_folio(ci, folio, entry, shadow);
> -               spin_unlock(&ci->lock);
> -               folio_unlock(folio);
> -               /* nr_pages refs from swap cache, 1 from allocation */
> -               folio_put_refs(folio, nr_pages + 1);
> +               err = -ENOMEM;
>                 count_mthp_stat(order, MTHP_STAT_SWPIN_FALLBACK_CHARGE);
> -               return ERR_PTR(-ENOMEM);
> +               goto delete_folio;
>         }
>
>         if (order > 1 && folio_memcg_alloc_deferred(folio)) {
> @@ -492,6 +498,15 @@ static struct folio *__swap_cache_alloc(struct swap_cluster_info *ci,
>         /* Caller will initiate read into locked new_folio */
>         folio_add_lru(folio);
>         return folio;
> +
> +delete_folio:
> +       spin_lock(&ci->lock);
> +       __swap_cache_do_del_folio(ci, folio, entry, shadow);
> +       spin_unlock(&ci->lock);
> +       folio_unlock(folio);
> +       /* nr_pages refs from swap cache, 1 from allocation */
> +       folio_put_refs(folio, nr_pages + 1);
> +       return ERR_PTR(err);
>  }
>
>  /**
> diff --git a/mm/zswap.c b/mm/zswap.c
> index 37f34e406c8e3..32671dc2bf84d 100644
> --- a/mm/zswap.c
> +++ b/mm/zswap.c
> @@ -1571,6 +1571,23 @@ bool zswap_store(struct folio *folio)
>         return ret;
>  }
>
> +/**
> + * zswap_is_present() - is any slot in [entry, entry + nr) in zswap?
> + * @entry: base swap entry of the range
> + * @nr: number of contiguous slots to check (pass 1 for a single-slot query)
> + */
> +bool zswap_is_present(swp_entry_t entry, unsigned int nr)
> +{
> +       pgoff_t offset = swp_offset(entry);
> +       struct xarray *tree = swap_zswap_tree(entry);
> +       unsigned long index = offset;
> +
> +       if (!nr || zswap_never_enabled())
> +               return false;
> +
> +       return xa_find(tree, &index, offset + nr - 1, XA_PRESENT);
> +}
> +
>  /**
>   * zswap_load() - load a folio from zswap
>   * @folio: folio to load
> @@ -1578,13 +1595,9 @@ bool zswap_store(struct folio *folio)
>   * Return: 0 on success, with the folio unlocked and marked up-to-date, or one
>   * of the following error codes:
>   *
> - *  -EIO: if the swapped out content was in zswap, but could not be loaded
> - *  into the page due to a decompression failure. The folio is unlocked, but
> - *  NOT marked up-to-date, so that an IO error is emitted (e.g. do_swap_page()
> - *  will SIGBUS).
> - *
> - *  -EINVAL: if the swapped out content was in zswap, but the page belongs
> - *  to a large folio, which is not supported by zswap. The folio is unlocked,
> + *  -EIO: if the swapped out content was in zswap but could not be handed
> + *  back, either because decompression failed or because a slot in a
> + *  large-folio range is unexpectedly still in zswap. The folio is unlocked,
>   *  but NOT marked up-to-date, so that an IO error is emitted (e.g.
>   *  do_swap_page() will SIGBUS).
>   *
> @@ -1605,13 +1618,20 @@ int zswap_load(struct folio *folio)
>                 return -ENOENT;
>
>         /*
> -        * Large folios should not be swapped in while zswap is being used, as
> -        * they are not properly handled. Zswap does not properly load large
> -        * folios, and a large folio may only be partially in zswap.
> +        * A large folio reaches zswap_load() only when its whole range is
> +        * expected to be on disk: PMD swap-entry consumers split before
> +        * calling into PMD-order swapin whenever any slot is still in zswap.
> +        * Confirm the range is entirely absent from zswap and return -ENOENT
> +        * so the caller reads it from disk; if a slot is unexpectedly still in
> +        * zswap, fail the read rather than return partially-initialized data.
>          */
> -       if (WARN_ON_ONCE(folio_test_large(folio))) {
> -               folio_unlock(folio);
> -               return -EINVAL;
> +       if (folio_test_large(folio)) {
> +               if (WARN_ON_ONCE(zswap_is_present(swp,
> +                                                 folio_nr_pages(folio)))) {
> +                       folio_unlock(folio);
> +                       return -EIO;
> +               }
> +               return -ENOENT;
>         }
>
>         entry = xa_load(tree, offset);
> --
> 2.53.0-Meta
>
>

^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH 1/1] mm/zswap: enable static key after runtime pool recovery
  2026-09-07 11:34       ` Yosry Ahmed
@ 2026-09-07 16:22         ` Usama Arif
  0 siblings, 0 replies; 11+ messages in thread
From: Usama Arif @ 2026-09-07 16:22 UTC (permalink / raw)
  To: Yosry Ahmed
  Cc: Andrew Morton, Longlong Xia, hannes, nphamcs, chengming.zhou,
	linux-mm, linux-kernel, stable, Longlong Xia, Alexandre Ghiti



On 07/09/2026 12:34, Yosry Ahmed wrote:
>>>> And thanks.  Sashiko might have found another issue in this zswap code:
>>>>         https://sashiko.dev/#/patchset/20260905125101.2970456-1-xialonglong2025@163.com
>>>
>>> Hmm I think this might be fixed by Alexandre's patch (in Usama's
>>> series): https://lore.kernel.org/linux-mm/20260818131202.494754-6-usama.arif@linux.dev/.
>>>
>>> Instead of always returning -EINVAL for large folios we only do so if
>>> they are actually in zswap. Usama/Alexandre, assuming I got this
>>> right, can I interest you in sending the zswap bits of that patch as a
>>> standalone fix? :)
>>
>>
>>
>> Hello!
>>
>> Below is what the patch looks like in my tree now. It can be sent independently of
>> the PMD swap series. Yosry if you are happy with it, will send it on the list
> 
> Would you be able to send the zswap bits only as a standalone (and
> hopefully backportable) fix to the issue Sashiko surfaced?


Yes, have sent it in https://lore.kernel.org/all/20260907161938.1932355-1-usama.arif@linux.dev/.

> 
>>
>> From a5b70b6d72eda15bd7a16b8fc51db1d271b08520 Mon Sep 17 00:00:00 2001
>> From: Alexandre Ghiti <alexghiti@fb.com>
>> Date: Wed, 22 Jul 2026 08:19:35 -0700
>> Subject: [PATCH 01/19] mm: zswap: add range lookup for large-folio swapin
>>
>> A large folio reaches zswap_load() only when the caller expects
>> the whole range to be on disk. Zswap still stores large folios as
>> independent order-0 entries, so reconstructing a large folio from
>> zswap entries would risk returning partially initialized data.
>>
>> Teach zswap_load() to scan the covered range. If no slot is in zswap,
>> return -ENOENT so swap_read_folio() reads the backing device. If any
>> slot is still in zswap, fail the large-folio read so the caller can
>> fall back to per-page swapin.
>>
>> Return -EIO rather than -EINVAL for that conflict. Large-folio loads
>> are now valid requests; the error means zswap cannot safely satisfy
>> the request from partial per-page compressed state, not that the
>> request is unsupported. Existing callers only distinguish -ENOENT,
>> so this is a semantic clarification rather than a behavioral change.
>>
>> Add zswap_is_present() so PMD swap-entry consumers can make the same
>> range decision before attempting PMD-order swapin. For high-order swap
>> cache allocations, check the zswap range after inserting the folio into
>> swap cache. The insertion stabilizes the range against zswap store and
>> writeback; if pre-existing per-page zswap entries are found, remove the
>> folio through the existing allocation rollback path and return -EBUSY.
>>
>> Signed-off-by: Alexandre Ghiti <alexghiti@fb.com>
>> Signed-off-by: Usama Arif <usama.arif@linux.dev>
>> ---
>>  include/linux/zswap.h |  6 ++++++
>>  mm/swap_state.c       | 29 ++++++++++++++++++++-------
>>  mm/zswap.c            | 46 +++++++++++++++++++++++++++++++------------
>>  3 files changed, 61 insertions(+), 20 deletions(-)
>>
>> diff --git a/include/linux/zswap.h b/include/linux/zswap.h
>> index 30c193a1207e1..cd9efcf9dec94 100644
>> --- a/include/linux/zswap.h
>> +++ b/include/linux/zswap.h
>> @@ -35,6 +35,7 @@ void zswap_lruvec_state_init(struct lruvec *lruvec);
>>  void zswap_folio_swapin(struct folio *folio);
>>  bool zswap_is_enabled(void);
>>  bool zswap_never_enabled(void);
>> +bool zswap_is_present(swp_entry_t entry, unsigned int nr);
>>  #else
>>
>>  struct zswap_lruvec_state {};
>> @@ -69,6 +70,11 @@ static inline bool zswap_never_enabled(void)
>>         return true;
>>  }
>>
>> +static inline bool zswap_is_present(swp_entry_t entry, unsigned int nr)
>> +{
>> +       return false;
>> +}
>> +
>>  #endif
>>
>>  #endif /* _LINUX_ZSWAP_H */
>> diff --git a/mm/swap_state.c b/mm/swap_state.c
>> index b76eb3d876fd7..103ae7ae8a4a6 100644
>> --- a/mm/swap_state.c
>> +++ b/mm/swap_state.c
>> @@ -12,6 +12,7 @@
>>  #include <linux/kernel_stat.h>
>>  #include <linux/mempolicy.h>
>>  #include <linux/swap.h>
>> +#include <linux/zswap.h>
>>  #include <linux/leafops.h>
>>  #include <linux/init.h>
>>  #include <linux/pagemap.h>
>> @@ -459,16 +460,21 @@ static struct folio *__swap_cache_alloc(struct swap_cluster_info *ci,
>>         __swap_cache_do_add_folio(ci, folio, entry);
>>         spin_unlock(&ci->lock);
>>
>> +       /*
>> +        * Once the folio is in swap cache, zswap cannot start storing or
>> +        * writing back any slot in the range. Reject high-order allocations
>> +        * that raced with pre-existing per-page zswap entries.
>> +        */
>> +       if (order && zswap_is_present(entry, nr_pages)) {
>> +               err = -EBUSY;
>> +               goto delete_folio;
>> +       }
>> +
>>         if (mem_cgroup_swapin_charge_folio(folio, memcg_id,
>>                                            vmf ? vmf->vma->vm_mm : NULL, gfp)) {
>> -               spin_lock(&ci->lock);
>> -               __swap_cache_do_del_folio(ci, folio, entry, shadow);
>> -               spin_unlock(&ci->lock);
>> -               folio_unlock(folio);
>> -               /* nr_pages refs from swap cache, 1 from allocation */
>> -               folio_put_refs(folio, nr_pages + 1);
>> +               err = -ENOMEM;
>>                 count_mthp_stat(order, MTHP_STAT_SWPIN_FALLBACK_CHARGE);
>> -               return ERR_PTR(-ENOMEM);
>> +               goto delete_folio;
>>         }
>>
>>         if (order > 1 && folio_memcg_alloc_deferred(folio)) {
>> @@ -492,6 +498,15 @@ static struct folio *__swap_cache_alloc(struct swap_cluster_info *ci,
>>         /* Caller will initiate read into locked new_folio */
>>         folio_add_lru(folio);
>>         return folio;
>> +
>> +delete_folio:
>> +       spin_lock(&ci->lock);
>> +       __swap_cache_do_del_folio(ci, folio, entry, shadow);
>> +       spin_unlock(&ci->lock);
>> +       folio_unlock(folio);
>> +       /* nr_pages refs from swap cache, 1 from allocation */
>> +       folio_put_refs(folio, nr_pages + 1);
>> +       return ERR_PTR(err);
>>  }
>>
>>  /**
>> diff --git a/mm/zswap.c b/mm/zswap.c
>> index 37f34e406c8e3..32671dc2bf84d 100644
>> --- a/mm/zswap.c
>> +++ b/mm/zswap.c
>> @@ -1571,6 +1571,23 @@ bool zswap_store(struct folio *folio)
>>         return ret;
>>  }
>>
>> +/**
>> + * zswap_is_present() - is any slot in [entry, entry + nr) in zswap?
>> + * @entry: base swap entry of the range
>> + * @nr: number of contiguous slots to check (pass 1 for a single-slot query)
>> + */
>> +bool zswap_is_present(swp_entry_t entry, unsigned int nr)
>> +{
>> +       pgoff_t offset = swp_offset(entry);
>> +       struct xarray *tree = swap_zswap_tree(entry);
>> +       unsigned long index = offset;
>> +
>> +       if (!nr || zswap_never_enabled())
>> +               return false;
>> +
>> +       return xa_find(tree, &index, offset + nr - 1, XA_PRESENT);
>> +}
>> +
>>  /**
>>   * zswap_load() - load a folio from zswap
>>   * @folio: folio to load
>> @@ -1578,13 +1595,9 @@ bool zswap_store(struct folio *folio)
>>   * Return: 0 on success, with the folio unlocked and marked up-to-date, or one
>>   * of the following error codes:
>>   *
>> - *  -EIO: if the swapped out content was in zswap, but could not be loaded
>> - *  into the page due to a decompression failure. The folio is unlocked, but
>> - *  NOT marked up-to-date, so that an IO error is emitted (e.g. do_swap_page()
>> - *  will SIGBUS).
>> - *
>> - *  -EINVAL: if the swapped out content was in zswap, but the page belongs
>> - *  to a large folio, which is not supported by zswap. The folio is unlocked,
>> + *  -EIO: if the swapped out content was in zswap but could not be handed
>> + *  back, either because decompression failed or because a slot in a
>> + *  large-folio range is unexpectedly still in zswap. The folio is unlocked,
>>   *  but NOT marked up-to-date, so that an IO error is emitted (e.g.
>>   *  do_swap_page() will SIGBUS).
>>   *
>> @@ -1605,13 +1618,20 @@ int zswap_load(struct folio *folio)
>>                 return -ENOENT;
>>
>>         /*
>> -        * Large folios should not be swapped in while zswap is being used, as
>> -        * they are not properly handled. Zswap does not properly load large
>> -        * folios, and a large folio may only be partially in zswap.
>> +        * A large folio reaches zswap_load() only when its whole range is
>> +        * expected to be on disk: PMD swap-entry consumers split before
>> +        * calling into PMD-order swapin whenever any slot is still in zswap.
>> +        * Confirm the range is entirely absent from zswap and return -ENOENT
>> +        * so the caller reads it from disk; if a slot is unexpectedly still in
>> +        * zswap, fail the read rather than return partially-initialized data.
>>          */
>> -       if (WARN_ON_ONCE(folio_test_large(folio))) {
>> -               folio_unlock(folio);
>> -               return -EINVAL;
>> +       if (folio_test_large(folio)) {
>> +               if (WARN_ON_ONCE(zswap_is_present(swp,
>> +                                                 folio_nr_pages(folio)))) {
>> +                       folio_unlock(folio);
>> +                       return -EIO;
>> +               }
>> +               return -ENOENT;
>>         }
>>
>>         entry = xa_load(tree, offset);
>> --
>> 2.53.0-Meta
>>
>>


^ permalink raw reply	[flat|nested] 11+ messages in thread

end of thread, other threads:[~2026-09-07 16:22 UTC | newest]

Thread overview: 11+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-05 12:50 [PATCH 1/1] mm/zswap: enable static key after runtime pool recovery Longlong Xia
2026-09-05 23:09 ` Andrew Morton
2026-09-06  0:31   ` Longlong Xia
2026-09-06  9:19   ` Yosry Ahmed
2026-09-07 11:00     ` Usama Arif
2026-09-07 11:34       ` Yosry Ahmed
2026-09-07 16:22         ` Usama Arif
2026-09-06  9:09 ` Yosry Ahmed
2026-09-06 13:36   ` [PATCH v2 1/1] mm/zswap: enable zswap_ever_enabled in zswap_pool_create() Longlong Xia
2026-09-06 13:43     ` Yosry Ahmed
2026-09-06 13:59       ` [PATCH v3 " Longlong Xia

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®