* [PATCH v3 1/2] mm, swap: distinguish a malformed swap entry from a dying device
2026-08-18 10:06 [PATCH v3 0/2] mm, swap: don't spin on a bad swap entry Breno Leitao
@ 2026-08-18 10:06 ` Breno Leitao
2026-08-18 16:18 ` Nhat Pham
2026-08-18 18:17 ` David Hildenbrand (Arm)
2026-08-18 10:06 ` [PATCH v3 2/2] mm: fail the fault on a malformed swap entry instead of retrying it Breno Leitao
2026-08-30 0:42 ` [PATCH v3 0/2] mm, swap: don't spin on a bad swap entry Andrew Morton
2 siblings, 2 replies; 8+ messages in thread
From: Breno Leitao @ 2026-08-18 10:06 UTC (permalink / raw)
To: Andrew Morton, Chris Li, Kairui Song, Kemeng Shi, Nhat Pham,
Baoquan He, Barry Song, Youngjun Park, David Hildenbrand,
Lorenzo Stoakes, Liam R. Howlett, Vlastimil Babka, Mike Rapoport,
Suren Baghdasaryan, Michal Hocko, Jann Horn, Pedro Falcato,
Hugh Dickins, Baolin Wang, Peter Xu, Johannes Weiner,
Yosry Ahmed, Chengming Zhou
Cc: linux-mm, linux-kernel, kernel-team, Breno Leitao
get_swap_device() returns NULL for two different things: an entry whose
type names no swap device or whose offset is past the end of one, and a
device that swapoff is taking away. The first never becomes valid, the
second does, and callers cannot tell them apart.
Return ERR_PTR(-EIO) for the two malformed cases and keep NULL for
swapoff. copy_nonpresent_pte() already reports -EIO for an entry whose
type names no device.
Callers bail out on failure either way, so switch them to
IS_ERR_OR_NULL(), and let the two paths that drop the reference skip an
error pointer. No functional change.
Reviewed-by: Barry Song <baohua@kernel.org>
Acked-by: Kairui Song <kasong@tencent.com>
Signed-off-by: Breno Leitao <leitao@debian.org>
---
mm/memory.c | 6 +++---
mm/mincore.c | 2 +-
mm/shmem.c | 2 +-
mm/swap_state.c | 4 ++--
mm/swapfile.c | 14 +++++++++-----
mm/userfaultfd.c | 4 ++--
mm/zswap.c | 2 +-
7 files changed, 19 insertions(+), 15 deletions(-)
diff --git a/mm/memory.c b/mm/memory.c
index d9cf941967cf0..03d8cf111d0be 100644
--- a/mm/memory.c
+++ b/mm/memory.c
@@ -4954,9 +4954,9 @@ vm_fault_t do_swap_page(struct vm_fault *vmf)
goto out;
}
- /* Prevent swapoff from happening to us. */
+ /* Prevent swapoff from happening to us, and reject a bad entry. */
si = get_swap_device(entry);
- if (unlikely(!si))
+ if (IS_ERR_OR_NULL(si))
goto out;
folio = swap_cache_get_folio(entry);
@@ -5266,7 +5266,7 @@ vm_fault_t do_swap_page(struct vm_fault *vmf)
if (vmf->pte)
pte_unmap_unlock(vmf->pte, vmf->ptl);
out:
- if (si)
+ if (!IS_ERR_OR_NULL(si))
put_swap_device(si);
return ret;
out_nomap:
diff --git a/mm/mincore.c b/mm/mincore.c
index ff4ac82817683..c086836bc4bcc 100644
--- a/mm/mincore.c
+++ b/mm/mincore.c
@@ -71,7 +71,7 @@ static unsigned char mincore_swap(swp_entry_t entry, bool shmem)
*/
if (shmem) {
si = get_swap_device(entry);
- if (!si)
+ if (IS_ERR_OR_NULL(si))
return 0;
}
folio = swap_cache_get_folio(entry);
diff --git a/mm/shmem.c b/mm/shmem.c
index 65572cbf1bd3c..d0a9f52bfed71 100644
--- a/mm/shmem.c
+++ b/mm/shmem.c
@@ -2276,7 +2276,7 @@ static int shmem_swapin_folio(struct inode *inode, pgoff_t index,
si = get_swap_device(index_entry);
order = shmem_confirm_swap(mapping, index, index_entry);
- if (unlikely(!si)) {
+ if (IS_ERR_OR_NULL(si)) {
if (order < 0)
return -EEXIST;
else
diff --git a/mm/swap_state.c b/mm/swap_state.c
index 4b7a3303c463b..f2e86d6626ecc 100644
--- a/mm/swap_state.c
+++ b/mm/swap_state.c
@@ -715,7 +715,7 @@ struct folio *read_swap_cache_async(struct swap_io_ctx *ctx, swp_entry_t entry,
struct folio *folio;
si = get_swap_device(entry);
- if (!si)
+ if (IS_ERR_OR_NULL(si))
return NULL;
mpol = get_vma_policy(vma, addr, 0, &ilx);
@@ -951,7 +951,7 @@ static struct folio *swap_vma_readahead(swp_entry_t targ_entry, gfp_t gfp_mask,
*/
if (swp_type(entry) != swp_type(targ_entry)) {
si = get_swap_device(entry);
- if (!si)
+ if (IS_ERR_OR_NULL(si))
continue;
}
folio = swap_cache_read_folio(&ctx, entry, gfp_mask, mpol, ilx,
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 4d4e3e3059f6b..3b1883930e943 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -1504,7 +1504,7 @@ int swap_retry_table_alloc(swp_entry_t entry, gfp_t gfp)
unsigned long offset = swp_offset(entry);
si = get_swap_device(entry);
- if (!si)
+ if (IS_ERR_OR_NULL(si))
return 0;
ci = __swap_offset_to_cluster(si, offset);
@@ -1859,7 +1859,10 @@ void folio_put_swap(struct folio *folio, struct page *page)
* Check whether swap entry is valid in the swap device. If so,
* return pointer to swap_info_struct, and keep the swap entry valid
* via preventing the swap device from being swapoff, until
- * put_swap_device() is called. Otherwise return NULL.
+ * put_swap_device() is called. Return NULL for an empty entry or a
+ * device that is going away, and ERR_PTR(-EIO) if the entry's type
+ * names no swap device or its offset is past the end of one. These EIOs
+ * are preceded by pr_err().
*
* Notice that swapoff or swapoff+swapon can still happen before the
* percpu_ref_tryget_live() in get_swap_device() or after the
@@ -1900,12 +1903,13 @@ struct swap_info_struct *get_swap_device(swp_entry_t entry)
return si;
bad_nofile:
pr_err("%s: %s%08lx\n", __func__, Bad_file, entry.val);
+ return ERR_PTR(-EIO);
out:
return NULL;
put_out:
pr_err("%s: %s%08lx\n", __func__, Bad_offset, entry.val);
percpu_ref_put(&si->users);
- return NULL;
+ return ERR_PTR(-EIO);
}
/*
@@ -2001,7 +2005,7 @@ int swp_swapcount(swp_entry_t entry)
int count;
si = get_swap_device(entry);
- if (!si)
+ if (IS_ERR_OR_NULL(si))
return 0;
ci = swap_cluster_lock(si, swp_offset(entry));
@@ -2127,7 +2131,7 @@ void swap_put_entries_direct(swp_entry_t entry, int nr)
struct swap_info_struct *si;
si = get_swap_device(entry);
- if (WARN_ON_ONCE(!si))
+ if (WARN_ON_ONCE(IS_ERR_OR_NULL(si)))
return;
if (WARN_ON_ONCE(end_offset > si->max))
goto out;
diff --git a/mm/userfaultfd.c b/mm/userfaultfd.c
index 24a4d92ffa3c2..cba5e20a641ed 100644
--- a/mm/userfaultfd.c
+++ b/mm/userfaultfd.c
@@ -1700,7 +1700,7 @@ static long move_pages_ptes(struct mm_struct *mm, pmd_t *dst_pmd, pmd_t *src_pmd
}
si = get_swap_device(entry);
- if (unlikely(!si)) {
+ if (IS_ERR_OR_NULL(si)) {
ret = -EAGAIN;
goto out;
}
@@ -1757,7 +1757,7 @@ static long move_pages_ptes(struct mm_struct *mm, pmd_t *dst_pmd, pmd_t *src_pmd
if (dst_pte)
pte_unmap(dst_pte);
mmu_notifier_invalidate_range_end(&range);
- if (si)
+ if (!IS_ERR_OR_NULL(si))
put_swap_device(si);
return ret;
diff --git a/mm/zswap.c b/mm/zswap.c
index f7c9c89f6449c..bc9b931d6f447 100644
--- a/mm/zswap.c
+++ b/mm/zswap.c
@@ -997,7 +997,7 @@ static int zswap_writeback_entry(struct zswap_entry *entry,
/* try to allocate swap cache folio */
si = get_swap_device(swpentry);
- if (!si)
+ if (IS_ERR_OR_NULL(si))
return -EEXIST;
mpol = get_task_policy(current);
--
2.53.0-Meta
^ permalink raw reply [flat|nested] 8+ messages in thread* [PATCH v3 2/2] mm: fail the fault on a malformed swap entry instead of retrying it
2026-08-18 10:06 [PATCH v3 0/2] mm, swap: don't spin on a bad swap entry Breno Leitao
2026-08-18 10:06 ` [PATCH v3 1/2] mm, swap: distinguish a malformed swap entry from a dying device Breno Leitao
@ 2026-08-18 10:06 ` Breno Leitao
2026-08-18 16:18 ` Nhat Pham
2026-08-18 18:18 ` David Hildenbrand (Arm)
2026-08-30 0:42 ` [PATCH v3 0/2] mm, swap: don't spin on a bad swap entry Andrew Morton
2 siblings, 2 replies; 8+ messages in thread
From: Breno Leitao @ 2026-08-18 10:06 UTC (permalink / raw)
To: Andrew Morton, Chris Li, Kairui Song, Kemeng Shi, Nhat Pham,
Baoquan He, Barry Song, Youngjun Park, David Hildenbrand,
Lorenzo Stoakes, Liam R. Howlett, Vlastimil Babka, Mike Rapoport,
Suren Baghdasaryan, Michal Hocko, Jann Horn, Pedro Falcato,
Hugh Dickins, Baolin Wang, Peter Xu, Johannes Weiner,
Yosry Ahmed, Chengming Zhou
Cc: linux-mm, linux-kernel, kernel-team, Breno Leitao
do_swap_page() returns 0 when get_swap_device() fails, which the fault
handler reads as "handled". For an entry that can never become valid
the retry takes the same fault again, so the thread spins forever,
retrying on the same fault.
Return VM_FAULT_SIGBUS (Bad access) for a malformed entry (pr_err() was
called at get_swap_device()).
Acked-by: Kairui Song <kasong@tencent.com>
Reviewed-by: Barry Song <baohua@kernel.org>
Signed-off-by: Breno Leitao <leitao@debian.org>
---
mm/memory.c | 5 ++++-
1 file changed, 4 insertions(+), 1 deletion(-)
diff --git a/mm/memory.c b/mm/memory.c
index 03d8cf111d0be..d9db4f1ae6f8a 100644
--- a/mm/memory.c
+++ b/mm/memory.c
@@ -4956,8 +4956,11 @@ vm_fault_t do_swap_page(struct vm_fault *vmf)
/* Prevent swapoff from happening to us, and reject a bad entry. */
si = get_swap_device(entry);
- if (IS_ERR_OR_NULL(si))
+ if (IS_ERR_OR_NULL(si)) {
+ if (IS_ERR(si))
+ ret = VM_FAULT_SIGBUS;
goto out;
+ }
folio = swap_cache_get_folio(entry);
if (folio)
--
2.53.0-Meta
^ permalink raw reply [flat|nested] 8+ messages in thread* Re: [PATCH v3 0/2] mm, swap: don't spin on a bad swap entry
2026-08-18 10:06 [PATCH v3 0/2] mm, swap: don't spin on a bad swap entry Breno Leitao
2026-08-18 10:06 ` [PATCH v3 1/2] mm, swap: distinguish a malformed swap entry from a dying device Breno Leitao
2026-08-18 10:06 ` [PATCH v3 2/2] mm: fail the fault on a malformed swap entry instead of retrying it Breno Leitao
@ 2026-08-30 0:42 ` Andrew Morton
2 siblings, 0 replies; 8+ messages in thread
From: Andrew Morton @ 2026-08-30 0:42 UTC (permalink / raw)
To: Breno Leitao
Cc: Chris Li, Kairui Song, Kemeng Shi, Nhat Pham, Baoquan He,
Barry Song, Youngjun Park, David Hildenbrand, Lorenzo Stoakes,
Liam R. Howlett, Vlastimil Babka, Mike Rapoport,
Suren Baghdasaryan, Michal Hocko, Jann Horn, Pedro Falcato,
Hugh Dickins, Baolin Wang, Peter Xu, Johannes Weiner,
Yosry Ahmed, Chengming Zhou, linux-mm, linux-kernel, kernel-team
On Tue, 18 Aug 2026 03:06:22 -0700 Breno Leitao <leitao@debian.org> wrote:
> I've seen some machines at Meta fleet that show the following type of
> problem:
>
> 1) It gets some weird warning:
>
> BUG: Bad page map in process khugepaged pte:f000eef300000017 pmd:00000067
> addr:00007f57c0a01000 vm_flags:20200073 anon_vma:ffff88829af7c340 mapping:0000000000000000 index:7f57c0a01
>
> The corruption is most likely the collapse/PT_RECLAIM race fixed by
> commit 366a4532d96f ("mm: fix the race between collapse and PT_RECLAIM
> under per-vma lock"). But this series is not about this one.
>
> 2) Then the fault never makes progress. do_swap_page() returns 0 when
> get_swap_device() fails, so the fault is retried, reads the same
> entry and faults again. Nothing in the round trip changes the PTE,
> and the same line comes out on every pass:
>
> get_swap_device: Bad swap offset entry 3ffffffc043c5
>
> Patch 1 makes get_swap_device() return ERR_PTR(-EIO) for a malformed
> entry, keeping NULL for a device swapoff is taking away, and converts
> the callers. No functional change expected.
>
> Patch 2 uses that to return VM_FAULT_SIGBUS instead of retrying.
>
> The rate limiting patch that used to open this series was split out and
> posted on its own as a backportable hotfix [1], per Andrew's request. It
> should land first: patch 1 here touches the lines next to it in
> get_swap_device(). This patch will probably conflict with [1], but the
> merge should be trivial, given the only change in [1] is the
> addition of the __ratelimited() suffix.
>
> pr_err("%s: %s%08lx\n", __func__, Bad_file, entry.val);
> pr_err_ratelimited("%s: %s%08lx\n", __func__, Bad_file, entry.val);
>
Thanks. The patches have bitrotted a little - get_swap_device() was an
easy fixup but please double-check that I didn't miss anything.
Or perhaps just refresh-retest-resend if there's any doubt.
> The rate limiting patch that used to open this series was split out and
> posted on its own as a backportable hotfix [1], per Andrew's request. It
> should land first: patch 1 here touches the lines next to it in
> get_swap_device(). This patch will probably conflict with [1], but the
> merge should be trivial, given the only change in [1] is the
> addition of the __ratelimited() suffix.
>
> pr_err("%s: %s%08lx\n", __func__, Bad_file, entry.val);
> pr_err_ratelimited("%s: %s%08lx\n", __func__, Bad_file, entry.val);
oh, you already said that.
As usual when inspecting our error-path code, Sashiko said "you all suck":
https://sashiko.dev/#/patchset/20260818-swap-v3-0-d3fa52598a59@debian.org
These things do seem on-topic for the changes you're proposing here, so
please take a look?
^ permalink raw reply [flat|nested] 8+ messages in thread