* [PATCH 0/2] mm: don't migrate folios containing hwpoisoned pages
@ 2026-09-28 10:58 Yuanhe Shu
2026-09-28 10:58 ` [PATCH 1/2] mm/compaction: skip " Yuanhe Shu
2026-09-28 10:58 ` [PATCH 2/2] mm/migrate: refuse to migrate " Yuanhe Shu
0 siblings, 2 replies; 4+ messages in thread
From: Yuanhe Shu @ 2026-09-28 10:58 UTC (permalink / raw)
To: akpm
Cc: xiangzao, david, vbabka, linmiaohe, nao.horiguchi, ziy,
ying.huang, wangkefeng.wang, tujinjiang, mgorman, kaitao.cheng,
chengkaitao, muchun.song, linux-mm, linux-kernel
Hello,
memory_failure() can end up abandoning a folio while PG_hwpoison is
already set on it. Two ways to get there:
1. Race: any isolator (compaction, NUMA balancing, move_pages(),
mbind(), CMA, long-term GUP) clears PG_lru from under
memory_failure(), which then fails the HWPoisonHandlable()
check, exhausts the get_any_page() retries and bails out with
MF_IGNORED - without holding any reference, but with the flag
set.
2. No race at all: the THP-split-failure path
(kill_procs_now() + MF_MSG_UNSPLIT_THP) leaves a large folio
on the LRU with a PG_hwpoisoned subpage by design.
In both cases the migrator still holds the folio on its migration
list and migrates it like any other: the corrupted data is copied
into a fresh folio with no poison marker, and no hwpoison PTE entry
is installed on the way back, so the userspace mapping is silently
redirected to the corrupted copy - no SIGBUS is ever delivered.
This was reproduced on v7.3-rc4-75-g62f4c998b297 (the base of
this series): an order-0 stress workload (madvise(MADV_HWPOISON)
racing move_pages()) left the poison flag set on 8426 abandoned
folios in a 180-second run; 698 of them were subsequently
migrated, and in every case the address then served the copied
corrupted data with no SIGBUS.
folio_mc_copy() only catches the corruption at copy time, and only
where ARCH_HAS_COPY_MC is implemented (x86_64, PPC64; elsewhere
copy_mc_highpage() degrades to a plain copy). Software-injected
poison - what MADV_HWPOISON and hwpoison-inject produce, and what
tests and fuzzers exercise - never traps during the copy on any
architecture. Checking before the copy avoids the migration
entirely, on all architectures, for real and simulated poison
alike.
Consumers of such a folio already refuse it elsewhere,
through the folio_contain_hwpoisoned_page() helper introduced by
commit 5f5ee52d4f58 ("mm/hwpoison: introduce folio_contain_hwpoisoned_page() helper"):
shmem fails writes to a poisoned folio; reclaim unmaps the
order-0 poisoned page and skips the hwpoisoned large folios it
cannot safely unmap
commit 1b0449544c64 ("mm/vmscan: don't try to reclaim hwpoison folio")
commit 9f1e8cd0b7c4 ("mm/vmscan: fix hwpoisoned large folio handling in shrink_folio_list")
(skipping, not unmapping, after the unmap variant turned out to
crash on large folios); memory hotplug refuses to migrate it in
do_migrate_range(); and the THP split path refuses to read
poisoned subpages
commit 841a8bfcbad9 ("mm: prevent poison consumption when splitting THP").
This series gives the remaining migrators
the same protection - skipping and refusing only, never unmapping:
patch 1: compaction, checked at the isolation point, failing
fast - patch 2's core check would also stop compaction's
migrations at migrate_pages() time, but only after the
avoidable isolation work
patch 2: the migration core, covering move_pages(2), mbind(),
NUMA balancing, alloc_contig_range() (CMA) and long-term
GUP - a check at the entry of migrate_folio_unmap() (fail
fast, before allocating the destination) plus a second
check in migrate_folio_move() immediately before the
copy. The pre-copy check catches poison that lands
between the two checks: a GUP-based injection path
takes its reference before the unmap, and
memory_failure() sets PG_hwpoison before releasing it,
so a racing software poison is visible by the time
migration is about to copy.
This series was developed independently; while preparing to
send it, we noticed that Kaitao Cheng had already proposed the
same check for the hugetlb migration path, and in v2 for
move_to_new_folio() [1]. That discussion stalled over the
concern that a check can still race with memory_failure().
Two points complete it here: the unsplit-THP state above is not
a race at all - the entry check catches it deterministically -
and for the race itself, the pre-copy check shrinks the window
to the final instructions before the copy; the measurement
below found zero leaks under adversarial load, and on copy_mc
architectures folio_mc_copy() catches real-DRAM corruption even
there. The reproducer and the userspace-visible effects are
what that thread's review asked for.
[1] https://lore.kernel.org/r/20260707090136.52904-1-kaitao.cheng@linux.dev/
Measured results:
- order-0 race against move_pages(): the unpatched kernel
propagated 698 of 8426 abandoned folios; with patch 2, 0 of
9711 (the abandoned-folio count varies between runs with the
race rate).
- deterministic compaction A/B (order-4 mTHP folios brought into
the stable unsplit state, clean neighbors as controls): the
unpatched kernel migrated 16/16 poisoned folios and 16/16
controls; with patch 1, 0/16 poisoned folios were migrated
while the controls still migrated 16/16.
The two patches touch different files, have no build dependency,
and each stands on its own.
No single Fixes: tag applies: the race path is as old as the
memory_failure() retry logic meeting page isolation, and the
unsplit-THP path has existed since commit
5d1fd5dc877b ("mm,hwpoison: introduce MF_MSG_UNSPLIT_THP")
(2020) - this series closes a long-standing gap rather than fixing
a regression. No cc:stable is proposed for the same reason;
backports can be prepared on request if wanted.
Soft offline is unaffected by both patches: it sets the poison
marker only after its own migration has succeeded. Out of reach
of any check is poison landing in the final instructions before a
copy - natively for paths that take no reference, such as a
hardware MCE or debugfs corrupt-pfn; for a GUP-based path only if
its injecting thread is delayed between taking the reference and
setting the flag - or after the copy has already happened, which
is a lost race regardless; on x86_64 folio_mc_copy()
additionally catches real-DRAM corruption at copy time.
A deterministic reproducer exists for the compaction part: fault
order-4 mTHP folios, pin a sibling subpage of each with vmsplice,
and inject MADV_HWPOISON on another subpage, so memory_failure()
takes the THP-split-failure path and leaves the folio on the LRU
with a poisoned subpage; then trigger compaction and check
/proc/kpageflags for whether each poisoned folio moved. It needs
a boot configuration where order-4, not PMD-sized, folios are
what compaction migrates. Both this and the race workload above
can be shared on request.
Yuanhe Shu (2):
mm/compaction: skip folios containing hwpoisoned pages
mm/migrate: refuse to migrate folios containing hwpoisoned pages
mm/compaction.c | 8 ++++++++
mm/migrate.c | 41 +++++++++++++++++++++++++++++++++++++++--
2 files changed, 47 insertions(+), 2 deletions(-)
--
2.43.7
^ permalink raw reply [flat|nested] 4+ messages in thread
* [PATCH 1/2] mm/compaction: skip folios containing hwpoisoned pages
2026-09-28 10:58 [PATCH 0/2] mm: don't migrate folios containing hwpoisoned pages Yuanhe Shu
@ 2026-09-28 10:58 ` Yuanhe Shu
2026-09-28 12:12 ` David Hildenbrand (Arm)
2026-09-28 10:58 ` [PATCH 2/2] mm/migrate: refuse to migrate " Yuanhe Shu
1 sibling, 1 reply; 4+ messages in thread
From: Yuanhe Shu @ 2026-09-28 10:58 UTC (permalink / raw)
To: akpm
Cc: xiangzao, david, vbabka, linmiaohe, nao.horiguchi, ziy,
ying.huang, wangkefeng.wang, tujinjiang, mgorman, kaitao.cheng,
chengkaitao, muchun.song, linux-mm, linux-kernel
memory_failure() sets PG_hwpoison on the error page before it manages
to pin the folio via get_hwpoison_page(). Compaction running on
another CPU can isolate that same folio in the meantime, so
HWPoisonHandlable() then fails on !PageLRU(), the retry loop in
get_any_page() exhausts itself and memory_failure() gives up with
MF_IGNORED - while compaction still has the folio on its
migratepages list, flagged PG_hwpoison.
Nothing stops compaction from migrating such a folio afterwards.
The corrupted data is copied into a fresh folio that carries no
poison marker, and no hwpoison PTE entry is installed on the way
back: the userspace mapping is silently redirected to the corrupted
copy and no SIGBUS is ever delivered.
Note that a large folio with a PG_hwpoisoned subpage can also sit on
the LRU without any race at all, via the THP-split-failure path of
memory_failure() (kill_procs_now() + MF_MSG_UNSPLIT_THP). Guarding
compaction is therefore needed regardless of how the race in
memory_failure() itself might be addressed.
folio_mc_copy() does not close this hole: it only detects corruption
at copy time and only where ARCH_HAS_COPY_MC is implemented (x86_64
and PPC64; elsewhere copy_mc_highpage() degrades to a plain copy).
Software-injected poison - what MADV_HWPOISON and the hwpoison-inject
interface produce, and what tests and fuzzers exercise - never traps
during the copy on any architecture. An isolation-time check avoids
the migration entirely, on all architectures and for both real and
simulated poison.
The check itself is two flag tests on the head page (PG_hwpoison
and the folio-level has_hwpoisoned marker), not a walk over
subpages.
Neither page reclaim nor memory hotplug lets such a folio be
copied into a fresh one: reclaim unmaps the poisoned order-0 page,
and skips the hwpoisoned large folios it cannot safely unmap, in
shrink_folio_list()
commit 1b0449544c64 ("mm/vmscan: don't try to reclaim hwpoison folio")
commit 9f1e8cd0b7c4 ("mm/vmscan: fix hwpoisoned large folio handling in shrink_folio_list")
while memory hotplug refuses to migrate them in do_migrate_range()
commit 5f5ee52d4f58 ("mm/hwpoison: introduce folio_contain_hwpoisoned_page() helper").
Skip them in compaction as well, for the same reason reclaim settled
on skipping: the UCE is rare and a race with compaction is rarer
still, so skipping is enough, and a later memory_failure() will
handle the folio if the UCE is triggered again - while a migrated
folio would silently propagate the corruption instead.
Deterministic validation on v7.3-rc4-75-g62f4c998b297: order-4
mTHP folios were brought into the stable "large, on LRU,
PG_hwpoisoned subpage" state by pinning a sibling subpage with
vmsplice and injecting MADV_HWPOISON on another subpage, which
makes memory_failure() take the THP-split-failure path; after
triggering compaction, /proc/kpageflags shows which poisoned
folios moved. The unpatched kernel migrated 16/16 poisoned
folios and 16/16 clean controls in the same pageblocks; with
this patch, 0/16 poisoned folios were migrated while their
controls still migrated 16/16.
Soft offline is not affected: it only sets the poison marker after
its own migration has succeeded.
Signed-off-by: Yuanhe Shu <xiangzao@linux.alibaba.com>
---
mm/compaction.c | 8 ++++++++
1 file changed, 8 insertions(+)
diff --git a/mm/compaction.c b/mm/compaction.c
index a049415512c6..491adcbc5313 100644
--- a/mm/compaction.c
+++ b/mm/compaction.c
@@ -1093,6 +1093,14 @@ isolate_migratepages_block(struct compact_control *cc, unsigned long low_pfn,
if (unlikely(!folio))
goto isolate_fail;
+ /*
+ * Migrating would copy the corrupted data into a fresh
+ * folio with no poison marker; skip, as memory hotplug
+ * refuses to migrate such folios for the same reason.
+ */
+ if (folio_contain_hwpoisoned_page(folio))
+ goto isolate_fail_put;
+
/*
* Migration will fail if an anonymous page is pinned in memory,
* so avoid taking lru_lock and isolating it unnecessarily in an
--
2.43.7
^ permalink raw reply [flat|nested] 4+ messages in thread
* [PATCH 2/2] mm/migrate: refuse to migrate folios containing hwpoisoned pages
2026-09-28 10:58 [PATCH 0/2] mm: don't migrate folios containing hwpoisoned pages Yuanhe Shu
2026-09-28 10:58 ` [PATCH 1/2] mm/compaction: skip " Yuanhe Shu
@ 2026-09-28 10:58 ` Yuanhe Shu
1 sibling, 0 replies; 4+ messages in thread
From: Yuanhe Shu @ 2026-09-28 10:58 UTC (permalink / raw)
To: akpm
Cc: xiangzao, david, vbabka, linmiaohe, nao.horiguchi, ziy,
ying.huang, wangkefeng.wang, tujinjiang, mgorman, kaitao.cheng,
chengkaitao, muchun.song, linux-mm, linux-kernel
All callers of migrate_pages() without their own poison handling -
move_pages(2), mbind(), NUMA balancing, alloc_contig_range() (CMA)
and long-term GUP - can race memory_failure() the same way compaction
does (see patch 1): the migration isolates a folio and clears
PG_lru, memory_failure() sets PG_hwpoison, then fails to pin it and
gives up with MF_IGNORED, and the pending migration copies the
corrupted data into a fresh folio that carries no poison marker;
the userspace mapping ends up silently serving the corrupted copy
with no SIGBUS.
This was reproduced on v7.3-rc4-75-g62f4c998b297: an order-0 stress
workload running madvise(MADV_HWPOISON) concurrently with
move_pages() left the poison flag set on 8426 abandoned folios in
a 180-second run, identified via /proc/kpageflags; 698 of them
were subsequently migrated, and in every case the address then
served the copied corrupted data with no SIGBUS.
Refuse such migrations in the migration core, with two checks:
- migrate_folio_unmap() checks at entry, so a folio that is already
poisoned when the batch reaches it fails immediately, without
allocating a destination folio or unmapping it.
- migrate_folio_move() checks again immediately before the copy,
because migrate_pages_batch() unmaps a whole chunk of folios
before copying any of them, so a poison can land in between. A
GUP-based path such as madvise(MADV_HWPOISON) takes its
reference while the folio is still mapped, i.e. before the
unmap, and memory_failure() sets PG_hwpoison before releasing
that reference, so a racing poison is visible by the time
migration reaches the copy, unless the injecting thread is
delayed between taking the reference and setting the flag. A
hardware MCE can likewise land in the final instructions
between the check and the copy; on x86_64 folio_mc_copy()
additionally catches real-DRAM corruption at copy time.
Both checks fail with -EHWPOISON rather than -EAGAIN: -EAGAIN would
leave the folio on the migration list and retry into the same check,
while -EHWPOISON is a permanent failure, so the folio is put back on
the LRU with the poison flag intact, for reclaim or a later
memory_failure() to deal with - the same reasoning reclaim settled
on in
commit 9f1e8cd0b7c4 ("mm/vmscan: fix hwpoisoned large folio handling in shrink_folio_list").
The hugetlb path gets the same two checks for symmetry and defense
in depth: its unmap-to-copy window is single-folio and much narrower
than the batched path, but the checks cost two bit tests.
Measured with the workload above (180-second runs; the count of
abandoned folios varies between runs with the race rate): the
unpatched kernel propagated 698 of 8426 abandoned folios; with
this patch, 0 of 9711 were propagated.
Soft offline is unaffected: it sets the poison marker only after its
own migration has succeeded. Memory hotplug offlining is
unaffected: it filters poisoned folios out before calling
migrate_pages(). Reclaim is unaffected: poisoned folios go through
its own unmap and skip paths, not migrate_pages().
A check at the copy site was previously proposed by Kaitao
Cheng, for the hugetlb path and for move_to_new_folio(); that
discussion stalled over the remaining race with memory_failure().
The GUP-reference argument and the measurement above answer
that concern for injection paths.
Signed-off-by: Yuanhe Shu <xiangzao@linux.alibaba.com>
Link: https://lore.kernel.org/r/20260707090136.52904-1-kaitao.cheng@linux.dev/
---
mm/migrate.c | 41 +++++++++++++++++++++++++++++++++++++++--
1 file changed, 39 insertions(+), 2 deletions(-)
diff --git a/mm/migrate.c b/mm/migrate.c
index 15b45832bcfa..bdfd490dd8e8 100644
--- a/mm/migrate.c
+++ b/mm/migrate.c
@@ -1222,6 +1222,18 @@ static int migrate_folio_unmap(new_folio_t get_new_folio,
bool locked = false;
bool dst_locked = false;
+ if (unlikely(folio_contain_hwpoisoned_page(src))) {
+ /*
+ * The copy would propagate the corrupted data into a
+ * fresh folio with no poison marker. Fail permanently
+ * so the folio is put back on the LRU for reclaim or a
+ * later memory_failure(); it was not unmapped yet, so
+ * there is nothing to restore.
+ */
+ migrate_folio_undo_src(src, 0, NULL, false, ret);
+ return -EHWPOISON;
+ }
+
dst = get_new_folio(src, private);
if (!dst)
return -ENOMEM;
@@ -1376,6 +1388,20 @@ static int migrate_folio_move(free_folio_t put_new_folio, unsigned long private,
prev = dst->lru.prev;
list_del(&dst->lru);
+ if (unlikely(folio_contain_hwpoisoned_page(src))) {
+ /*
+ * Recheck before the copy: a poison can land while the
+ * batch is still unmapping the rest of its chunk. A
+ * GUP-based path takes its reference before the unmap,
+ * and memory_failure() sets PG_hwpoison before
+ * releasing it, so a racing poison is visible by now;
+ * a hardware MCE can still land in the final
+ * instructions.
+ */
+ rc = -EHWPOISON;
+ goto out;
+ }
+
if (unlikely(page_has_movable_ops(&src->page))) {
rc = migrate_movable_ops_page(&dst->page, &src->page, mode);
if (rc)
@@ -1521,6 +1547,12 @@ static int unmap_and_move_hugetlb_folio(new_folio_t get_new_folio,
goto out_unlock;
}
+ if (unlikely(folio_contain_hwpoisoned_page(src))) {
+ /* Same reasoning as in migrate_folio_unmap(). */
+ rc = -EHWPOISON;
+ goto out_unlock;
+ }
+
if (folio_test_anon(src))
anon_vma = folio_get_anon_vma(src);
@@ -1546,8 +1578,13 @@ static int unmap_and_move_hugetlb_folio(new_folio_t get_new_folio,
was_mapped = 1;
}
- if (!folio_mapped(src))
- rc = move_to_new_folio(dst, src, mode);
+ if (!folio_mapped(src)) {
+ /* Same reasoning as in migrate_folio_move(). */
+ if (unlikely(folio_contain_hwpoisoned_page(src)))
+ rc = -EHWPOISON;
+ else
+ rc = move_to_new_folio(dst, src, mode);
+ }
if (was_mapped)
remove_migration_ptes(src, !rc ? dst : src, ttu);
--
2.43.7
^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [PATCH 1/2] mm/compaction: skip folios containing hwpoisoned pages
2026-09-28 10:58 ` [PATCH 1/2] mm/compaction: skip " Yuanhe Shu
@ 2026-09-28 12:12 ` David Hildenbrand (Arm)
0 siblings, 0 replies; 4+ messages in thread
From: David Hildenbrand (Arm) @ 2026-09-28 12:12 UTC (permalink / raw)
To: Yuanhe Shu, akpm
Cc: vbabka, linmiaohe, nao.horiguchi, ziy, ying.huang,
wangkefeng.wang, tujinjiang, mgorman, kaitao.cheng, chengkaitao,
muchun.song, linux-mm, linux-kernel
On 9/28/26 12:58, Yuanhe Shu wrote:
> memory_failure() sets PG_hwpoison on the error page before it manages
> to pin the folio via get_hwpoison_page(). Compaction running on
> another CPU can isolate that same folio in the meantime, so
> HWPoisonHandlable() then fails on !PageLRU(), the retry loop in
> get_any_page() exhausts itself and memory_failure() gives up with
> MF_IGNORED - while compaction still has the folio on its
> migratepages list, flagged PG_hwpoison.
>
> Nothing stops compaction from migrating such a folio afterwards.
> The corrupted data is copied into a fresh folio that carries no
> poison marker, and no hwpoison PTE entry is installed on the way
> back: the userspace mapping is silently redirected to the corrupted
> copy and no SIGBUS is ever delivered.
>
> Note that a large folio with a PG_hwpoisoned subpage can also sit on
> the LRU without any race at all, via the THP-split-failure path of
> memory_failure() (kill_procs_now() + MF_MSG_UNSPLIT_THP). Guarding
> compaction is therefore needed regardless of how the race in
> memory_failure() itself might be addressed.
>
> folio_mc_copy() does not close this hole: it only detects corruption
> at copy time and only where ARCH_HAS_COPY_MC is implemented (x86_64
> and PPC64; elsewhere copy_mc_highpage() degrades to a plain copy).
Then they should implement it.
> Software-injected poison - what MADV_HWPOISON and the hwpoison-inject
> interface produce, and what tests and fuzzers exercise - never traps
> during the copy on any architecture.
And these are debug interfaces, why do we care?
> An isolation-time check avoids
> the migration entirely, on all architectures and for both real and
> simulated poison.
It's racy. See the link below.
> The check itself is two flag tests on the head page (PG_hwpoison
> and the folio-level has_hwpoisoned marker), not a walk over
> subpages.
>
> Neither page reclaim nor memory hotplug lets such a folio be
> copied into a fresh one: reclaim unmaps the poisoned order-0 page,
> and skips the hwpoisoned large folios it cannot safely unmap, in
> shrink_folio_list()
> commit 1b0449544c64 ("mm/vmscan: don't try to reclaim hwpoison folio")
> commit 9f1e8cd0b7c4 ("mm/vmscan: fix hwpoisoned large folio handling in shrink_folio_list")
> while memory hotplug refuses to migrate them in do_migrate_range()
> commit 5f5ee52d4f58 ("mm/hwpoison: introduce folio_contain_hwpoisoned_page() helper").
> Skip them in compaction as well, for the same reason reclaim settled
> on skipping: the UCE is rare and a race with compaction is rarer
> still, so skipping is enough, and a later memory_failure() will
> handle the folio if the UCE is triggered again - while a migrated
> folio would silently propagate the corruption instead.
>
> Deterministic validation on v7.3-rc4-75-g62f4c998b297: order-4
> mTHP folios were brought into the stable "large, on LRU,
> PG_hwpoisoned subpage" state by pinning a sibling subpage with
> vmsplice and injecting MADV_HWPOISON on another subpage, which
> makes memory_failure() take the THP-split-failure path; after
> triggering compaction, /proc/kpageflags shows which poisoned
> folios moved. The unpatched kernel migrated 16/16 poisoned
> folios and 16/16 clean controls in the same pageblocks; with
> this patch, 0/16 poisoned folios were migrated while their
> controls still migrated 16/16.
Was any of this written by an LLM?
>
> Soft offline is not affected: it only sets the poison marker after
> its own migration has succeeded.
>
> Signed-off-by: Yuanhe Shu <xiangzao@linux.alibaba.com>
> ---
> mm/compaction.c | 8 ++++++++
> 1 file changed, 8 insertions(+)
>
> diff --git a/mm/compaction.c b/mm/compaction.c
> index a049415512c6..491adcbc5313 100644
> --- a/mm/compaction.c
> +++ b/mm/compaction.c
> @@ -1093,6 +1093,14 @@ isolate_migratepages_block(struct compact_control *cc, unsigned long low_pfn,
> if (unlikely(!folio))
> goto isolate_fail;
>
> + /*
> + * Migrating would copy the corrupted data into a fresh
> + * folio with no poison marker; skip, as memory hotplug
> + * refuses to migrate such folios for the same reason.
> + */
> + if (folio_contain_hwpoisoned_page(folio))
> + goto isolate_fail_put;
> +
> /*
> * Migration will fail if an anonymous page is pinned in memory,
> * so avoid taking lru_lock and isolating it unnecessarily in an
See
https://lore.kernel.org/linux-mm/20260707090136.52904-1-kaitao.cheng@linux.dev/
--
Cheers,
David
^ permalink raw reply [flat|nested] 4+ messages in thread
end of thread, other threads:[~2026-09-28 12:12 UTC | newest]
Thread overview: 4+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-28 10:58 [PATCH 0/2] mm: don't migrate folios containing hwpoisoned pages Yuanhe Shu
2026-09-28 10:58 ` [PATCH 1/2] mm/compaction: skip " Yuanhe Shu
2026-09-28 12:12 ` David Hildenbrand (Arm)
2026-09-28 10:58 ` [PATCH 2/2] mm/migrate: refuse to migrate " Yuanhe Shu
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®