mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Yuanhe Shu <xiangzao@linux.alibaba.com>
To: akpm@linux-foundation.org
Cc: xiangzao@linux.alibaba.com, david@kernel.org, vbabka@kernel.org,
	linmiaohe@huawei.com, nao.horiguchi@gmail.com, ziy@nvidia.com,
	ying.huang@linux.alibaba.com, wangkefeng.wang@huawei.com,
	tujinjiang@huawei.com, mgorman@techsingularity.net,
	kaitao.cheng@linux.dev, chengkaitao@kylinos.cn,
	muchun.song@linux.dev, linux-mm@kvack.org,
	linux-kernel@vger.kernel.org
Subject: [PATCH 2/2] mm/migrate: refuse to migrate folios containing hwpoisoned pages
Date: Mon, 28 Sep 2026 18:58:05 +0800	[thread overview]
Message-ID: <20260928105805.1215770-3-xiangzao@linux.alibaba.com> (raw)
In-Reply-To: <20260928105805.1215770-1-xiangzao@linux.alibaba.com>

All callers of migrate_pages() without their own poison handling -
move_pages(2), mbind(), NUMA balancing, alloc_contig_range() (CMA)
and long-term GUP - can race memory_failure() the same way compaction
does (see patch 1): the migration isolates a folio and clears
PG_lru, memory_failure() sets PG_hwpoison, then fails to pin it and
gives up with MF_IGNORED, and the pending migration copies the
corrupted data into a fresh folio that carries no poison marker;
the userspace mapping ends up silently serving the corrupted copy
with no SIGBUS.

This was reproduced on v7.3-rc4-75-g62f4c998b297: an order-0 stress
workload running madvise(MADV_HWPOISON) concurrently with
move_pages() left the poison flag set on 8426 abandoned folios in
a 180-second run, identified via /proc/kpageflags; 698 of them
were subsequently migrated, and in every case the address then
served the copied corrupted data with no SIGBUS.

Refuse such migrations in the migration core, with two checks:

- migrate_folio_unmap() checks at entry, so a folio that is already
  poisoned when the batch reaches it fails immediately, without
  allocating a destination folio or unmapping it.

- migrate_folio_move() checks again immediately before the copy,
  because migrate_pages_batch() unmaps a whole chunk of folios
  before copying any of them, so a poison can land in between.  A
  GUP-based path such as madvise(MADV_HWPOISON) takes its
  reference while the folio is still mapped, i.e. before the
  unmap, and memory_failure() sets PG_hwpoison before releasing
  that reference, so a racing poison is visible by the time
  migration reaches the copy, unless the injecting thread is
  delayed between taking the reference and setting the flag.  A
  hardware MCE can likewise land in the final instructions
  between the check and the copy; on x86_64 folio_mc_copy()
  additionally catches real-DRAM corruption at copy time.

Both checks fail with -EHWPOISON rather than -EAGAIN: -EAGAIN would
leave the folio on the migration list and retry into the same check,
while -EHWPOISON is a permanent failure, so the folio is put back on
the LRU with the poison flag intact, for reclaim or a later
memory_failure() to deal with - the same reasoning reclaim settled
on in
commit 9f1e8cd0b7c4 ("mm/vmscan: fix hwpoisoned large folio handling in shrink_folio_list").
The hugetlb path gets the same two checks for symmetry and defense
in depth: its unmap-to-copy window is single-folio and much narrower
than the batched path, but the checks cost two bit tests.

Measured with the workload above (180-second runs; the count of
abandoned folios varies between runs with the race rate): the
unpatched kernel propagated 698 of 8426 abandoned folios; with
this patch, 0 of 9711 were propagated.

Soft offline is unaffected: it sets the poison marker only after its
own migration has succeeded.  Memory hotplug offlining is
unaffected: it filters poisoned folios out before calling
migrate_pages().  Reclaim is unaffected: poisoned folios go through
its own unmap and skip paths, not migrate_pages().

A check at the copy site was previously proposed by Kaitao
Cheng, for the hugetlb path and for move_to_new_folio(); that
discussion stalled over the remaining race with memory_failure().
The GUP-reference argument and the measurement above answer
that concern for injection paths.

Signed-off-by: Yuanhe Shu <xiangzao@linux.alibaba.com>
Link: https://lore.kernel.org/r/20260707090136.52904-1-kaitao.cheng@linux.dev/
---
 mm/migrate.c | 41 +++++++++++++++++++++++++++++++++++++++--
 1 file changed, 39 insertions(+), 2 deletions(-)

diff --git a/mm/migrate.c b/mm/migrate.c
index 15b45832bcfa..bdfd490dd8e8 100644
--- a/mm/migrate.c
+++ b/mm/migrate.c
@@ -1222,6 +1222,18 @@ static int migrate_folio_unmap(new_folio_t get_new_folio,
 	bool locked = false;
 	bool dst_locked = false;
 
+	if (unlikely(folio_contain_hwpoisoned_page(src))) {
+		/*
+		 * The copy would propagate the corrupted data into a
+		 * fresh folio with no poison marker.  Fail permanently
+		 * so the folio is put back on the LRU for reclaim or a
+		 * later memory_failure(); it was not unmapped yet, so
+		 * there is nothing to restore.
+		 */
+		migrate_folio_undo_src(src, 0, NULL, false, ret);
+		return -EHWPOISON;
+	}
+
 	dst = get_new_folio(src, private);
 	if (!dst)
 		return -ENOMEM;
@@ -1376,6 +1388,20 @@ static int migrate_folio_move(free_folio_t put_new_folio, unsigned long private,
 	prev = dst->lru.prev;
 	list_del(&dst->lru);
 
+	if (unlikely(folio_contain_hwpoisoned_page(src))) {
+		/*
+		 * Recheck before the copy: a poison can land while the
+		 * batch is still unmapping the rest of its chunk.  A
+		 * GUP-based path takes its reference before the unmap,
+		 * and memory_failure() sets PG_hwpoison before
+		 * releasing it, so a racing poison is visible by now;
+		 * a hardware MCE can still land in the final
+		 * instructions.
+		 */
+		rc = -EHWPOISON;
+		goto out;
+	}
+
 	if (unlikely(page_has_movable_ops(&src->page))) {
 		rc = migrate_movable_ops_page(&dst->page, &src->page, mode);
 		if (rc)
@@ -1521,6 +1547,12 @@ static int unmap_and_move_hugetlb_folio(new_folio_t get_new_folio,
 		goto out_unlock;
 	}
 
+	if (unlikely(folio_contain_hwpoisoned_page(src))) {
+		/* Same reasoning as in migrate_folio_unmap(). */
+		rc = -EHWPOISON;
+		goto out_unlock;
+	}
+
 	if (folio_test_anon(src))
 		anon_vma = folio_get_anon_vma(src);
 
@@ -1546,8 +1578,13 @@ static int unmap_and_move_hugetlb_folio(new_folio_t get_new_folio,
 		was_mapped = 1;
 	}
 
-	if (!folio_mapped(src))
-		rc = move_to_new_folio(dst, src, mode);
+	if (!folio_mapped(src)) {
+		/* Same reasoning as in migrate_folio_move(). */
+		if (unlikely(folio_contain_hwpoisoned_page(src)))
+			rc = -EHWPOISON;
+		else
+			rc = move_to_new_folio(dst, src, mode);
+	}
 
 	if (was_mapped)
 		remove_migration_ptes(src, !rc ? dst : src, ttu);
-- 
2.43.7


      parent reply	other threads:[~2026-09-28 10:58 UTC|newest]

Thread overview: 5+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-28 10:58 [PATCH 0/2] mm: don't " Yuanhe Shu
2026-09-28 10:58 ` [PATCH 1/2] mm/compaction: skip " Yuanhe Shu
2026-09-28 12:12   ` David Hildenbrand (Arm)
2026-09-29  3:41     ` Yuanhe Shu
2026-09-28 10:58 ` Yuanhe Shu [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260928105805.1215770-3-xiangzao@linux.alibaba.com \
    --to=xiangzao@linux.alibaba.com \
    --cc=akpm@linux-foundation.org \
    --cc=chengkaitao@kylinos.cn \
    --cc=david@kernel.org \
    --cc=kaitao.cheng@linux.dev \
    --cc=linmiaohe@huawei.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=mgorman@techsingularity.net \
    --cc=muchun.song@linux.dev \
    --cc=nao.horiguchi@gmail.com \
    --cc=tujinjiang@huawei.com \
    --cc=vbabka@kernel.org \
    --cc=wangkefeng.wang@huawei.com \
    --cc=ying.huang@linux.alibaba.com \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®