mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Kaitao Cheng <kaitao.cheng@linux.dev>
To: "David Hildenbrand (Arm)" <david@kernel.org>,
	Yuanhe Shu <xiangzao@linux.alibaba.com>,
	akpm@linux-foundation.org
Cc: vbabka@kernel.org, linmiaohe@huawei.com, nao.horiguchi@gmail.com,
	ziy@nvidia.com, ying.huang@linux.alibaba.com,
	wangkefeng.wang@huawei.com, tujinjiang@huawei.com,
	mgorman@techsingularity.net, chengkaitao@kylinos.cn,
	muchun.song@linux.dev, linux-mm@kvack.org,
	linux-kernel@vger.kernel.org
Subject: Re: [PATCH 1/2] mm/compaction: skip folios containing hwpoisoned pages
Date: Sat, 10 Oct 2026 16:35:53 +0800	[thread overview]
Message-ID: <331b4e40-8d0c-4333-b2b6-b81e4c4b2e17@linux.dev> (raw)
In-Reply-To: <99cbd960-439d-4672-aba9-6b49697db8b3@kernel.org>



在 2026/9/28 20:12, David Hildenbrand (Arm) 写道:
> On 9/28/26 12:58, Yuanhe Shu wrote:
>> memory_failure() sets PG_hwpoison on the error page before it manages
>> to pin the folio via get_hwpoison_page().  Compaction running on
>> another CPU can isolate that same folio in the meantime, so
>> HWPoisonHandlable() then fails on !PageLRU(), the retry loop in
>> get_any_page() exhausts itself and memory_failure() gives up with
>> MF_IGNORED - while compaction still has the folio on its
>> migratepages list, flagged PG_hwpoison.
>>
>> Nothing stops compaction from migrating such a folio afterwards.
>> The corrupted data is copied into a fresh folio that carries no
>> poison marker, and no hwpoison PTE entry is installed on the way
>> back: the userspace mapping is silently redirected to the corrupted
>> copy and no SIGBUS is ever delivered.
>>
>> Note that a large folio with a PG_hwpoisoned subpage can also sit on
>> the LRU without any race at all, via the THP-split-failure path of
>> memory_failure() (kill_procs_now() + MF_MSG_UNSPLIT_THP).  Guarding
>> compaction is therefore needed regardless of how the race in
>> memory_failure() itself might be addressed.
>>
>> folio_mc_copy() does not close this hole: it only detects corruption
>> at copy time and only where ARCH_HAS_COPY_MC is implemented (x86_64
>> and PPC64; elsewhere copy_mc_highpage() degrades to a plain copy).
> 
> Then they should implement it.
> 
>> Software-injected poison - what MADV_HWPOISON and the hwpoison-inject
>> interface produce, and what tests and fuzzers exercise - never traps
>> during the copy on any architecture.  
> 
> And these are debug interfaces, why do we care?
> 
>> An isolation-time check avoids
>> the migration entirely, on all architectures and for both real and
>> simulated poison.
> 
> It's racy. See the link below.

Could we recheck PG_hwpoison on the source page after folio_mc_copy()
completes, but before the mapping is transferred, and abort migration
if the flag is set? This would prevent PG_hwpoison from being lost
after migration.

Completely closing the race window by modifying the arch-specific
copy_mc_to_kernel() implementation would presumably also require
corresponding hardware and architectural capabilities, which does
not seem very practical.

>> The check itself is two flag tests on the head page (PG_hwpoison
>> and the folio-level has_hwpoisoned marker), not a walk over
>> subpages.
>>
>> Neither page reclaim nor memory hotplug lets such a folio be
>> copied into a fresh one: reclaim unmaps the poisoned order-0 page,
>> and skips the hwpoisoned large folios it cannot safely unmap, in
>> shrink_folio_list()
>> commit 1b0449544c64 ("mm/vmscan: don't try to reclaim hwpoison folio")
>> commit 9f1e8cd0b7c4 ("mm/vmscan: fix hwpoisoned large folio handling in shrink_folio_list")
>> while memory hotplug refuses to migrate them in do_migrate_range()
>> commit 5f5ee52d4f58 ("mm/hwpoison: introduce folio_contain_hwpoisoned_page() helper").
>> Skip them in compaction as well, for the same reason reclaim settled
>> on skipping: the UCE is rare and a race with compaction is rarer
>> still, so skipping is enough, and a later memory_failure() will
>> handle the folio if the UCE is triggered again - while a migrated
>> folio would silently propagate the corruption instead.
>>
>> Deterministic validation on v7.3-rc4-75-g62f4c998b297: order-4
>> mTHP folios were brought into the stable "large, on LRU,
>> PG_hwpoisoned subpage" state by pinning a sibling subpage with
>> vmsplice and injecting MADV_HWPOISON on another subpage, which
>> makes memory_failure() take the THP-split-failure path; after
>> triggering compaction, /proc/kpageflags shows which poisoned
>> folios moved.  The unpatched kernel migrated 16/16 poisoned
>> folios and 16/16 clean controls in the same pageblocks; with
>> this patch, 0/16 poisoned folios were migrated while their
>> controls still migrated 16/16.
> 
> 
> Was any of this written by an LLM?
> 
>>
>> Soft offline is not affected: it only sets the poison marker after
>> its own migration has succeeded.
>>
>> Signed-off-by: Yuanhe Shu <xiangzao@linux.alibaba.com>
>> ---
>>  mm/compaction.c | 8 ++++++++
>>  1 file changed, 8 insertions(+)
>>
>> diff --git a/mm/compaction.c b/mm/compaction.c
>> index a049415512c6..491adcbc5313 100644
>> --- a/mm/compaction.c
>> +++ b/mm/compaction.c
>> @@ -1093,6 +1093,14 @@ isolate_migratepages_block(struct compact_control *cc, unsigned long low_pfn,
>>  		if (unlikely(!folio))
>>  			goto isolate_fail;
>>  
>> +		/*
>> +		 * Migrating would copy the corrupted data into a fresh
>> +		 * folio with no poison marker; skip, as memory hotplug
>> +		 * refuses to migrate such folios for the same reason.
>> +		 */
>> +		if (folio_contain_hwpoisoned_page(folio))
>> +			goto isolate_fail_put;
>> +
>>  		/*
>>  		 * Migration will fail if an anonymous page is pinned in memory,
>>  		 * so avoid taking lru_lock and isolating it unnecessarily in an
> 
> See
> 
> https://lore.kernel.org/linux-mm/20260707090136.52904-1-kaitao.cheng@linux.dev/
> 

-- 
Thanks
Kaitao Cheng


  parent reply	other threads:[~2026-10-10  8:36 UTC|newest]

Thread overview: 6+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-28 10:58 [PATCH 0/2] mm: don't migrate " Yuanhe Shu
2026-09-28 10:58 ` [PATCH 1/2] mm/compaction: skip " Yuanhe Shu
2026-09-28 12:12   ` David Hildenbrand (Arm)
2026-09-29  3:41     ` Yuanhe Shu
2026-10-10  8:35     ` Kaitao Cheng [this message]
2026-09-28 10:58 ` [PATCH 2/2] mm/migrate: refuse to migrate " Yuanhe Shu

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=331b4e40-8d0c-4333-b2b6-b81e4c4b2e17@linux.dev \
    --to=kaitao.cheng@linux.dev \
    --cc=akpm@linux-foundation.org \
    --cc=chengkaitao@kylinos.cn \
    --cc=david@kernel.org \
    --cc=linmiaohe@huawei.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=mgorman@techsingularity.net \
    --cc=muchun.song@linux.dev \
    --cc=nao.horiguchi@gmail.com \
    --cc=tujinjiang@huawei.com \
    --cc=vbabka@kernel.org \
    --cc=wangkefeng.wang@huawei.com \
    --cc=xiangzao@linux.alibaba.com \
    --cc=ying.huang@linux.alibaba.com \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®