mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Yuanhe Shu <xiangzao@linux.alibaba.com>
To: david@kernel.org, akpm@linux-foundation.org
Cc: xiangzao@linux.alibaba.com, vbabka@kernel.org,
	linmiaohe@huawei.com, nao.horiguchi@gmail.com, ziy@nvidia.com,
	ying.huang@linux.alibaba.com, wangkefeng.wang@huawei.com,
	tujinjiang@huawei.com, mgorman@techsingularity.net,
	kaitao.cheng@linux.dev, chengkaitao@kylinos.cn,
	muchun.song@linux.dev, linux-mm@kvack.org,
	linux-kernel@vger.kernel.org
Subject: Re: [PATCH 1/2] mm/compaction: skip folios containing hwpoisoned pages
Date: Tue, 29 Sep 2026 11:41:23 +0800	[thread overview]
Message-ID: <20260929034123.2705745-1-xiangzao@linux.alibaba.com> (raw)
In-Reply-To: <99cbd960-439d-4672-aba9-6b49697db8b3@kernel.org>

On 9/28/26 20:12, David Hildenbrand wrote:
> On 9/28/26 12:58, Yuanhe Shu wrote:
>> folio_mc_copy() does not close this hole: it only detects corruption
>> at copy time and only where ARCH_HAS_COPY_MC is implemented (x86_64
>> and PPC64; elsewhere copy_mc_highpage() degrades to a plain copy).
>
> Then they should implement it.

Agreed, and nothing here replaces copy_mc.  Two things the check adds
while that work is pending: on archs without copy_mc the plain copy
does not fail, it consumes the poison - the same poison consumption
commit 841a8bfcbad9 ("mm: prevent poison consumption when splitting
THP") refused to risk on the split path, at your suggestion; and the
refusal happens before the copy is attempted at all, while copy_mc
recovers only during the read.  The check and copy_mc compose.

>> Software-injected poison - what MADV_HWPOISON and the hwpoison-inject
>>  interface produce, and what tests and fuzzers exercise - never traps
>>  during the copy on any architecture.
>
> And these are debug interfaces, why do we care?

Because they are how this bug class gets found and reached in
practice: the reclaim-side check this series builds on (1b0449544c64)
exists because of a syzkaller report through these interfaces, and
carries Cc: stable - as does its large-folio follow-up (9f1e8cd0b7c4),
which you Acked.  You were also cc'd when Andrew raised the
urgency on the hugetlb sibling of this fix for the same reason -
"perhaps Bad Guys can find a way of triggering this, so the urgency
becomes higher".

The window itself is not debug-only either: a real UCE landing in it
behaves identically, and on a non-copy_mc arch that copy consumes real
poison instead of failing.

>> An isolation-time check avoids
>>  the migration entirely, on all architectures and for both real and
>>  simulated poison.
>
> It's racy. See the link below.

In isolation, yes - and dealing with that race is what patch 2 is
for, which is why this is a series.  The steady state (split failure
leaves the folio on the LRU with the flag already set) is not a race
at all, and that is what the isolation-time check catches
deterministically: 16/16 such folios migrated unpatched, 0/16 with
it.  For poison racing the batch, the load-bearing check is the one
patch 2 adds, immediately before the copy.  In patch 2's racing
workload, the entry check alone still propagated 693 of 9606 - that
datapoint is exactly why the second check is there.  A GUP-based
path takes its reference before the unmap, and memory_failure()
sets PG_hwpoison before releasing it, so a racing injection is
visible by copy time, unless the injecting thread is delayed
between taking the reference and setting the flag.  A real MCE can
likewise land in the final instructions before the copy; on x86_64
copy_mc catches that one at the read.

That thread is linked from patch 2's changelog precisely because the
concern you raised there is what the pre-copy check tries to answer.
The patch 1 sentence you quoted overstates the isolation check on
its own - I'll scope it to the already-flagged case in a v2.

> Was any of this written by an LLM?

The patches, the reproducer and the measurements are mine.  I used
an LLM mainly for the prose, to keep it accurate and unambiguous,
and also to look up related patches and past discussions on the
mailing lists.  I verified every technical claim against the code
and the test logs.  Fair point on the missing tag - v2 will carry
Assisted-by: LLM.

-- 
Yuanhe

  reply	other threads:[~2026-09-29  3:41 UTC|newest]

Thread overview: 5+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-28 10:58 [PATCH 0/2] mm: don't migrate " Yuanhe Shu
2026-09-28 10:58 ` [PATCH 1/2] mm/compaction: skip " Yuanhe Shu
2026-09-28 12:12   ` David Hildenbrand (Arm)
2026-09-29  3:41     ` Yuanhe Shu [this message]
2026-09-28 10:58 ` [PATCH 2/2] mm/migrate: refuse to migrate " Yuanhe Shu

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260929034123.2705745-1-xiangzao@linux.alibaba.com \
    --to=xiangzao@linux.alibaba.com \
    --cc=akpm@linux-foundation.org \
    --cc=chengkaitao@kylinos.cn \
    --cc=david@kernel.org \
    --cc=kaitao.cheng@linux.dev \
    --cc=linmiaohe@huawei.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=mgorman@techsingularity.net \
    --cc=muchun.song@linux.dev \
    --cc=nao.horiguchi@gmail.com \
    --cc=tujinjiang@huawei.com \
    --cc=vbabka@kernel.org \
    --cc=wangkefeng.wang@huawei.com \
    --cc=ying.huang@linux.alibaba.com \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®