From: "Zi Yan" <ziy@nvidia.com>
To: "Lance Yang" <lance.yang@linux.dev>,
"David Hildenbrand (Arm)" <david@kernel.org>
Cc: "Kyle Zeng" <kylebot@openai.com>, <linux-mm@kvack.org>,
<linux-kernel@vger.kernel.org>,
"Andrew Morton" <akpm@linux-foundation.org>,
"Baolin Wang" <baolin.wang@linux.alibaba.com>,
"Nico Pache" <nico.pache@linux.dev>,
"Ryan Roberts" <ryan.roberts@arm.com>,
"Dev Jain" <dev.jain@arm.com>, "Barry Song" <baohua@kernel.org>,
"Usama Arif" <usama.arif@linux.dev>,
"Kiryl Shutsemau" <kas@kernel.org>, <stable@vger.kernel.org>
Subject: Re: [PATCH] mm/huge_memory: avoid transient none PMDs during lazyfree reclaim
Date: Fri, 09 Oct 2026 21:56:17 -0400 [thread overview]
Message-ID: <DM0S9L075B57.1FCNOI7LA1FJW@nvidia.com> (raw)
In-Reply-To: <0ec81c65-fbf5-405f-948b-0fb1b2e083e8@linux.dev>
On Fri Oct 9, 2026 at 9:12 PM EDT, Lance Yang wrote:
>
>
> On 2026/10/10 04:44, David Hildenbrand (Arm) wrote:
>> On 10/9/26 22:27, Zi Yan wrote:
>>> On 9 Oct 2026, at 16:23, David Hildenbrand (Arm) wrote:
>>>
>>>> On 10/9/26 20:17, Zi Yan wrote:
>>>>>
>>>>> There are some gaps we need to close before getting this fix in:
>>>>>
>>>>> 1. GUP-fast cannot follow the invalidated PMD: riscv and LoongArch need
>>>>> pmd_access_permitted() that requires _PAGE_PRESENT; s390's
>>>>> pmdp_invalidate() needs a change.
>>>>
>>>> Ugh. We really need pmdp_invalidate() to have reasonable semantics. This whole
>>>> PMD locking is a mess :(
>>>>
>>>>>
>>>>> 2. sparc64's thp_pte_count can be imbalanced with this change (based on
>>>>> Lance's offlist feedback).
>>>>
>>>> Ack.
>>>>
>>>>>
>>>>> In addition, powerpc has a page table check issue similr to riscv, where riscv
>>>>> fixed it with commit 9f4a88b8d01a6. This is not related to this issue
>>>>> but discovered along with the investigation.
>>>>
>>>> Zi, do you have the capacity to take over this patch?
>>>
>>> Yes, I can take over it. My plan is to send a series including
>>> patches for 1 and this patch as is. powerpc fix can be a separate one.
>>>
>>> Best Regards,
>>> Yan, Zi
>>
>> If we need a quick stable fix we could temporarily disable the whole thing until
>> it is fixed, just a thought.
>
> +1 we could do that for now. A quick fix with minimal churn (correctness
> comes first).
How about the patch below? unmap_huge_pmd_locked() and
__discard_anon_folio_pmd_locked() will be dead code to keep the patch
small. Later, with arch fixes, the function can be re-enabled.
From 2204b9d08530796847a586428b1f782b987d071c Mon Sep 17 00:00:00 2001
From: Zi Yan <ziy@nvidia.com>
Date: Fri, 9 Oct 2026 21:28:10 -0400
Subject: [PATCH] mm/rmap: don't discard lazyfree THPs at PMD level
__discard_anon_folio_pmd_locked() clears a lazyfree THP PMD before it knows
whether the folio can be discarded, and restores it if the folio was
redirtied or has extra references. A concurrent munmap() or
MREMAP_DONTUNMAP skips the temporary none PMD and unlinks the VMA from its
anon_vma, so the folio stays mapped after the anon_vma is freed and a later
rmap walk uses the freed anon_vma.
Using an invalidated PMD instead of a cleared one requires additional arch
code fixes. Instead, disable the PMD level discard of lazyfree THPs, as
before commit 735ecdfaf4e8 ("mm/vmscan: avoid split lazyfree THP during
shrink_folio_list()").
Fixes: 735ecdfaf4e8 ("mm/vmscan: avoid split lazyfree THP during shrink_folio_list()")
Reported-by: Kyle Zeng <kylebot@openai.com>
Closes: https://lore.kernel.org/r/20261009165214.40212-2-kylebot@openai.com
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Zi Yan <ziy@nvidia.com>
---
mm/rmap.c | 11 -----------
1 file changed, 11 deletions(-)
diff --git a/mm/rmap.c b/mm/rmap.c
index 805db93fe0428..1131b76bbbc28 100644
--- a/mm/rmap.c
+++ b/mm/rmap.c
@@ -2275,17 +2275,6 @@ static bool try_to_unmap_one(struct folio *folio, struct vm_area_struct *vma,
}
if (!pvmw.pte) {
- if (folio_test_lazyfree(folio)) {
- if (unmap_huge_pmd_locked(vma, pvmw.address, pvmw.pmd, folio))
- goto walk_done;
- /*
- * unmap_huge_pmd_locked has either already marked
- * the folio as swap-backed or decided to retain it
- * due to GUP or speculative references.
- */
- goto walk_abort;
- }
-
if (flags & TTU_SPLIT_HUGE_PMD) {
/*
* We temporarily have to drop the PTL and
--
2.53.0
--
Best Regards,
Yan, Zi
next prev parent reply other threads:[~2026-10-10 1:56 UTC|newest]
Thread overview: 9+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-09 16:52 Kyle Zeng
2026-10-09 18:00 ` David Hildenbrand (Arm)
2026-10-09 18:17 ` Zi Yan
2026-10-09 20:23 ` David Hildenbrand (Arm)
2026-10-09 20:27 ` Zi Yan
2026-10-09 20:44 ` David Hildenbrand (Arm)
2026-10-10 1:12 ` Lance Yang
2026-10-10 1:56 ` Zi Yan [this message]
2026-10-10 2:04 ` Lance Yang
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=DM0S9L075B57.1FCNOI7LA1FJW@nvidia.com \
--to=ziy@nvidia.com \
--cc=akpm@linux-foundation.org \
--cc=baohua@kernel.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=david@kernel.org \
--cc=dev.jain@arm.com \
--cc=kas@kernel.org \
--cc=kylebot@openai.com \
--cc=lance.yang@linux.dev \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=nico.pache@linux.dev \
--cc=ryan.roberts@arm.com \
--cc=stable@vger.kernel.org \
--cc=usama.arif@linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®