From: Kyle Zeng <kylebot@openai.com>
To: linux-mm@kvack.org
Cc: linux-kernel@vger.kernel.org,
Andrew Morton <akpm@linux-foundation.org>,
David Hildenbrand <david@kernel.org>, Zi Yan <ziy@nvidia.com>,
Baolin Wang <baolin.wang@linux.alibaba.com>,
Nico Pache <nico.pache@linux.dev>,
Ryan Roberts <ryan.roberts@arm.com>, Dev Jain <dev.jain@arm.com>,
Barry Song <baohua@kernel.org>, Lance Yang <lance.yang@linux.dev>,
Usama Arif <usama.arif@linux.dev>,
Kiryl Shutsemau <kas@kernel.org>, Kyle Zeng <kylebot@openai.com>,
stable@vger.kernel.org
Subject: [PATCH] mm/huge_memory: avoid transient none PMDs during lazyfree reclaim
Date: Fri, 9 Oct 2026 09:52:15 -0700 [thread overview]
Message-ID: <20261009165214.40212-2-kylebot@openai.com> (raw)
__discard_anon_folio_pmd_locked() clears and flushes a huge PMD before
checking whether the folio can be discarded. If the folio was redirtied
or has unexpected references, it restores the original PMD. Reclaim holds
the PMD lock and the anon_vma read lock, but not the owning mmap_lock.
Both zap_pmd_range() and mremap's get_old_pmd() can skip a none PMD
without taking its lock. A concurrent munmap() or whole-VMA
MREMAP_DONTUNMAP can therefore miss the PMD and later unlink the source
VMA from its anon_vma, leaving a restored mapping behind. The anon_vma
write lock taken by unlink_anon_vmas() waits for reclaim to finish, but
does not repeat the skipped page-table walk. Once the remaining VMA
links are removed, the folio's positive mapcount no longer guarantees a
live anon_vma. Racing lazyfree reclaim against MREMAP_DONTUNMAP as an
unprivileged user reproduces a KASAN use-after-free in
folio_lock_anon_vma_read().
Use pmdp_invalidate() to keep a recognizable, non-none huge PMD while
the discard can still fail. Concurrent unmap and move operations then
have to synchronize on the PMD lock instead of skipping the mapping.
Clear the invalidated PMD only after the dirty and reference checks have
succeeded, before removing the rmap and withdrawing the deposited page
table. The full invalidation retains the TLB flush required for the
dirty and GUP-fast checks.
Fixes: 735ecdfaf4e8 ("mm/vmscan: avoid split lazyfree THP during shrink_folio_list()")
Cc: stable@vger.kernel.org
Assisted-by: Codex:gpt-6-astra
Signed-off-by: Kyle Zeng <kylebot@openai.com>
---
mm/huge_memory.c | 11 +++++++++--
1 file changed, 9 insertions(+), 2 deletions(-)
diff --git a/mm/huge_memory.c b/mm/huge_memory.c
index 1e5d68acf62a..20262195385a 100644
--- a/mm/huge_memory.c
+++ b/mm/huge_memory.c
@@ -3569,11 +3569,17 @@ static bool __discard_anon_folio_pmd_locked(struct vm_area_struct *vma,
return false;
}
- orig_pmd = pmdp_huge_clear_flush(vma, addr, pmdp);
+ /*
+ * The discard may fail, so keep the PMD non-none until we're
+ * committed to discarding it. Otherwise, concurrent munmap() or
+ * mremap() can skip the PMD without taking the PTL and later unlink
+ * the VMA from its anon_vma despite a restored mapping.
+ */
+ orig_pmd = pmdp_invalidate(vma, addr, pmdp);
/*
* Syncing against concurrent GUP-fast:
- * - clear PMD; barrier; read refcount
+ * - invalidate PMD; barrier; read refcount
* - inc refcount; barrier; read PMD
*/
smp_mb();
@@ -3607,6 +3613,7 @@ static bool __discard_anon_folio_pmd_locked(struct vm_area_struct *vma,
return false;
}
+ pmdp_huge_get_and_clear(mm, addr, pmdp);
folio_remove_rmap_pmd(folio, pmd_page(orig_pmd), vma);
zap_deposited_table(mm, pmdp);
add_mm_counter(mm, MM_ANONPAGES, -HPAGE_PMD_NR);
next reply other threads:[~2026-10-09 16:52 UTC|newest]
Thread overview: 9+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-09 16:52 Kyle Zeng [this message]
2026-10-09 18:00 ` David Hildenbrand (Arm)
2026-10-09 18:17 ` Zi Yan
2026-10-09 20:23 ` David Hildenbrand (Arm)
2026-10-09 20:27 ` Zi Yan
2026-10-09 20:44 ` David Hildenbrand (Arm)
2026-10-10 1:12 ` Lance Yang
2026-10-10 1:56 ` Zi Yan
2026-10-10 2:04 ` Lance Yang
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261009165214.40212-2-kylebot@openai.com \
--to=kylebot@openai.com \
--cc=akpm@linux-foundation.org \
--cc=baohua@kernel.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=david@kernel.org \
--cc=dev.jain@arm.com \
--cc=kas@kernel.org \
--cc=lance.yang@linux.dev \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=nico.pache@linux.dev \
--cc=ryan.roberts@arm.com \
--cc=stable@vger.kernel.org \
--cc=usama.arif@linux.dev \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®