From: James Houghton <jthoughton@google.com>
To: Andrew Morton <akpm@linux-foundation.org>
Cc: David Hildenbrand <david@kernel.org>,
Lorenzo Stoakes <ljs@kernel.org>, Zi Yan <ziy@nvidia.com>,
Baolin Wang <baolin.wang@linux.alibaba.com>,
liam@infradead.org, Nico Pache <nico.pache@linux.dev>,
Ryan Roberts <ryan.roberts@arm.com>, Dev Jain <dev.jain@arm.com>,
Barry Song <baohua@kernel.org>,
Lance Yang <lance.yang@linux.dev>,
Usama Arif <usama.arif@linux.dev>,
Yang Shi <shy828301@gmail.com>,
zokeefe@google.com, hughd@google.com,
Kiryl Shutsemau <kas@kernel.org>,
jthoughton@google.com, linux-mm@kvack.org,
linux-kernel@vger.kernel.org
Subject: [PATCH v4 1/2] mm/khugepaged: Never install PMDs in uffd-minor-registered VMAs
Date: Sat, 3 Oct 2026 00:18:58 +0000 [thread overview]
Message-ID: <20261003001859.502725-1-jthoughton@google.com> (raw)
Userfaultfd minor faults provides userspace with the ability to manually
install PTEs with UFFDIO_CONTINUE. Right now, MADV_COLLAPSE can map
holes in the VMA when a naturally-aligned THP is present. This is not
true for khugepaged collapse: the PTEs will be retracted, but a PMD will
not be installed.
When MADV_COLLAPSE installs a PMD that mapped holes in the VMA,
userspace is likely to expect UFFDIO_CONTINUE to succeed on the
should-be holes. UFFDIO_CONTINUE will fail and return EEXIST.
This is not inherently a problem, as MADV_COLLAPSE is an explicit
userspace action. But, especially because MADV_COLLAPSE can be invoked
by an external process via process_madvise(), a rogue caller could break
a userfaultfd-minor resolver thread.
Userspace cannot generally use MADV_COLLAPSE to resolve userfault minor
faults, as MADV_COLLAPSE will only resolve such faults if a
naturally-aligned THP is present, so this is not a functional
regression for userspace.
The naturally-aligned THP case is the only case where this quirk exists.
Collapsing otherwise requires all PTEs to be present for
userfaultfd-registered VMAs (i.e., max none PTEs is 0), which is
correct. This check is essentially bypassed for naturally-aligned THPs.
Suggested-by: Lance Yang <lance.yang@linux.dev>
Tested-by: Lance Yang <lance.yang@linux.dev>
Signed-off-by: James Houghton <jthoughton@google.com>
---
Changes since v3:
- Prevent page table retraction in patch 1. Please see the comment
next to it. So this patch now has two hunks instead of one.
- Reworded the comment and combined the userfaultfd checks in patch 1.
(Thanks David)
- Simplified the selftest a bit given the behavior change in patch 1.
- Undid a change in v3's selftest that incorrectly handled the case
where MADV_COLLAPSE was not supported entirely.
- Rebased on top of mm-unstable, which includes Kiryl's changes.
v3: https://lore.kernel.org/linux-mm/20260910023411.514987-1-jthoughton@google.com/
v2: https://lore.kernel.org/linux-mm/20260828222640.1638457-1-jthoughton@google.com/
v1: https://lore.kernel.org/linux-mm/20260828005004.2870750-1-jthoughton@google.com/
---
mm/khugepaged.c | 14 ++++++++++----
1 file changed, 10 insertions(+), 4 deletions(-)
diff --git a/mm/khugepaged.c b/mm/khugepaged.c
index 913086eaf17b..438c4af1c058 100644
--- a/mm/khugepaged.c
+++ b/mm/khugepaged.c
@@ -1868,10 +1868,11 @@ static enum scan_result try_collapse_pte_mapped_thp(struct mm_struct *mm, unsign
return SCAN_VMA_CHECK;
/*
- * Keep pmd pgtable while the uffd bit is in use; see comment in
- * retract_page_tables().
+ * Don't collapse if there might be PTE markers for userfaultfd-based
+ * access protection or if collapsing might bypass userfaultfd minor
+ * faults.
*/
- if (userfaultfd_protected(vma))
+ if (userfaultfd_protected(vma) || userfaultfd_minor(vma))
return SCAN_PTE_UFFD;
folio = filemap_lock_folio(vma->vm_file->f_mapping,
@@ -2093,8 +2094,13 @@ static bool file_backed_vma_is_retractable(struct vm_area_struct *vma)
* and cannot be recycled to a shared PMD. Other vmas can still
* have the same file mapped hugely, but skip this one: it will
* always be mapped in small page size for these registrations.
+ *
+ * Userfaultfd-minor-registered VMAs should also not be retracted.
+ * PMDs will not be installed, as doing so can suppress minor faults.
+ * If retraction were allowed, khugepaged might continually cause
+ * unnecessary userfaultfd minor faults for already-CONTINUE'd pages.
*/
- if (userfaultfd_protected(vma))
+ if (userfaultfd_protected(vma) || userfaultfd_minor(vma))
return false;
/*
base-commit: b2b4b29b76dabdee576eba66953a66ca61c5fca0
--
2.56.0.rc1.315.gc6ed9934b7-goog
next reply other threads:[~2026-10-03 0:19 UTC|newest]
Thread overview: 2+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-03 0:18 James Houghton [this message]
2026-10-03 0:18 ` [PATCH v4 2/2] mm: selftests: Adjust the MADV_COLLAPSE selftests for uffd-minor James Houghton
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261003001859.502725-1-jthoughton@google.com \
--to=jthoughton@google.com \
--cc=akpm@linux-foundation.org \
--cc=baohua@kernel.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=david@kernel.org \
--cc=dev.jain@arm.com \
--cc=hughd@google.com \
--cc=kas@kernel.org \
--cc=lance.yang@linux.dev \
--cc=liam@infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=nico.pache@linux.dev \
--cc=ryan.roberts@arm.com \
--cc=shy828301@gmail.com \
--cc=usama.arif@linux.dev \
--cc=ziy@nvidia.com \
--cc=zokeefe@google.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®