From: Lance Yang <lance.yang@linux.dev>
To: jthoughton@google.com
Cc: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org,
ziy@nvidia.com, baolin.wang@linux.alibaba.com,
liam@infradead.org, nico.pache@linux.dev, ryan.roberts@arm.com,
dev.jain@arm.com, baohua@kernel.org, lance.yang@linux.dev,
usama.arif@linux.dev, shy828301@gmail.com, zokeefe@google.com,
hughd@google.com, kas@kernel.org, linux-mm@kvack.org,
linux-kernel@vger.kernel.org, stable@vger.kernel.org
Subject: Re: [PATCH] mm/khugepaged: Don't collapse uffd-minor-registered VMAs
Date: Fri, 28 Aug 2026 17:47:03 +0800 [thread overview]
Message-ID: <20260828094703.11081-1-lance.yang@linux.dev> (raw)
In-Reply-To: <20260828005004.2870750-1-jthoughton@google.com>
On Fri, Aug 28, 2026 at 12:50:04AM +0000, James Houghton wrote:
>Userfaultfd minor faults provides userspace with the ability to manually
>install PTEs with UFFDIO_CONTINUE. Right now, khugepaged collapse can
>map holes in the VMA when a naturally-aligned THP is present without
>explicit action from userspace.
>
>This is a problem, as it bypasses userfaultfd minor faults that
>userspace is expecting to handle.
One basic question first. Should MADV_COLLAPSE refuse to collapse a
UFFD-minor-registered VMA, regardless of whether all PTEs are present?
I'd leave that to the maintainers :D
Anyway, assuming the answer is yes, I wonder whether the new check is
sufficient. See below.
>
>If userspace implements post-copy live migration using userfaultfd minor
>faults, this situation is currently possible:
>1. The VMA for guest memory is userfaultfd-minor-registered and nothing
> is mapped in the page tables.
>2. A stale copy of a page is present in a naturally-aligned THP (from
> pre-copy live migration).
>3. khugepaged collapses the mapping of the THP, installs a PMD.
>4. The VM now has access to the stale contents => VM is broken.
>5. After installing the correct contents, userspace attempts to map the
> page with UFFDIO_CONTINUE; it gets EEXIST, indicating that something
> unexpectedly mapped the page.
>
>The naturally-aligned THP case is the only case where this is a problem.
>khugepaged otherwise requires all PTEs to be present for
>userfaultfd-registered VMAs (i.e., max none PTEs is 0), which is
>correct. This check is essentially bypassed for naturally-aligned THPs.
>
>To deal with this issue, completely disallow collapsing in
>userfaultfd-minor-registered VMAs. This is slightly pessimistic; it
>would be nice to allow MADV_COLLAPSE to work if all PTEs are in fact
>present, but that seems more complex than it is worth.
>
>Fixes: 58ac9a8993a1 ("mm/khugepaged: attempt to map file/shmem-backed pte-mapped THPs by pmds")
>Cc: <stable@vger.kernel.org> # 6.1
>Signed-off-by: James Houghton <jthoughton@google.com>
>---
>This was caught with manual review while diagnosing a related issue
>that came up with in Google's live migration testing.
>
>I've uploaded a mostly-AI-generated reproducer here[1]. As long as
>/sys/kernel/mm/transparent_hugepage/shmem_enabled is not set to 'deny',
>the repro should work.
>
>[1] https://gist.github.com/48ca/d399bf534158e80241fb4937ef1ff664
>---
> mm/khugepaged.c | 9 +++++++++
> 1 file changed, 9 insertions(+)
>
>diff --git a/mm/khugepaged.c b/mm/khugepaged.c
>index b237f6e7662a..66f956d3dd67 100644
>--- a/mm/khugepaged.c
>+++ b/mm/khugepaged.c
>@@ -2804,6 +2804,15 @@ static enum scan_result collapse_single_pmd(unsigned long addr,
> goto end;
> }
>
>+ /*
>+ * Userfaultfd-minor-registered VMAs should not be collapsed, as
>+ * userspace is expecting to explicitly install PTEs.
>+ */
>+ if (userfaultfd_minor(vma)) {
>+ result = SCAN_PTE_UFFD;
>+ goto end;
>+ }
Assume UFFDIO_REGISTER_MODE_MINOR completes after collapse_single_pmd()
drops the mmap read lock and before it reacquires it.
Doesn't this still leave a registration race, no?
int madvise_collapse(struct vm_area_struct *vma, unsigned long start,
unsigned long end, bool *lock_dropped)
{
...
cc->is_khugepaged = false;
...
result = collapse_single_pmd(addr, vma, &mmap_unlocked, cc);
...
}
static enum scan_result collapse_single_pmd(unsigned long addr,
struct vm_area_struct *vma, bool *lock_dropped,
struct collapse_control *cc)
{
...
if (userfaultfd_minor(vma)) {
result = SCAN_PTE_UFFD;
goto end;
}
...
mmap_read_unlock(mm);
*lock_dropped = true;
...
if (result == SCAN_PTE_MAPPED_HUGEPAGE) {
mmap_read_lock(mm);
if (collapse_test_exit_or_disable(mm))
result = SCAN_ANY_PROCESS;
else
result = try_collapse_pte_mapped_thp(mm, addr,
!cc->is_khugepaged);
...
mmap_read_unlock(mm);
}
...
}
static enum scan_result try_collapse_pte_mapped_thp(struct mm_struct *mm, unsigned long addr,
bool install_pmd)
{
...
struct vm_area_struct *vma = vma_lookup(mm, haddr);
...
if (!vma || !vma->vm_file ||
!range_in_vma(vma, haddr, haddr + HPAGE_PMD_SIZE))
return SCAN_VMA_CHECK;
...
if (userfaultfd_protected(vma))
return SCAN_PTE_UFFD;
...
result = find_pmd_or_thp_or_none(mm, haddr, &pmd);
switch (result) {
case SCAN_SUCCEED:
break;
case SCAN_NO_PTE_TABLE:
...
goto maybe_install_pmd;
default:
goto drop_folio;
}
...
maybe_install_pmd:
/* step 5: install pmd entry */
result = install_pmd
? set_huge_pmd(vma, haddr, pmd, folio, &folio->page)
: SCAN_SUCCEED;
...
}
static inline bool userfaultfd_minor(struct vm_area_struct *vma)
{
return vma_test_any_mask(vma, VMA_UFFD_MINOR);
}
static inline bool userfaultfd_protected(struct vm_area_struct *vma)
{
return userfaultfd_wp(vma) || userfaultfd_rwp(vma);
}
Emm ... userfaultfd_protected() only covers WP and RWP. MADV_COLLAPSE
passes install_pmd=true, so the SCAN_NO_PTE_TABLE case can still reach
set_huge_pmd() after UFFDIO_REGISTER_MODE_MINOR has completed ...
Maybe:
---8<---
diff --git a/mm/khugepaged.c b/mm/khugepaged.c
index 33c41bc32af8..0eada7265d59 100644
--- a/mm/khugepaged.c
+++ b/mm/khugepaged.c
@@ -1893,6 +1893,8 @@ static enum scan_result try_collapse_pte_mapped_thp(struct mm_struct *mm, unsign
*/
if (userfaultfd_protected(vma))
return SCAN_PTE_UFFD;
+ if (userfaultfd_minor(vma))
+ return SCAN_PTE_UFFD;
folio = filemap_lock_folio(vma->vm_file->f_mapping,
linear_page_index(vma, haddr));
--
With that, LGTM.
Tested-by: Lance Yang <lance.yang@linux.dev>
Cheers, Lance
next prev parent reply other threads:[~2026-08-28 9:47 UTC|newest]
Thread overview: 4+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-28 0:50 James Houghton
2026-08-28 9:47 ` Lance Yang [this message]
2026-08-28 13:07 ` Kiryl Shutsemau
2026-08-28 19:07 ` James Houghton
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260828094703.11081-1-lance.yang@linux.dev \
--to=lance.yang@linux.dev \
--cc=akpm@linux-foundation.org \
--cc=baohua@kernel.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=david@kernel.org \
--cc=dev.jain@arm.com \
--cc=hughd@google.com \
--cc=jthoughton@google.com \
--cc=kas@kernel.org \
--cc=liam@infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=nico.pache@linux.dev \
--cc=ryan.roberts@arm.com \
--cc=shy828301@gmail.com \
--cc=stable@vger.kernel.org \
--cc=usama.arif@linux.dev \
--cc=ziy@nvidia.com \
--cc=zokeefe@google.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®