mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Lance Yang <lance.yang@linux.dev>
To: usama.arif@linux.dev
Cc: akpm@linux-foundation.org, david@kernel.org, chrisl@kernel.org,
	kasong@tencent.com, ljs@kernel.org, ziy@nvidia.com,
	linux-mm@kvack.org, ying.huang@linux.alibaba.com,
	baoquan.he@linux.dev, willy@infradead.org, youngjun.park@lge.com,
	hannes@cmpxchg.org, riel@surriel.com, shakeel.butt@linux.dev,
	alex@ghiti.fr, kas@kernel.org, baohua@kernel.org,
	dev.jain@arm.com, baolin.wang@linux.alibaba.com,
	nico.pache@linux.dev, liam@infradead.org, ryan.roberts@arm.com,
	vbabka@kernel.org, lance.yang@linux.dev,
	linux-kernel@vger.kernel.org, nphamcs@gmail.com,
	shikemeng@huaweicloud.com, yosry@kernel.org, qi.zheng@linux.dev,
	luizcap@redhat.com, kernel-team@meta.com
Subject: Re: [PATCH v8 26/30] mm: handle PMD swap entries in UFFDIO_MOVE
Date: Sun,  4 Oct 2026 16:19:17 +0800	[thread overview]
Message-ID: <20261004081917.20111-1-lance.yang@linux.dev> (raw)
In-Reply-To: <20261002095503.3585565-27-usama.arif@linux.dev>

Hey Usama,

On Fri, Oct 02, 2026 at 02:52:40AM -0700, Usama Arif wrote:
[...]
>+static int move_swap_pmd(struct mm_struct *mm, struct vm_area_struct *dst_vma,
>+			 unsigned long dst_addr, unsigned long src_addr,
>+			 pmd_t *dst_pmd, pmd_t *src_pmd,
>+			 pmd_t orig_dst_pmd, pmd_t orig_src_pmd,
>+			 spinlock_t *dst_ptl, spinlock_t *src_ptl,
>+			 struct folio *src_folio, swp_entry_t entry)
>+{
[...]
>+
>+	moved_pmd = pmdp_huge_get_and_clear(mm, src_addr, src_pmd);
>+	if (pgtable_supports_soft_dirty())
>+		moved_pmd = pmd_swp_mksoft_dirty(moved_pmd);

Shouldn't we clear the source UFFD marker here?

	moved_pmd = pmd_swp_clear_uffd(moved_pmd);

just like move_swap_pte(), no?

>+	/* Re-arm RWP on the moved swap entry if dst_vma is RWP-registered. */
>+	if (userfaultfd_rwp(dst_vma))
>+		moved_pmd = pmd_swp_mkuffd(moved_pmd);
>+	set_pmd_at(mm, dst_addr, dst_pmd, moved_pmd);

static int move_swap_pte(struct mm_struct *mm, struct vm_area_struct *dst_vma,
			 unsigned long dst_addr, unsigned long src_addr,
			 pte_t *dst_pte, pte_t *src_pte,
			 pte_t orig_dst_pte, pte_t orig_src_pte,
			 pmd_t *dst_pmd, pmd_t dst_pmdval,
			 spinlock_t *dst_ptl, spinlock_t *src_ptl,
			 struct folio *src_folio,
			 struct swap_info_struct *si, swp_entry_t entry)
{
...
	orig_src_pte = ptep_get_and_clear(mm, src_addr, src_pte);
...
	orig_src_pte = pte_swp_clear_uffd(orig_src_pte);
	/* Re-arm RWP on the moved swap entry if dst_vma is RWP-registered. */
	if (userfaultfd_rwp(dst_vma))
		orig_src_pte = pte_swp_mkuffd(orig_src_pte);
	set_pte_at(mm, dst_addr, dst_pte, orig_src_pte);
...
}

Suppose the source is an exclusive swap PMD carrying a UFFD-WP marker, and
the destination is an empty hole registered for synchronous UFFD-WP without
RWP. We haven't write-protected the destination, but a successful whole-PMD
MOVE carries the source's UFFD-WP marker into it.

On swap-in, that marker is restored in the present PMD, leaving it
write-protected. The write fault then reaches
handle_userfault(vmf, VM_UFFD_WP), even though we never write-protected
the destination ...

That would be quite a surprise for userspace, and could leave the writer
stuck if the unexpected WP event goes unhandled :)

vm_fault_t do_huge_pmd_swap_page(struct vm_fault *vmf)
{
	struct vm_area_struct *vma = vmf->vma;
...
	bool exclusive, stable_writes, rwp_restore = false;
	bool write = vmf->flags & FAULT_FLAG_WRITE;
...
	pmd = folio_mk_pmd(folio, vma->vm_page_prot);
...
	if (pmd_swp_uffd(vmf->orig_pmd))
		pmd = pmd_mkuffd(pmd);
	if (pmd_swp_uffd(vmf->orig_pmd) && userfaultfd_rwp(vma)) {
		pmd = pmd_modify(pmd, PAGE_NONE);
		rwp_restore = true;
	}
...
	if (exclusive) {
		if (!rwp_restore && (vma->vm_flags & VM_WRITE) &&
		    !userfaultfd_huge_pmd_wp(vma, pmd) &&
		    !pmd_needs_soft_dirty_wp(vma, pmd)) {
			pmd = pmd_mkwrite(pmd, vma);
			if (write)
				pmd = pmd_mkdirty(pmd);
		}
		rmap_flags |= RMAP_EXCLUSIVE;
	}
...
	vmf->orig_pmd = pmd;
...
	if (write && !pmd_write(pmd) && !rwp_restore) {
		vm_fault_t wp_ret = wp_huge_pmd(vmf);
...
	}
...
}

vm_fault_t wp_huge_pmd(struct vm_fault *vmf)
{
	struct vm_area_struct *vma = vmf->vma;
	const bool unshare = vmf->flags & FAULT_FLAG_UNSHARE;
...
	if (vma_is_anonymous(vma)) {
		if (likely(!unshare) &&
		    userfaultfd_huge_pmd_wp(vma, vmf->orig_pmd)) {
			if (userfaultfd_wp_async(vmf->vma))
				goto split;
			return handle_userfault(vmf, VM_UFFD_WP);
		}
		return do_huge_pmd_wp_page(vmf);
	}
...
}

So, clearing the source UFFD-WP marker before the destination RWP check
would match the PTE helper and keep the RWP re-arm :)

Am I missing something?

>+	src_pgtable = pgtable_trans_huge_withdraw(mm, src_pmd);
>+	pgtable_trans_huge_deposit(mm, dst_pmd, src_pgtable);
>+
>+	double_pt_unlock(dst_ptl, src_ptl);
>+	return 0;
>+}
>+#endif /* CONFIG_THP_SWAP */
>+

[...]

Cheers, Lance

  reply	other threads:[~2026-10-04  8:19 UTC|newest]

Thread overview: 39+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-02  9:52 [PATCH v8 00/30] mm: PMD-level swap entries for anonymous THPs Usama Arif
2026-10-02  9:52 ` [PATCH v8 01/30] mm: rename pmd_to_softleaf_folio() to pmd_softleaf_to_folio() Usama Arif
2026-10-02  9:52 ` [PATCH v8 02/30] arm64: mm: add PMD swap-exclusive helpers Usama Arif
2026-10-02  9:52 ` [PATCH v8 03/30] loongarch: " Usama Arif
2026-10-02 14:17   ` Huacai Chen
2026-10-02  9:52 ` [PATCH v8 04/30] powerpc: " Usama Arif
2026-10-05  4:55   ` LEROY Christophe
2026-10-02  9:52 ` [PATCH v8 05/30] riscv: " Usama Arif
2026-10-02  9:52 ` [PATCH v8 06/30] s390: " Usama Arif
2026-10-02  9:52 ` [PATCH v8 07/30] x86: " Usama Arif
2026-10-02  9:52 ` [PATCH v8 08/30] mm: recognize PMD swap entries in the softleaf layer Usama Arif
2026-10-02  9:52 ` [PATCH v8 09/30] mm/debug_vm_pgtable: test PMD swap-exclusive helpers Usama Arif
2026-10-02  9:52 ` [PATCH v8 10/30] mm: make PMD migration-entry splitting explicit Usama Arif
2026-10-02  9:52 ` [PATCH v8 11/30] mm: split PMD swap entries into PTE swap entries Usama Arif
2026-10-02  9:52 ` [PATCH v8 12/30] mm/swap: allow duplicating a range of " Usama Arif
2026-10-02  9:52 ` [PATCH v8 13/30] mm: handle PMD swap entries in fork path Usama Arif
2026-10-02  9:52 ` [PATCH v8 14/30] mm: zswap: reject high-order swap cache allocations backed by zswap Usama Arif
2026-10-02  9:52 ` [PATCH v8 15/30] mm: swap in PMD swap entries as whole THPs during swapoff Usama Arif
2026-10-02  9:52 ` [PATCH v8 16/30] fs/proc: account PMD swap entries in smaps Usama Arif
2026-10-02  9:52 ` [PATCH v8 17/30] mm: handle soft-dirty and uffd-wp on PMD swap entries Usama Arif
2026-10-02  9:52 ` [PATCH v8 18/30] mm/hmm: fault PMD swap entries on demand Usama Arif
2026-10-02  9:52 ` [PATCH v8 19/30] mm: free PMD swap entries in zap_huge_pmd() Usama Arif
2026-10-02  9:52 ` [PATCH v8 20/30] mm/madvise: free PMD swap entries with MADV_FREE Usama Arif
2026-10-02  9:52 ` [PATCH v8 21/30] mm/madvise: skip PMD swap entries for MADV_COLD and MADV_PAGEOUT Usama Arif
2026-10-02  9:52 ` [PATCH v8 22/30] mm/madvise: keep PMD swap entries whole for MADV_GUARD_INSTALL/REMOVE Usama Arif
2026-10-02  9:52 ` [PATCH v8 23/30] mm/mincore: report PMD swap-cache residency Usama Arif
2026-10-02  9:52 ` [PATCH v8 24/30] mm/khugepaged: treat PMD swap entries as mapped THPs Usama Arif
2026-10-02  9:52 ` [PATCH v8 25/30] mm: handle PMD swap entries in MADV_WILLNEED Usama Arif
2026-10-02  9:52 ` [PATCH v8 26/30] mm: handle PMD swap entries in UFFDIO_MOVE Usama Arif
2026-10-04  8:19   ` Lance Yang [this message]
2026-10-02  9:52 ` [PATCH v8 27/30] mm: don't PTE-batch a swap-in over a hardware-poisoned subpage Usama Arif
2026-10-02  9:52 ` [PATCH v8 28/30] mm: handle PMD swap entry faults on swap-in Usama Arif
2026-10-04 10:19   ` Lance Yang
2026-10-02  9:52 ` [PATCH v8 29/30] mm: install PMD swap entries on swap-out Usama Arif
2026-10-02  9:52 ` [PATCH v8 30/30] selftests/mm: add PMD swap entry tests Usama Arif
2026-10-02 14:28 ` [PATCH v8 00/30] mm: PMD-level swap entries for anonymous THPs David Hildenbrand (Arm)
2026-10-02 15:13   ` Zi Yan
2026-10-04 12:39     ` Usama Arif
2026-10-04  3:08 ` Lance Yang

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261004081917.20111-1-lance.yang@linux.dev \
    --to=lance.yang@linux.dev \
    --cc=akpm@linux-foundation.org \
    --cc=alex@ghiti.fr \
    --cc=baohua@kernel.org \
    --cc=baolin.wang@linux.alibaba.com \
    --cc=baoquan.he@linux.dev \
    --cc=chrisl@kernel.org \
    --cc=david@kernel.org \
    --cc=dev.jain@arm.com \
    --cc=hannes@cmpxchg.org \
    --cc=kas@kernel.org \
    --cc=kasong@tencent.com \
    --cc=kernel-team@meta.com \
    --cc=liam@infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=luizcap@redhat.com \
    --cc=nico.pache@linux.dev \
    --cc=nphamcs@gmail.com \
    --cc=qi.zheng@linux.dev \
    --cc=riel@surriel.com \
    --cc=ryan.roberts@arm.com \
    --cc=shakeel.butt@linux.dev \
    --cc=shikemeng@huaweicloud.com \
    --cc=usama.arif@linux.dev \
    --cc=vbabka@kernel.org \
    --cc=willy@infradead.org \
    --cc=ying.huang@linux.alibaba.com \
    --cc=yosry@kernel.org \
    --cc=youngjun.park@lge.com \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®