* [PATCH v3 0/2] mm: stop calling pmd_folio() on special PMDs
@ 2026-09-26 10:51 Gregory Price
2026-09-26 10:51 ` [PATCH v3 1/2] mm/mempolicy: use vm_normal_folio_pmd() in queue_folios_pmd() Gregory Price
` (2 more replies)
0 siblings, 3 replies; 4+ messages in thread
From: Gregory Price @ 2026-09-26 10:51 UTC (permalink / raw)
To: linux-mm
Cc: linux-kernel, kernel-team, akpm, liam, ljs, david, vbabka, jannh,
gourry, ziy, joshua.hahnjy, rakie.kim, ying.huang, peterx, jgg,
sashiko-bot
Two page table walkers resolve the folio behind a PMD with pmd_folio(),
which is only valid for a PMD mapping a refcounted struct page:
madvise_cold_or_pageout_pte_range() mm/madvise.c
queue_folios_pmd() mm/mempolicy.c
vmf_insert_pfn_pmd() installs special PMDs holding a raw pfn that need not
have a memmap entry at all. Both walkers can reach one and fault on the
first folio field read. The PTE halves of both already use
vm_normal_folio(); these two patches make the PMD halves match.
The four callers of vmf_insert_pfn_pmd(), and which walker each reaches:
drivers/vfio/pci/vfio_pci_core.c VM_PFNMAP mempolicy
drivers/gpu/drm/drm_gem_shmem_helper.c VM_PFNMAP mempolicy
drivers/gpu/drm/panthor/panthor_gem.c VM_PFNMAP mempolicy
drivers/hv/mshv_vtl_main.c VM_MIXEDMAP both
can_madv_lru_vma() rejects VM_PFNMAP, so only mshv_vtl_low reaches the
madvise walker, and that needs CAP_SYS_ADMIN. queue_pages_walk_ops
supplies its own ->test_walk, so walk_page_test()'s generic VM_PFNMAP skip
never runs and vfio-pci is reachable by any process holding the device fd.
Hence the different stable tags.
One behaviour change: mbind(MPOL_MF_STRICT) over a PMD mapped VM_PFNMAP
region now returns 0 rather than -EIO. The PTE loop already returned 0
there. drm_gem_shmem and panthor are where this is observable, since they
PMD map pages that do have a memmap entry and so never faulted.
Reproducer
==========
No hardware needed. An out of tree module stands in for the drivers above:
three misc devices, each with a ->huge_fault calling vmf_insert_pfn_pmd(),
plus VM_HUGEPAGE so the fault path takes the PMD branch.
/dev/pmdspec_mixed VM_MIXEDMAP, pfn at the 1 TiB mark, no memmap
/dev/pmdspec_pfnmap VM_PFNMAP, pfn at the 1 TiB mark, no memmap
/dev/pmdspec_real VM_PFNMAP, real alloc_pages(PMD_ORDER) on node 0
Userspace maps the device into a PMD aligned window, reads one byte to
fault the PMD in, checks a module parameter to confirm it went in, then
issues the operation.
vng --run <bzImage> --user root --memory 4G --verbose \
--append "numa=fake=2" \
--exec "insmod pmdspec.ko && ./pmdspec_test <subtest>"
numa=fake=2 gives a node 1 to bind to; the module allocates its real page
on node 0, which is what makes queue_folio_required() true.
subtest operation parent series
--------------------------------------------------------------------
madv_cold madvise(MADV_COLD) oops ret=0
madv_pageout madvise(MADV_PAGEOUT) oops ret=0
mbind_mixed mbind(MPOL_BIND, n1, MPOL_MF_MOVE) oops ret=0
mbind_pfnmap mbind(MPOL_BIND, n1, MPOL_MF_STRICT) oops ret=0
mbind_real mbind(MPOL_BIND, n1, MPOL_MF_STRICT) -EIO ret=0
Two things the table shows that are easy to miss in the code:
- mbind_mixed passes only MPOL_MF_MOVE. MPOL_MF_STRICT is not needed for
a VM_MIXEDMAP vma: walk_page_test() only skips VM_PFNMAP, and
vma_migratable() is true for VM_MIXEDMAP.
- mbind_real demonstrates the user visible change (-EIO -> 0)
Reported-by: sashiko-bot <sashiko-bot@kernel.org>
Closes: https://sashiko.dev/#/patchset/20260817220810.1175596-1-gourry%40gourry.net
Assisted-by: LLM
---
v3: improvements from Lorenzo
v2: improvements from David
Gregory Price (2):
mm/mempolicy: use vm_normal_folio_pmd() in queue_folios_pmd()
mm/madvise: use vm_normal_folio_pmd() in cold/pageout PMD range
mm/madvise.c | 7 +++----
mm/mempolicy.c | 13 +++++--------
2 files changed, 8 insertions(+), 12 deletions(-)
--
2.55.0
^ permalink raw reply [flat|nested] 4+ messages in thread
* [PATCH v3 1/2] mm/mempolicy: use vm_normal_folio_pmd() in queue_folios_pmd()
2026-09-26 10:51 [PATCH v3 0/2] mm: stop calling pmd_folio() on special PMDs Gregory Price
@ 2026-09-26 10:51 ` Gregory Price
2026-09-26 10:51 ` [PATCH v3 2/2] mm/madvise: use vm_normal_folio_pmd() in cold/pageout PMD range Gregory Price
2026-09-27 21:45 ` [PATCH v3 0/2] mm: stop calling pmd_folio() on special PMDs Andrew Morton
2 siblings, 0 replies; 4+ messages in thread
From: Gregory Price @ 2026-09-26 10:51 UTC (permalink / raw)
To: linux-mm
Cc: linux-kernel, kernel-team, akpm, liam, ljs, david, vbabka, jannh,
gourry, ziy, joshua.hahnjy, rakie.kim, ying.huang, peterx, jgg,
sashiko-bot, stable
mmap a VM_PFNMAP region whose ->huge_fault installs a PMD through
vmf_insert_pfn_pmd() - a vfio-pci MMIO BAR does this - then
mbind(p, len, MPOL_BIND, &mask, maxnode, MPOL_MF_STRICT);
With a stand-in module for the driver:
BUG: unable to handle page fault for address: fffff96dc0000008
RIP: 0010:queue_folios_pte_range+0xaf/0x440
walk_pgd_range+0x52b/0xaf0
__walk_page_range+0x6a/0x1d0
walk_page_range_mm_unsafe+0x193/0x230
queue_pages_range+0x64/0xa0
do_mbind+0x25e/0x640
queue_folios_pmd(), inlined above, calls pmd_folio() on that PMD. The pfn
is raw MMIO with no memmap entry, so the folio lands in unpopulated
vmemmap. Neither guard stops the walk:
walk_page_test() skips VM_PFNMAP, but queue_pages_walk_ops
supplies ->test_walk, so it never runs
queue_pages_test_walk() honours vma_migratable(), but only while
MPOL_MF_STRICT is clear
A VM_MIXEDMAP vma needs neither flag, being vma_migratable(), so plain
mbind(MPOL_MF_MOVE) reaches this too - and there the bad folio carries on
into migrate_folio_add() and folio_isolate_lru(). mshv_vtl_low is such a
mapping.
Use vm_normal_folio_pmd() and skip on NULL, as the PTE loop in
queue_folios_pte_range() already does with vm_normal_folio(). This also
filters the huge zero PMD, so its separate check is no longer needed.
mbind(MPOL_MF_STRICT) over a PMD mapped VM_PFNMAP region now returns 0
rather than -EIO. The PTE loop already returned 0 there.
Fixes: 3c8e44c9b369 ("mm: mark special bits for huge pfn mappings when inject")
Reported-by: sashiko-bot <sashiko-bot@kernel.org>
Closes: https://sashiko.dev/#/patchset/20260817220810.1175596-1-gourry%40gourry.net
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Gregory Price (Meta) <gourry@gourry.net>
Acked-by: David Hildenbrand (Arm) <david@kernel.org>
Reviewed-by: Zi Yan <ziy@nvidia.com>
---
mm/mempolicy.c | 13 +++++--------
1 file changed, 5 insertions(+), 8 deletions(-)
diff --git a/mm/mempolicy.c b/mm/mempolicy.c
index 2063ab7577d7..40744658483b 100644
--- a/mm/mempolicy.c
+++ b/mm/mempolicy.c
@@ -667,7 +667,8 @@ static inline bool queue_folio_required(struct folio *folio,
return node_isset(nid, *qp->nmask) == !(flags & MPOL_MF_INVERT);
}
-static void queue_folios_pmd(pmd_t *pmd, struct mm_walk *walk)
+static void queue_folios_pmd(pmd_t *pmd, unsigned long addr,
+ struct mm_walk *walk)
{
struct folio *folio;
struct queue_pages *qp = walk->private;
@@ -678,13 +679,9 @@ static void queue_folios_pmd(pmd_t *pmd, struct mm_walk *walk)
qp->nr_failed++;
return;
}
- folio = pmd_folio(pmdval);
- if (folio_is_zone_device(folio))
+ folio = vm_normal_folio_pmd(walk->vma, addr, pmdval);
+ if (!folio || folio_is_zone_device(folio))
return;
- if (is_huge_zero_folio(folio)) {
- walk->action = ACTION_CONTINUE;
- return;
- }
if (!queue_folio_required(folio, qp))
return;
if (!(qp->flags & (MPOL_MF_MOVE | MPOL_MF_MOVE_ALL)) ||
@@ -717,7 +714,7 @@ static int queue_folios_pte_range(pmd_t *pmd, unsigned long addr,
ptl = pmd_trans_huge_lock(pmd, vma);
if (ptl) {
- queue_folios_pmd(pmd, walk);
+ queue_folios_pmd(pmd, addr, walk);
spin_unlock(ptl);
goto out;
}
--
2.55.0
^ permalink raw reply [flat|nested] 4+ messages in thread
* [PATCH v3 2/2] mm/madvise: use vm_normal_folio_pmd() in cold/pageout PMD range
2026-09-26 10:51 [PATCH v3 0/2] mm: stop calling pmd_folio() on special PMDs Gregory Price
2026-09-26 10:51 ` [PATCH v3 1/2] mm/mempolicy: use vm_normal_folio_pmd() in queue_folios_pmd() Gregory Price
@ 2026-09-26 10:51 ` Gregory Price
2026-09-27 21:45 ` [PATCH v3 0/2] mm: stop calling pmd_folio() on special PMDs Andrew Morton
2 siblings, 0 replies; 4+ messages in thread
From: Gregory Price @ 2026-09-26 10:51 UTC (permalink / raw)
To: linux-mm
Cc: linux-kernel, kernel-team, akpm, liam, ljs, david, vbabka, jannh,
gourry, ziy, joshua.hahnjy, rakie.kim, ying.huang, peterx, jgg,
sashiko-bot, stable
mmap a VM_MIXEDMAP region whose ->huge_fault installs a PMD through
vmf_insert_pfn_pmd() - mshv_vtl_low does this, and needs CAP_SYS_ADMIN
to open - then:
madvise(p, PMD_SIZE, MADV_PAGEOUT);
With a stand-in module for the driver:
BUG: unable to handle page fault for address: fffff587c0000008
RIP: 0010:madvise_cold_or_pageout_pte_range+0x410/0x9b0
walk_pgd_range+0x52b/0xaf0
__walk_page_range+0x6a/0x1d0
walk_page_range_vma_unsafe+0x8e/0x120
madvise_pageout+0xb2/0x180
madvise_vma_behavior+0x46b/0xa90
do_madvise+0x108/0x190
__x64_sys_madvise+0x26/0x30
Nothing validates the pfn on the way in:
can_madv_lru_vma() rejects VM_PFNMAP, but not VM_MIXEDMAP
can_fault() *pfn = vmf->pgoff & ~(mask >> PAGE_SHIFT);
vmf_insert_pfn_pmd() no pfn_valid() check
pmd_folio() pfn_to_page() -> unpopulated vmemmap
Even with a valid pfn the path is wrong. The mapping carries no rmap, so
folio_maybe_mapped_shared() sees mapcount 0, and the walker goes on to
folio_deactivate(), or folio_isolate_lru() plus reclaim_pages(), against a
folio this mapping does not own.
Use vm_normal_folio_pmd() and skip on NULL, as the PTE half of this same
walker already does with vm_normal_folio(). This also filters the huge
zero PMD, so its separate check is no longer needed.
Fixes: 3c8e44c9b369 ("mm: mark special bits for huge pfn mappings when inject")
Reported-by: sashiko-bot <sashiko-bot@kernel.org>
Closes: https://sashiko.dev/#/patchset/20260817220810.1175596-1-gourry%40gourry.net
Cc: stable@vger.kernel.org # v6.19+
Assisted-by: LLM
Signed-off-by: Gregory Price (Meta) <gourry@gourry.net>
Reviewed-by: Lorenzo Stoakes (ARM) <ljs@kernel.org>
Acked-by: David Hildenbrand (Arm) <david@kernel.org>
Reviewed-by: Zi Yan <ziy@nvidia.com>
---
mm/madvise.c | 7 +++----
1 file changed, 3 insertions(+), 4 deletions(-)
diff --git a/mm/madvise.c b/mm/madvise.c
index 61e15d9502ca..85ab9bc76e06 100644
--- a/mm/madvise.c
+++ b/mm/madvise.c
@@ -395,16 +395,15 @@ static int madvise_cold_or_pageout_pte_range(pmd_t *pmd,
return 0;
orig_pmd = *pmd;
- if (is_huge_zero_pmd(orig_pmd))
- goto huge_unlock;
-
if (unlikely(!pmd_present(orig_pmd))) {
VM_WARN_ON_ONCE(!pmd_is_migration_entry(orig_pmd) &&
!pmd_is_device_private_entry(orig_pmd));
goto huge_unlock;
}
- folio = pmd_folio(orig_pmd);
+ folio = vm_normal_folio_pmd(vma, addr, orig_pmd);
+ if (!folio)
+ goto huge_unlock;
if (folio_is_zone_device(folio))
goto huge_unlock;
--
2.55.0
^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [PATCH v3 0/2] mm: stop calling pmd_folio() on special PMDs
2026-09-26 10:51 [PATCH v3 0/2] mm: stop calling pmd_folio() on special PMDs Gregory Price
2026-09-26 10:51 ` [PATCH v3 1/2] mm/mempolicy: use vm_normal_folio_pmd() in queue_folios_pmd() Gregory Price
2026-09-26 10:51 ` [PATCH v3 2/2] mm/madvise: use vm_normal_folio_pmd() in cold/pageout PMD range Gregory Price
@ 2026-09-27 21:45 ` Andrew Morton
2 siblings, 0 replies; 4+ messages in thread
From: Andrew Morton @ 2026-09-27 21:45 UTC (permalink / raw)
To: Gregory Price
Cc: linux-mm, linux-kernel, kernel-team, liam, ljs, david, vbabka,
jannh, ziy, joshua.hahnjy, rakie.kim, ying.huang, peterx, jgg,
sashiko-bot
On Sat, 26 Sep 2026 06:51:08 -0400 Gregory Price <gourry@gourry.net> wrote:
> Two page table walkers resolve the folio behind a PMD with pmd_folio(),
> which is only valid for a PMD mapping a refcounted struct page:
>
> madvise_cold_or_pageout_pte_range() mm/madvise.c
> queue_folios_pmd() mm/mempolicy.c
>
> vmf_insert_pfn_pmd() installs special PMDs holding a raw pfn that need not
> have a memmap entry at all. Both walkers can reach one and fault on the
> first folio field read. The PTE halves of both already use
> vm_normal_folio(); these two patches make the PMD halves match.
>
> The four callers of vmf_insert_pfn_pmd(), and which walker each reaches:
>
> drivers/vfio/pci/vfio_pci_core.c VM_PFNMAP mempolicy
> drivers/gpu/drm/drm_gem_shmem_helper.c VM_PFNMAP mempolicy
> drivers/gpu/drm/panthor/panthor_gem.c VM_PFNMAP mempolicy
> drivers/hv/mshv_vtl_main.c VM_MIXEDMAP both
>
> can_madv_lru_vma() rejects VM_PFNMAP, so only mshv_vtl_low reaches the
> madvise walker, and that needs CAP_SYS_ADMIN. queue_pages_walk_ops
> supplies its own ->test_walk, so walk_page_test()'s generic VM_PFNMAP skip
> never runs and vfio-pci is reachable by any process holding the device fd.
> Hence the different stable tags.
>
> One behaviour change: mbind(MPOL_MF_STRICT) over a PMD mapped VM_PFNMAP
> region now returns 0 rather than -EIO. The PTE loop already returned 0
> there. drm_gem_shmem and panthor are where this is observable, since they
> PMD map pages that do have a memmap entry and so never faulted.
Thanks, I've updated mm-unstable to this version.
> v3: improvements from Lorenzo
Here's how v3 altered mm.git:
mm/mempolicy.c | 9 ++-------
1 file changed, 2 insertions(+), 7 deletions(-)
--- a/mm/mempolicy.c~b
+++ a/mm/mempolicy.c
@@ -668,7 +668,7 @@ static inline bool queue_folio_required(
}
static void queue_folios_pmd(pmd_t *pmd, unsigned long addr,
- struct mm_walk *walk)
+ struct mm_walk *walk)
{
struct folio *folio;
struct queue_pages *qp = walk->private;
@@ -680,12 +680,7 @@ static void queue_folios_pmd(pmd_t *pmd,
return;
}
folio = vm_normal_folio_pmd(walk->vma, addr, pmdval);
- if (!folio) {
- if (is_huge_zero_pmd(pmdval))
- walk->action = ACTION_CONTINUE;
- return;
- }
- if (folio_is_zone_device(folio))
+ if (!folio || folio_is_zone_device(folio))
return;
if (!queue_folio_required(folio, qp))
return;
_
^ permalink raw reply [flat|nested] 4+ messages in thread
end of thread, other threads:[~2026-09-27 21:45 UTC | newest]
Thread overview: 4+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-26 10:51 [PATCH v3 0/2] mm: stop calling pmd_folio() on special PMDs Gregory Price
2026-09-26 10:51 ` [PATCH v3 1/2] mm/mempolicy: use vm_normal_folio_pmd() in queue_folios_pmd() Gregory Price
2026-09-26 10:51 ` [PATCH v3 2/2] mm/madvise: use vm_normal_folio_pmd() in cold/pageout PMD range Gregory Price
2026-09-27 21:45 ` [PATCH v3 0/2] mm: stop calling pmd_folio() on special PMDs Andrew Morton
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®