mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [syzbot] [kernel?] WARNING in __ioremap_caller
@ 2026-02-22 11:33 syzbot
  2026-09-04 22:02 ` syzbot
  2026-10-01 15:25 ` [PATCH] mm: don't ioremap COWed anon pages in generic_access_phys() Nguyen Ngoc Thang
  0 siblings, 2 replies; 13+ messages in thread
From: syzbot @ 2026-02-22 11:33 UTC (permalink / raw)
  To: bp, dave.hansen, hpa, linux-kernel, luto, mingo, peterz,
	syzkaller-bugs, tglx, x86

Hello,

syzbot found the following issue on:

HEAD commit:    2961f841b025 Merge tag 'turbostat-2026.02.14' of git://git..
git tree:       upstream
console output: https://syzkaller.appspot.com/x/log.txt?x=17674594580000
kernel config:  https://syzkaller.appspot.com/x/.config?x=6259cfbe2d15cac4
dashboard link: https://syzkaller.appspot.com/bug?extid=49b1021becba70c1f3f6
compiler:       gcc (Debian 14.2.0-19) 14.2.0, GNU ld (GNU Binutils for Debian) 2.44

Unfortunately, I don't have any reproducer for this issue yet.

Downloadable assets:
disk image: https://storage.googleapis.com/syzbot-assets/4db266f4e980/disk-2961f841.raw.xz
vmlinux: https://storage.googleapis.com/syzbot-assets/9ec134f45048/vmlinux-2961f841.xz
kernel image: https://storage.googleapis.com/syzbot-assets/5ba4f405d712/bzImage-2961f841.xz

IMPORTANT: if you fix the issue, please add the following tag to the commit:
Reported-by: syzbot+49b1021becba70c1f3f6@syzkaller.appspotmail.com

ioremap: invalid physical address a729678abbd17117
------------[ cut here ]------------
1
WARNING: arch/x86/mm/ioremap.c:206 at __ioremap_caller.isra.0.cold+0x59/0xa4 arch/x86/mm/ioremap.c:206, CPU#0: syz.1.4191/26162
Modules linked in:
CPU: 0 UID: 0 PID: 26162 Comm: syz.1.4191 Tainted: G     U       L      syzkaller #0 PREEMPT(full) 
Tainted: [U]=USER, [L]=SOFTLOCKUP
Hardware name: Google Google Compute Engine/Google Compute Engine, BIOS Google 02/12/2026
RIP: 0010:__ioremap_caller.isra.0.cold+0x59/0xa4 arch/x86/mm/ioremap.c:206
Code: 48 8b 34 24 48 c7 c7 e0 fc ab 8b e8 30 1d 01 00 e9 61 c2 97 00 e8 76 ca e6 00 4c 89 ee 48 c7 c7 40 fb ab 8b e8 17 1d 01 00 90 <0f> 0b 90 e9 41 c2 97 00 e8 59 ca e6 00 41 0f b6 d7 4c 89 ee 48 c7
RSP: 0018:ffffc90004d1f718 EFLAGS: 00010286
RAX: 0000000000000032 RBX: 1ffff920009a3ee7 RCX: 0000000000000000
RDX: 0000000000000032 RSI: ffffffff81e6f929 RDI: fffff520009a3ed4
RBP: 0000000000000040 R08: 0000000000000005 R09: 0000000000000000
R10: 0000000080000000 R11: 0000000000038c68 R12: a729678abbd17156
R13: a729678abbd17117 R14: 0000000000000000 R15: 0000000000029ca5
FS:  00007fc5be1016c0(0000) GS:ffff888124351000(0000) knlGS:0000000000000000
CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
CR2: 0000001b32306ff8 CR3: 00000000831e0000 CR4: 00000000003526f0
Call Trace:
 <TASK>
 ioremap_cache arch/x86/mm/ioremap.c:436 [inline]
 arch_memremap_wb+0x23/0x40 arch/x86/mm/ioremap.c:508
 memremap+0x1cb/0x7d0 kernel/iomem.c:95
 pcibios_device_add+0x101/0x600 arch/x86/pci/common.c:652
 pci_device_add+0xd4f/0x1810 drivers/pci/probe.c:2761
 pci_scan_single_device drivers/pci/probe.c:2793 [inline]
 pci_scan_single_device+0x1d0/0x240 drivers/pci/probe.c:2779
 pci_scan_slot+0x1c9/0x7c0 drivers/pci/probe.c:2876
 pci_scan_child_bus_extend+0x6b/0x7b0 drivers/pci/probe.c:3095
 pci_scan_child_bus drivers/pci/probe.c:3208 [inline]
 pci_rescan_bus+0x18/0x40 drivers/pci/probe.c:3499
 rescan_store+0xfb/0x130 drivers/pci/pci-sysfs.c:472
 bus_attr_store+0x74/0xb0 drivers/base/bus.c:172
 sysfs_kf_write+0xf2/0x150 fs/sysfs/file.c:142
 kernfs_fop_write_iter+0x3e0/0x5f0 fs/kernfs/file.c:352
 new_sync_write fs/read_write.c:595 [inline]
 vfs_write+0x6ac/0x1070 fs/read_write.c:688
 ksys_write+0x12a/0x250 fs/read_write.c:740
 do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
 do_syscall_64+0x106/0xf80 arch/x86/entry/syscall_64.c:94
 entry_SYSCALL_64_after_hwframe+0x77/0x7f
RIP: 0033:0x7fc5bd19c629
Code: ff c3 66 2e 0f 1f 84 00 00 00 00 00 0f 1f 44 00 00 48 89 f8 48 89 f7 48 89 d6 48 89 ca 4d 89 c2 4d 89 c8 4c 8b 4c 24 08 0f 05 <48> 3d 01 f0 ff ff 73 01 c3 48 c7 c1 e8 ff ff ff f7 d8 64 89 01 48
RSP: 002b:00007fc5be101028 EFLAGS: 00000246 ORIG_RAX: 0000000000000001
RAX: ffffffffffffffda RBX: 00007fc5bd415fa0 RCX: 00007fc5bd19c629
RDX: 0000000000000001 RSI: 0000200000000200 RDI: 0000000000000003
RBP: 00007fc5bd232b39 R08: 0000000000000000 R09: 0000000000000000
R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000000
R13: 00007fc5bd416038 R14: 00007fc5bd415fa0 R15: 00007ffc328057a8
 </TASK>


---
This report is generated by a bot. It may contain errors.
See https://goo.gl/tpsmEJ for more information about syzbot.
syzbot engineers can be reached at syzkaller@googlegroups.com.

syzbot will keep track of this issue. See:
https://goo.gl/tpsmEJ#status for how to communicate with syzbot.

If the report is already addressed, let syzbot know by replying with:
#syz fix: exact-commit-title

If you want to overwrite report's subsystems, reply with:
#syz set subsystems: new-subsystem
(See the list of subsystem names on the web dashboard)

If the report is a duplicate of another one, reply with:
#syz dup: exact-subject-of-another-report

If you want to undo deduplication, reply with:
#syz undup

^ permalink raw reply	[flat|nested] 13+ messages in thread

* Re: [syzbot] [kernel?] WARNING in __ioremap_caller
  2026-02-22 11:33 [syzbot] [kernel?] WARNING in __ioremap_caller syzbot
@ 2026-09-04 22:02 ` syzbot
  2026-09-04 23:20   ` Dave Hansen
  2026-10-01 15:25 ` [PATCH] mm: don't ioremap COWed anon pages in generic_access_phys() Nguyen Ngoc Thang
  1 sibling, 1 reply; 13+ messages in thread
From: syzbot @ 2026-09-04 22:02 UTC (permalink / raw)
  To: bp, dave.hansen, hpa, linux-kernel, luto, mingo, peterz,
	syzkaller-bugs, tglx, x86

syzbot has found a reproducer for the following issue on:

HEAD commit:    bc35965f6940 Merge tag 'mm-hotfixes-stable-2026-09-03-17-4..
git tree:       upstream
console output: https://syzkaller.appspot.com/x/log.txt?x=164e88f9580000
kernel config:  https://syzkaller.appspot.com/x/.config?x=85bc5cc2fc7394d9
dashboard link: https://syzkaller.appspot.com/bug?extid=49b1021becba70c1f3f6
compiler:       gcc (Debian 14.2.0-19) 14.2.0, GNU ld (GNU Binutils for Debian) 2.44
userspace arch: i386
syz repro:      https://syzkaller.appspot.com/x/repro.syz?x=16bca8f9580000
C reproducer:   https://syzkaller.appspot.com/x/repro.c?x=16a09215580000

Downloadable assets:
disk image (non-bootable): https://storage.googleapis.com/syzbot-assets/d900f083ada3/non_bootable_disk-bc35965f.raw.xz
vmlinux: https://storage.googleapis.com/syzbot-assets/9eaaad19cc52/vmlinux-bc35965f.xz
kernel image: https://storage.googleapis.com/syzbot-assets/b6466cfeb6f0/bzImage-bc35965f.xz

IMPORTANT: if you fix the issue, please add the following tag to the commit:
Reported-by: syzbot+49b1021becba70c1f3f6@syzkaller.appspotmail.com

------------[ cut here ]------------
ioremap on RAM at 0x000000001bf67000 - 0x000000001bf67fff
WARNING: arch/x86/mm/ioremap.c:216 at __ioremap_caller.isra.0+0x4c2/0x5f0 arch/x86/mm/ioremap.c:216, CPU#2: syz.0.17/5915
Modules linked in:
CPU: 2 UID: 0 PID: 5915 Comm: syz.0.17 Not tainted syzkaller #0 PREEMPT(full) 
Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 1.16.3-debian-1.16.3-2 04/01/2014
RIP: 0010:__ioremap_caller.isra.0+0x4d3/0x5f0 arch/x86/mm/ioremap.c:216
Code: 41 5c 41 5d 41 5e 41 5f c3 cc cc cc cc e8 45 cb 50 00 48 8d 3d 2e 85 8d 0f 48 8d 84 24 c0 00 00 00 48 8d 54 24 60 48 8d 70 c0 <67> 48 0f b9 3a eb 89 41 83 ff 04 75 32 e8 1b cb 50 00 bf 04 00 00
RSP: 0018:ffffc900036478a0 EFLAGS: 00010293
RAX: ffffc90003647960 RBX: 1ffff920006c8f18 RCX: 0000000000000000
RDX: ffffc90003647900 RSI: ffffc90003647920 RDI: ffffffff9147f620
RBP: 0000000000001000 R08: 0000000000000005 R09: 0000000000000000
R10: 0000000000000001 R11: 0000000000000001 R12: 0000000000000001
R13: 000000001bf67000 R14: 0000000000000002 R15: 000000001bf67000
FS:  0000000000000000(0000) GS:ffff888096b64000(0063) knlGS:0000000056c88480
CS:  0010 DS: 002b ES: 002b CR0: 0000000080050033
CR2: 0000000000000018 CR3: 000000005eb62000 CR4: 0000000000352ef0
Call Trace:
 <TASK>
 generic_access_phys+0x130/0x4d0 mm/memory.c:7178
 kernfs_vma_access+0x1ce/0x280 fs/kernfs/file.c:437
 __access_remote_vm+0x58f/0x890 mm/memory.c:7256
 mem_rw+0x2a1/0x670 fs/proc/base.c:912
 do_loop_readv_writev fs/read_write.c:848 [inline]
 do_loop_readv_writev fs/read_write.c:836 [inline]
 vfs_readv+0x5d8/0x8d0 fs/read_write.c:1021
 do_readv+0x13e/0x340 fs/read_write.c:1081
 do_syscall_32_irqs_on arch/x86/entry/syscall_32.c:79 [inline]
 __do_fast_syscall_32+0x13a/0x8b0 arch/x86/entry/syscall_32.c:291
 do_fast_syscall_32+0x32/0x70 arch/x86/entry/syscall_32.c:316
 entry_SYSENTER_compat_after_hwframe+0x84/0x8e
RIP: 0023:0xf7f14fec
Code: Unable to access opcode bytes at 0xf7f14fc2.
RSP: 002b:00000000ffc02c5c EFLAGS: 00000296 ORIG_RAX: 0000000000000091
RAX: ffffffffffffffda RBX: 0000000000000003 RCX: 0000000080000100
RDX: 0000000000000002 RSI: 0000000000000000 RDI: 0000000000000000
RBP: 0000000000000000 R08: 0000000000000000 R09: 0000000000000000
R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000000
R13: 0000000000000000 R14: 0000000000000000 R15: 0000000000000000
 </TASK>
----------------
Code disassembly (best guess):
   0:	41 5c                	pop    %r12
   2:	41 5d                	pop    %r13
   4:	41 5e                	pop    %r14
   6:	41 5f                	pop    %r15
   8:	c3                   	ret
   9:	cc                   	int3
   a:	cc                   	int3
   b:	cc                   	int3
   c:	cc                   	int3
   d:	e8 45 cb 50 00       	call   0x50cb57
  12:	48 8d 3d 2e 85 8d 0f 	lea    0xf8d852e(%rip),%rdi        # 0xf8d8547
  19:	48 8d 84 24 c0 00 00 	lea    0xc0(%rsp),%rax
  20:	00
  21:	48 8d 54 24 60       	lea    0x60(%rsp),%rdx
  26:	48 8d 70 c0          	lea    -0x40(%rax),%rsi
* 2a:	67 48 0f b9 3a       	ud1    (%edx),%rdi <-- trapping instruction
  2f:	eb 89                	jmp    0xffffffba
  31:	41 83 ff 04          	cmp    $0x4,%r15d
  35:	75 32                	jne    0x69
  37:	e8 1b cb 50 00       	call   0x50cb57
  3c:	bf                   	.byte 0xbf
  3d:	04 00                	add    $0x0,%al


---
If you want syzbot to run the reproducer, reply with:
#syz test: git://repo/address.git branch-or-commit-hash
If you attach or paste a git patch, syzbot will apply it before testing.

^ permalink raw reply	[flat|nested] 13+ messages in thread

* Re: [syzbot] [kernel?] WARNING in __ioremap_caller
  2026-09-04 22:02 ` syzbot
@ 2026-09-04 23:20   ` Dave Hansen
  0 siblings, 0 replies; 13+ messages in thread
From: Dave Hansen @ 2026-09-04 23:20 UTC (permalink / raw)
  To: syzbot, bp, dave.hansen, hpa, linux-kernel, luto, mingo, peterz,
	syzkaller-bugs, tglx, x86

On 9/4/26 15:02, syzbot wrote:
> C reproducer:   https://syzkaller.appspot.com/x/repro.c?x=16a09215580000

Heh, that's a fun reproducer as usual. Looks like it's mmap()'ing some
PCI device resource the accessing the mmap() with /proc/$pid/mem. The
thing that's weird is how a PCI device resource ends up pointing at
system RAM.

/proc/iomem might be enlightening as would 'lspci -vvv'. But it makes me
wonder if the PCI device enumeration inside the VM is broken or something.

^ permalink raw reply	[flat|nested] 13+ messages in thread

* [PATCH] mm: don't ioremap COWed anon pages in generic_access_phys()
  2026-02-22 11:33 [syzbot] [kernel?] WARNING in __ioremap_caller syzbot
  2026-09-04 22:02 ` syzbot
@ 2026-10-01 15:25 ` Nguyen Ngoc Thang
  2026-10-01 20:22   ` David Hildenbrand (Arm)
  1 sibling, 1 reply; 13+ messages in thread
From: Nguyen Ngoc Thang @ 2026-10-01 15:25 UTC (permalink / raw)
  To: akpm, david
  Cc: Nguyen Ngoc Thang, ljs, liam, vbabka, rppt, surenb, mhocko,
	peterx, dave.hansen, linux-mm, linux-kernel,
	syzbot+49b1021becba70c1f3f6

A MAP_PRIVATE mapping of iomem (e.g. a PCI sysfs resourceN file) is a
COW pfnmap: a write fault replaces the pfn with an anonymous page, which
remap_pfn_range() allows by keeping vm_pgoff equal to the base pfn.

generic_access_phys() does not tell those COWed pages apart from the
original pfns and ioremaps whatever the PTE points to. Reading such an
address via /proc/pid/mem or ptrace then ioremaps RAM:

  ioremap on RAM at 0x0000000045623000 - 0x0000000045623fff
  WARNING: arch/x86/mm/ioremap.c:216 at __ioremap_caller.isra.0+0x4c2/0x5f0
  Call Trace:
   generic_access_phys+0x130/0x4d0 mm/memory.c:7178
   kernfs_vma_access+0x1ce/0x280 fs/kernfs/file.c:437
   __access_remote_vm+0x58f/0x890 mm/memory.c:7256
   mem_rw+0x2a1/0x670 fs/proc/base.c:912

Reject COWed pages using the same linearity rule vm_normal_page() uses.
All generic_access_phys() users set up the mapping with a full-vma
(io_)remap_pfn_range(), so the rule holds for them. The access already
fails on x86 and arm64, whose ioremap refuses RAM; elsewhere it no longer
reads the anon page through a device-memory alias.

Fixes: 28b2ee20c7cb ("access_process_vm device memory infrastructure")
Reported-by: syzbot+49b1021becba70c1f3f6@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=49b1021becba70c1f3f6
Signed-off-by: Nguyen Ngoc Thang <ngocthang2710.1999@gmail.com>
---
This answers Dave's question on the syzbot thread of how a PCI BAR can
point at RAM: it doesn't. The repro mmap()s resource1 MAP_PRIVATE with
PROT_WRITE at address 0, then (via syz_ublk_add_dev() with NULL args)
writes to address 0. That write COWs the BAR page into an anonymous
page, and the later /proc/self/mem read hands its pfn to ioremap.

Forbidding MAP_PRIVATE on BARs would not cover /dev/mem, uio, cdx or
dfl-afu, which share this ->access hook, and private pfnmap COW is
deliberately supported (see get_remap_pgoff()).

Tested in QEMU (q35, virtio-net at 00:03.0) with a small program that
maps resource1 shared and private, COWs the private page and reads both
via pread(/proc/self/mem):

                  before              after
  shared          4, 0xfee01000       4, 0xfee01000
  private         4, 0xfee01000       4, 0xfee01000
  private-cowed   -1 EIO + WARN       -1 EIO, no WARN

(pread() sees EIO either way; internally generic_access_phys() went from
ioremap WARN + -ENOMEM to -EINVAL.)

plus 4 x 500 parallel runs on the patched kernel, no warnings.

Note: the syzbot report also groups an unrelated signature under the
same title (pcibios_device_add() -> memremap() of setup_data on PCI
rescan); this patch does not address that one.

 mm/memory.c | 11 +++++++++++
 1 file changed, 11 insertions(+)

diff --git a/mm/memory.c b/mm/memory.c
index 8b0c2c735d3d..8551b144028d 100644
--- a/mm/memory.c
+++ b/mm/memory.c
@@ -7141,6 +7141,14 @@ void follow_pfnmap_end(struct follow_pfnmap_args *args)
 EXPORT_SYMBOL_GPL(follow_pfnmap_end);
 
 #ifdef CONFIG_HAVE_IOREMAP_PROT
+/* A write fault on a private pfnmap replaces the pfn with an anon page. */
+static bool pfnmap_pfn_is_cowed(struct vm_area_struct *vma, unsigned long addr,
+				unsigned long pfn)
+{
+	return (vma->vm_flags & VM_PFNMAP) && vma_is_cow_mapping(vma) &&
+	       pfn != linear_page_index(vma, addr);
+}
+
 /**
  * generic_access_phys - generic implementation for iomem mmap access
  * @vma: the vma to access
@@ -7175,6 +7183,9 @@ int generic_access_phys(struct vm_area_struct *vma, unsigned long addr,
 	if ((write & FOLL_WRITE) && !writable)
 		return -EINVAL;
 
+	if (pfnmap_pfn_is_cowed(vma, addr, phys_addr >> PAGE_SHIFT))
+		return -EINVAL;
+
 	maddr = ioremap_prot(phys_addr, PAGE_ALIGN(len + offset), prot);
 	if (!maddr)
 		return -ENOMEM;
-- 
2.43.0


^ permalink raw reply	[flat|nested] 13+ messages in thread

* Re: [PATCH] mm: don't ioremap COWed anon pages in generic_access_phys()
  2026-10-01 15:25 ` [PATCH] mm: don't ioremap COWed anon pages in generic_access_phys() Nguyen Ngoc Thang
@ 2026-10-01 20:22   ` David Hildenbrand (Arm)
  2026-10-02  8:38     ` Lorenzo Stoakes (ARM)
  0 siblings, 1 reply; 13+ messages in thread
From: David Hildenbrand (Arm) @ 2026-10-01 20:22 UTC (permalink / raw)
  To: Nguyen Ngoc Thang, akpm
  Cc: ljs, liam, vbabka, rppt, surenb, mhocko, peterx, dave.hansen,
	linux-mm, linux-kernel, syzbot+49b1021becba70c1f3f6

On 10/1/26 17:25, Nguyen Ngoc Thang wrote:
> A MAP_PRIVATE mapping of iomem (e.g. a PCI sysfs resourceN file) is a
> COW pfnmap: a write fault replaces the pfn with an anonymous page, which
> remap_pfn_range() allows by keeping vm_pgoff equal to the base pfn.
> 
> generic_access_phys() does not tell those COWed pages apart from the
> original pfns and ioremaps whatever the PTE points to. Reading such an
> address via /proc/pid/mem or ptrace then ioremaps RAM:
> 
>   ioremap on RAM at 0x0000000045623000 - 0x0000000045623fff
>   WARNING: arch/x86/mm/ioremap.c:216 at __ioremap_caller.isra.0+0x4c2/0x5f0
>   Call Trace:
>    generic_access_phys+0x130/0x4d0 mm/memory.c:7178
>    kernfs_vma_access+0x1ce/0x280 fs/kernfs/file.c:437
>    __access_remote_vm+0x58f/0x890 mm/memory.c:7256
>    mem_rw+0x2a1/0x670 fs/proc/base.c:912
> 

Ack, apart from that, nothing bad should happen.

> Reject COWed pages using the same linearity rule vm_normal_page() uses.

Hm I don't enjoy that.

What about letting follow_pfnmap_start() just perform a vm_normal_page()
etc and indicate that information to the caller in the
follow_pfnmap_args() ?

We kind-of have that information in the form of follow_pfnmap_args.special
... but it's not available on all architectures. And it shouldn't exist.

I was thinking a while ago about disallowing follow_pfnmap_start() entirely
on anon folios, but the VM_IO thingy made me assume that there are some
odd users (kvm/vfio) that actually need CoWed folios.

Anyhow, this is what I think we should do for now:


From 0b75f30b137612f18cdee28c8c0dfa96870868c2 Mon Sep 17 00:00:00 2001
From: "David Hildenbrand (Arm)" <david@kernel.org>
Date: Thu, 1 Oct 2026 22:19:57 +0200
Subject: [PATCH] tmp

Signed-off-by: David Hildenbrand (Arm) <david@kernel.org>
---
 include/linux/mm.h |  4 +++-
 mm/memory.c        | 48 ++++++++++++++++++++++++++++++++++++++++------
 2 files changed, 45 insertions(+), 7 deletions(-)

diff --git a/include/linux/mm.h b/include/linux/mm.h
index c49ef99b4413..b90e547797a9 100644
--- a/include/linux/mm.h
+++ b/include/linux/mm.h
@@ -3193,6 +3193,8 @@ struct folio *vm_normal_folio_pmd(struct vm_area_struct *vma,
 				  unsigned long addr, pmd_t pmd);
 struct page *vm_normal_page_pmd(struct vm_area_struct *vma, unsigned long addr,
 				pmd_t pmd);
+struct folio *vm_normal_folio_pud(struct vm_area_struct *vma,
+		unsigned long addr, pud_t pud);
 struct page *vm_normal_page_pud(struct vm_area_struct *vma, unsigned long addr,
 		pud_t pud);
 
@@ -3245,7 +3247,7 @@ struct follow_pfnmap_args {
 	unsigned long addr_mask;
 	pgprot_t pgprot;
 	bool writable;
-	bool special;
+	bool cow;
 };
 int follow_pfnmap_start(struct follow_pfnmap_args *args);
 void follow_pfnmap_end(struct follow_pfnmap_args *args);
diff --git a/mm/memory.c b/mm/memory.c
index 926276d41920..764b79204ab7 100644
--- a/mm/memory.c
+++ b/mm/memory.c
@@ -843,7 +843,7 @@ struct page *vm_normal_page_pmd(struct vm_area_struct *vma, unsigned long addr,
  * @addr: The address where the @pmd is mapped.
  * @pmd: The PMD.
  *
- * Get the "struct folio" associated with a PTE. See __vm_normal_page()
+ * Get the "struct folio" associated with a PMD. See __vm_normal_page()
  * for details on "normal" and "special" mappings.
  *
  * Return: Returns the "struct folio" if this is a "normal" mapping. Returns
@@ -879,6 +879,28 @@ struct page *vm_normal_page_pud(struct vm_area_struct *vma,
 	return __vm_normal_page(vma, addr, pud_pfn(pud), pud_special(pud),
 				&entry, sizeof(entry), PGTABLE_LEVEL_PUD);
 }
+
+/**
+ * vm_normal_folio_pud() - Get the "struct folio" associated with a PUD
+ * @vma: The VMA mapping the @pud.
+ * @addr: The address where the @pud is mapped.
+ * @pud: The PUD.
+ *
+ * Get the "struct folio" associated with a PUD. See __vm_normal_page()
+ * for details on "normal" and "special" mappings.
+ *
+ * Return: Returns the "struct folio" if this is a "normal" mapping. Returns
+ *	   NULL if this is a "special" mapping.
+ */
+struct folio *vm_normal_folio_pud(struct vm_area_struct *vma,
+				  unsigned long addr, pud_t pud)
+{
+	struct page *page = vm_normal_page_pud(vma, addr, pud);
+
+	if (page)
+		return page_folio(page);
+	return NULL;
+}
 #endif
 
 /**
@@ -6947,7 +6969,7 @@ static inline void pfnmap_args_setup(struct follow_pfnmap_args *args,
 				     spinlock_t *lock, pte_t *ptep,
 				     pgprot_t pgprot, unsigned long pfn_base,
 				     unsigned long addr_mask, bool writable,
-				     bool special)
+				     struct folio *folio)
 {
 	args->lock = lock;
 	args->ptep = ptep;
@@ -6955,7 +6977,7 @@ static inline void pfnmap_args_setup(struct follow_pfnmap_args *args,
 	args->addr_mask = addr_mask;
 	args->pgprot = pgprot;
 	args->writable = writable;
-	args->special = special;
+	args->cow = folio && folio_test_anon(folio);
 }
 
 static inline void pfnmap_lockdep_assert(struct vm_area_struct *vma)
@@ -7008,6 +7030,7 @@ int follow_pfnmap_start(struct follow_pfnmap_args *args)
 	struct vm_area_struct *vma = args->vma;
 	unsigned long address = args->address;
 	struct mm_struct *mm = vma->vm_mm;
+	struct folio *folio = NULL;
 	spinlock_t *lock;
 	pgd_t *pgdp;
 	p4d_t *p4dp, p4d;
@@ -7047,9 +7070,12 @@ int follow_pfnmap_start(struct follow_pfnmap_args *args)
 			spin_unlock(lock);
 			goto retry;
 		}
+
+		if (vma_is_cow_mapping(vma))
+			folio = vm_normal_folio_pud(vma, address, pud);
 		pfnmap_args_setup(args, lock, NULL, pud_pgprot(pud),
 				  pud_pfn(pud), PUD_MASK, pud_write(pud),
-				  pud_special(pud));
+				  folio);
 		return 0;
 	}
 
@@ -7068,9 +7094,12 @@ int follow_pfnmap_start(struct follow_pfnmap_args *args)
 			spin_unlock(lock);
 			goto retry;
 		}
+
+		if (vma_is_cow_mapping(vma))
+			folio = vm_normal_folio_pmd(vma, address, pmd);
 		pfnmap_args_setup(args, lock, NULL, pmd_pgprot(pmd),
 				  pmd_pfn(pmd), PMD_MASK, pmd_write(pmd),
-				  pmd_special(pmd));
+				  folio);
 		return 0;
 	}
 
@@ -7080,9 +7109,12 @@ int follow_pfnmap_start(struct follow_pfnmap_args *args)
 	pte = ptep_get(ptep);
 	if (!pte_present(pte))
 		goto unlock;
+
+	if (vma_is_cow_mapping(vma))
+		folio = vm_normal_folio(vma, address, pte);
 	pfnmap_args_setup(args, lock, ptep, pte_pgprot(pte),
 			  pte_pfn(pte), PAGE_MASK, pte_write(pte),
-			  pte_special(pte));
+			  folio);
 	return 0;
 unlock:
 	pte_unmap_unlock(ptep, lock);
@@ -7145,6 +7177,10 @@ int generic_access_phys(struct vm_area_struct *vma, unsigned long addr,
 	writable = args.writable;
 	follow_pfnmap_end(&args);
 
+	/* Reject CoW'ed anonymous folios. */
+	if (args.cow)
+		return -EINVAL;
+
 	if ((write & FOLL_WRITE) && !writable)
 		return -EINVAL;
 
-- 
2.43.0


-- 
Cheers,

David

^ permalink raw reply	[flat|nested] 13+ messages in thread

* Re: [PATCH] mm: don't ioremap COWed anon pages in generic_access_phys()
  2026-10-01 20:22   ` David Hildenbrand (Arm)
@ 2026-10-02  8:38     ` Lorenzo Stoakes (ARM)
  2026-10-02 10:02       ` David Hildenbrand (Arm)
  0 siblings, 1 reply; 13+ messages in thread
From: Lorenzo Stoakes (ARM) @ 2026-10-02  8:38 UTC (permalink / raw)
  To: David Hildenbrand (Arm)
  Cc: Nguyen Ngoc Thang, akpm, liam, vbabka, rppt, surenb, mhocko,
	peterx, dave.hansen, linux-mm, linux-kernel,
	syzbot+49b1021becba70c1f3f6

On Thu, Oct 01, 2026 at 10:22:29PM +0200, David Hildenbrand (Arm) wrote:
> On 10/1/26 17:25, Nguyen Ngoc Thang wrote:
> > A MAP_PRIVATE mapping of iomem (e.g. a PCI sysfs resourceN file) is a
> > COW pfnmap: a write fault replaces the pfn with an anonymous page, which
> > remap_pfn_range() allows by keeping vm_pgoff equal to the base pfn.
> >
> > generic_access_phys() does not tell those COWed pages apart from the
> > original pfns and ioremaps whatever the PTE points to. Reading such an
> > address via /proc/pid/mem or ptrace then ioremaps RAM:
> >
> >   ioremap on RAM at 0x0000000045623000 - 0x0000000045623fff
> >   WARNING: arch/x86/mm/ioremap.c:216 at __ioremap_caller.isra.0+0x4c2/0x5f0
> >   Call Trace:
> >    generic_access_phys+0x130/0x4d0 mm/memory.c:7178
> >    kernfs_vma_access+0x1ce/0x280 fs/kernfs/file.c:437
> >    __access_remote_vm+0x58f/0x890 mm/memory.c:7256
> >    mem_rw+0x2a1/0x670 fs/proc/base.c:912
> >
>

Sorry to say this looks schlopped.

This guy has sent 10 series across 6 subsystems over ~21 hrs:

https://lore.kernel.org/all/?q=f%3Angocthang2710.1999%40gmail.com

Nguyen - please do not flood the kernel with patches, and please use the
Assisted-by tag for generated content.

The original code you submitted is really not great even if the issue may
be valid.

So I'd say somebody from the core team should take over this if we want to
come up with a patch.

> Ack, apart from that, nothing bad should happen.
>
> > Reject COWed pages using the same linearity rule vm_normal_page() uses.
>
> Hm I don't enjoy that.
>
> What about letting follow_pfnmap_start() just perform a vm_normal_page()
> etc and indicate that information to the caller in the
> follow_pfnmap_args() ?
>
> We kind-of have that information in the form of follow_pfnmap_args.special
> ... but it's not available on all architectures. And it shouldn't exist.
>
> I was thinking a while ago about disallowing follow_pfnmap_start() entirely
> on anon folios, but the VM_IO thingy made me assume that there are some
> odd users (kvm/vfio) that actually need CoWed folios.

Ugh.

>
> Anyhow, this is what I think we should do for now:
>
>
> From 0b75f30b137612f18cdee28c8c0dfa96870868c2 Mon Sep 17 00:00:00 2001
> From: "David Hildenbrand (Arm)" <david@kernel.org>
> Date: Thu, 1 Oct 2026 22:19:57 +0200
> Subject: [PATCH] tmp
>
> Signed-off-by: David Hildenbrand (Arm) <david@kernel.org>

This looks reasonable but I hate that we have 'special' CoW overrides like
this :)

> ---
>  include/linux/mm.h |  4 +++-
>  mm/memory.c        | 48 ++++++++++++++++++++++++++++++++++++++++------
>  2 files changed, 45 insertions(+), 7 deletions(-)
>
> diff --git a/include/linux/mm.h b/include/linux/mm.h
> index c49ef99b4413..b90e547797a9 100644
> --- a/include/linux/mm.h
> +++ b/include/linux/mm.h
> @@ -3193,6 +3193,8 @@ struct folio *vm_normal_folio_pmd(struct vm_area_struct *vma,
>  				  unsigned long addr, pmd_t pmd);
>  struct page *vm_normal_page_pmd(struct vm_area_struct *vma, unsigned long addr,
>  				pmd_t pmd);
> +struct folio *vm_normal_folio_pud(struct vm_area_struct *vma,
> +		unsigned long addr, pud_t pud);

Hmm, if not defined before why would this need a new PUD handler? Do we
even have PUD-leaf PFN mappings?

>  struct page *vm_normal_page_pud(struct vm_area_struct *vma, unsigned long addr,
>  		pud_t pud);
>
> @@ -3245,7 +3247,7 @@ struct follow_pfnmap_args {
>  	unsigned long addr_mask;
>  	pgprot_t pgprot;
>  	bool writable;
> -	bool special;
> +	bool cow;
>  };
>  int follow_pfnmap_start(struct follow_pfnmap_args *args);
>  void follow_pfnmap_end(struct follow_pfnmap_args *args);
> diff --git a/mm/memory.c b/mm/memory.c
> index 926276d41920..764b79204ab7 100644
> --- a/mm/memory.c
> +++ b/mm/memory.c
> @@ -843,7 +843,7 @@ struct page *vm_normal_page_pmd(struct vm_area_struct *vma, unsigned long addr,
>   * @addr: The address where the @pmd is mapped.
>   * @pmd: The PMD.
>   *
> - * Get the "struct folio" associated with a PTE. See __vm_normal_page()
> + * Get the "struct folio" associated with a PMD. See __vm_normal_page()
>   * for details on "normal" and "special" mappings.
>   *
>   * Return: Returns the "struct folio" if this is a "normal" mapping. Returns
> @@ -879,6 +879,28 @@ struct page *vm_normal_page_pud(struct vm_area_struct *vma,
>  	return __vm_normal_page(vma, addr, pud_pfn(pud), pud_special(pud),
>  				&entry, sizeof(entry), PGTABLE_LEVEL_PUD);
>  }
> +
> +/**
> + * vm_normal_folio_pud() - Get the "struct folio" associated with a PUD
> + * @vma: The VMA mapping the @pud.
> + * @addr: The address where the @pud is mapped.
> + * @pud: The PUD.
> + *
> + * Get the "struct folio" associated with a PUD. See __vm_normal_page()
> + * for details on "normal" and "special" mappings.
> + *
> + * Return: Returns the "struct folio" if this is a "normal" mapping. Returns
> + *	   NULL if this is a "special" mapping.
> + */
> +struct folio *vm_normal_folio_pud(struct vm_area_struct *vma,
> +				  unsigned long addr, pud_t pud)
> +{
> +	struct page *page = vm_normal_page_pud(vma, addr, pud);
> +
> +	if (page)
> +		return page_folio(page);
> +	return NULL;
> +}
>  #endif
>
>  /**
> @@ -6947,7 +6969,7 @@ static inline void pfnmap_args_setup(struct follow_pfnmap_args *args,
>  				     spinlock_t *lock, pte_t *ptep,
>  				     pgprot_t pgprot, unsigned long pfn_base,
>  				     unsigned long addr_mask, bool writable,
> -				     bool special)
> +				     struct folio *folio)
>  {
>  	args->lock = lock;
>  	args->ptep = ptep;
> @@ -6955,7 +6977,7 @@ static inline void pfnmap_args_setup(struct follow_pfnmap_args *args,
>  	args->addr_mask = addr_mask;
>  	args->pgprot = pgprot;
>  	args->writable = writable;
> -	args->special = special;
> +	args->cow = folio && folio_test_anon(folio);

I actually wonder if this should be changed to folio_has_anon_rmap() at
some point :)

Since a folio being 'anon' is vague, because we stupidly made 'anon' vague
in general.

Anyway I was going to ask does this suffice for CoW but having an anon rmap
implies CoW so it does.

>  }
>
>  static inline void pfnmap_lockdep_assert(struct vm_area_struct *vma)
> @@ -7008,6 +7030,7 @@ int follow_pfnmap_start(struct follow_pfnmap_args *args)
>  	struct vm_area_struct *vma = args->vma;
>  	unsigned long address = args->address;
>  	struct mm_struct *mm = vma->vm_mm;
> +	struct folio *folio = NULL;
>  	spinlock_t *lock;
>  	pgd_t *pgdp;
>  	p4d_t *p4dp, p4d;
> @@ -7047,9 +7070,12 @@ int follow_pfnmap_start(struct follow_pfnmap_args *args)
>  			spin_unlock(lock);
>  			goto retry;
>  		}
> +
> +		if (vma_is_cow_mapping(vma))
> +			folio = vm_normal_folio_pud(vma, address, pud);
>  		pfnmap_args_setup(args, lock, NULL, pud_pgprot(pud),
>  				  pud_pfn(pud), PUD_MASK, pud_write(pud),
> -				  pud_special(pud));
> +				  folio);
>  		return 0;
>  	}
>
> @@ -7068,9 +7094,12 @@ int follow_pfnmap_start(struct follow_pfnmap_args *args)
>  			spin_unlock(lock);
>  			goto retry;
>  		}
> +
> +		if (vma_is_cow_mapping(vma))
> +			folio = vm_normal_folio_pmd(vma, address, pmd);
>  		pfnmap_args_setup(args, lock, NULL, pmd_pgprot(pmd),
>  				  pmd_pfn(pmd), PMD_MASK, pmd_write(pmd),
> -				  pmd_special(pmd));
> +				  folio);
>  		return 0;
>  	}
>
> @@ -7080,9 +7109,12 @@ int follow_pfnmap_start(struct follow_pfnmap_args *args)
>  	pte = ptep_get(ptep);
>  	if (!pte_present(pte))
>  		goto unlock;
> +
> +	if (vma_is_cow_mapping(vma))
> +		folio = vm_normal_folio(vma, address, pte);
>  	pfnmap_args_setup(args, lock, ptep, pte_pgprot(pte),
>  			  pte_pfn(pte), PAGE_MASK, pte_write(pte),
> -			  pte_special(pte));
> +			  folio);

This all seems reasonable, I guess, insofar as the horrible 'special' CoW
handling is reasonable, but this function is gross to have this degree of
duplication across page table levels.

But I guess one for the future.

>  	return 0;
>  unlock:
>  	pte_unmap_unlock(ptep, lock);
> @@ -7145,6 +7177,10 @@ int generic_access_phys(struct vm_area_struct *vma, unsigned long addr,
>  	writable = args.writable;
>  	follow_pfnmap_end(&args);
>
> +	/* Reject CoW'ed anonymous folios. */
> +	if (args.cow)
> +		return -EINVAL;
> +
>  	if ((write & FOLL_WRITE) && !writable)
>  		return -EINVAL;
>
> --
> 2.43.0
>
>
> --
> Cheers,
>
> David

--
Cheers, Lorenzo

^ permalink raw reply	[flat|nested] 13+ messages in thread

* Re: [PATCH] mm: don't ioremap COWed anon pages in generic_access_phys()
  2026-10-02  8:38     ` Lorenzo Stoakes (ARM)
@ 2026-10-02 10:02       ` David Hildenbrand (Arm)
  2026-10-02 10:04         ` David Hildenbrand (Arm)
  2026-10-02 10:07         ` Lorenzo Stoakes (ARM)
  0 siblings, 2 replies; 13+ messages in thread
From: David Hildenbrand (Arm) @ 2026-10-02 10:02 UTC (permalink / raw)
  To: Lorenzo Stoakes (ARM)
  Cc: Nguyen Ngoc Thang, akpm, liam, vbabka, rppt, surenb, mhocko,
	peterx, dave.hansen, linux-mm, linux-kernel,
	syzbot+49b1021becba70c1f3f6

On 10/2/26 10:38, Lorenzo Stoakes (ARM) wrote:
> On Thu, Oct 01, 2026 at 10:22:29PM +0200, David Hildenbrand (Arm) wrote:
>> On 10/1/26 17:25, Nguyen Ngoc Thang wrote:
>>> A MAP_PRIVATE mapping of iomem (e.g. a PCI sysfs resourceN file) is a
>>> COW pfnmap: a write fault replaces the pfn with an anonymous page, which
>>> remap_pfn_range() allows by keeping vm_pgoff equal to the base pfn.
>>>
>>> generic_access_phys() does not tell those COWed pages apart from the
>>> original pfns and ioremaps whatever the PTE points to. Reading such an
>>> address via /proc/pid/mem or ptrace then ioremaps RAM:
>>>
>>>   ioremap on RAM at 0x0000000045623000 - 0x0000000045623fff
>>>   WARNING: arch/x86/mm/ioremap.c:216 at __ioremap_caller.isra.0+0x4c2/0x5f0
>>>   Call Trace:
>>>    generic_access_phys+0x130/0x4d0 mm/memory.c:7178
>>>    kernfs_vma_access+0x1ce/0x280 fs/kernfs/file.c:437
>>>    __access_remote_vm+0x58f/0x890 mm/memory.c:7256
>>>    mem_rw+0x2a1/0x670 fs/proc/base.c:912
>>>
>>
> 
> Sorry to say this looks schlopped.
> 
> This guy has sent 10 series across 6 subsystems over ~21 hrs:
> 
> https://lore.kernel.org/all/?q=f%3Angocthang2710.1999%40gmail.com
> 
> Nguyen - please do not flood the kernel with patches, and please use the
> Assisted-by tag for generated content.
> 
> The original code you submitted is really not great even if the issue may
> be valid.
> 
> So I'd say somebody from the core team should take over this if we want to
> come up with a patch.

Yes, I'll take care of it.

[...]
>>
>> Signed-off-by: David Hildenbrand (Arm) <david@kernel.org>
> 
> This looks reasonable but I hate that we have 'special' CoW overrides like
> this :)

After sending this yesterday, I concluded that we can do this cleaner: just have

bool normal_page;

(naming suggestions?)

that express that this is something refcounted with a struct page, like
documented for vm_normal_page().

Then we can just refuse all of these.

> 
>> ---
>>  include/linux/mm.h |  4 +++-
>>  mm/memory.c        | 48 ++++++++++++++++++++++++++++++++++++++++------
>>  2 files changed, 45 insertions(+), 7 deletions(-)
>>
>> diff --git a/include/linux/mm.h b/include/linux/mm.h
>> index c49ef99b4413..b90e547797a9 100644
>> --- a/include/linux/mm.h
>> +++ b/include/linux/mm.h
>> @@ -3193,6 +3193,8 @@ struct folio *vm_normal_folio_pmd(struct vm_area_struct *vma,
>>  				  unsigned long addr, pmd_t pmd);
>>  struct page *vm_normal_page_pmd(struct vm_area_struct *vma, unsigned long addr,
>>  				pmd_t pmd);
>> +struct folio *vm_normal_folio_pud(struct vm_area_struct *vma,
>> +		unsigned long addr, pud_t pud);
> 
> Hmm, if not defined before why would this need a new PUD handler? Do we
> even have PUD-leaf PFN mappings?

Yes we do. In any case, good for consistency.

[...]

>> +	args->cow = folio && folio_test_anon(folio);
> 
> I actually wonder if this should be changed to folio_has_anon_rmap() at
> some point :)

Not a fan.

> 
> Since a folio being 'anon' is vague, because we stupidly made 'anon' vague
> in general.

For folios it's an established term :)

> 
> Anyway I was going to ask does this suffice for CoW but having an anon rmap
> implies CoW so it does.


The downside of using "bool normal_page;" is that we should check
vm_normal_page() for any mapping, not just cow mappings. I suspect
performance-wise we don't really care.

-- 
Cheers,

David

^ permalink raw reply	[flat|nested] 13+ messages in thread

* Re: [PATCH] mm: don't ioremap COWed anon pages in generic_access_phys()
  2026-10-02 10:02       ` David Hildenbrand (Arm)
@ 2026-10-02 10:04         ` David Hildenbrand (Arm)
  2026-10-02 10:09           ` Lorenzo Stoakes (ARM)
  2026-10-02 10:07         ` Lorenzo Stoakes (ARM)
  1 sibling, 1 reply; 13+ messages in thread
From: David Hildenbrand (Arm) @ 2026-10-02 10:04 UTC (permalink / raw)
  To: Lorenzo Stoakes (ARM)
  Cc: Nguyen Ngoc Thang, akpm, liam, vbabka, rppt, surenb, mhocko,
	peterx, dave.hansen, linux-mm, linux-kernel,
	syzbot+49b1021becba70c1f3f6

On 10/2/26 12:02, David Hildenbrand (Arm) wrote:
> On 10/2/26 10:38, Lorenzo Stoakes (ARM) wrote:
>> On Thu, Oct 01, 2026 at 10:22:29PM +0200, David Hildenbrand (Arm) wrote:
>>>
>>
>> Sorry to say this looks schlopped.
>>
>> This guy has sent 10 series across 6 subsystems over ~21 hrs:
>>
>> https://lore.kernel.org/all/?q=f%3Angocthang2710.1999%40gmail.com
>>
>> Nguyen - please do not flood the kernel with patches, and please use the
>> Assisted-by tag for generated content.
>>
>> The original code you submitted is really not great even if the issue may
>> be valid.
>>
>> So I'd say somebody from the core team should take over this if we want to
>> come up with a patch.
> 
> Yes, I'll take care of it.
> 
> [...]
>>>
>>> Signed-off-by: David Hildenbrand (Arm) <david@kernel.org>
>>
>> This looks reasonable but I hate that we have 'special' CoW overrides like
>> this :)
> 
> After sending this yesterday, I concluded that we can do this cleaner: just have
> 
> bool normal_page;
> 
> (naming suggestions?)
> 
> that express that this is something refcounted with a struct page, like
> documented for vm_normal_page().

Hmm, have to think about that once more, regarding VM_IO and if there are some
cases that would actually have to work in generic_access_phys().

-- 
Cheers,

David

^ permalink raw reply	[flat|nested] 13+ messages in thread

* Re: [PATCH] mm: don't ioremap COWed anon pages in generic_access_phys()
  2026-10-02 10:02       ` David Hildenbrand (Arm)
  2026-10-02 10:04         ` David Hildenbrand (Arm)
@ 2026-10-02 10:07         ` Lorenzo Stoakes (ARM)
  1 sibling, 0 replies; 13+ messages in thread
From: Lorenzo Stoakes (ARM) @ 2026-10-02 10:07 UTC (permalink / raw)
  To: David Hildenbrand (Arm)
  Cc: Nguyen Ngoc Thang, akpm, liam, vbabka, rppt, surenb, mhocko,
	peterx, dave.hansen, linux-mm, linux-kernel,
	syzbot+49b1021becba70c1f3f6

On Fri, Oct 02, 2026 at 12:02:15PM +0200, David Hildenbrand (Arm) wrote:
> On 10/2/26 10:38, Lorenzo Stoakes (ARM) wrote:
> > On Thu, Oct 01, 2026 at 10:22:29PM +0200, David Hildenbrand (Arm) wrote:

> > So I'd say somebody from the core team should take over this if we want to
> > come up with a patch.
>
> Yes, I'll take care of it.

Thanks!

> >> Signed-off-by: David Hildenbrand (Arm) <david@kernel.org>
> >
> > This looks reasonable but I hate that we have 'special' CoW overrides like
> > this :)
>
> After sending this yesterday, I concluded that we can do this cleaner: just have
>
> bool normal_page;
>
> (naming suggestions?)
>
> that express that this is something refcounted with a struct page, like
> documented for vm_normal_page().
>
> Then we can just refuse all of these.

Yeah it's all a bit tricky.

Maybe is_vm_normal_page ?

Or do we want to default to referring to a folio... but then pfnmap not a
folio... ugh.

is_normal?

With a comment explaining in sense of vm_normal_page/folio()?

> >> diff --git a/include/linux/mm.h b/include/linux/mm.h
> >> index c49ef99b4413..b90e547797a9 100644
> >> --- a/include/linux/mm.h
> >> +++ b/include/linux/mm.h

> >>  				  unsigned long addr, pmd_t pmd);
> >>  struct page *vm_normal_page_pmd(struct vm_area_struct *vma, unsigned long addr,
> >>  				pmd_t pmd);
> >> +struct folio *vm_normal_folio_pud(struct vm_area_struct *vma,
> >> +		unsigned long addr, pud_t pud);
> >
> > Hmm, if not defined before why would this need a new PUD handler? Do we
> > even have PUD-leaf PFN mappings?
>
> Yes we do. In any case, good for consistency.

In that case then we should definitely have it!

> > Since a folio being 'anon' is vague, because we stupidly made 'anon' vague
> > in general.
>
> For folios it's an established term :)

Yeah, fair enough, that is true.

And makes the anon folio -> tracked as such in my change consistent end-to-end.

I guess we have swapbacked for the shmem stuff as a clear delineation too.

>
> >
> > Anyway I was going to ask does this suffice for CoW but having an anon rmap
> > implies CoW so it does.
>
>
> The downside of using "bool normal_page;" is that we should check
> vm_normal_page() for any mapping, not just cow mappings. I suspect
> performance-wise we don't really care.

Yep, I'm sure it's fine!

--
Cheers, Lorenzo

^ permalink raw reply	[flat|nested] 13+ messages in thread

* Re: [PATCH] mm: don't ioremap COWed anon pages in generic_access_phys()
  2026-10-02 10:04         ` David Hildenbrand (Arm)
@ 2026-10-02 10:09           ` Lorenzo Stoakes (ARM)
  2026-10-02 12:19             ` David Hildenbrand (Arm)
  0 siblings, 1 reply; 13+ messages in thread
From: Lorenzo Stoakes (ARM) @ 2026-10-02 10:09 UTC (permalink / raw)
  To: David Hildenbrand (Arm)
  Cc: Nguyen Ngoc Thang, akpm, liam, vbabka, rppt, surenb, mhocko,
	peterx, dave.hansen, linux-mm, linux-kernel,
	syzbot+49b1021becba70c1f3f6

On Fri, Oct 02, 2026 at 12:04:27PM +0200, David Hildenbrand (Arm) wrote:
> On 10/2/26 12:02, David Hildenbrand (Arm) wrote:
> > On 10/2/26 10:38, Lorenzo Stoakes (ARM) wrote:
> >> On Thu, Oct 01, 2026 at 10:22:29PM +0200, David Hildenbrand (Arm) wrote:
> >>>
> >>
> >> Sorry to say this looks schlopped.
> >>
> >> This guy has sent 10 series across 6 subsystems over ~21 hrs:
> >>
> >> https://lore.kernel.org/all/?q=f%3Angocthang2710.1999%40gmail.com
> >>
> >> Nguyen - please do not flood the kernel with patches, and please use the
> >> Assisted-by tag for generated content.
> >>
> >> The original code you submitted is really not great even if the issue may
> >> be valid.
> >>
> >> So I'd say somebody from the core team should take over this if we want to
> >> come up with a patch.
> >
> > Yes, I'll take care of it.
> >
> > [...]
> >>>
> >>> Signed-off-by: David Hildenbrand (Arm) <david@kernel.org>
> >>
> >> This looks reasonable but I hate that we have 'special' CoW overrides like
> >> this :)
> >
> > After sending this yesterday, I concluded that we can do this cleaner: just have
> >
> > bool normal_page;
> >
> > (naming suggestions?)
> >
> > that express that this is something refcounted with a struct page, like
> > documented for vm_normal_page().
>
> Hmm, have to think about that once more, regarding VM_IO and if there are some
> cases that would actually have to work in generic_access_phys().

Well, my series at least makes it easier to reason about VMA_IO_BIT!

Though not sure if it really touches PFN map cases specifically.

>
> --
> Cheers,
>
> David

--
Cheers, Lorenzo

^ permalink raw reply	[flat|nested] 13+ messages in thread

* Re: [PATCH] mm: don't ioremap COWed anon pages in generic_access_phys()
  2026-10-02 10:09           ` Lorenzo Stoakes (ARM)
@ 2026-10-02 12:19             ` David Hildenbrand (Arm)
  2026-10-02 14:03               ` Lorenzo Stoakes (ARM)
  0 siblings, 1 reply; 13+ messages in thread
From: David Hildenbrand (Arm) @ 2026-10-02 12:19 UTC (permalink / raw)
  To: Lorenzo Stoakes (ARM)
  Cc: Nguyen Ngoc Thang, akpm, liam, vbabka, rppt, surenb, mhocko,
	peterx, dave.hansen, linux-mm, linux-kernel,
	syzbot+49b1021becba70c1f3f6

On 10/2/26 12:09, Lorenzo Stoakes (ARM) wrote:
> On Fri, Oct 02, 2026 at 12:04:27PM +0200, David Hildenbrand (Arm) wrote:
>> On 10/2/26 12:02, David Hildenbrand (Arm) wrote:
>>>
>>> Yes, I'll take care of it.
>>>
>>> [...]
>>>
>>> After sending this yesterday, I concluded that we can do this cleaner: just have
>>>
>>> bool normal_page;
>>>
>>> (naming suggestions?)
>>>
>>> that express that this is something refcounted with a struct page, like
>>> documented for vm_normal_page().
>>
>> Hmm, have to think about that once more, regarding VM_IO and if there are some
>> cases that would actually have to work in generic_access_phys().
> 
> Well, my series at least makes it easier to reason about VMA_IO_BIT!
> 
> Though not sure if it really touches PFN map cases specifically.

I'm more concerned about someone using this function on VM_MIXEDMAP | VM_IO with
a memory page that has a struct page but is actually not memory. So we could get
something that vm_normal_page() would flag but generic_access_phys() could
actually read ... I'll have to explore the generic_access_phys() users once more.

All way to complicated (and you series improves things).

-- 
Cheers,

David

^ permalink raw reply	[flat|nested] 13+ messages in thread

* Re: [PATCH] mm: don't ioremap COWed anon pages in generic_access_phys()
  2026-10-02 12:19             ` David Hildenbrand (Arm)
@ 2026-10-02 14:03               ` Lorenzo Stoakes (ARM)
  2026-10-02 14:26                 ` David Hildenbrand (Arm)
  0 siblings, 1 reply; 13+ messages in thread
From: Lorenzo Stoakes (ARM) @ 2026-10-02 14:03 UTC (permalink / raw)
  To: David Hildenbrand (Arm)
  Cc: Nguyen Ngoc Thang, akpm, liam, vbabka, rppt, surenb, mhocko,
	peterx, dave.hansen, linux-mm, linux-kernel,
	syzbot+49b1021becba70c1f3f6

On Fri, Oct 02, 2026 at 02:19:00PM +0200, David Hildenbrand (Arm) wrote:
> On 10/2/26 12:09, Lorenzo Stoakes (ARM) wrote:
> > On Fri, Oct 02, 2026 at 12:04:27PM +0200, David Hildenbrand (Arm) wrote:
> >> On 10/2/26 12:02, David Hildenbrand (Arm) wrote:
> >>>
> >>> Yes, I'll take care of it.
> >>>
> >>> [...]
> >>>
> >>> After sending this yesterday, I concluded that we can do this cleaner: just have
> >>>
> >>> bool normal_page;
> >>>
> >>> (naming suggestions?)
> >>>
> >>> that express that this is something refcounted with a struct page, like
> >>> documented for vm_normal_page().
> >>
> >> Hmm, have to think about that once more, regarding VM_IO and if there are some
> >> cases that would actually have to work in generic_access_phys().
> >
> > Well, my series at least makes it easier to reason about VMA_IO_BIT!
> >
> > Though not sure if it really touches PFN map cases specifically.
>
> I'm more concerned about someone using this function on VM_MIXEDMAP | VM_IO with
> a memory page that has a struct page but is actually not memory. So we could get
> something that vm_normal_page() would flag but generic_access_phys() could
> actually read ... I'll have to explore the generic_access_phys() users once more.

Hmm that could be a problem also for struct page's that are there but you're not
supposed to access for other reasons to as well?

Definitely need to be careful about that.

>
> All way to complicated (and you series improves things).

:) well there's still a lot of complexity in there, one battle at a time...

>
> --
> Cheers,
>
> David

--
Cheers, Lorenzo

^ permalink raw reply	[flat|nested] 13+ messages in thread

* Re: [PATCH] mm: don't ioremap COWed anon pages in generic_access_phys()
  2026-10-02 14:03               ` Lorenzo Stoakes (ARM)
@ 2026-10-02 14:26                 ` David Hildenbrand (Arm)
  0 siblings, 0 replies; 13+ messages in thread
From: David Hildenbrand (Arm) @ 2026-10-02 14:26 UTC (permalink / raw)
  To: Lorenzo Stoakes (ARM)
  Cc: Nguyen Ngoc Thang, akpm, liam, vbabka, rppt, surenb, mhocko,
	peterx, dave.hansen, linux-mm, linux-kernel,
	syzbot+49b1021becba70c1f3f6

On 10/2/26 16:03, Lorenzo Stoakes (ARM) wrote:
> On Fri, Oct 02, 2026 at 02:19:00PM +0200, David Hildenbrand (Arm) wrote:
>> On 10/2/26 12:09, Lorenzo Stoakes (ARM) wrote:
>>>
>>> Well, my series at least makes it easier to reason about VMA_IO_BIT!
>>>
>>> Though not sure if it really touches PFN map cases specifically.
>>
>> I'm more concerned about someone using this function on VM_MIXEDMAP | VM_IO with
>> a memory page that has a struct page but is actually not memory. So we could get
>> something that vm_normal_page() would flag but generic_access_phys() could
>> actually read ... I'll have to explore the generic_access_phys() users once more.
> 
> Hmm that could be a problem also for struct page's that are there but you're not
> supposed to access for other reasons to as well?
> 
> Definitely need to be careful about that.
> 
>>
>> All way to complicated (and you series improves things).
> 
> :) well there's still a lot of complexity in there, one battle at a time...

My conclusion so far is: indicating is_normal should work when using it in
generic_access_phys(), as it is never used on non-VM_PFNMAP VMAs.

follow_pfnmap_start(), however, can be called on some non-VM_PFNMAP-but-VM_IO
VMAs from KVM as it seems.

So we cannot easily restrict follow_pfnmap_start() to VM_PFNMAP only (weird,
right?) without risking breaking some weird KVM use case. Gah.

Anyhow, I'll prepare a fix for this one and send it out likely later today.

-- 
Cheers,

David

^ permalink raw reply	[flat|nested] 13+ messages in thread

end of thread, other threads:[~2026-10-02 14:26 UTC | newest]

Thread overview: 13+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-02-22 11:33 [syzbot] [kernel?] WARNING in __ioremap_caller syzbot
2026-09-04 22:02 ` syzbot
2026-09-04 23:20   ` Dave Hansen
2026-10-01 15:25 ` [PATCH] mm: don't ioremap COWed anon pages in generic_access_phys() Nguyen Ngoc Thang
2026-10-01 20:22   ` David Hildenbrand (Arm)
2026-10-02  8:38     ` Lorenzo Stoakes (ARM)
2026-10-02 10:02       ` David Hildenbrand (Arm)
2026-10-02 10:04         ` David Hildenbrand (Arm)
2026-10-02 10:09           ` Lorenzo Stoakes (ARM)
2026-10-02 12:19             ` David Hildenbrand (Arm)
2026-10-02 14:03               ` Lorenzo Stoakes (ARM)
2026-10-02 14:26                 ` David Hildenbrand (Arm)
2026-10-02 10:07         ` Lorenzo Stoakes (ARM)

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®