* [RFC 0/1] x86/vmemmap: Add missing update of PML4 table / PML5 table entry
@ 2025-02-14 19:51 Gwan-gyeong Mun
2025-02-14 19:51 ` [RFC 1/1] " Gwan-gyeong Mun
0 siblings, 1 reply; 6+ messages in thread
From: Gwan-gyeong Mun @ 2025-02-14 19:51 UTC (permalink / raw)
To: linux-kernel
Cc: osalvador, 42.hyeyoo, byungchul, dave.hansen, luto, peterz, akpm,
max.byungchul.park, max.byungchul.park
When performing test of loading XE GPU drive module after applying the
GPU SVM and Xe SVM patch series[1] and the Dept patch series[2],
unexpected pagefault occur.
Through identifying the reported callstack[3] and the entry value of the
PML4 table/PML5 table corresponding to the virtual address where the page
fault occurred was empty, I wrote this rfc.
But this is a temporary solution to prevent page fault problems, and it
requires improvement of the routine that updates the missing entry in
the PML4 table or PML5 table.
[1] https://lore.kernel.org/dri-devel/20250213021112.1228481-1-matthew.brost@intel.com/
[2] https://lore.kernel.org/lkml/20240508094726.35754-1-byungchul@sk.com/
[3]
[ 49.103630] xe 0000:00:04.0: [drm] Available VRAM: 0x0000000800000000, 0x00000002fb800000
[ 49.116710] BUG: unable to handle page fault for address: ffffeb3ff1200000
[ 49.117175] #PF: supervisor write access in kernel mode
[ 49.117511] #PF: error_code(0x0002) - not-present page
[ 49.117835] PGD 0 P4D 0
[ 49.118015] Oops: Oops: 0002 [#1] PREEMPT SMP NOPTI
[ 49.118366] CPU: 3 UID: 0 PID: 302 Comm: modprobe Tainted: G W 6.13.0-drm-tip-test+ #62
[ 49.118976] Tainted: [W]=WARN
[ 49.119179] Hardware name: QEMU Standard PC (i440FX + PIIX, 1996), BIOS 1.15.0-1 04/01/2014
[ 49.119710] RIP: 0010:vmemmap_set_pmd+0xff/0x230
[ 49.120011] Code: 77 22 02 a9 ff ff 1f 00 74 58 48 8b 3d 62 77 22 02 48 85 ff 0f 85 9a 00 00 00 48 8d 7d 08 48 89 e9 31 c0 48 89 ea 48 83 e7 f8 <48> c7 45 00 00 00 00 00 48 29 f9 48 c7 45 48 00 00 00 00 83 c1 50
[ 49.121158] RSP: 0018:ffffc900016d37a8 EFLAGS: 00010282
[ 49.121502] RAX: 0000000000000000 RBX: ffff888164000000 RCX: ffffeb3ff1200000
[ 49.121966] RDX: ffffeb3ff1200000 RSI: 80000000000001e3 RDI: ffffeb3ff1200008
[ 49.122499] RBP: ffffeb3ff1200000 R08: ffffeb3ff1280000 R09: 0000000000000000
[ 49.123032] R10: ffff88817b94dc48 R11: 0000000000000003 R12: ffffeb3ff1280000
[ 49.123566] R13: 0000000000000000 R14: ffff88817b94dc48 R15: 8000000163e001e3
[ 49.124096] FS: 00007f53ae71d740(0000) GS:ffff88843fd80000(0000) knlGS:0000000000000000
[ 49.124698] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[ 49.125129] CR2: ffffeb3ff1200000 CR3: 000000017c7d2000 CR4: 0000000000750ef0
[ 49.125662] PKRU: 55555554
[ 49.125880] Call Trace:
[ 49.126078] <TASK>
[ 49.126252] ? __die_body.cold+0x19/0x26
[ 49.126509] ? page_fault_oops+0xa2/0x240
[ 49.126736] ? preempt_count_add+0x47/0xa0
[ 49.126968] ? search_module_extables+0x4a/0x80
[ 49.127224] ? exc_page_fault+0x206/0x230
[ 49.127454] ? asm_exc_page_fault+0x22/0x30
[ 49.127691] ? vmemmap_set_pmd+0xff/0x230
[ 49.127919] vmemmap_populate_hugepages+0x176/0x180
[ 49.128194] vmemmap_populate+0x34/0x80
[ 49.128416] __populate_section_memmap+0x41/0x90
[ 49.128676] sparse_add_section+0x121/0x3e0
[ 49.128914] __add_pages+0xba/0x150
[ 49.129116] add_pages+0x1d/0x70
[ 49.129305] memremap_pages+0x3dc/0x810
[ 49.129529] devm_memremap_pages+0x1c/0x60
[ 49.129762] xe_devm_add+0x8b/0x100 [xe]
[ 49.130072] xe_tile_init_noalloc+0x6a/0x70 [xe]
[ 49.130408] xe_device_probe+0x48c/0x740 [xe]
[ 49.130714] ? __pfx___drmm_mutex_release+0x10/0x10
[ 49.130982] ? __drmm_add_action+0x85/0xd0
[ 49.131208] ? __pfx___drmm_mutex_release+0x10/0x10
[ 49.131478] xe_pci_probe+0x7ef/0xd90 [xe]
[ 49.131777] ? _raw_spin_unlock_irqrestore+0x66/0x90
[ 49.132049] ? lockdep_hardirqs_on+0xba/0x140
[ 49.132290] pci_device_probe+0x99/0x110
[ 49.132510] really_probe+0xdb/0x340
[ 49.132710] ? pm_runtime_barrier+0x50/0x90
[ 49.132941] ? __pfx___driver_attach+0x10/0x10
[ 49.133190] __driver_probe_device+0x78/0x110
[ 49.133433] driver_probe_device+0x1f/0xa0
[ 49.133661] __driver_attach+0xba/0x1c0
[ 49.133874] bus_for_each_dev+0x7a/0xd0
[ 49.134089] bus_add_driver+0x114/0x200
[ 49.134302] driver_register+0x6e/0xc0
[ 49.134515] xe_init+0x1e/0x50 [xe]
[ 49.134827] ? __pfx_xe_init+0x10/0x10 [xe]
[ 49.134926] xe 0000:00:04.0: [drm:process_one_work] GT1: GuC CT safe-mode canceled
[ 49.135112] do_one_initcall+0x5b/0x2b0
[ 49.135734] ? rcu_is_watching+0xd/0x40
[ 49.135995] ? __kmalloc_cache_noprof+0x231/0x310
[ 49.136315] do_init_module+0x60/0x210
[ 49.136572] init_module_from_file+0x86/0xc0
[ 49.136863] idempotent_init_module+0x12b/0x340
[ 49.137156] __x64_sys_finit_module+0x61/0xc0
[ 49.137437] do_syscall_64+0x69/0x140
[ 49.137681] entry_SYSCALL_64_after_hwframe+0x76/0x7e
[ 49.137953] RIP: 0033:0x7f53ae1261fd
[ 49.138153] Code: ff c3 66 2e 0f 1f 84 00 00 00 00 00 90 f3 0f 1e fa 48 89 f8 48 89 f7 48 89 d6 48 89 ca 4d 89 c2 4d 89 c8 4c 8b 4c 24 08 0f 05 <48> 3d 01 f0 ff ff 73 01 c3 48 8b 0d e3 fa 0c 00 f7 d8 64 89 01 48
[ 49.139117] RSP: 002b:00007ffd0e9021e8 EFLAGS: 00000246 ORIG_RAX: 0000000000000139
[ 49.139525] RAX: ffffffffffffffda RBX: 000055c02951ee50 RCX: 00007f53ae1261fd
[ 49.139905] RDX: 0000000000000000 RSI: 000055bfff125478 RDI: 0000000000000010
[ 49.140282] RBP: 000055bfff125478 R08: 00007f53ae1f6b20 R09: 00007ffd0e902230
[ 49.140663] R10: 000055c029522000 R11: 0000000000000246 R12: 0000000000040000
[ 49.141040] R13: 000055c02951ef80 R14: 0000000000000000 R15: 000055c029521fc0
[ 49.141424] </TASK>
[ 49.141552] Modules linked in: xe(+) drm_ttm_helper gpu_sched drm_suballoc_helper drm_gpuvm drm_exec drm_gpusvm i2c_algo_bit drm_buddy video wmi ttm drm_display_helper drm_kms_helper crct10dif_pclmul crc32_pclmul i2c_piix4 e1000 ghash_clmulni_intel i2c_smbus fuse
[ 49.142824] CR2: ffffeb3ff1200000
[ 49.143010] ---[ end trace 0000000000000000 ]---
[ 49.143268] RIP: 0010:vmemmap_set_pmd+0xff/0x230
[ 49.143523] Code: 77 22 02 a9 ff ff 1f 00 74 58 48 8b 3d 62 77 22 02 48 85 ff 0f 85 9a 00 00 00 48 8d 7d 08 48 89 e9 31 c0 48 89 ea 48 83 e7 f8 <48> c7 45 00 00 00 00 00 48 29 f9 48 c7 45 48 00 00 00 00 83 c1 50
[ 49.144489] RSP: 0018:ffffc900016d37a8 EFLAGS: 00010282
[ 49.144775] RAX: 0000000000000000 RBX: ffff888164000000 RCX: ffffeb3ff1200000
[ 49.145154] RDX: ffffeb3ff1200000 RSI: 80000000000001e3 RDI: ffffeb3ff1200008
[ 49.145536] RBP: ffffeb3ff1200000 R08: ffffeb3ff1280000 R09: 0000000000000000
[ 49.145914] R10: ffff88817b94dc48 R11: 0000000000000003 R12: ffffeb3ff1280000
[ 49.146292] R13: 0000000000000000 R14: ffff88817b94dc48 R15: 8000000163e001e3
[ 49.146671] FS: 00007f53ae71d740(0000) GS:ffff88843fd80000(0000) knlGS:0000000000000000
[ 49.147097] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[ 49.147407] CR2: ffffeb3ff1200000 CR3: 000000017c7d2000 CR4: 0000000000750ef0
[ 49.147786] PKRU: 55555554
[ 49.147941] note: modprobe[302] exited with irqs disabled
Gwan-gyeong Mun (1):
x86/vmemmap: Add missing update of PML4 table / PML5 table entry
arch/x86/mm/init_64.c | 1 +
1 file changed, 1 insertion(+)
--
2.48.1
^ permalink raw reply [flat|nested] 6+ messages in thread
* [RFC 1/1] x86/vmemmap: Add missing update of PML4 table / PML5 table entry
2025-02-14 19:51 [RFC 0/1] x86/vmemmap: Add missing update of PML4 table / PML5 table entry Gwan-gyeong Mun
@ 2025-02-14 19:51 ` Gwan-gyeong Mun
2025-02-14 19:57 ` Dave Hansen
0 siblings, 1 reply; 6+ messages in thread
From: Gwan-gyeong Mun @ 2025-02-14 19:51 UTC (permalink / raw)
To: linux-kernel
Cc: osalvador, 42.hyeyoo, byungchul, dave.hansen, luto, peterz, akpm,
max.byungchul.park, max.byungchul.park
when performing vmemmap populate, if the entry of the PML4 table/PML5 table
pointing to the target virtual address has never been updated, a page fault
occurs when the memset(start) called from the vmemmap_use_new_sub_pmd()
execution flow.
This fixes the problem of using the virtual address without updating the
entry in the PML4 table or PML5 table. But this is a temporary solution to
prevent page fault problems, and it requires improvement of the routine
that updates the missing entry in the PML4 table or PML5 table.
Fixes: faf1c0008a33 ("x86/vmemmap: optimize for consecutive sections in partial populated PMDs")
Signed-off-by: Gwan-gyeong Mun <gwan-gyeong.mun@intel.com>
Cc: Oscar Salvador <osalvador@suse.de>
Cc: Hyeonggon Yoo <42.hyeyoo@gmail.com>
Cc: Byungchul Park <byungchul@sk.com>
Cc: Dave Hansen <dave.hansen@linux.intel.com>
Cc: Andy Lutomirski <luto@kernel.org>
Cc: Peter Zijlstra <peterz@infradead.org>
Cc: Andrew Morton <akpm@linux-foundation.org>
---
arch/x86/mm/init_64.c | 1 +
1 file changed, 1 insertion(+)
diff --git a/arch/x86/mm/init_64.c b/arch/x86/mm/init_64.c
index 01ea7c6df303..7a4d8cea1a2e 100644
--- a/arch/x86/mm/init_64.c
+++ b/arch/x86/mm/init_64.c
@@ -912,6 +912,7 @@ static void __meminit vmemmap_use_new_sub_pmd(unsigned long start, unsigned long
{
const unsigned long page = ALIGN_DOWN(start, PMD_SIZE);
+ sync_global_pgds(start, end - 1);
vmemmap_flush_unused_pmd();
/*
--
2.48.1
^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [RFC 1/1] x86/vmemmap: Add missing update of PML4 table / PML5 table entry
2025-02-14 19:51 ` [RFC 1/1] " Gwan-gyeong Mun
@ 2025-02-14 19:57 ` Dave Hansen
2025-02-15 0:20 ` Harry (Hyeonggon) Yoo
0 siblings, 1 reply; 6+ messages in thread
From: Dave Hansen @ 2025-02-14 19:57 UTC (permalink / raw)
To: Gwan-gyeong Mun, linux-kernel
Cc: osalvador, 42.hyeyoo, byungchul, dave.hansen, luto, peterz, akpm,
max.byungchul.park, max.byungchul.park
On 2/14/25 11:51, Gwan-gyeong Mun wrote:
> when performing vmemmap populate, if the entry of the PML4 table/PML5 table
> pointing to the target virtual address has never been updated, a page fault
> occurs when the memset(start) called from the vmemmap_use_new_sub_pmd()
> execution flow.
"Page fault" meaning oops? Or something that we manage to handle and
return from without oopsing?
> This fixes the problem of using the virtual address without updating the
> entry in the PML4 table or PML5 table. But this is a temporary solution to
> prevent page fault problems, and it requires improvement of the routine
> that updates the missing entry in the PML4 table or PML5 table.
Can we please skip past the band-aid and go to the real fix?
^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [RFC 1/1] x86/vmemmap: Add missing update of PML4 table / PML5 table entry
2025-02-14 19:57 ` Dave Hansen
@ 2025-02-15 0:20 ` Harry (Hyeonggon) Yoo
2025-02-15 0:29 ` Dave Hansen
0 siblings, 1 reply; 6+ messages in thread
From: Harry (Hyeonggon) Yoo @ 2025-02-15 0:20 UTC (permalink / raw)
To: Dave Hansen
Cc: Gwan-gyeong Mun, linux-kernel, osalvador, byungchul, dave.hansen,
luto, peterz, akpm, max.byungchul.park, max.byungchul.park
On Fri, Feb 14, 2025 at 11:57:50AM -0800, Dave Hansen wrote:
> On 2/14/25 11:51, Gwan-gyeong Mun wrote:
> > when performing vmemmap populate, if the entry of the PML4 table/PML5 table
> > pointing to the target virtual address has never been updated, a page fault
> > occurs when the memset(start) called from the vmemmap_use_new_sub_pmd()
> > execution flow.
>
> "Page fault" meaning oops? Or something that we manage to handle and
> return from without oopsing?
It means oops, because the kernel accesses part of vmemmap that's not
populated (yet) in current process's page table.
This oops was observed after increasing the size of struct page (as a part of
developing a debug feature), but the real cause is that page table entries are
only installed in init_mm's page table and then sync'd later, but in the mean
time the process that triggered hot-plug accesses new portion of vmemmap.
If the process does not directly use the page table of init_mm (like swapper)
this oops can occur (e.g., I was able to trigger with `sudo modprobe hmm_test`
after increasing the size of struct page).
> > This fixes the problem of using the virtual address without updating the
> > entry in the PML4 table or PML5 table. But this is a temporary solution to
> > prevent page fault problems, and it requires improvement of the routine
> > that updates the missing entry in the PML4 table or PML5 table.
>
> Can we please skip past the band-aid and go to the real fix?
Yes, of course it'd best to skip a temporary fix.
The intention is to report/discuss the problem and a fix as a starting point.
--
Harry
^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [RFC 1/1] x86/vmemmap: Add missing update of PML4 table / PML5 table entry
2025-02-15 0:20 ` Harry (Hyeonggon) Yoo
@ 2025-02-15 0:29 ` Dave Hansen
2025-02-17 11:40 ` Gwan-gyeong Mun
0 siblings, 1 reply; 6+ messages in thread
From: Dave Hansen @ 2025-02-15 0:29 UTC (permalink / raw)
To: Harry (Hyeonggon) Yoo
Cc: Gwan-gyeong Mun, linux-kernel, osalvador, byungchul, dave.hansen,
luto, peterz, akpm, max.byungchul.park, max.byungchul.park
On 2/14/25 16:20, Harry (Hyeonggon) Yoo wrote:
> On Fri, Feb 14, 2025 at 11:57:50AM -0800, Dave Hansen wrote:
>> On 2/14/25 11:51, Gwan-gyeong Mun wrote:
>>> when performing vmemmap populate, if the entry of the PML4 table/PML5 table
>>> pointing to the target virtual address has never been updated, a page fault
>>> occurs when the memset(start) called from the vmemmap_use_new_sub_pmd()
>>> execution flow.
>>
>> "Page fault" meaning oops? Or something that we manage to handle and
>> return from without oopsing?
>
> It means oops, because the kernel accesses part of vmemmap that's not
> populated (yet) in current process's page table.
Your 0/1 cover letter got to me after this mail did. I see the oops
there clear as day now.
> This oops was observed after increasing the size of struct page (as a part of
> developing a debug feature), but the real cause is that page table entries are
> only installed in init_mm's page table and then sync'd later, but in the mean
> time the process that triggered hot-plug accesses new portion of vmemmap.
>
> If the process does not directly use the page table of init_mm (like swapper)
> this oops can occur (e.g., I was able to trigger with `sudo modprobe hmm_test`
> after increasing the size of struct page).
Makes sense. Thanks for the explanation.
>>> This fixes the problem of using the virtual address without updating the
>>> entry in the PML4 table or PML5 table. But this is a temporary solution to
>>> prevent page fault problems, and it requires improvement of the routine
>>> that updates the missing entry in the PML4 table or PML5 table.
>>
>> Can we please skip past the band-aid and go to the real fix?
>
> Yes, of course it'd best to skip a temporary fix.
> The intention is to report/discuss the problem and a fix as a starting point.
Do you have a better fix in mind?
^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [RFC 1/1] x86/vmemmap: Add missing update of PML4 table / PML5 table entry
2025-02-15 0:29 ` Dave Hansen
@ 2025-02-17 11:40 ` Gwan-gyeong Mun
0 siblings, 0 replies; 6+ messages in thread
From: Gwan-gyeong Mun @ 2025-02-17 11:40 UTC (permalink / raw)
To: Dave Hansen, Harry (Hyeonggon) Yoo
Cc: linux-kernel, osalvador, byungchul, dave.hansen, luto, peterz,
akpm, max.byungchul.park, max.byungchul.park
On 2/15/25 2:29 AM, Dave Hansen wrote:
> On 2/14/25 16:20, Harry (Hyeonggon) Yoo wrote:
>> On Fri, Feb 14, 2025 at 11:57:50AM -0800, Dave Hansen wrote:
>>> On 2/14/25 11:51, Gwan-gyeong Mun wrote:
>>>> when performing vmemmap populate, if the entry of the PML4 table/PML5 table
>>>> pointing to the target virtual address has never been updated, a page fault
>>>> occurs when the memset(start) called from the vmemmap_use_new_sub_pmd()
>>>> execution flow.
>>>
>>> "Page fault" meaning oops? Or something that we manage to handle and
>>> return from without oopsing?
>>
>> It means oops, because the kernel accesses part of vmemmap that's not
>> populated (yet) in current process's page table.
>
> Your 0/1 cover letter got to me after this mail did. I see the oops
> there clear as day now.
>
>> This oops was observed after increasing the size of struct page (as a part of
>> developing a debug feature), but the real cause is that page table entries are
>> only installed in init_mm's page table and then sync'd later, but in the mean
>> time the process that triggered hot-plug accesses new portion of vmemmap.
>>
>> If the process does not directly use the page table of init_mm (like swapper)
>> this oops can occur (e.g., I was able to trigger with `sudo modprobe hmm_test`
>> after increasing the size of struct page).
>
> Makes sense. Thanks for the explanation.
>
>>>> This fixes the problem of using the virtual address without updating the
>>>> entry in the PML4 table or PML5 table. But this is a temporary solution to
>>>> prevent page fault problems, and it requires improvement of the routine
>>>> that updates the missing entry in the PML4 table or PML5 table.
>>>
>>> Can we please skip past the band-aid and go to the real fix?
>>
>> Yes, of course it'd best to skip a temporary fix.
>> The intention is to report/discuss the problem and a fix as a starting point.
>
> Do you have a better fix in mind?
>
Yes, first what comes to mind right now to safely access the virtual
address is; translating vmemmap-based virtual address to direct-mapped
virtual address and use it, if the current top-level page table is not
init_mm's page table when accessing a vmemmap-based virtual address
before page table sync.
I will send a patch first with this idea.
If you have any better ideas, please let me know.
Br,
G.G.
^ permalink raw reply [flat|nested] 6+ messages in thread
end of thread, other threads:[~2025-02-17 11:42 UTC | newest]
Thread overview: 6+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2025-02-14 19:51 [RFC 0/1] x86/vmemmap: Add missing update of PML4 table / PML5 table entry Gwan-gyeong Mun
2025-02-14 19:51 ` [RFC 1/1] " Gwan-gyeong Mun
2025-02-14 19:57 ` Dave Hansen
2025-02-15 0:20 ` Harry (Hyeonggon) Yoo
2025-02-15 0:29 ` Dave Hansen
2025-02-17 11:40 ` Gwan-gyeong Mun
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®