* [PATCH 1/1] device-dax: Correct pgoff align in dax_set_mapping()
@ 2024-09-27 7:45 Kun(llfl)
2024-09-27 17:46 ` Andrew Morton
0 siblings, 1 reply; 4+ messages in thread
From: Kun(llfl) @ 2024-09-27 7:45 UTC (permalink / raw)
To: Andrew Morton, linux-kernel
pgoff should be aligned using ALIGN_DOWN() instead of ALIGN(). Otherwise,
vmf->address not aligned to fault_size will be aligned to the next
alignment, that can result in memory failure getting the wrong address.
Fixes: b9b5777f09be ("device-dax: use ALIGN() for determining pgoff")
Signed-off-by: Kun(llfl) <llfl@linux.alibaba.com>
Tested-by: JianXiong Zhao <zhaojianxiong.zjx@alibaba-inc.com>
---
drivers/dax/device.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/drivers/dax/device.c b/drivers/dax/device.c
index 9c1a729cd77e..6d74e62bbee0 100644
--- a/drivers/dax/device.c
+++ b/drivers/dax/device.c
@@ -86,7 +86,7 @@ static void dax_set_mapping(struct vm_fault *vmf, pfn_t pfn,
nr_pages = 1;
pgoff = linear_page_index(vmf->vma,
- ALIGN(vmf->address, fault_size));
+ ALIGN_DOWN(vmf->address, fault_size));
for (i = 0; i < nr_pages; i++) {
struct page *page = pfn_to_page(pfn_t_to_pfn(pfn) + i);
--
2.43.0
^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [PATCH 1/1] device-dax: Correct pgoff align in dax_set_mapping()
2024-09-27 7:45 [PATCH 1/1] device-dax: Correct pgoff align in dax_set_mapping() Kun(llfl)
@ 2024-09-27 17:46 ` Andrew Morton
2024-09-29 3:00 ` Kun(llfl)
2024-09-30 11:06 ` Joao Martins
0 siblings, 2 replies; 4+ messages in thread
From: Andrew Morton @ 2024-09-27 17:46 UTC (permalink / raw)
To: Kun(llfl); +Cc: linux-kernel, Dan Williams, Joao Martins
(cc's added)
On Fri, 27 Sep 2024 15:45:09 +0800 "Kun(llfl)" <llfl@linux.alibaba.com> wrote:
> pgoff should be aligned using ALIGN_DOWN() instead of ALIGN(). Otherwise,
> vmf->address not aligned to fault_size will be aligned to the next
> alignment, that can result in memory failure getting the wrong address.
>
> Fixes: b9b5777f09be ("device-dax: use ALIGN() for determining pgoff")
That's quite an old change. Can you suggest why it took this long to
be discovered? What is your userspace doing to trigger this?
> Signed-off-by: Kun(llfl) <llfl@linux.alibaba.com>
> Tested-by: JianXiong Zhao <zhaojianxiong.zjx@alibaba-inc.com>
> ---
> drivers/dax/device.c | 2 +-
> 1 file changed, 1 insertion(+), 1 deletion(-)
>
> diff --git a/drivers/dax/device.c b/drivers/dax/device.c
> index 9c1a729cd77e..6d74e62bbee0 100644
> --- a/drivers/dax/device.c
> +++ b/drivers/dax/device.c
> @@ -86,7 +86,7 @@ static void dax_set_mapping(struct vm_fault *vmf, pfn_t pfn,
> nr_pages = 1;
>
> pgoff = linear_page_index(vmf->vma,
> - ALIGN(vmf->address, fault_size));
> + ALIGN_DOWN(vmf->address, fault_size));
>
> for (i = 0; i < nr_pages; i++) {
> struct page *page = pfn_to_page(pfn_t_to_pfn(pfn) + i);
> --
> 2.43.0
^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [PATCH 1/1] device-dax: Correct pgoff align in dax_set_mapping()
2024-09-27 17:46 ` Andrew Morton
@ 2024-09-29 3:00 ` Kun(llfl)
2024-09-30 11:06 ` Joao Martins
1 sibling, 0 replies; 4+ messages in thread
From: Kun(llfl) @ 2024-09-29 3:00 UTC (permalink / raw)
To: Andrew Morton; +Cc: linux-kernel, Dan Williams, Joao Martins
That's a subtle situation that only can be observed in
page_mapped_in_vma() after the page is page fault handled by
dev_dax_huge_fault. Generally, there is little chance to perform
page_mapped_in_vma in dev-dax's page unless in specific error injection
to the dax device to trigger an MCE - memory-failure. In that case,
page_mapped_in_vma() will be triggered to determine which task is
accessing the failure address and kill that task in the end.
We used self-developed dax device (which is 2M aligned mapping) , to
perform error injection to random address. It turned out that error
injected to non-2M-aligned address was causing endless MCE until panic.
Because page_mapped_in_vma() kept resulting wrong address and the task
accessing the failure address was never killed properly:
[ 3783.719419] Memory failure: 0x200c9742: recovery action for dax page:
Recovered
[ 3784.049006] mce: Uncorrected hardware memory error in user-access at
200c9742380
[ 3784.049190] Memory failure: 0x200c9742: recovery action for dax page:
Recovered
[ 3784.448042] mce: Uncorrected hardware memory error in user-access at
200c9742380
[ 3784.448186] Memory failure: 0x200c9742: recovery action for dax page:
Recovered
[ 3784.792026] mce: Uncorrected hardware memory error in user-access at
200c9742380
[ 3784.792179] Memory failure: 0x200c9742: recovery action for dax page:
Recovered
[ 3785.162502] mce: Uncorrected hardware memory error in user-access at
200c9742380
[ 3785.162633] Memory failure: 0x200c9742: recovery action for dax page:
Recovered
[ 3785.461116] mce: Uncorrected hardware memory error in user-access at
200c9742380
[ 3785.461247] Memory failure: 0x200c9742: recovery action for dax page:
Recovered
[ 3785.764730] mce: Uncorrected hardware memory error in user-access at
200c9742380
[ 3785.764859] Memory failure: 0x200c9742: recovery action for dax page:
Recovered
[ 3786.042128] mce: Uncorrected hardware memory error in user-access at
200c9742380
[ 3786.042259] Memory failure: 0x200c9742: recovery action for dax page:
Recovered
[ 3786.464293] mce: Uncorrected hardware memory error in user-access at
200c9742380
[ 3786.464423] Memory failure: 0x200c9742: recovery action for dax page:
Recovered
[ 3786.818090] mce: Uncorrected hardware memory error in user-access at
200c9742380
[ 3786.818217] Memory failure: 0x200c9742: recovery action for dax page:
Recovered
[ 3787.085297] mce: Uncorrected hardware memory error in user-access at
200c9742380
[ 3787.085424] Memory failure: 0x200c9742: recovery action for dax page:
Recovered
It took us several weeks to pinpoint this problem, but we eventually
used bpftrace to trace the page fault and mce address and successfully
identified the issue.
On 9/28/24 1:46 AM, Andrew Morton wrote:
> (cc's added)
>
> On Fri, 27 Sep 2024 15:45:09 +0800 "Kun(llfl)" <llfl@linux.alibaba.com> wrote:
>
>> pgoff should be aligned using ALIGN_DOWN() instead of ALIGN(). Otherwise,
>> vmf->address not aligned to fault_size will be aligned to the next
>> alignment, that can result in memory failure getting the wrong address.
>>
>> Fixes: b9b5777f09be ("device-dax: use ALIGN() for determining pgoff")
> That's quite an old change. Can you suggest why it took this long to
> be discovered? What is your userspace doing to trigger this?
>
>> Signed-off-by: Kun(llfl) <llfl@linux.alibaba.com>
>> Tested-by: JianXiong Zhao <zhaojianxiong.zjx@alibaba-inc.com>
>> ---
>> drivers/dax/device.c | 2 +-
>> 1 file changed, 1 insertion(+), 1 deletion(-)
>>
>> diff --git a/drivers/dax/device.c b/drivers/dax/device.c
>> index 9c1a729cd77e..6d74e62bbee0 100644
>> --- a/drivers/dax/device.c
>> +++ b/drivers/dax/device.c
>> @@ -86,7 +86,7 @@ static void dax_set_mapping(struct vm_fault *vmf, pfn_t pfn,
>> nr_pages = 1;
>>
>> pgoff = linear_page_index(vmf->vma,
>> - ALIGN(vmf->address, fault_size));
>> + ALIGN_DOWN(vmf->address, fault_size));
>>
>> for (i = 0; i < nr_pages; i++) {
>> struct page *page = pfn_to_page(pfn_t_to_pfn(pfn) + i);
>> --
>> 2.43.0
--
Best,
KUN
^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [PATCH 1/1] device-dax: Correct pgoff align in dax_set_mapping()
2024-09-27 17:46 ` Andrew Morton
2024-09-29 3:00 ` Kun(llfl)
@ 2024-09-30 11:06 ` Joao Martins
1 sibling, 0 replies; 4+ messages in thread
From: Joao Martins @ 2024-09-30 11:06 UTC (permalink / raw)
To: Andrew Morton, Kun(llfl); +Cc: linux-kernel, Dan Williams
On 27/09/2024 18:46, Andrew Morton wrote:
> (cc's added)
>
> On Fri, 27 Sep 2024 15:45:09 +0800 "Kun(llfl)" <llfl@linux.alibaba.com> wrote:
>
>> pgoff should be aligned using ALIGN_DOWN() instead of ALIGN(). Otherwise,
>> vmf->address not aligned to fault_size will be aligned to the next
>> alignment, that can result in memory failure getting the wrong address.
>>
>> Fixes: b9b5777f09be ("device-dax: use ALIGN() for determining pgoff")
>
> That's quite an old change. Can you suggest why it took this long to
> be discovered? What is your userspace doing to trigger this?
>
Likely we never reproduce in production because we always pin device-dax regions
in the region align they provide (Qemu does similarly with prealloc in
hugetlb/file backed memory). I think this bug requires that we touch *unpinned*
device-dax regions unaligned to the device-dax selected alignment (page size
i.e. 4K/2M/1G)
>> Signed-off-by: Kun(llfl) <llfl@linux.alibaba.com>
>> Tested-by: JianXiong Zhao <zhaojianxiong.zjx@alibaba-inc.com>
Thanks a lot for the catch and sorry for the pain that it may have caused you:
Reviewed-by: Joao Martins <joao.m.martins@oracle.com>
>> ---
>> drivers/dax/device.c | 2 +-
>> 1 file changed, 1 insertion(+), 1 deletion(-)
>>
>> diff --git a/drivers/dax/device.c b/drivers/dax/device.c
>> index 9c1a729cd77e..6d74e62bbee0 100644
>> --- a/drivers/dax/device.c
>> +++ b/drivers/dax/device.c
>> @@ -86,7 +86,7 @@ static void dax_set_mapping(struct vm_fault *vmf, pfn_t pfn,
>> nr_pages = 1;
>>
>> pgoff = linear_page_index(vmf->vma,
>> - ALIGN(vmf->address, fault_size));
>> + ALIGN_DOWN(vmf->address, fault_size));
>>
>> for (i = 0; i < nr_pages; i++) {
>> struct page *page = pfn_to_page(pfn_t_to_pfn(pfn) + i);
>> --
>> 2.43.0
^ permalink raw reply [flat|nested] 4+ messages in thread
end of thread, other threads:[~2024-09-30 11:06 UTC | newest]
Thread overview: 4+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2024-09-27 7:45 [PATCH 1/1] device-dax: Correct pgoff align in dax_set_mapping() Kun(llfl)
2024-09-27 17:46 ` Andrew Morton
2024-09-29 3:00 ` Kun(llfl)
2024-09-30 11:06 ` Joao Martins
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®