mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: "Kun(llfl)" <llfl@linux.alibaba.com>
To: Andrew Morton <akpm@linux-foundation.org>
Cc: linux-kernel@vger.kernel.org,
	Dan Williams <dan.j.williams@intel.com>,
	Joao Martins <joao.m.martins@oracle.com>
Subject: Re: [PATCH 1/1] device-dax: Correct pgoff align in dax_set_mapping()
Date: Sun, 29 Sep 2024 11:00:17 +0800	[thread overview]
Message-ID: <0e574409-cf12-47fd-b107-664e7f1b9cb6@linux.alibaba.com> (raw)
In-Reply-To: <20240927104646.3a0b777b5114ec62becd7f47@linux-foundation.org>

That's a subtle situation that only can be observed in 
page_mapped_in_vma() after the page is page fault handled by 
dev_dax_huge_fault. Generally, there is little chance to perform 
page_mapped_in_vma in dev-dax's page unless in specific error injection 
to the dax device to trigger an MCE - memory-failure. In that case, 
page_mapped_in_vma() will be triggered to determine which task is 
accessing the failure address and kill that task in the end.


We used self-developed dax device (which is 2M aligned mapping) , to 
perform error injection to random address. It turned out that error 
injected to non-2M-aligned address was causing endless MCE until panic. 
Because page_mapped_in_vma() kept resulting wrong address and the task 
accessing the failure address was never killed properly:


[ 3783.719419] Memory failure: 0x200c9742: recovery action for dax page: 
Recovered
[ 3784.049006] mce: Uncorrected hardware memory error in user-access at 
200c9742380
[ 3784.049190] Memory failure: 0x200c9742: recovery action for dax page: 
Recovered
[ 3784.448042] mce: Uncorrected hardware memory error in user-access at 
200c9742380
[ 3784.448186] Memory failure: 0x200c9742: recovery action for dax page: 
Recovered
[ 3784.792026] mce: Uncorrected hardware memory error in user-access at 
200c9742380
[ 3784.792179] Memory failure: 0x200c9742: recovery action for dax page: 
Recovered
[ 3785.162502] mce: Uncorrected hardware memory error in user-access at 
200c9742380
[ 3785.162633] Memory failure: 0x200c9742: recovery action for dax page: 
Recovered
[ 3785.461116] mce: Uncorrected hardware memory error in user-access at 
200c9742380
[ 3785.461247] Memory failure: 0x200c9742: recovery action for dax page: 
Recovered
[ 3785.764730] mce: Uncorrected hardware memory error in user-access at 
200c9742380
[ 3785.764859] Memory failure: 0x200c9742: recovery action for dax page: 
Recovered
[ 3786.042128] mce: Uncorrected hardware memory error in user-access at 
200c9742380
[ 3786.042259] Memory failure: 0x200c9742: recovery action for dax page: 
Recovered
[ 3786.464293] mce: Uncorrected hardware memory error in user-access at 
200c9742380
[ 3786.464423] Memory failure: 0x200c9742: recovery action for dax page: 
Recovered
[ 3786.818090] mce: Uncorrected hardware memory error in user-access at 
200c9742380
[ 3786.818217] Memory failure: 0x200c9742: recovery action for dax page: 
Recovered
[ 3787.085297] mce: Uncorrected hardware memory error in user-access at 
200c9742380
[ 3787.085424] Memory failure: 0x200c9742: recovery action for dax page: 
Recovered

It took us several weeks to pinpoint this problem,  but we eventually 
used bpftrace to trace the page fault and mce address and successfully 
identified the issue.

On 9/28/24 1:46 AM, Andrew Morton wrote:
> (cc's added)
>
> On Fri, 27 Sep 2024 15:45:09 +0800 "Kun(llfl)" <llfl@linux.alibaba.com> wrote:
>
>> pgoff should be aligned using ALIGN_DOWN() instead of ALIGN(). Otherwise,
>> vmf->address not aligned to fault_size will be aligned to the next
>> alignment, that can result in memory failure getting the wrong address.
>>
>> Fixes: b9b5777f09be ("device-dax: use ALIGN() for determining pgoff")
> That's quite an old change.  Can you suggest why it took this long to
> be discovered?  What is your userspace doing to trigger this?
>
>> Signed-off-by: Kun(llfl) <llfl@linux.alibaba.com>
>> Tested-by: JianXiong Zhao <zhaojianxiong.zjx@alibaba-inc.com>
>> ---
>>   drivers/dax/device.c | 2 +-
>>   1 file changed, 1 insertion(+), 1 deletion(-)
>>
>> diff --git a/drivers/dax/device.c b/drivers/dax/device.c
>> index 9c1a729cd77e..6d74e62bbee0 100644
>> --- a/drivers/dax/device.c
>> +++ b/drivers/dax/device.c
>> @@ -86,7 +86,7 @@ static void dax_set_mapping(struct vm_fault *vmf, pfn_t pfn,
>>   		nr_pages = 1;
>>   
>>   	pgoff = linear_page_index(vmf->vma,
>> -			ALIGN(vmf->address, fault_size));
>> +			ALIGN_DOWN(vmf->address, fault_size));
>>   
>>   	for (i = 0; i < nr_pages; i++) {
>>   		struct page *page = pfn_to_page(pfn_t_to_pfn(pfn) + i);
>> -- 
>> 2.43.0

-- 
Best,
KUN


  reply	other threads:[~2024-09-29  3:00 UTC|newest]

Thread overview: 4+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2024-09-27  7:45 Kun(llfl)
2024-09-27 17:46 ` Andrew Morton
2024-09-29  3:00   ` Kun(llfl) [this message]
2024-09-30 11:06   ` Joao Martins

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=0e574409-cf12-47fd-b107-664e7f1b9cb6@linux.alibaba.com \
    --to=llfl@linux.alibaba.com \
    --cc=akpm@linux-foundation.org \
    --cc=dan.j.williams@intel.com \
    --cc=joao.m.martins@oracle.com \
    --cc=linux-kernel@vger.kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®