From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1751940AbdENEMs (ORCPT ); Sun, 14 May 2017 00:12:48 -0400 Received: from mx0a-001b2d01.pphosted.com ([148.163.156.1]:54792 "EHLO mx0a-001b2d01.pphosted.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1751105AbdENEMr (ORCPT ); Sun, 14 May 2017 00:12:47 -0400 Subject: Re: [PATCH V2] mm/madvise: Enable (soft|hard) offline of HugeTLB pages at PGD level To: Andrew Morton , Anshuman Khandual References: <20170426035731.6924-1-khandual@linux.vnet.ibm.com> <20170512143503.81e0de2ae3d88a53168c601a@linux-foundation.org> Cc: linux-kernel@vger.kernel.org, linux-mm@kvack.org, n-horiguchi@ah.jp.nec.com, aneesh.kumar@linux.vnet.ibm.com From: Anshuman Khandual Date: Sun, 14 May 2017 09:41:49 +0530 User-Agent: Mozilla/5.0 (X11; Linux x86_64; rv:45.0) Gecko/20100101 Thunderbird/45.5.1 MIME-Version: 1.0 In-Reply-To: <20170512143503.81e0de2ae3d88a53168c601a@linux-foundation.org> Content-Type: text/plain; charset=windows-1252 Content-Transfer-Encoding: 7bit X-TM-AS-MML: disable x-cbid: 17051404-0008-0000-0000-0000012C3235 X-IBM-AV-DETECTION: SAVI=unused REMOTE=unused XFE=unused x-cbparentid: 17051404-0009-0000-0000-0000095AD427 Message-Id: X-Proofpoint-Virus-Version: vendor=fsecure engine=2.50.10432:,, definitions=2017-05-13_12:,, signatures=0 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 spamscore=0 suspectscore=2 malwarescore=0 phishscore=0 adultscore=0 bulkscore=0 classifier=spam adjust=0 reason=mlx scancount=1 engine=8.0.1-1703280000 definitions=main-1705140059 Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org On 05/13/2017 03:05 AM, Andrew Morton wrote: > On Wed, 26 Apr 2017 09:27:31 +0530 Anshuman Khandual wrote: > >> Though migrating gigantic HugeTLB pages does not sound much like real >> world use case, they can be affected by memory errors. Hence migration >> at the PGD level HugeTLB pages should be supported just to enable soft >> and hard offline use cases. >> >> While allocating the new gigantic HugeTLB page, it should not matter >> whether new page comes from the same node or not. There would be very >> few gigantic pages on the system afterall, we should not be bothered >> about node locality when trying to save a big page from crashing. >> >> This introduces a new HugeTLB allocator called alloc_huge_page_nonid() >> which will scan over all online nodes on the system and allocate a >> single HugeTLB page. >> >> ... >> >> --- a/mm/hugetlb.c >> +++ b/mm/hugetlb.c >> @@ -1669,6 +1669,23 @@ struct page *__alloc_buddy_huge_page_with_mpol(struct hstate *h, >> return __alloc_buddy_huge_page(h, vma, addr, NUMA_NO_NODE); >> } >> >> +struct page *alloc_huge_page_nonid(struct hstate *h) >> +{ >> + struct page *page = NULL; >> + int nid = 0; >> + >> + spin_lock(&hugetlb_lock); >> + if (h->free_huge_pages - h->resv_huge_pages > 0) { >> + for_each_online_node(nid) { >> + page = dequeue_huge_page_node(h, nid); >> + if (page) >> + break; >> + } >> + } >> + spin_unlock(&hugetlb_lock); >> + return page; >> +} >> + >> /* >> * This allocation function is useful in the context where vma is irrelevant. >> * E.g. soft-offlining uses this function because it only cares physical >> diff --git a/mm/memory-failure.c b/mm/memory-failure.c >> index fe64d7729a8e..d4f5710cf3f7 100644 >> --- a/mm/memory-failure.c >> +++ b/mm/memory-failure.c >> @@ -1481,11 +1481,15 @@ EXPORT_SYMBOL(unpoison_memory); >> static struct page *new_page(struct page *p, unsigned long private, int **x) >> { >> int nid = page_to_nid(p); >> - if (PageHuge(p)) >> + if (PageHuge(p)) { >> + if (hstate_is_gigantic(page_hstate(compound_head(p)))) >> + return alloc_huge_page_nonid(page_hstate(compound_head(p))); >> + >> return alloc_huge_page_node(page_hstate(compound_head(p)), >> nid); >> - else >> + } else { >> return __alloc_pages_node(nid, GFP_HIGHUSER_MOVABLE, 0); >> + } >> } > > Rather than adding alloc_huge_page_nonid(), would it be neater to teach > alloc_huge_page_node() (actually dequeue_huge_page_node()) to understand > nid==NUMA_NO_NODE? Sure, will change dequeue_huge_page_node() to accommodate NUMA_NO_NODE and let soft offline call with NUMA_NO_NODE in case of gigantic huge pages.