From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from canpmsgout05.his.huawei.com (canpmsgout05.his.huawei.com [113.46.200.220]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E29A126A0A7 for ; Wed, 20 May 2026 02:40:28 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=113.46.200.220 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1779244832; cv=none; b=ZCINC+0aPnlM56pZQYgQzV0PnKVUAjQZstkyqp16e/ZcTudU0y0rkXEsyKtkzxYLa6ECm3WiisX4Zylcd12UE3QkVUCEq46udA5w45Wdx8odPOAvX1a0bsF2hJ/6QAr4PJbu2pCiCx/IwjHnkPVE/kbVG5q2wc6ijMc+zagqN/s= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1779244832; c=relaxed/simple; bh=7MicLkA5y2DiPJFSbR1uRaxS1UkVaDOQZSx3ybWoP1o=; h=Message-ID:Date:MIME-Version:Subject:To:CC:References:From: In-Reply-To:Content-Type; b=f3FmbWkREUp3vuV/Bq16Jz0WzwwvUvo19CIOLlU7FEefwwus9wDWtF7md8HdiqRgHaUKOO1xTKAIkfX6Z5GXgJfr/ZUVHq8S2mscP1CJisguB9OtEaihGA5o1rS8VKHBrjg2FYEwE1URCKv0+d6hPEdsCMp7v5uYqNtpUEKiN3w= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=huawei.com; spf=pass smtp.mailfrom=huawei.com; dkim=pass (1024-bit key) header.d=huawei.com header.i=@huawei.com header.b=WTAaxDc4; arc=none smtp.client-ip=113.46.200.220 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=huawei.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=huawei.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=huawei.com header.i=@huawei.com header.b="WTAaxDc4" dkim-signature: v=1; a=rsa-sha256; d=huawei.com; s=dkim; c=relaxed/relaxed; q=dns/txt; h=From; bh=FaSlZq7+rot967gNRvfkEumv8IhDHP9pWNaSpmceWiI=; b=WTAaxDc4eokgnuVAhzn6gLrz6PgN76SppCJNHJTneYy5V2uecBmKnrz1d4f26AG6Q9xTXvP2F MNeKRoJB2/Q4pZT+2OebHErmhPR3vOmMRkcqDFBI1t6utrzZliqwX587ca7DXiXvyiui42KMmd/ dK/q3JDXmaZ5kLCmstlronE= Received: from mail.maildlp.com (unknown [172.19.163.0]) by canpmsgout05.his.huawei.com (SkyGuard) with ESMTPS id 4gKwZW1T2Yz12LfK; Wed, 20 May 2026 10:33:23 +0800 (CST) Received: from dggpemf100008.china.huawei.com (unknown [7.185.36.138]) by mail.maildlp.com (Postfix) with ESMTPS id 689A64056B; Wed, 20 May 2026 10:40:24 +0800 (CST) Received: from [10.174.177.243] (10.174.177.243) by dggpemf100008.china.huawei.com (7.185.36.138) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.1544.11; Wed, 20 May 2026 10:40:22 +0800 Message-ID: <07873271-6671-4092-becb-7b2cef4217a1@huawei.com> Date: Wed, 20 May 2026 10:40:20 +0800 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH v3] mm/memory-failure: fix hugetlb_lock AA deadlock in get_huge_page_for_hwpoison To: Wupeng Ma , , , , , , , , , , , , , CC: , References: <20260520020128.3506168-1-mawupeng1@huawei.com> Content-Language: en-US From: Kefeng Wang In-Reply-To: <20260520020128.3506168-1-mawupeng1@huawei.com> Content-Type: text/plain; charset="UTF-8"; format=flowed Content-Transfer-Encoding: 7bit X-ClientProxiedBy: kwepems200001.china.huawei.com (7.221.188.67) To dggpemf100008.china.huawei.com (7.185.36.138) You should remove the v3 since this is the first version for mainline. On 5/20/2026 10:01 AM, Wupeng Ma wrote: > madvise(MADV_HWPOISON) can trigger a recursive spinlock self-deadlock > (AA deadlock) on hugetlb_lock due to a race with concurrent folio > unmapping. The race scenario: > > Thread 1 (madvise MADV_HWPOISON) Thread 2 (unmap) > ------------------------------- ----------------- > madvise_inject_error() > get_user_pages_fast() <- refcount++ > memory_failure(MF_COUNT_INCREASED) > get_huge_page_for_hwpoison() > spin_lock_irq(&hugetlb_lock) > // refcount == 2 (gup + map) > // MF_COUNT_INCREASED path: > count_increased = true > zap_pte_range() > page_remove_rmap() > put_page() <- drops map ref > // refcount: 2 -> 1 > hugetlb_update_hwpoison() > -> MF_HUGETLB_FOLIO_PRE_POISONED > -> goto out > out: > folio_put(folio) <- drops gup ref > // refcount: 1 -> 0 > free_huge_folio() > spin_lock_irq(&hugetlb_lock) <- AA DEADLOCK > > When Thread 2's put_page() drops the mapping reference while Thread 1 > holds hugetlb_lock, the folio refcount drops to 1. The subsequent > folio_put() at the out: label frees the folio, and free_huge_folio() > attempts to re-acquire the non-recursive hugetlb_lock on the same CPU, > resulting in an AA self-deadlock. > > The same deadlock can also occur on the folio_try_get() path: when a > migratable folio is found and folio_try_get() succeeds (refcount rises > to refcount+1), a concurrent unmap and a hugetlb_update_hwpoison() > returning pre-poisoned status will land at out: where folio_put() again > may free the folio under hugetlb_lock. > > Fix this by removing the hugetlb_lock wrapper from hugetlb.c and > moving the lock acquisition directly into get_huge_page_for_hwpoison() > (formerly __get_huge_page_for_hwpoison) in memory-failure.c. Place > spin_unlock_irq() before folio_put() at the out: label so that the > folio is always released outside the lock, preventing any recursive > lock acquisition. The following comment is not useful and correct for this version. Remove the now-incorrect "Called from hugetlb code > with hugetlb_lock held" comment, and update the stale > __get_huge_page_for_hwpoison declarations in include/linux/mm.h. > The change is lgtm, Reviewed-by: Kefeng Wang > Fixes: 405ce051236c ("mm/hwpoison: fix race between hugetlb free/demotion and memory_failure_hugetlb()") > Signed-off-by: Wupeng Ma > --- > include/linux/hugetlb.h | 8 -------- > include/linux/mm.h | 8 -------- > mm/hugetlb.c | 11 ----------- > mm/memory-failure.c | 8 ++++---- > 4 files changed, 4 insertions(+), 31 deletions(-) > > diff --git a/include/linux/hugetlb.h b/include/linux/hugetlb.h > index 65910437be1ca..aa3eb42e0a01a 100644 > --- a/include/linux/hugetlb.h > +++ b/include/linux/hugetlb.h > @@ -153,8 +153,6 @@ long hugetlb_unreserve_pages(struct inode *inode, long start, long end, > long freed); > bool folio_isolate_hugetlb(struct folio *folio, struct list_head *list); > int get_hwpoison_hugetlb_folio(struct folio *folio, bool *hugetlb, bool unpoison); > -int get_huge_page_for_hwpoison(unsigned long pfn, int flags, > - bool *migratable_cleared); > void folio_putback_hugetlb(struct folio *folio); > void move_hugetlb_state(struct folio *old_folio, struct folio *new_folio, int reason); > void hugetlb_fix_reserve_counts(struct inode *inode); > @@ -422,12 +420,6 @@ static inline int get_hwpoison_hugetlb_folio(struct folio *folio, bool *hugetlb, > return 0; > } > > -static inline int get_huge_page_for_hwpoison(unsigned long pfn, int flags, > - bool *migratable_cleared) > -{ > - return 0; > -} > - > static inline void folio_putback_hugetlb(struct folio *folio) > { > } > diff --git a/include/linux/mm.h b/include/linux/mm.h > index abb4963c1f064..46e5936dabaa8 100644 > --- a/include/linux/mm.h > +++ b/include/linux/mm.h > @@ -4602,8 +4602,6 @@ extern int soft_offline_page(unsigned long pfn, int flags); > */ > extern const struct attribute_group memory_failure_attr_group; > extern void memory_failure_queue(unsigned long pfn, int flags); > -extern int __get_huge_page_for_hwpoison(unsigned long pfn, int flags, > - bool *migratable_cleared); > void num_poisoned_pages_inc(unsigned long pfn); > void num_poisoned_pages_sub(unsigned long pfn, long i); > #else > @@ -4611,12 +4609,6 @@ static inline void memory_failure_queue(unsigned long pfn, int flags) > { > } > > -static inline int __get_huge_page_for_hwpoison(unsigned long pfn, int flags, > - bool *migratable_cleared) > -{ > - return 0; > -} > - > static inline void num_poisoned_pages_inc(unsigned long pfn) > { > } > diff --git a/mm/hugetlb.c b/mm/hugetlb.c > index 327eaa4074d39..4c99bb868ad08 100644 > --- a/mm/hugetlb.c > +++ b/mm/hugetlb.c > @@ -7170,17 +7170,6 @@ int get_hwpoison_hugetlb_folio(struct folio *folio, bool *hugetlb, bool unpoison > return ret; > } > > -int get_huge_page_for_hwpoison(unsigned long pfn, int flags, > - bool *migratable_cleared) > -{ > - int ret; > - > - spin_lock_irq(&hugetlb_lock); > - ret = __get_huge_page_for_hwpoison(pfn, flags, migratable_cleared); > - spin_unlock_irq(&hugetlb_lock); > - return ret; > -} > - > /** > * folio_putback_hugetlb - unisolate a hugetlb folio > * @folio: the isolated hugetlb folio > diff --git a/mm/memory-failure.c b/mm/memory-failure.c > index ee42d43613097..28522180cf7f8 100644 > --- a/mm/memory-failure.c > +++ b/mm/memory-failure.c > @@ -1966,10 +1966,7 @@ void folio_clear_hugetlb_hwpoison(struct folio *folio) > folio_free_raw_hwp(folio, true); > } > > -/* > - * Called from hugetlb code with hugetlb_lock held. > - */ > -int __get_huge_page_for_hwpoison(unsigned long pfn, int flags, > +static int get_huge_page_for_hwpoison(unsigned long pfn, int flags, > bool *migratable_cleared) > { > struct page *page = pfn_to_page(pfn); > @@ -1977,6 +1974,7 @@ int __get_huge_page_for_hwpoison(unsigned long pfn, int flags, > bool count_increased = false; > int ret, rc; > > + spin_lock_irq(&hugetlb_lock); > if (!folio_test_hugetlb(folio)) { > ret = MF_HUGETLB_NON_HUGEPAGE; > goto out; > @@ -2013,8 +2011,10 @@ int __get_huge_page_for_hwpoison(unsigned long pfn, int flags, > *migratable_cleared = true; > } > > + spin_unlock_irq(&hugetlb_lock); > return ret; > out: > + spin_unlock_irq(&hugetlb_lock); > if (count_increased) > folio_put(folio); > return ret;