From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pj1-f74.google.com (mail-pj1-f74.google.com [209.85.216.74]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id EB28E34D4EA for ; Sun, 5 Jul 2026 18:07:17 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.74 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783274839; cv=none; b=Cq+fcCL3q1puci3mhFk8PpA58LP/2k/5VQ76lh3s5ybg9W7860erpc4zBpVws3LxiHy+2kGWc3/xG0sNSSlKCDRDSI/TbiewngkuxXAU0ptcKwl0fCRPUDnhXYDBuRDGXZmCFkpREyXfHs9MPq612SdojahJo+l1JXUNzB/wgt8= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783274839; c=relaxed/simple; bh=k8mwuvdv2NDIMYl/Rn9cZOQRcWwgg4VoUHVh2cH4854=; h=Date:Mime-Version:Message-ID:Subject:From:To:Cc:Content-Type; b=L1BHanXqbcZcMaKADLlYjyhMNagko/blKcqeVLN/W5iK2O8ZQkowd2X8xOuJKQOtAKYGqV9RfRTfJ5t+Sc6JBI91bOGEBLan+WE8KqEyao0m7l3BIuDEVY73/3qUte8FtxmO6ZAC5PfFnt/Mbw7hwrF9D6m1ANjGjQ8o401edpI= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--jiaqiyan.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=D+HiYDYL; arc=none smtp.client-ip=209.85.216.74 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--jiaqiyan.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="D+HiYDYL" Received: by mail-pj1-f74.google.com with SMTP id 98e67ed59e1d1-380b630c505so2735947a91.1 for ; Sun, 05 Jul 2026 11:07:17 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1783274837; x=1783879637; darn=vger.kernel.org; h=cc:to:from:subject:message-id:mime-version:date:from:to:cc:subject :date:message-id:reply-to; bh=W0l/T3c1tBTPd59g2vWHPm6URn5LD4/gNmBayGnnWek=; b=D+HiYDYLvbnW1cUzDXuVmLXmKv223fzHwfemqlT3JgKoemBea+WGr4CcOJC0BEeaIK 52ddt1qzmf7uI3Z9ayKSN+Osh0+Yi6StdxJbesUHMY0QXVrZQgmZ2gcYrDyVAlIf2IOS g+p30BE/FGeAy+aDrgt/bT5zTTX3Wr9A0TZsu4snudHLTv/h0ItlyQP05hOuaFeKg8k0 qrolYyLo0Wpywe8q0PRZy33uXhMls/tk3ZWo4UXqvN4DUiKam0q28nnET87HTmQTGlsI yuXHb/QoeEhqCOrCKKK5lNdqvhQ98YTap76ptf7xkBu/l5zNUlvcr75XC5dP3iSFM08v 95dg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1783274837; x=1783879637; h=cc:to:from:subject:message-id:mime-version:date:x-gm-message-state :from:to:cc:subject:date:message-id:reply-to; bh=W0l/T3c1tBTPd59g2vWHPm6URn5LD4/gNmBayGnnWek=; b=mFSthFtmxg4sVDiSfCdOXo3kxoL5JTxDCZye/fmWzwOX9NtH4zi4zbTwUFrSvSh8VY XIo3v+nG/9L7kZ2sHLWTiY212ZVMrz5lVZdwmkLLLXG1po3vGt+V9y/zLmRrfCzAEvPw Lif01xGTfRpv/9G8ra50lXkRmRseSXKvHFsFUh5GcipHIEJeEp/rPcmOJUw/OHeVYpNG lAzPmm5OMYRVNXgNt0pY6TfiPAPUS4PpiAEU5kXVtCRvSQHePNEwQjFdR6phvS2gLaLb nqMgZ6WhF4HzCN9flqwq/sX7Kvh34QIEee4qgbo6fXe2YpYqS+RfhKJplu5OWs7oXrd9 Vi6g== X-Forwarded-Encrypted: i=1; AHgh+RpHjn5Dedrvc8kj4BKD4lM/Ssv9PNwHKOz/k1mw0PrOmQt/KQgD/JAI6tvU3UaoSDFIUiwA4dcFzeKA0h4=@vger.kernel.org X-Gm-Message-State: AOJu0YzO5mrsd/c64PMtiHTg3lJP//CsbmlpFMQZMNpZ98MzkDOzNuoc Yy0E6s9KgAWMRT6jpOQEQ49dA2NEqrCE0eks8f0Rbr3hi45T2locD5I88zxKJC4OghJoaKjSrNJ K+Rt7mpKxLZUgzw== X-Received: from pjoz16.prod.google.com ([2002:a17:90a:9810:b0:384:f6e1:ff81]) (user=jiaqiyan job=prod-delivery.src-stubby-dispatcher) by 2002:a17:90b:48d1:b0:37f:e326:6557 with SMTP id 98e67ed59e1d1-382803b8446mr7223203a91.4.1783274836929; Sun, 05 Jul 2026 11:07:16 -0700 (PDT) Date: Sun, 5 Jul 2026 18:07:09 +0000 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 X-Mailer: git-send-email 2.55.0.rc0.799.gd6f94ed593-goog Message-ID: <20260705180714.3708947-1-jiaqiyan@google.com> Subject: [PATCH v6 0/5] Only free healthy pages in high-order has_hwpoisoned folio From: Jiaqi Yan To: linmiaohe@huawei.com, ljs@kernel.org, ziy@nvidia.com, vbabka@kernel.org Cc: osalvador@kernel.org, harry.yoo@oracle.com, willy@infradead.org, osalvador@suse.de, jackmanb@google.com, hannes@cmpxchg.org, nao.horiguchi@gmail.com, david@kernel.org, william.roche@oracle.com, tony.luck@intel.com, wangkefeng.wang@huawei.com, jane.chu@oracle.com, akpm@linux-foundation.org, muchun.song@linux.dev, liam@infradead.org, rientjes@google.com, duenwen@google.com, jthoughton@google.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, vbabka@suse.cz, rppt@kernel.org, shuah@kernel.org, surenb@google.com, mhocko@suse.com, boudewijn@delta-utec.com, Jiaqi Yan Content-Type: text/plain; charset="UTF-8" At the end of dissolve_free_hugetlb_folio(), a free HugeTLB folio becomes non-HugeTLB and is released to buddy allocator as a high-order folio, e.g. a folio that contains 262144 pages if the folio was a 1G HugeTLB hugepage. This is problematic if the HugeTLB hugepage contained HWPoison subpages. In that case, since buddy allocator does not check HWPoison for non-zero-order folio, the raw HWPoison page can be given out with its buddy page and be re-used by either kernel or userspace. Memory failure recovery (MFR) code (mm/memory-failure.c) does attempt to take raw HWPoison page off buddy allocator after dissolve_free_hugetlb_folio(). However, there is always a time window between dissolve_free_hugetlb_folio() frees a HWPoison high-order folio to buddy allocator and MFR takes HWPoison raw page off buddy allocator. Another similar situation is when a transparent huge page (THP) is handled by MFR but splitting failed. Such THP will eventually be released to buddy allocator when owning userspace processes are gone, but with certain subpages having HWPoison [9]. One obvious way to avoid both problems is to add page sanity checks in page allocate or free path. However, it is against the past efforts to reduce sanity check overhead [1,2,3]. Introduce free_has_hwpoisoned() to only free the healthy pages and exclude the HWPoison ones in the high-order folio. free_has_hwpoisoned() happens at the end of free_pages_prepare(), which already deals with both decomposing the original compound page, updating page metadata like alloc tag and page owner. It is also only applied when PG_has_hwpoisoned indicates folio contains certain HWPoison page(s) for performance reason. Its idea is to iterate through the sub-pages of the folio to identify contiguous ranges of healthy pages. Instead of freeing pages one by one, free_has_hwpoisoned() then re-use free_prepared_contig_range() [11] to decompose healthy ranges into the largest possible chunks of different orders. Every chunk is freed via __free_frozen_pages(). free_has_hwpoisoned() has linear time complexity wrt the number of pages in the folio. While the power-of-two decomposition ensures that the number of calls to the buddy allocator is logarithmic for each contiguous healthy range, the mandatory linear scan of pages to identify PageHWPoison() defines the overall time complexity. I tested with some test-only code [4] and hugetlb-mfr [5], by checking the status of pcplist and freelist immediately after dissolve_free_hugetlb_folio() a free 2M or 1G HugeTLB page that contains 1~8 HWPoison raw pages: - HWPoison pages are excluded by free_has_hwpoisoned(). - Some healthy pages can be in zone->per_cpu_pageset (pcplist) because pcp_count is not high enough. Many healthy pages are in some order's zone->free_area[order].free_list (freelist). - In rare cases, some healthy pages are in neither pcplist nor freelist. My best guest is they are allocated before the test checks. To illustrate the latency free_has_hwpoisoned() added to the memory freeing path, I tested its time cost with 8 HWPoison pages with instrument code in [4] for 20 sample runs on a machine having 56 Intel Skylake physical cores and 768GB memory: - Has HWPoison path: mean=1030us, stdev=21us - No HWPoison path: mean=66us, stdev=6us free_has_hwpoisoned() is around 15x the baseline. Its cost is nontrivial, but still far from triggering soft lockup, and fair for handling exceptional hardware memory errors. Now that free_has_hwpoisoned() ensures HWPoison pages never made into buddy allocator, MFR don't need to take_page_off_buddy() anymore after disovling HWPoison hugepages. So replace __page_handle_poison() with new __hugepage_handle_poison() for HugeTLB specific call sites. It may worthy to note that this patchset doesn't affect the soft offline behavior in MFR. This is because soft offline does not folio_set_hwpoison() upfront, and for HugeTLB case, doesn't involve get_huge_page_for_hwpoison(). To provide test coverage for the new __hugepage_handle_poison() in me_huge_page(), the last commit adds a MADV_HARD test variant for anonymous HugeTLB pages. It also cover the code path that frees a HugeTLB page that contains 1 raw HWPoison page. Based on commit cfb8731f5396 ("mm: fix CONFIG_STACK_GROWSUP typo in tools/testing/vma/include/dup.h") Changelog v5 [11] -> v6 - Rebase to recent akpm/mm-unstable and address comments from Zi Yan, Miaohe Lin, Vlastimil Babka. - Extract free_pages_sanitize() to avoid touching HWPoison page(s) at the end of __free_pages_prepare(). - Introduce FPI_SANITIZE and add a fpi_t argument to the freeing path free_has_hwpoisoned() -> __free_prepared_contig_range(), so that healthy page blocks are sanitized before __free_frozen_pages(). - Make free_has_hwpoisoned() be compatible with order==0. - Repeat the previous test done on both X86 and ARM64 machines. CONFIG_KASAN + CONFIG_KASAN_SW_TAGS + kasan.fault=report are enabled on the ARM64 machine. v4 [10] -> v5 - Rebase to very recent akpm/mm-unstable. - Re-use free_prepared_contig_range() introduced by [11], and remove free_contiguous_pages() in v4. - Instead of using struct page pointer, iterate over pfn in free_has_hwpoison(). - Re-ested and re-evaluated free_has_hwpoison()'s time cost. - Add memory failure recovery test for anonymous 1G HugeTLB page to gain test coverage for __hugepage_handle_poison() and for freeing 1G HugeTLB page that has 1 HWPoison page. v3 [8] -> v4 - Address comments from Zi Yan, Miaohe Lin, Harry Yoo. - Set has_hwpoisoned flag after introducing free_has_hwpoisoned(). - Unwrap free_pages_prepare_has_hwpoisoned() into free_pages_prepare(). - If folio has HWPoison, its healthy pages will be freed with FPI_NONE right in free_pages_prepare(), who returns false to indicate caller should not proceeding its own freeing action. - Rework the commit on __page_handle_poison. Only change the handling for HWPoison HugeTLB page, leaving free buddy page and soft offline handling alone. v2 [7] -> v3: - Address comments from Mathew Wilcox, Harry Hoo, Miaohe Lin. - Let free_has_hwpoisoned() happen after free_pages_prepare(), which help to deal with decomposing the original compound page, and with page metadata like alloc tag and page owner. - Tested with "page_owner=on" and CONFIG_MEM_ALLOC_PROFILING*=y. - Wrap checking PG_has_hwpoisoned and free_has_hwpoisoned() into free_pages_prepare_has_hwpoisoned(), which replaces free_pages_prepare() calls in free_frozen_pages(). - Rename free_has_hwpoison_page() to free_has_hwpoisoned(). - Measure latency added by free_has_hwpoisoned(). - Ensure struct page *end is only used for pointer arithmetic, instead of accessed as page. - Refactor page_handl_poison instead of just __page_handle_poison(). v1 [6] -> v2: - Total reimplementation based on discussions with Mathew Wilcox, Harry Hoo, Zi Yan etc - hugetlb_free_hwpoison_folio() => free_has_hwpoison_pages(). - Utilize has_hwpoisoned flag to tell buddy allocator a high-order folio contains HWPoison. - Simplify __page_handle_poison() given that the HWPoison page(s) won't be freed within high-order folio. [1] https://lore.kernel.org/linux-mm/1460711275-1130-15-git-send-email-mgorman@techsingularity.net [2] https://lore.kernel.org/linux-mm/1460711275-1130-16-git-send-email-mgorman@techsingularity.net [3] https://lore.kernel.org/all/20230216095131.17336-1-vbabka@suse.cz [4] https://drive.google.com/file/d/1CzJn1Cc4wCCm183Y77h244fyZIkTLzCt/view?usp=sharing [5] https://lore.kernel.org/linux-mm/20251116013223.1557158-3-jiaqiyan@google.com [6] https://lore.kernel.org/linux-mm/20251116014721.1561456-1-jiaqiyan@google.com [7] https://lore.kernel.org/linux-mm/20251219183346.3627510-1-jiaqiyan@google.com [8] https://lore.kernel.org/linux-mm/20260112004923.888429-1-jiaqiyan@google.com [9] https://lore.kernel.org/linux-mm/20260113205441.506897-1-boudewijn@delta-utec.com [10] https://lore.kernel.org/linux-mm/20260202194125.2191216-1-jiaqiyan@google.com [11] https://lore.kernel.org/all/20260401101634.2868165-2-usama.anjum@arm.com [12] https://lore.kernel.org/linux-mm/20260531055829.3636554-1-jiaqiyan@google.com Jiaqi Yan (5): mm/page_alloc: introduce __free_prepared_contig_range() with fpi_t mm/page_alloc: only free healthy pages in high-order has_hwpoisoned folio mm/memory-failure: set has_hwpoisoned flags on dissolved HugeTLB folio mm/memory-failure: skip take_page_off_buddy after dissolving HWPoison HugeTLB page selftests/mm: add hard memory failure anonymous HugeTLB test include/linux/page-flags.h | 2 +- mm/memory-failure.c | 37 +++- mm/page_alloc.c | 176 ++++++++++++++++---- tools/testing/selftests/mm/memory-failure.c | 70 +++++++- 4 files changed, 247 insertions(+), 38 deletions(-) -- 2.55.0.rc0.799.gd6f94ed593-goog