From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from casper.infradead.org (casper.infradead.org [90.155.50.34]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C0E3E40862A for ; Mon, 27 Jul 2026 18:21:08 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=90.155.50.34 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785176470; cv=none; b=O4NSecxEEMRBgXnbIsnCmSaOdoKIABOwU7orKG1WsXNByLzsWkC6VvyVMsLHHHbT/RFZR4tGsH5p4wQY+cH3XKKidwu6oSybPRst+v8uMV7BtyxTvGrQ1f1THJWBCvgByLdYf1PqZbcb4z/GXA+FU9i77VfrzsfuoSaJbtaLb5c= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785176470; c=relaxed/simple; bh=us7+PaTsjNCOj1hesmbVtrOeT+s+uuSK/gI1Eyd/zGk=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=nRxeSX2EEytu4v+SigWfCYek+dKXizpI1huBSky4xCxfzvJPzpa4nUsupklMHxBIZ7S1LPv8Vk3loMSrvBXxprMlx9hvOJXiqZ2HdwQZ/KbmhiMvOo/RgwMd7dN10lm9SZzYFNAQx6UXmScV2gMoBFEuAkbjXmMrjQaE4awQnu4= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=infradead.org; spf=pass smtp.mailfrom=infradead.org; dkim=pass (2048-bit key) header.d=infradead.org header.i=@infradead.org header.b=sFCBoWUU; arc=none smtp.client-ip=90.155.50.34 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=infradead.org Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=infradead.org Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=infradead.org header.i=@infradead.org header.b="sFCBoWUU" DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=infradead.org; s=casper.20170209; h=In-Reply-To:Content-Type:MIME-Version: References:Message-ID:Subject:Cc:To:From:Date:Sender:Reply-To: Content-Transfer-Encoding:Content-ID:Content-Description; bh=bQ4rvhkya2fSTiIGocjlzq+705xNxIlgjGPvKqk8WCE=; b=sFCBoWUUjmQrn9tNM8L/Mt9RbL ErJLZbBhPPBL1oOZoxwgY8VXtJxOqD+XgRxB6FiP5wQgyo4aRZToporXWy7WDzWQ3zd+IkoW7ufiD Wcz5oPsLv6jD2DecxMLS/eqwv4D/Fok2KEDa7+LBIhOkTcwzckx7rck3ZO59PrZy6I89x5Kuz2Acu ExrN7adRmfscEVOu+CKiFdPZgFg62uY2sVb2zGZtHdi8s1ovWuP7HxjQBXjyKQAlPomqvdsywgY8k gni1QFD40YBqOLOZzpFqvubGDtDmy3AEOwOH8cX0jdqIvtYHiedEbAlnAk4Qy16sKpyLLExY/oDj3 QFQl8Dlw==; Received: from willy by casper.infradead.org with local (Exim 4.99.1 #2 (Red Hat Linux)) id 1woPwY-0000000BDxl-3gF9; Mon, 27 Jul 2026 18:20:50 +0000 Date: Mon, 27 Jul 2026 19:20:50 +0100 From: Matthew Wilcox To: "David Hildenbrand (Arm)" Cc: "Vlastimil Babka (SUSE)" , William Roche , Jiaqi Yan , linmiaohe@huawei.com, ljs@kernel.org, ziy@nvidia.com, osalvador@kernel.org, harry.yoo@oracle.com, osalvador@suse.de, jackmanb@google.com, hannes@cmpxchg.org, nao.horiguchi@gmail.com, tony.luck@intel.com, wangkefeng.wang@huawei.com, jane.chu@oracle.com, akpm@linux-foundation.org, muchun.song@linux.dev, liam@infradead.org, rientjes@google.com, duenwen@google.com, jthoughton@google.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, rppt@kernel.org, shuah@kernel.org, surenb@google.com, mhocko@suse.com, boudewijn@delta-utec.com Subject: Re: [PATCH v6 0/5] Only free healthy pages in high-order has_hwpoisoned folio Message-ID: References: <20260705180714.3708947-1-jiaqiyan@google.com> <85cb7ea8-8116-4092-8310-69b61eb8602c@kernel.org> <1dea7b3c-7740-474d-b9d4-cd2baf47f181@kernel.org> <168ad406-f6b5-4623-adac-9f894e410e48@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <168ad406-f6b5-4623-adac-9f894e410e48@kernel.org> On Mon, Jul 27, 2026 at 04:20:54PM +0200, David Hildenbrand (Arm) wrote: > >> Just adding a comment about this aspect: > >> The check_new_pages() mechanism used by the __rmqueue functions should > >> filter these pages out, but this has been disabled by default in 2023 > >> with: > >> [PATCH] mm, page_alloc: reduce page alloc/free sanity checks > >> https://lore.kernel.org/all/20230216095131.17336-1-vbabka@suse.cz > >> > >> So it would need to be enabled back, taking some of the performance hit. > >> (and I personally think that it has to be done) > > > > Would it truly fix the issue, or rather there would still be a race window > > left where we check that there's no hwpoison flag in the re-enabled check, > > and only then someone sets it? > > Why are we checking PageHWPoison at all then in check_new_page()? > > I think we created a mess. > > The PageHWPoison check is not just a "nice to have" sanity check for kernel bugs. > > So it should never have been optimized out that way before reworking the bigger > picture. There's A Lot Going On (and I don't think I understand it all yet). We can soft-poison pages while they're in Buddy, for example. And then soft-unpoison them again. Is it handled properly? I doubt it. Looks to me like it's full of races. > > Also, can the hardware actually detect a problem with a page that nobody > > accesses? I guess if yes, it's only in some corner cases. > > Yes, quite frequently I think. > > > > > So I'm wary about penalizing the allocator paths again. > > I get the feeling that we don't have a proper plan on how to handle HWPoisoned > pages. We should take a step back and discuss how we actually want to handle > them instead of optimizing here and there and creating more of a mess. I feel like there's a general lack of understanding of how hwpoison works amongst those of us who work on the core of memory handling, and the memory-failure code has not been kept up to date with how we think about page/folio/memdesc handling. It might be good to have a BOF at Plumbers to share our (mis)understandings of how all of this works yesterday/today/tomorrow.