From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-dl1-f54.google.com (mail-dl1-f54.google.com [74.125.82.54]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1C30A38BF8D for ; Sat, 10 Oct 2026 16:21:33 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.82.54 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1791649295; cv=none; b=c3pvzB1+dkF6PrMS1+HzHE4UlA2AAKOhz6gsLanV44w8SvLpnjeJX0AQI5UfG4ECEe86tU1d8+vuzfRJibJj9F3yRLyDoYmPaP2MAkMWnCDtEsw9uKlydKCoawB2F5OJxfRqGoomAausrmYieG1bLWJTWgKbTU9boFpocb8enZg= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1791649295; c=relaxed/simple; bh=Jf9uxSOi0/xM7Jz6cv13LtM+3Zv+s2pvZFEqWS1Sh2g=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=nYrFjg70jVVsCy6RzAmHl6OkNWoKBlCvQaPVRSRpeyZ+YcAcH/4kboZ4LlVIc3JFk+ddDt7SBj2r6x0zCFBLfGlXEI96q1TLuICgNL0fWiqSupPZTAbjF2c4Yy0fkDT/bP1FGExSZEZxjMTmGhcAr2RhmULEfOdqqgzP77FYqME= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=amcIm2kT; arc=none smtp.client-ip=74.125.82.54 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="amcIm2kT" Received: by mail-dl1-f54.google.com with SMTP id a92af1059eb24-16079d54c17so1000842c88.0 for ; Sat, 10 Oct 2026 09:21:33 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1791649293; x=1792254093; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=EOFRoJt6U85Tsj+Ev9WdukczofMEitvTriF7KGly0Fw=; b=amcIm2kTOo/7Cy4qE08bncn+JjF7EGQQfqdFcaq91K9B46gYi7rw7ZzhfmNb6w83bS nCalWLB2harh4jzJ4JgNNkC69Hk3NQKdqmQ04QH2toQ+JsYSsVN5zVEpoR43ywSgffqb 6v49iv6xHJCDNQdTaqiDfhgYyM1kml3ejp/+V1F6Hr8rXhgDXh/3LfZrS2YIELGeTOTx JHuxZBgPQYCQE+am9Ei+K3rELdexUqIQ/MiVuSwRDmCw5rHYdAuUFfsfFF1bJNudpNFH SD1Rwc8a7pXmGOLz3vRXOCdLFzB67DsUaOkw1x5UF8l3C6D+i3si/x7tiT0IH+ymAVEM Nabg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1791649293; x=1792254093; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=EOFRoJt6U85Tsj+Ev9WdukczofMEitvTriF7KGly0Fw=; b=clph/HleXLe2XKYwTxuQrT3p7pzio6nI90VZGnO78jLOZI5bqQhb22c1NSZMfArJwr G63Y5LXamZ5SluFTtyJXVreewzJud0NGuQ9wixKTEQLNJy1RMXur7zggSKLasztiaETO CW+Jb9pG/nKlyLxVpLSmrjhX2ULQiMAq0F4R6cmG5EEK4YmMKPpK/81ZBdN6jYFrptGW ly4stirUJzx/jvRC00+tore1MqHkiOOXahE4da3osMS2UFROrenf/7HYBjv4zsO0niJd 5/XCC9RhYQSbcKfJBw8AhvsNumfbqoTOr7oJViOODBR1fkPhY2elobiKKtJpJ9jZ4T9/ j7Sw== X-Forwarded-Encrypted: i=1; AKwUvBxuukxiYOXEtTo0UyDrNYhNA0UQbywjexrGb8GXjCvfBK9xCPy5AagzayMQ/B1MUbSCZ/vjWn5f8Ni7T7Q=@vger.kernel.org X-Gm-Message-State: AFq9FYKdvuFjhxJYUTNPWgk94iWXD1Gkrnf/49nJ/AcNqtU6NCiZhBfW /XoR3Dafu/hY1okKqp72OH6feQ5e13QaWEAi/0mQZqZFXa9uMfjopkzb X-Gm-Gg: AYBFou3nGLem2B9qVtrVoeiTDjj1gWnzoHpdecS7vIouxK6xA5vOIUfubrkiC63xB90 /iYwZTUFt8g7juLkZ6uj7tN6xRCWnqt5/Uh0POSOQLaX2h2iGuxEM5zBDcJjLGanPkBwSaJUijg ZrpyfSmJY2th77BdA3UfAp0yTR9OK1B7GqPY1UF51+6mdTXOZPUPV/oTv8e+k+aiTHyDZn2PYhO bbrNKVscjEbXFvhAb7gnHXThbjEPmNCNatwkGLVmrOUaaaRLQky5pZWwm0NRJnyzpYWTTxWyBcw ziMwBxjxAN3BdpthWnyxUnN94yYGnwhU4KdbJ8Iy2jppQdsDoTkfx3MoK/y+9wxrn3VKUQ4ZyJz ZEGCUMa1SqPr4toZNVtGKIHB4ivRAXSa7ole9H4qmPqayrFeyhciOkEMG9T4572VteO8lXaKFFI i3Aua0RCnp/BUg2PkdB/qOBmjI7bZ9sGfBc8bUmyIlhimqLF84DcokxjUWr7GzrlNwRFWuMGHZy ylF/ufFNpdApLay1w== X-Received: by 2002:a05:701b:2212:b0:144:c124:279 with SMTP id a92af1059eb24-16a5f7c79c8mr7463252c88.35.1791649292992; Sat, 10 Oct 2026 09:21:32 -0700 (PDT) Received: from acer-nitro-anv15-41.. ([115.96.176.39]) by smtp.gmail.com with ESMTPSA id a92af1059eb24-169a3f6186fsm15778622c88.14.2026.10.10.09.21.23 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sat, 10 Oct 2026 09:21:32 -0700 (PDT) From: Shaikh Kamaluddin To: Davidlohr Bueso , Jonathan Cameron , Dave Jiang , Alison Schofield , Vishal Verma , Dan Williams , Miaohe Lin , Andrew Morton , Vlastimil Babka , Breno Leitao Cc: Ira Weiny , Li Ming , Richard Cheng , Naoya Horiguchi , Suren Baghdasaryan , Michal Hocko , Brendan Jackman , Johannes Weiner , Zi Yan , linux-cxl@vger.kernel.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org, David Hildenbrand , Oscar Salvador , Kiryl Shutsemau , Harry Yoo , Shaikh Kamaluddin Subject: [RFC PATCH 3/5] mm/memory-failure: Add a pre-online poison registry Date: Sat, 10 Oct 2026 21:50:11 +0530 Message-ID: <20261010162017.62506-4-shaikhkamal2012@gmail.com> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20261010162017.62506-1-shaikhkamal2012@gmail.com> References: <20261010162017.62506-1-shaikhkamal2012@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit A frame can be known bad before its backing memory is onlined. A CXL Type-3 device retains poison as DPA-based Media Error Records and can report those records through GET_POISON_LIST after a reboot, kexec, or device hotplug. CXL poison can be retrieved while a region is being prepared, before dax/kmem calls add_memory_driver_managed() for the range. At that point there is no initialized struct page on which memory_failure() can operate. The poison information must therefore be retained until the corresponding pages pass through memory initialization. Add a registry of sorted, nonoverlapping poisoned PFN ranges. Each registration represents one owner and its independently managed memory. For example, a CXL region registers the poisoned PFN ranges translated from its device poison records and receives an opaque handle for that range set. A second CXL region receives a separate handle for its own ranges. The owner uses its handle to remove exactly its registration after the associated memory can no longer enter the allocator. Removing one CXL region's registration does not affect poison ranges registered by another region. Keep each registered range set immutable and protect the registration list with RCU. This keeps allocator-side lookups lockless while allowing a registration to be removed safely after an RCU grace period. A later patch will consult the registry before initialized pages are released to the buddy allocator. Signed-off-by: Shaikh Kamaluddin --- include/linux/memory-failure.h | 92 +++++++++++++++++ mm/memory-failure.c | 175 +++++++++++++++++++++++++++++++++ 2 files changed, 267 insertions(+) diff --git a/include/linux/memory-failure.h b/include/linux/memory-failure.h index d333dcdbeae7..ebb988c2fe12 100644 --- a/include/linux/memory-failure.h +++ b/include/linux/memory-failure.h @@ -11,9 +11,73 @@ struct pfn_address_space { unsigned long pfn, pgoff_t *pgoff); }; +struct preonline_hwpoison; + +/** + * struct preonline_hwpoison_range - Pre-online poisoned PFN range + * @start_pfn: First PFN in the range + * @nr_pages: Number of consecutive PFNs in the range + */ +struct preonline_hwpoison_range { + unsigned long start_pfn; + unsigned long nr_pages; +}; + #ifdef CONFIG_MEMORY_FAILURE int register_pfn_address_space(struct pfn_address_space *pfn_space); void unregister_pfn_address_space(struct pfn_address_space *pfn_space); + +/** + * preonline_hwpoison_register - Register poison before memory is onlined + * @ranges: Sorted, nonoverlapping PFN ranges in ascending start PFN order. + * Adjacent ranges are accepted, but callers should coalesce them. + * The array is copied and need not outlive this call. + * @nr_ranges: Number of entries in @ranges. Must be nonzero. + * @handle: Registration handle returned on success. Set to NULL on failure. + * + * Register PFN ranges that must be withheld when their backing memory is + * initialized. Pass the returned handle to + * preonline_hwpoison_unregister() only after the memory can no longer enter + * the allocator. + * + * Return: 0 on success, -EINVAL if @handle is NULL or @ranges is empty, + * unsorted, overlapping, or contains a zero-length entry, -EOVERFLOW on PFN + * wraparound, and -ENOMEM on allocation failure. + */ +int preonline_hwpoison_register(const struct preonline_hwpoison_range *ranges, + unsigned int nr_ranges, + struct preonline_hwpoison **handle); + +/** + * preonline_hwpoison_unregister - Remove a pre-online poison registration + * @handle: Handle returned by preonline_hwpoison_register(), or NULL + * + * Remove @handle from future pre-online poison lookups. The caller must first + * ensure that the associated memory can no longer be initialized or handed + * to the page allocator. + * + * Pages already marked as hwpoison remain marked. + */ +void preonline_hwpoison_unregister(struct preonline_hwpoison *handle); + +/** + * preonline_hwpoison_intersects - Does a PFN range cover a known-bad frame? + * @start_pfn: First PFN of the range to test + * @nr_pages: Length of the range in frames + * + * Return: %true if any registered range overlaps [@start_pfn, @start_pfn + + * @nr_pages), %false otherwise. + */ +bool preonline_hwpoison_intersects(unsigned long start_pfn, + unsigned long nr_pages); + +/** + * preonline_hwpoison_contains - Is a single frame known bad? + * @pfn: Frame to test + * + * Return: %true if @pfn falls in a registered range. + */ +bool preonline_hwpoison_contains(unsigned long pfn); #else static inline int register_pfn_address_space(struct pfn_address_space *pfn_space) { @@ -23,6 +87,34 @@ static inline int register_pfn_address_space(struct pfn_address_space *pfn_space static inline void unregister_pfn_address_space(struct pfn_address_space *pfn_space) { } + +static inline int +preonline_hwpoison_register(const struct preonline_hwpoison_range *ranges, + unsigned int nr_ranges, + struct preonline_hwpoison **handle) +{ + if (handle) + *handle = NULL; + + return -EOPNOTSUPP; +} + +static inline void +preonline_hwpoison_unregister(struct preonline_hwpoison *handle) +{ +} + +static inline bool preonline_hwpoison_intersects(unsigned long start_pfn, + unsigned long nr_pages) +{ + return false; +} + +static inline bool preonline_hwpoison_contains(unsigned long pfn) +{ + return false; +} + #endif /* CONFIG_MEMORY_FAILURE */ #endif /* _LINUX_MEMORY_FAILURE_H */ diff --git a/mm/memory-failure.c b/mm/memory-failure.c index 60e968243470..e89b400c3fa2 100644 --- a/mm/memory-failure.c +++ b/mm/memory-failure.c @@ -2291,6 +2291,181 @@ static void kill_procs_now(struct page *p, unsigned long pfn, int flags, kill_procs(&tokill, true, pfn, flags); } +/** + * struct preonline_hwpoison - Owner of pre-online poisoned PFN ranges + * @list: Entry in the global owner list + * @rcu: Used to defer the free past a grace period + * @nr_ranges: Number of ranges in @ranges + * @ranges: Sorted, nonoverlapping PFN ranges + * + * The ranges are immutable after the object is added to the owner list, which is + * what lets readers walk them under RCU with no lock. + */ +struct preonline_hwpoison { + struct list_head list; + struct rcu_head rcu; + unsigned int nr_ranges; + struct preonline_hwpoison_range ranges[]; +}; + +static LIST_HEAD(preonline_hwpoison_owners); +/* Serialises writers only. Readers use RCU. */ +static DEFINE_MUTEX(preonline_hwpoison_lock); + +static int +preonline_hwpoison_validate(const struct preonline_hwpoison_range *ranges, + unsigned int nr_ranges) +{ + unsigned long previous_end = 0; + + if (!ranges || !nr_ranges) + return -EINVAL; + + for (unsigned int i = 0; i < nr_ranges; i++) { + unsigned long end; + + if (!ranges[i].nr_pages) + return -EINVAL; + + if (check_add_overflow(ranges[i].start_pfn, + ranges[i].nr_pages - 1, &end)) + return -EOVERFLOW; + + /* + * Require sorted, nonoverlapping ranges. Adjacent ranges + * remain valid, although callers may coalesce them to reduce + * storage and lookup overhead. + */ + if (i && ranges[i].start_pfn <= previous_end) + return -EINVAL; + + previous_end = end; + } + + return 0; +} + +int preonline_hwpoison_register(const struct preonline_hwpoison_range *ranges, + unsigned int nr_ranges, + struct preonline_hwpoison **handle) +{ + struct preonline_hwpoison *owner; + int rc; + + if (!handle) + return -EINVAL; + + *handle = NULL; + + rc = preonline_hwpoison_validate(ranges, nr_ranges); + if (rc) + return rc; + + /* + * The range array may be larger than a page, so do not require + * physically contiguous memory. + */ + owner = kvmalloc(struct_size(owner, ranges, nr_ranges), GFP_KERNEL); + if (!owner) + return -ENOMEM; + + owner->nr_ranges = nr_ranges; + memcpy(owner->ranges, ranges, + flex_array_size(owner, ranges, nr_ranges)); + + /* + * list_add_tail_rcu() publishes with a release barrier, so a reader + * that observes the entry also observes the fully initialised + * ranges[] above. + */ + scoped_guard(mutex, &preonline_hwpoison_lock) { + list_add_tail_rcu(&owner->list, &preonline_hwpoison_owners); + } + + *handle = owner; + + return 0; +} +EXPORT_SYMBOL_GPL(preonline_hwpoison_register); + +void preonline_hwpoison_unregister(struct preonline_hwpoison *owner) +{ + if (!owner) + return; + + scoped_guard(mutex, &preonline_hwpoison_lock) { + list_del_rcu(&owner->list); + } + + /* + * A reader may still be walking this entry. list_del_rcu() leaves its + * forward pointer intact so that walk completes into the live list; + * the free waits for a grace period. + */ + kvfree_rcu(owner, rcu); +} +EXPORT_SYMBOL_GPL(preonline_hwpoison_unregister); + +static bool preonline_owner_intersects(const struct preonline_hwpoison *owner, + unsigned long start_pfn, + unsigned long end_pfn) +{ + unsigned int low = 0; + unsigned int high = owner->nr_ranges; + + /* + * Find the first range whose inclusive end is greater than or equal + * to start_pfn. + */ + while (low < high) { + const struct preonline_hwpoison_range *range; + unsigned long range_end; + unsigned int mid; + + mid = low + (high - low) / 2; + range = &owner->ranges[mid]; + range_end = range->start_pfn + range->nr_pages - 1; + + if (range_end < start_pfn) + low = mid + 1; + else + high = mid; + } + + if (low == owner->nr_ranges) + return false; + + return owner->ranges[low].start_pfn <= end_pfn; +} + +bool preonline_hwpoison_intersects(unsigned long start_pfn, + unsigned long nr_pages) +{ + struct preonline_hwpoison *owner; + unsigned long end_pfn; + + if (!nr_pages) + return false; + + /* Fail safe: withhold a range we cannot reason about. */ + if (WARN_ON_ONCE(check_add_overflow(start_pfn, nr_pages - 1, &end_pfn))) + return true; + + guard(rcu)(); + + list_for_each_entry_rcu(owner, &preonline_hwpoison_owners, list) { + if (preonline_owner_intersects(owner, start_pfn, end_pfn)) + return true; + } + + return false; +} + +bool preonline_hwpoison_contains(unsigned long pfn) +{ + return preonline_hwpoison_intersects(pfn, 1); +} + int register_pfn_address_space(struct pfn_address_space *pfn_space) { guard(mutex)(&pfn_space_lock); -- 2.43.0