From: Shaikh Kamaluddin <shaikhkamal2012@gmail.com>
To: Davidlohr Bueso <dave@stgolabs.net>,
Jonathan Cameron <jic23@kernel.org>,
Dave Jiang <dave.jiang@intel.com>,
Alison Schofield <alison.schofield@intel.com>,
Vishal Verma <vishal.l.verma@intel.com>,
Dan Williams <djbw@kernel.org>, Miaohe Lin <linmiaohe@huawei.com>,
Andrew Morton <akpm@linux-foundation.org>,
Vlastimil Babka <vbabka@kernel.org>,
Breno Leitao <leitao@debian.org>
Cc: Ira Weiny <iweiny@kernel.org>, Li Ming <ming.li@zohomail.com>,
Richard Cheng <icheng@nvidia.com>,
Naoya Horiguchi <nao.horiguchi@gmail.com>,
Suren Baghdasaryan <surenb@google.com>,
Michal Hocko <mhocko@suse.com>,
Brendan Jackman <brendan.jackman@linux.dev>,
Johannes Weiner <hannes@cmpxchg.org>, Zi Yan <ziy@nvidia.com>,
linux-cxl@vger.kernel.org, linux-kernel@vger.kernel.org,
linux-mm@kvack.org, David Hildenbrand <david@kernel.org>,
Oscar Salvador <osalvador@suse.de>,
Kiryl Shutsemau <kas@kernel.org>, Harry Yoo <harry@kernel.org>,
Shaikh Kamaluddin <shaikhkamal2012@gmail.com>
Subject: [RFC PATCH 3/5] mm/memory-failure: Add a pre-online poison registry
Date: Sat, 10 Oct 2026 21:50:11 +0530 [thread overview]
Message-ID: <20261010162017.62506-4-shaikhkamal2012@gmail.com> (raw)
In-Reply-To: <20261010162017.62506-1-shaikhkamal2012@gmail.com>
A frame can be known bad before its backing memory is onlined. A CXL
Type-3 device retains poison as DPA-based Media Error Records and can
report those records through GET_POISON_LIST after a reboot, kexec, or
device hotplug.
CXL poison can be retrieved while a region is being prepared, before
dax/kmem calls add_memory_driver_managed() for the range. At that point
there is no initialized struct page on which memory_failure() can
operate. The poison information must therefore be retained until the
corresponding pages pass through memory initialization.
Add a registry of sorted, nonoverlapping poisoned PFN ranges. Each
registration represents one owner and its independently managed memory.
For example, a CXL region registers the poisoned PFN ranges translated
from its device poison records and receives an opaque handle for that
range set. A second CXL region receives a separate handle for its own
ranges.
The owner uses its handle to remove exactly its registration after the
associated memory can no longer enter the allocator. Removing one CXL
region's registration does not affect poison ranges registered by
another region.
Keep each registered range set immutable and protect the registration
list with RCU. This keeps allocator-side lookups lockless while allowing
a registration to be removed safely after an RCU grace period. A later
patch will consult the registry before initialized pages are released
to the buddy allocator.
Signed-off-by: Shaikh Kamaluddin <shaikhkamal2012@gmail.com>
---
include/linux/memory-failure.h | 92 +++++++++++++++++
mm/memory-failure.c | 175 +++++++++++++++++++++++++++++++++
2 files changed, 267 insertions(+)
diff --git a/include/linux/memory-failure.h b/include/linux/memory-failure.h
index d333dcdbeae7..ebb988c2fe12 100644
--- a/include/linux/memory-failure.h
+++ b/include/linux/memory-failure.h
@@ -11,9 +11,73 @@ struct pfn_address_space {
unsigned long pfn, pgoff_t *pgoff);
};
+struct preonline_hwpoison;
+
+/**
+ * struct preonline_hwpoison_range - Pre-online poisoned PFN range
+ * @start_pfn: First PFN in the range
+ * @nr_pages: Number of consecutive PFNs in the range
+ */
+struct preonline_hwpoison_range {
+ unsigned long start_pfn;
+ unsigned long nr_pages;
+};
+
#ifdef CONFIG_MEMORY_FAILURE
int register_pfn_address_space(struct pfn_address_space *pfn_space);
void unregister_pfn_address_space(struct pfn_address_space *pfn_space);
+
+/**
+ * preonline_hwpoison_register - Register poison before memory is onlined
+ * @ranges: Sorted, nonoverlapping PFN ranges in ascending start PFN order.
+ * Adjacent ranges are accepted, but callers should coalesce them.
+ * The array is copied and need not outlive this call.
+ * @nr_ranges: Number of entries in @ranges. Must be nonzero.
+ * @handle: Registration handle returned on success. Set to NULL on failure.
+ *
+ * Register PFN ranges that must be withheld when their backing memory is
+ * initialized. Pass the returned handle to
+ * preonline_hwpoison_unregister() only after the memory can no longer enter
+ * the allocator.
+ *
+ * Return: 0 on success, -EINVAL if @handle is NULL or @ranges is empty,
+ * unsorted, overlapping, or contains a zero-length entry, -EOVERFLOW on PFN
+ * wraparound, and -ENOMEM on allocation failure.
+ */
+int preonline_hwpoison_register(const struct preonline_hwpoison_range *ranges,
+ unsigned int nr_ranges,
+ struct preonline_hwpoison **handle);
+
+/**
+ * preonline_hwpoison_unregister - Remove a pre-online poison registration
+ * @handle: Handle returned by preonline_hwpoison_register(), or NULL
+ *
+ * Remove @handle from future pre-online poison lookups. The caller must first
+ * ensure that the associated memory can no longer be initialized or handed
+ * to the page allocator.
+ *
+ * Pages already marked as hwpoison remain marked.
+ */
+void preonline_hwpoison_unregister(struct preonline_hwpoison *handle);
+
+/**
+ * preonline_hwpoison_intersects - Does a PFN range cover a known-bad frame?
+ * @start_pfn: First PFN of the range to test
+ * @nr_pages: Length of the range in frames
+ *
+ * Return: %true if any registered range overlaps [@start_pfn, @start_pfn +
+ * @nr_pages), %false otherwise.
+ */
+bool preonline_hwpoison_intersects(unsigned long start_pfn,
+ unsigned long nr_pages);
+
+/**
+ * preonline_hwpoison_contains - Is a single frame known bad?
+ * @pfn: Frame to test
+ *
+ * Return: %true if @pfn falls in a registered range.
+ */
+bool preonline_hwpoison_contains(unsigned long pfn);
#else
static inline int register_pfn_address_space(struct pfn_address_space *pfn_space)
{
@@ -23,6 +87,34 @@ static inline int register_pfn_address_space(struct pfn_address_space *pfn_space
static inline void unregister_pfn_address_space(struct pfn_address_space *pfn_space)
{
}
+
+static inline int
+preonline_hwpoison_register(const struct preonline_hwpoison_range *ranges,
+ unsigned int nr_ranges,
+ struct preonline_hwpoison **handle)
+{
+ if (handle)
+ *handle = NULL;
+
+ return -EOPNOTSUPP;
+}
+
+static inline void
+preonline_hwpoison_unregister(struct preonline_hwpoison *handle)
+{
+}
+
+static inline bool preonline_hwpoison_intersects(unsigned long start_pfn,
+ unsigned long nr_pages)
+{
+ return false;
+}
+
+static inline bool preonline_hwpoison_contains(unsigned long pfn)
+{
+ return false;
+}
+
#endif /* CONFIG_MEMORY_FAILURE */
#endif /* _LINUX_MEMORY_FAILURE_H */
diff --git a/mm/memory-failure.c b/mm/memory-failure.c
index 60e968243470..e89b400c3fa2 100644
--- a/mm/memory-failure.c
+++ b/mm/memory-failure.c
@@ -2291,6 +2291,181 @@ static void kill_procs_now(struct page *p, unsigned long pfn, int flags,
kill_procs(&tokill, true, pfn, flags);
}
+/**
+ * struct preonline_hwpoison - Owner of pre-online poisoned PFN ranges
+ * @list: Entry in the global owner list
+ * @rcu: Used to defer the free past a grace period
+ * @nr_ranges: Number of ranges in @ranges
+ * @ranges: Sorted, nonoverlapping PFN ranges
+ *
+ * The ranges are immutable after the object is added to the owner list, which is
+ * what lets readers walk them under RCU with no lock.
+ */
+struct preonline_hwpoison {
+ struct list_head list;
+ struct rcu_head rcu;
+ unsigned int nr_ranges;
+ struct preonline_hwpoison_range ranges[];
+};
+
+static LIST_HEAD(preonline_hwpoison_owners);
+/* Serialises writers only. Readers use RCU. */
+static DEFINE_MUTEX(preonline_hwpoison_lock);
+
+static int
+preonline_hwpoison_validate(const struct preonline_hwpoison_range *ranges,
+ unsigned int nr_ranges)
+{
+ unsigned long previous_end = 0;
+
+ if (!ranges || !nr_ranges)
+ return -EINVAL;
+
+ for (unsigned int i = 0; i < nr_ranges; i++) {
+ unsigned long end;
+
+ if (!ranges[i].nr_pages)
+ return -EINVAL;
+
+ if (check_add_overflow(ranges[i].start_pfn,
+ ranges[i].nr_pages - 1, &end))
+ return -EOVERFLOW;
+
+ /*
+ * Require sorted, nonoverlapping ranges. Adjacent ranges
+ * remain valid, although callers may coalesce them to reduce
+ * storage and lookup overhead.
+ */
+ if (i && ranges[i].start_pfn <= previous_end)
+ return -EINVAL;
+
+ previous_end = end;
+ }
+
+ return 0;
+}
+
+int preonline_hwpoison_register(const struct preonline_hwpoison_range *ranges,
+ unsigned int nr_ranges,
+ struct preonline_hwpoison **handle)
+{
+ struct preonline_hwpoison *owner;
+ int rc;
+
+ if (!handle)
+ return -EINVAL;
+
+ *handle = NULL;
+
+ rc = preonline_hwpoison_validate(ranges, nr_ranges);
+ if (rc)
+ return rc;
+
+ /*
+ * The range array may be larger than a page, so do not require
+ * physically contiguous memory.
+ */
+ owner = kvmalloc(struct_size(owner, ranges, nr_ranges), GFP_KERNEL);
+ if (!owner)
+ return -ENOMEM;
+
+ owner->nr_ranges = nr_ranges;
+ memcpy(owner->ranges, ranges,
+ flex_array_size(owner, ranges, nr_ranges));
+
+ /*
+ * list_add_tail_rcu() publishes with a release barrier, so a reader
+ * that observes the entry also observes the fully initialised
+ * ranges[] above.
+ */
+ scoped_guard(mutex, &preonline_hwpoison_lock) {
+ list_add_tail_rcu(&owner->list, &preonline_hwpoison_owners);
+ }
+
+ *handle = owner;
+
+ return 0;
+}
+EXPORT_SYMBOL_GPL(preonline_hwpoison_register);
+
+void preonline_hwpoison_unregister(struct preonline_hwpoison *owner)
+{
+ if (!owner)
+ return;
+
+ scoped_guard(mutex, &preonline_hwpoison_lock) {
+ list_del_rcu(&owner->list);
+ }
+
+ /*
+ * A reader may still be walking this entry. list_del_rcu() leaves its
+ * forward pointer intact so that walk completes into the live list;
+ * the free waits for a grace period.
+ */
+ kvfree_rcu(owner, rcu);
+}
+EXPORT_SYMBOL_GPL(preonline_hwpoison_unregister);
+
+static bool preonline_owner_intersects(const struct preonline_hwpoison *owner,
+ unsigned long start_pfn,
+ unsigned long end_pfn)
+{
+ unsigned int low = 0;
+ unsigned int high = owner->nr_ranges;
+
+ /*
+ * Find the first range whose inclusive end is greater than or equal
+ * to start_pfn.
+ */
+ while (low < high) {
+ const struct preonline_hwpoison_range *range;
+ unsigned long range_end;
+ unsigned int mid;
+
+ mid = low + (high - low) / 2;
+ range = &owner->ranges[mid];
+ range_end = range->start_pfn + range->nr_pages - 1;
+
+ if (range_end < start_pfn)
+ low = mid + 1;
+ else
+ high = mid;
+ }
+
+ if (low == owner->nr_ranges)
+ return false;
+
+ return owner->ranges[low].start_pfn <= end_pfn;
+}
+
+bool preonline_hwpoison_intersects(unsigned long start_pfn,
+ unsigned long nr_pages)
+{
+ struct preonline_hwpoison *owner;
+ unsigned long end_pfn;
+
+ if (!nr_pages)
+ return false;
+
+ /* Fail safe: withhold a range we cannot reason about. */
+ if (WARN_ON_ONCE(check_add_overflow(start_pfn, nr_pages - 1, &end_pfn)))
+ return true;
+
+ guard(rcu)();
+
+ list_for_each_entry_rcu(owner, &preonline_hwpoison_owners, list) {
+ if (preonline_owner_intersects(owner, start_pfn, end_pfn))
+ return true;
+ }
+
+ return false;
+}
+
+bool preonline_hwpoison_contains(unsigned long pfn)
+{
+ return preonline_hwpoison_intersects(pfn, 1);
+}
+
int register_pfn_address_space(struct pfn_address_space *pfn_space)
{
guard(mutex)(&pfn_space_lock);
--
2.43.0
next prev parent reply other threads:[~2026-10-10 16:21 UTC|newest]
Thread overview: 6+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-10 16:20 [RFC PATCH 0/5] cxl, mm: Keep device-reported poison out of the page allocator Shaikh Kamaluddin
2026-10-10 16:20 ` [RFC PATCH 1/5] cxl/mbox: Add Get Poison List snapshot support Shaikh Kamaluddin
2026-10-10 16:20 ` [RFC PATCH 2/5] cxl/region: Factor decoder poison-list retrieval Shaikh Kamaluddin
2026-10-10 16:20 ` Shaikh Kamaluddin [this message]
2026-10-10 16:20 ` [RFC PATCH 4/5] mm/page_alloc: Keep pre-online poison out of the buddy allocator Shaikh Kamaluddin
2026-10-10 16:20 ` [RFC PATCH 5/5] cxl/region: Register device poison before exposing RAM Shaikh Kamaluddin
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261010162017.62506-4-shaikhkamal2012@gmail.com \
--to=shaikhkamal2012@gmail.com \
--cc=akpm@linux-foundation.org \
--cc=alison.schofield@intel.com \
--cc=brendan.jackman@linux.dev \
--cc=dave.jiang@intel.com \
--cc=dave@stgolabs.net \
--cc=david@kernel.org \
--cc=djbw@kernel.org \
--cc=hannes@cmpxchg.org \
--cc=harry@kernel.org \
--cc=icheng@nvidia.com \
--cc=iweiny@kernel.org \
--cc=jic23@kernel.org \
--cc=kas@kernel.org \
--cc=leitao@debian.org \
--cc=linmiaohe@huawei.com \
--cc=linux-cxl@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=mhocko@suse.com \
--cc=ming.li@zohomail.com \
--cc=nao.horiguchi@gmail.com \
--cc=osalvador@suse.de \
--cc=surenb@google.com \
--cc=vbabka@kernel.org \
--cc=vishal.l.verma@intel.com \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®