From: Srirangan Madhavan <smadhavan@nvidia.com>
To: Alison Schofield <alison.schofield@intel.com>,
Bjorn Helgaas <bhelgaas@google.com>,
Dave Jiang <dave.jiang@intel.com>,
Davidlohr Bueso <dave@stgolabs.net>,
Ira Weiny <ira.weiny@intel.com>,
Jonathan Cameron <jic23@kernel.org>,
Vishal Verma <vishal.l.verma@intel.com>,
linux-cxl@vger.kernel.org, linux-pci@vger.kernel.org,
linux-kernel@vger.kernel.org
Cc: Alex Williamson <alex.williamson@redhat.com>,
vsethi@nvidia.com, alwilliamson@nvidia.com,
Sai Yashwanth Reddy Kancherla <skancherla@nvidia.com>,
Vishal Aslot <vaslot@nvidia.com>,
Manish Honap <mhonap@nvidia.com>, Jiandi An <jan@nvidia.com>,
Richard Cheng <icheng@nvidia.com>,
linux-tegra@vger.kernel.org,
Srirangan Madhavan <smadhavan@nvidia.com>
Subject: [PATCH v15 15/16] PCI/CXL: Restore CXL state after CXL bus reset
Date: Sun, 11 Oct 2026 02:14:21 +0000 [thread overview]
Message-ID: <20261011021422.3428136-16-smadhavan@nvidia.com> (raw)
In-Reply-To: <20261011021422.3428136-1-smadhavan@nvidia.com>
CXL bus reset can clear HDM decoder programming and CXL Device DVSEC
control state. Restore the cached state after a successful cxl_bus reset
while IOMMU exclusion remains active.
Restore PCI configuration first so HDM MMIO is accessible, and preserve a
disabled state if restoration fails.
Copy the cached state before the reset, then refresh its decoder settings
from the cache under the CXL region quiesce once the reset completes. A
region decommit during the reset updates the cache, and restoring the
earlier decoders would leave the endpoint decoding an HPA range that the
CXL core considers free. Keep global and DVSEC control from before the
reset, since a refresh during the reset can read them as all-ones.
The prepare/restore interface and decoder-refresh structure follow Dave
Jiang's cxl-type2-reset reference branch.
Link: https://git.kernel.org/pub/scm/linux/kernel/git/djiang/linux.git/commit/?id=b93f0c8e67b198a63786fd3721b7177ab9df60be
Suggested-by: Dave Jiang <dave.jiang@intel.com>
Signed-off-by: Srirangan Madhavan <smadhavan@nvidia.com>
Assisted-by: LLM
Acked-by: Bjorn Helgaas <bhelgaas@google.com> # pci/pci.c
---
drivers/cxl/core/hdm_state.c | 31 ++++++++++++++++++
drivers/pci/cxl.c | 62 ++++++++++++++++++++++++++++++++++++
drivers/pci/pci.c | 14 ++++++++
drivers/pci/pci.h | 16 ++++++++++
include/cxl/hdm.h | 2 ++
5 files changed, 125 insertions(+)
diff --git a/drivers/cxl/core/hdm_state.c b/drivers/cxl/core/hdm_state.c
index cef604013350..95daeda0ef80 100644
--- a/drivers/cxl/core/hdm_state.c
+++ b/drivers/cxl/core/hdm_state.c
@@ -120,6 +120,37 @@ struct cxl_hdm_info *cxl_hdm_cache_snapshot(struct cxl_hdm_info *const *slot)
return copy;
}
+/**
+ * cxl_hdm_cache_copy_decoders() - refresh the decoders of a snapshot
+ * @slot: cache pointer owned by the device
+ * @dst: snapshot of the same device from cxl_hdm_cache_snapshot()
+ *
+ * Copy only the per-decoder settings, which region commit and teardown keep
+ * current. Leave the rest of @dst, including global and DVSEC control, as
+ * captured: a refresh during a reset may have read those registers from a
+ * device that returned all-ones.
+ *
+ * Return: 0 on success, -ENXIO when nothing is published, or -EINVAL when
+ * @dst holds a different number of decoders.
+ */
+int cxl_hdm_cache_copy_decoders(struct cxl_hdm_info *const *slot,
+ struct cxl_hdm_info *dst)
+{
+ struct cxl_hdm_info *info;
+
+ lockdep_assert_held_write(&cxl_rwsem.region);
+ guard(rwsem_read)(&cxl_rwsem.dpa);
+ info = *slot;
+ if (!info)
+ return -ENXIO;
+ if (info->decoder_count != dst->decoder_count)
+ return -EINVAL;
+
+ memcpy(dst->settings, info->settings,
+ flex_array_size(info, settings, info->decoder_count));
+ return 0;
+}
+
void cxl_region_quiesce_lock(void)
__acquires(cxl_region_quiesce)
__context_unsafe(token maps to cxl_rwsem.region)
diff --git a/drivers/pci/cxl.c b/drivers/pci/cxl.c
index 7a225e215ca5..ed0b6f619c70 100644
--- a/drivers/pci/cxl.c
+++ b/drivers/pci/cxl.c
@@ -372,6 +372,68 @@ static int cxl_reset_save_restored_state(struct pci_dev *pdev, u16 command)
return rc;
}
+/**
+ * pci_cxl_reset_prepare() - allocate HDM state storage ahead of a bus reset
+ * @pdev: device about to be reset
+ *
+ * Copy the cached HDM state before the reset, so that restoring cannot fail
+ * for lack of memory. The decoder settings are refreshed when the state is
+ * restored.
+ *
+ * Return: storage for cxl_restore_state_after_pci_reset(), to be freed with
+ * kfree(); NULL when @pdev has no cached HDM state; or an ERR_PTR() on
+ * failure.
+ */
+struct cxl_hdm_info *pci_cxl_reset_prepare(struct pci_dev *pdev)
+{
+ struct cxl_hdm_info *snapshot;
+
+ device_lock_assert(&pdev->dev);
+ snapshot = cxl_hdm_cache_snapshot(&pdev->hdm);
+ if (IS_ERR(snapshot) && PTR_ERR(snapshot) == -ENXIO)
+ return NULL;
+
+ return snapshot;
+}
+
+/**
+ * cxl_restore_state_after_pci_reset() - restore CXL state after a bus reset
+ * @pdev: device that was reset
+ * @snapshot: storage from pci_cxl_reset_prepare()
+ *
+ * Under the region quiesce, refresh the decoder settings in @snapshot from the
+ * cache and restore: first the PCI configuration needed to reach HDM MMIO,
+ * then the HDM decoder and CXL Device DVSEC state. Leave @pdev disabled on
+ * failure.
+ *
+ * Refresh the decoders now rather than rely on the copy taken before the
+ * reset: a region commit or teardown during the reset has updated the cache,
+ * and restoring older decoders would leave the device decoding ranges the
+ * CXL core considers free.
+ */
+int cxl_restore_state_after_pci_reset(struct pci_dev *pdev,
+ struct cxl_hdm_info *snapshot)
+{
+ u16 command;
+ int rc;
+
+ device_lock_assert(&pdev->dev);
+ cxl_region_quiesce_lock();
+
+ rc = cxl_hdm_cache_copy_decoders(&pdev->hdm, snapshot);
+ if (!rc) {
+ cxl_restore_pci_state_for_hdm_restore(pdev, &command);
+ rc = cxl_restore_state(pdev, snapshot);
+ }
+ if (rc)
+ cxl_reset_save_disabled_state(pdev);
+ else
+ rc = cxl_reset_save_restored_state(pdev, command);
+
+ cxl_region_quiesce_unlock();
+ return rc;
+}
+
/*
* CXL r4.0 sec 9.7.2 defines the reset completion timeout encodings.
* Sec 9.7.3 leaves config-space access behavior undefined for 100 ms after
diff --git a/drivers/pci/pci.c b/drivers/pci/pci.c
index b58a6a613311..cecf8e681e69 100644
--- a/drivers/pci/pci.c
+++ b/drivers/pci/pci.c
@@ -5014,6 +5014,7 @@ static int pci_reset_bus_function(struct pci_dev *dev, bool probe)
static int cxl_reset_bus_function(struct pci_dev *dev, bool probe)
{
+ struct cxl_hdm_info *snapshot __free(kfree) = NULL;
struct pci_dev *bridge;
u16 dvsec, reg, val;
int rc;
@@ -5036,6 +5037,16 @@ static int cxl_reset_bus_function(struct pci_dev *dev, bool probe)
if (rc)
return -ENOTTY;
+ snapshot = pci_cxl_reset_prepare(dev);
+ if (IS_ERR(snapshot))
+ return PTR_ERR(snapshot);
+
+ /* Raw reset callers may not have saved the PCI state needed for restore. */
+ if (snapshot && !dev->state_saved) {
+ pci_err(dev, "CXL bus reset requires saved PCI state\n");
+ return -EINVAL;
+ }
+
rc = pci_dev_reset_iommu_prepare(dev);
if (rc) {
pci_err(dev, "failed to stop IOMMU for a PCI reset: %d\n", rc);
@@ -5056,6 +5067,9 @@ static int cxl_reset_bus_function(struct pci_dev *dev, bool probe)
pci_write_config_word(bridge, dvsec + PCI_DVSEC_CXL_PORT_CTL,
reg);
+ if (!rc && snapshot)
+ rc = cxl_restore_state_after_pci_reset(dev, snapshot);
+
pci_dev_reset_iommu_done(dev);
return rc;
}
diff --git a/drivers/pci/pci.h b/drivers/pci/pci.h
index 743e65abd758..476e8eefafb2 100644
--- a/drivers/pci/pci.h
+++ b/drivers/pci/pci.h
@@ -730,10 +730,14 @@ static inline void pci_doe_destroy(struct pci_dev *pdev) { }
static inline void pci_doe_disconnected(struct pci_dev *pdev) { }
#endif
+struct cxl_hdm_info;
#ifdef CONFIG_CXL_RESET
void pci_cxl_hdm_cache_init(struct pci_dev *pdev);
void pci_cxl_hdm_cache_release(struct pci_dev *pdev);
int cxl_reset_function(struct pci_dev *pdev, bool probe);
+struct cxl_hdm_info *pci_cxl_reset_prepare(struct pci_dev *pdev);
+int cxl_restore_state_after_pci_reset(struct pci_dev *pdev,
+ struct cxl_hdm_info *snapshot);
#else
static inline void pci_cxl_hdm_cache_init(struct pci_dev *pdev) { }
static inline void pci_cxl_hdm_cache_release(struct pci_dev *pdev) { }
@@ -741,6 +745,18 @@ static inline int cxl_reset_function(struct pci_dev *pdev, bool probe)
{
return -ENOTTY;
}
+
+static inline struct cxl_hdm_info *pci_cxl_reset_prepare(struct pci_dev *pdev)
+{
+ return NULL;
+}
+
+static inline int
+cxl_restore_state_after_pci_reset(struct pci_dev *pdev,
+ struct cxl_hdm_info *snapshot)
+{
+ return 0;
+}
#endif
#ifdef CONFIG_PCI_NPEM
diff --git a/include/cxl/hdm.h b/include/cxl/hdm.h
index 445093cf0452..02cad2d74329 100644
--- a/include/cxl/hdm.h
+++ b/include/cxl/hdm.h
@@ -60,6 +60,8 @@ void cxl_hdm_cache_publish(struct cxl_hdm_info **slot,
void cxl_hdm_cache_release(struct cxl_hdm_info **slot);
bool cxl_hdm_cache_dvsec_decode(struct cxl_hdm_info *const *slot);
struct cxl_hdm_info *cxl_hdm_cache_snapshot(struct cxl_hdm_info *const *slot);
+int cxl_hdm_cache_copy_decoders(struct cxl_hdm_info *const *slot,
+ struct cxl_hdm_info *dst);
/*
* Region quiesce blocks CXL region commit and teardown, and with them decoder
--
2.43.0
next prev parent reply other threads:[~2026-10-11 2:15 UTC|newest]
Thread overview: 17+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-11 2:14 [PATCH v15 00/16] PCI/CXL: Add CXL reset support for Type 2 devices Srirangan Madhavan
2026-10-11 2:14 ` [PATCH v15 01/16] cxl: Drop stale decoder interleave limit comment Srirangan Madhavan
2026-10-11 2:14 ` [PATCH v15 02/16] cxl: Share CXL port upstream PCI device lookup Srirangan Madhavan
2026-10-11 2:14 ` [PATCH v15 03/16] cxl: Move decoder declarations to shared header Srirangan Madhavan
2026-10-11 2:14 ` [PATCH v15 04/16] cxl: Embed decoder configuration in a standalone structure Srirangan Madhavan
2026-10-11 2:14 ` [PATCH v15 05/16] cxl: Introduce endpoint HDM decoder settings Srirangan Madhavan
2026-10-11 2:14 ` [PATCH v15 06/16] cxl: Move HDM decoder helpers to built-in code Srirangan Madhavan
2026-10-11 2:14 ` [PATCH v15 07/16] cxl: Share HDM decoder register unpacking Srirangan Madhavan
2026-10-11 2:14 ` [PATCH v15 08/16] cxl: Refresh cached PCI HDM decoder settings Srirangan Madhavan
2026-10-11 2:14 ` [PATCH v15 09/16] PCI/CXL: Cache endpoint HDM state during PCI enumeration Srirangan Madhavan
2026-10-11 2:14 ` [PATCH v15 10/16] PCI/CXL: Add CXL Device Reset sequencing Srirangan Madhavan
2026-10-11 2:14 ` [PATCH v15 11/16] cxl: Validate and synchronize HDM ranges around reset Srirangan Madhavan
2026-10-11 2:14 ` [PATCH v15 12/16] PCI/CXL: Reject reset with unsafe function scope Srirangan Madhavan
2026-10-11 2:14 ` [PATCH v15 13/16] PCI/CXL: Restore CXL state after PCI reset Srirangan Madhavan
2026-10-11 2:14 ` [PATCH v15 14/16] PCI/CXL: Expose CXL Reset as a PCI reset method Srirangan Madhavan
2026-10-11 2:14 ` Srirangan Madhavan [this message]
2026-10-11 2:14 ` [PATCH v15 16/16] tools/testing/cxl: Add HDM decoder range checks Srirangan Madhavan
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261011021422.3428136-16-smadhavan@nvidia.com \
--to=smadhavan@nvidia.com \
--cc=alex.williamson@redhat.com \
--cc=alison.schofield@intel.com \
--cc=alwilliamson@nvidia.com \
--cc=bhelgaas@google.com \
--cc=dave.jiang@intel.com \
--cc=dave@stgolabs.net \
--cc=icheng@nvidia.com \
--cc=ira.weiny@intel.com \
--cc=jan@nvidia.com \
--cc=jic23@kernel.org \
--cc=linux-cxl@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-pci@vger.kernel.org \
--cc=linux-tegra@vger.kernel.org \
--cc=mhonap@nvidia.com \
--cc=skancherla@nvidia.com \
--cc=vaslot@nvidia.com \
--cc=vishal.l.verma@intel.com \
--cc=vsethi@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®