From: Ankit Soni <Ankit.Soni@amd.com>
To: <iommu@lists.linux.dev>, <joro@8bytes.org>, <will@kernel.org>,
<jgg@nvidia.com>
Cc: <suravee.suthikulpanit@amd.com>, <vasant.hegde@amd.com>,
<robin.murphy@arm.com>, <joao.m.martins@oracle.com>,
<alejandro.j.jimenez@oracle.com>, <pasha.tatashin@soleen.com>,
<rppt@kernel.org>, <pratyush@kernel.org>, <skhawaja@google.com>,
<praan@google.com>, <baolu.lu@linux.intel.com>,
<dwmw2@infradead.org>, <kevin.tian@intel.com>,
<dmatlack@google.com>, <vipinsh@google.com>,
<kexec@lists.infradead.org>, <linux-kernel@vger.kernel.org>
Subject: [RFC PATCH 7/7] iommu/amd: reattach preserved devices to their restored domains
Date: Mon, 5 Oct 2026 06:40:17 +0000 [thread overview]
Message-ID: <20261005064018.1558-8-Ankit.Soni@amd.com> (raw)
In-Reply-To: <20261005064018.1558-1-Ankit.Soni@amd.com>
A device that rode across the kexec is already translating through the
DTE this kernel adopted, so adopt its domain ID and GCR3 table instead
of programming the DTE again, and verify the adopted state against the
live DTE. A mismatch means the kernel and the hardware disagree about
how the device's translations are cached, so fail the attach and leave
the DTE alone.
Refuse every other attach for such a device, in all four attach ops,
because the blocking, identity, nested and freshly built paging domains
would each rewrite that DTE under a device that is still doing DMA. The
paging path compares the target against the domain recorded for this
device, so a second restored domain cannot be swapped in either, and
clone_alias() is skipped for the same reason.
On device removal the core skips the release domain for a preserved
device, so detach_device() never runs and dev_data->domain stays set.
Drop the software state and leave the DTE alone; the restored domain
outlives the device.
Signed-off-by: Ankit Soni <Ankit.Soni@amd.com>
---
drivers/iommu/amd/amd_iommu.h | 10 +++
drivers/iommu/amd/iommu.c | 130 ++++++++++++++++++++++++++++-
drivers/iommu/amd/liveupdate.c | 146 +++++++++++++++++++++++++++++++++
drivers/iommu/amd/nested.c | 8 ++
4 files changed, 292 insertions(+), 2 deletions(-)
diff --git a/drivers/iommu/amd/amd_iommu.h b/drivers/iommu/amd/amd_iommu.h
index ac6d7a17eb40..93fe65b28d53 100644
--- a/drivers/iommu/amd/amd_iommu.h
+++ b/drivers/iommu/amd/amd_iommu.h
@@ -239,6 +239,9 @@ void amd_iommu_unpreserve_device(struct device *dev,
void amd_iommu_clear_unpreserved_dtes(struct amd_iommu *iommu);
void amd_iommu_restore_dev_table(struct amd_iommu *iommu,
struct iommu_hw_ser *ser);
+int amd_iommu_reattach_device(struct device *dev,
+ struct protection_domain *domain,
+ struct iommu_device_ser *device_ser);
#else
static inline void amd_iommu_clear_unpreserved_dtes(struct amd_iommu *iommu)
{
@@ -248,5 +251,12 @@ static inline void amd_iommu_restore_dev_table(struct amd_iommu *iommu,
struct iommu_hw_ser *ser)
{
}
+
+static inline int amd_iommu_reattach_device(struct device *dev,
+ struct protection_domain *domain,
+ struct iommu_device_ser *device_ser)
+{
+ return -EOPNOTSUPP;
+}
#endif /* CONFIG_IOMMU_LIVEUPDATE */
#endif /* AMD_IOMMU_H */
diff --git a/drivers/iommu/amd/iommu.c b/drivers/iommu/amd/iommu.c
index 72df97b99589..0bdd93016814 100644
--- a/drivers/iommu/amd/iommu.c
+++ b/drivers/iommu/amd/iommu.c
@@ -444,6 +444,13 @@ static int clone_alias(struct pci_dev *pdev_origin, u16 alias, void *data)
ret = -EINVAL;
goto out;
}
+ /*
+ * A preserved device is still translating through the DTE this kernel
+ * adopted, so never clone another device's DTE over it.
+ */
+ if (alias_data->dev && dev_iommu_restored_state(alias_data->dev))
+ goto out;
+
update_dte256(iommu, alias_data, &new);
amd_iommu_set_rlookup_table(iommu, alias);
@@ -2376,6 +2383,55 @@ static void pdom_detach_iommu(struct amd_iommu *iommu,
spin_unlock_irqrestore(&pdom->lock, flags);
}
+/*
+ * True when @dom is the one domain this kernel restored for @dev. Attaching a
+ * preserved device to anything else has to be refused, including a different
+ * restored domain.
+ */
+static bool dev_restored_domain_matches(struct device *dev,
+ struct iommu_domain *dom)
+{
+ struct iommu_device_ser *device_ser = dev_iommu_restored_state(dev);
+ struct iommu_domain_ser *domain_ser;
+
+ if (!device_ser || !device_ser->domain_iommu_ser.domain_phys)
+ return false;
+
+ domain_ser = phys_to_virt(device_ser->domain_iommu_ser.domain_phys);
+
+ return domain_ser->restored_domain == dom;
+}
+
+static int reattach_verify_dte(struct amd_iommu *iommu,
+ struct iommu_dev_data *dev_data, u16 domid)
+{
+ struct dev_table_entry dte;
+
+ get_dte256(iommu, dev_data, &dte);
+
+ if (!(dte.data[0] & DTE_FLAG_V) ||
+ FIELD_GET(DTE_DOMID_MASK, dte.data[1]) != domid) {
+ dev_err(dev_data->dev,
+ "preserved DTE does not describe the adopted domain ID %u (DTE 0x%llx/0x%llx)\n",
+ domid, dte.data[0], dte.data[1]);
+ return -EINVAL;
+ }
+
+ /*
+ * If these two disagree, either the device caches translations
+ * nobody invalidates or the IOMMU waits for completions from a
+ * device that will not send them.
+ */
+ if (!!(dte.data[1] & DTE_FLAG_IOTLB) != !!dev_data->ats_enabled) {
+ dev_err(dev_data->dev,
+ "preserved DTE and restored ATS state disagree (DTE 0x%llx/0x%llx, ats_enabled %u)\n",
+ dte.data[0], dte.data[1], dev_data->ats_enabled);
+ return -EINVAL;
+ }
+
+ return 0;
+}
+
/*
* If a device is not yet associated with a domain, this function makes the
* device visible in the domain
@@ -2385,6 +2441,7 @@ static int attach_device(struct device *dev,
{
struct iommu_dev_data *dev_data = dev_iommu_priv_get(dev);
struct amd_iommu *iommu = get_amd_iommu_from_dev_data(dev_data);
+ struct iommu_device_ser *device_ser = NULL;
struct pci_dev *pdev;
unsigned long flags;
int ret = 0;
@@ -2401,8 +2458,17 @@ static int attach_device(struct device *dev,
if (ret)
goto out;
+ if (iommu_domain_restored_state(&domain->domain))
+ device_ser = dev_iommu_restored_state(dev);
+
/* Setup GCR3 table */
- if (pdom_is_sva_capable(domain)) {
+ if (device_ser) {
+ ret = amd_iommu_reattach_device(dev, domain, device_ser);
+ if (ret) {
+ pdom_detach_iommu(iommu, domain);
+ goto out;
+ }
+ } else if (pdom_is_sva_capable(domain)) {
ret = init_gcr3_table(dev_data, domain);
if (ret) {
pdom_detach_iommu(iommu, domain);
@@ -2432,12 +2498,33 @@ static int attach_device(struct device *dev,
spin_unlock_irqrestore(&domain->lock, flags);
/* Update device table */
- dev_update_dte(dev_data, true);
+ if (device_ser) {
+ ret = reattach_verify_dte(iommu, dev_data,
+ device_ser->domain_iommu_ser.attachment_id);
+ if (ret)
+ goto err_reattach;
+ } else {
+ dev_update_dte(dev_data, true);
+ }
out:
mutex_unlock(&dev_data->mutex);
return ret;
+
+ /*
+ * The DTE is the previous kernel's and the hardware is still walking
+ * it, so unwind the software state only and leave it alone.
+ */
+err_reattach:
+ spin_lock_irqsave(&domain->lock, flags);
+ list_del(&dev_data->list);
+ spin_unlock_irqrestore(&domain->lock, flags);
+ dev_data->domain = NULL;
+ pdom_detach_iommu(iommu, domain);
+ mutex_unlock(&dev_data->mutex);
+
+ return ret;
}
/*
@@ -2493,6 +2580,30 @@ static void detach_device(struct device *dev)
mutex_unlock(&dev_data->mutex);
}
+/*
+ * The core skips the release domain for a preserved device, so its DTE is still
+ * live here. Drop the software state only. The restored domain outlives the
+ * device and the hardware keeps walking the page tables we adopted.
+ */
+static void detach_restored_device(struct device *dev)
+{
+ struct iommu_dev_data *dev_data = dev_iommu_priv_get(dev);
+ struct amd_iommu *iommu = get_amd_iommu_from_dev_data(dev_data);
+ struct protection_domain *domain = dev_data->domain;
+ unsigned long flags;
+
+ mutex_lock(&dev_data->mutex);
+
+ spin_lock_irqsave(&domain->lock, flags);
+ list_del(&dev_data->list);
+ spin_unlock_irqrestore(&domain->lock, flags);
+
+ dev_data->domain = NULL;
+ pdom_detach_iommu(iommu, domain);
+
+ mutex_unlock(&dev_data->mutex);
+}
+
static struct iommu_device *amd_iommu_probe_device(struct device *dev)
{
struct iommu_device *iommu_dev;
@@ -2561,6 +2672,9 @@ static void amd_iommu_release_device(struct device *dev)
{
struct iommu_dev_data *dev_data = dev_iommu_priv_get(dev);
+ if (dev_iommu_restored_state(dev) && dev_data->domain)
+ detach_restored_device(dev);
+
WARN_ON(dev_data->domain);
/*
@@ -2935,6 +3049,9 @@ static int blocked_domain_attach_device(struct iommu_domain *domain,
{
struct iommu_dev_data *dev_data = dev_iommu_priv_get(dev);
+ if (dev_iommu_restored_state(dev))
+ return -EBUSY;
+
if (dev_data->domain)
detach_device(dev);
@@ -3018,6 +3135,15 @@ static int amd_iommu_attach_device(struct iommu_domain *dom, struct device *dev,
if (dom->dirty_ops && !amd_iommu_hd_support(iommu))
return -EINVAL;
+ /*
+ * A preserved device is still translating through the DTE this kernel
+ * adopted, so only the one domain restored for it may be attached.
+ * Refuse before the detach below, which would tear that DTE down.
+ */
+ if (dev_iommu_restored_state(dev) &&
+ !dev_restored_domain_matches(dev, dom))
+ return -EBUSY;
+
if (dev_data->domain)
detach_device(dev);
diff --git a/drivers/iommu/amd/liveupdate.c b/drivers/iommu/amd/liveupdate.c
index 2e9a1de6b1ff..3d09054be438 100644
--- a/drivers/iommu/amd/liveupdate.c
+++ b/drivers/iommu/amd/liveupdate.c
@@ -395,3 +395,149 @@ void amd_iommu_restore_dev_table(struct amd_iommu *iommu,
pci_seg->old_dev_tbl_cpy = __va(ser->amd.dev_table_phys);
}
+
+/* Inverse of pd_mode_to_ser(), for a value coming off the wire. */
+static enum protection_domain_mode ser_to_pd_mode(u32 mode)
+{
+ switch (mode) {
+ case IOMMU_AMD_SER_PD_MODE_V1:
+ return PD_MODE_V1;
+ case IOMMU_AMD_SER_PD_MODE_V2:
+ return PD_MODE_V2;
+ default:
+ return PD_MODE_NONE;
+ }
+}
+
+static void restore_gcr3_level(u64 *tbl, int level)
+{
+ int i;
+
+ iommu_restore_pages(__pa(tbl));
+
+ if (level == 0)
+ return;
+
+ for (i = 0; i < GCR3_ENTRIES_PER_LEVEL; i++) {
+ if (!(tbl[i] & GCR3_VALID))
+ continue;
+
+ restore_gcr3_level(iommu_phys_to_virt(tbl[i] & PAGE_MASK),
+ level - 1);
+ }
+}
+
+static u64 *gcr3_pasid0_entry(u64 *tbl, int level)
+{
+ while (level--) {
+ if (!(tbl[0] & GCR3_VALID))
+ return NULL;
+
+ tbl = iommu_phys_to_virt(tbl[0] & PAGE_MASK);
+ }
+
+ return &tbl[0];
+}
+
+/*
+ * Every preserved domain ID is reserved before the first fresh allocation (see
+ * reserve_dev_table_domain_ids()), so @attachment_id already matching means an
+ * earlier device of this domain swapped it rather than a collision.
+ */
+static int reattach_domain_id(struct device *dev,
+ struct protection_domain *domain,
+ u16 attachment_id)
+{
+ bool disagree = false;
+ unsigned long flags;
+ int fresh_id = -1;
+
+ spin_lock_irqsave(&domain->lock, flags);
+ if (domain->id != attachment_id) {
+ if (list_empty(&domain->dev_list)) {
+ fresh_id = domain->id;
+ domain->id = attachment_id;
+ } else {
+ disagree = true;
+ }
+ }
+ spin_unlock_irqrestore(&domain->lock, flags);
+
+ if (disagree) {
+ dev_err(dev, "preserved domain ID %u conflicts with %u already adopted for its domain\n",
+ attachment_id, domain->id);
+ return -EINVAL;
+ }
+
+ if (fresh_id >= 0)
+ amd_iommu_pdom_id_free(fresh_id);
+
+ return 0;
+}
+
+static int reattach_gcr3_table(struct device *dev,
+ struct protection_domain *domain,
+ struct iommu_device_ser *device_ser)
+{
+ struct iommu_dev_data *dev_data = dev_iommu_priv_get(dev);
+ struct gcr3_tbl_info *gcr3_info = &dev_data->gcr3_info;
+ struct pt_iommu_x86_64_hw_info pt_info;
+ u32 glx = device_ser->amd.gcr3_glx;
+ u64 *gcr3_tbl, *pte;
+
+ if (amd_iommu_max_glx_val < 0 || glx > (u32)amd_iommu_max_glx_val) {
+ dev_err(dev, "cannot adopt a %u-level preserved GCR3 tree, hardware supports %d\n",
+ glx, amd_iommu_max_glx_val);
+ return -EINVAL;
+ }
+
+ gcr3_tbl = phys_to_virt(device_ser->amd.gcr3_tbl_phys);
+ restore_gcr3_level(gcr3_tbl, glx);
+
+ gcr3_info->gcr3_tbl = gcr3_tbl;
+ gcr3_info->glx = glx;
+ gcr3_info->domid = device_ser->domain_iommu_ser.attachment_id;
+
+ if (domain->pd_mode != PD_MODE_V2)
+ return 0;
+
+ /* Double-check the hardware is walking the page tables we restored. */
+ pt_iommu_x86_64_hw_info(&domain->amdv2, &pt_info);
+ pte = gcr3_pasid0_entry(gcr3_tbl, gcr3_info->glx);
+ if (!pte || (__sme_clr(*pte) & PAGE_MASK) != (pt_info.gcr3_pt & PAGE_MASK)) {
+ dev_err(dev, "preserved GCR3[0] 0x%llx does not match restored v2 page-table root 0x%llx\n",
+ pte ? *pte : 0, pt_info.gcr3_pt);
+ return -EINVAL;
+ }
+
+ return 0;
+}
+
+/**
+ * amd_iommu_reattach_device - Adopt a device's preserved DTE state on LU boot
+ * @dev: Device being attached
+ * @domain: Domain this kernel rebuilt from the same preserved state
+ * @device_ser: Preserved per-device record handed over by the previous kernel
+ *
+ * Return: 0 on success, or a negative error code if the preserved state does
+ * not describe the domain this kernel restored.
+ */
+int amd_iommu_reattach_device(struct device *dev,
+ struct protection_domain *domain,
+ struct iommu_device_ser *device_ser)
+{
+ enum protection_domain_mode pd_mode;
+
+ pd_mode = ser_to_pd_mode(device_ser->amd.pd_mode);
+ if (pd_mode != domain->pd_mode) {
+ dev_err(dev, "preserved page-table mode %d does not match restored domain mode %d\n",
+ pd_mode, domain->pd_mode);
+ return -EINVAL;
+ }
+
+ if (!device_ser->amd.gcr3_tbl_phys)
+ return reattach_domain_id(dev, domain,
+ device_ser->domain_iommu_ser.attachment_id);
+
+ return reattach_gcr3_table(dev, domain, device_ser);
+}
diff --git a/drivers/iommu/amd/nested.c b/drivers/iommu/amd/nested.c
index 63b53b29e029..4b84b7b4b829 100644
--- a/drivers/iommu/amd/nested.c
+++ b/drivers/iommu/amd/nested.c
@@ -6,6 +6,7 @@
#define dev_fmt(fmt) "AMD-Vi: " fmt
#include <linux/iommu.h>
+#include <linux/iommu-liveupdate.h>
#include <linux/refcount.h>
#include <uapi/linux/iommufd.h>
@@ -248,6 +249,13 @@ static int nested_attach_device(struct iommu_domain *dom, struct device *dev,
if (WARN_ON(dev_data->pasid_enabled))
return -EINVAL;
+ /*
+ * A preserved device is still translating through the DTE this kernel
+ * adopted, and a nested domain is never the domain restored for it.
+ */
+ if (dev_iommu_restored_state(dev))
+ return -EBUSY;
+
mutex_lock(&dev_data->mutex);
set_dte_nested(iommu, dom, dev_data, &new);
--
2.43.0
prev parent reply other threads:[~2026-10-05 6:43 UTC|newest]
Thread overview: 8+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-05 6:40 [RFC PATCH 0/7] iommu/amd: Implement live update state preservation Ankit Soni
2026-10-05 6:40 ` [RFC PATCH 1/7] iommu/amd: defer device attach only on a kdump boot Ankit Soni
2026-10-05 6:40 ` [RFC PATCH 2/7] liveupdate: parse the incoming handover tree before late_time_init() Ankit Soni
2026-10-05 6:40 ` [RFC PATCH 3/7] iommu/kho/abi: add AMD IOMMU live-update serialisation structs Ankit Soni
2026-10-05 6:40 ` [RFC PATCH 4/7] iommu/amd: preserve IOMMU and device state for live update Ankit Soni
2026-10-05 6:40 ` [RFC PATCH 5/7] iommu/amd: clear unpreserved DTEs and quiesce logs at live-update shutdown Ankit Soni
2026-10-05 6:40 ` [RFC PATCH 6/7] iommu/amd: restore preserved state on a live-update boot Ankit Soni
2026-10-05 6:40 ` Ankit Soni [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261005064018.1558-8-Ankit.Soni@amd.com \
--to=ankit.soni@amd.com \
--cc=alejandro.j.jimenez@oracle.com \
--cc=baolu.lu@linux.intel.com \
--cc=dmatlack@google.com \
--cc=dwmw2@infradead.org \
--cc=iommu@lists.linux.dev \
--cc=jgg@nvidia.com \
--cc=joao.m.martins@oracle.com \
--cc=joro@8bytes.org \
--cc=kevin.tian@intel.com \
--cc=kexec@lists.infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=pasha.tatashin@soleen.com \
--cc=praan@google.com \
--cc=pratyush@kernel.org \
--cc=robin.murphy@arm.com \
--cc=rppt@kernel.org \
--cc=skhawaja@google.com \
--cc=suravee.suthikulpanit@amd.com \
--cc=vasant.hegde@amd.com \
--cc=vipinsh@google.com \
--cc=will@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®