* [RFC PATCH 0/7] iommu/amd: Implement live update state preservation
@ 2026-10-05 6:40 Ankit Soni
2026-10-05 6:40 ` [RFC PATCH 1/7] iommu/amd: defer device attach only on a kdump boot Ankit Soni
` (6 more replies)
0 siblings, 7 replies; 8+ messages in thread
From: Ankit Soni @ 2026-10-05 6:40 UTC (permalink / raw)
To: iommu, joro, will, jgg
Cc: suravee.suthikulpanit, vasant.hegde, robin.murphy,
joao.m.martins, alejandro.j.jimenez, pasha.tatashin, rppt,
pratyush, skhawaja, praan, baolu.lu, dwmw2, kevin.tian, dmatlack,
vipinsh, kexec, linux-kernel
This series adds AMD-Vi support for IOMMU state preservation across a
kexec-based live update, so a passed-through device keeps its translations
and keeps doing DMA while the kernel underneath is replaced.
It applies on top of the IOMMU live update core series v5 [1], currently in
review, and reuses the FLB preservation mechanism that series introduces.
This is an RFC. The reclaim path is not included; it depends on the phase 2
iommufd work and will follow.
What is preserved
=================
The Device Table, adopted by the incoming kernel rather than rebuilt, and for
each preserved device its domain ID, page-table mode, and GCR3 tree if the
device uses PASID.
What is not preserved
=====================
The command buffer, event log, PPR log, GA log, GA tail and cmd_sem, and the
interrupt remapping tables. The incoming kernel allocates all of them fresh.
Interrupts raised during the window are lost, but the data and the completion
records travel by DMA, so a driver finds the completed work when it next
reads its queues.
Disabling the PPR log also disables PPR, so a preserved device cannot
raise a page request across the window. That is safe only because an HWPT
with a fault queue cannot be marked for preservation today; if that is
relaxed for SVA, such a request would go unanswered and, with the event
log stopped, unlogged. Whoever adds PRI preservation needs to revisit
this.
Architectural overview
======================
At preserve time the driver pins the PCI segment's Device Table for KHO and
records each preserved device's domain ID, page-table mode and GCR3 tree.
At shutdown translation stays enabled on any unit carrying preserved devices,
so their DMA never stops. Everything else on that unit is made safe: DTEs of
unpreserved devices are reset to blocked, the interrupt fields of every DTE
are cleared because no DTE may point at a table the next kernel can recycle,
and the command, event, PPR and GA engines are stopped and polled until idle.
That last step matters because the incoming kernel may reuse the pages those
buffers occupied, and an engine still writing after the kexec would corrupt
them with nothing to attribute the damage to. If an engine does not go idle
the driver panics rather than complete the handover.
On the live update boot the driver finds its record by MMIO physical base,
adopts the Device Table, and reserves the domain IDs it describes so a new
domain cannot alias a preserved one. A unit handed over translating is left
translating. When the core probes the preserved devices their domain ID and
GCR3 tree are adopted on attach and verified against the live DTE; a mismatch
fails the attach. A restored domain is immutable, so a preserved device may
only attach to the domain restored for it.
One generic change
==================
Patch 2 moves the LUO handover-tree parse into start_kernel(), immediately
before late_time_init().
AMD-Vi needs this because it is the x86 interrupt remapping provider, so
amd_iommu_prepare() runs from late_time_init() via enable_IR_x2apic(). It
queries preserved state from three sites on that path. LUO parses the tree
from an early_initcall inside rest_init(), so those queries run before the
data exists and cannot tell "nothing was preserved" from "not parsed yet".
The driver then disables a unit the previous kernel left translating.
Note:
ATS and PASID state on the device itself is not adopted. For a preserved
device the attach path still runs the normal enable sequence, so an
ATS-capable device would have its ATS Control register rewritten while it
is doing DMA. Adopting that state is PCI core roadmap item #4 [2], so this
series leaves it alone rather than reprogramming the PCI side here. The
restored DTE's IOTLB bit is checked against the device's ATS state, which
catches a disagreement.
[1] Samiullah Khawaja, "iommu: Add live update state preservation" (v5)
https://lore.kernel.org/linux-iommu/20260921004834.2601285-1-skhawaja@google.com/
[2] David Matlack, "RFC: PCI core Live Update Roadmap"
https://lore.kernel.org/linux-pci/20261001232133.560284-1-dmatlack@google.com/
Ankit Soni (7):
iommu/amd: defer device attach only on a kdump boot
liveupdate: parse the incoming handover tree before late_time_init()
iommu/kho/abi: add AMD IOMMU live-update serialisation structs
iommu/amd: preserve IOMMU and device state for live update
iommu/amd: clear unpreserved DTEs and quiesce logs at live-update shutdown
iommu/amd: restore preserved state on a live-update boot
iommu/amd: reattach preserved devices to their restored domains
drivers/iommu/amd/Makefile | 1 +
drivers/iommu/amd/amd_iommu.h | 34 ++
drivers/iommu/amd/amd_iommu_types.h | 14 +
drivers/iommu/amd/init.c | 253 +++++++++++--
drivers/iommu/amd/iommu.c | 144 +++++++-
drivers/iommu/amd/liveupdate.c | 543 ++++++++++++++++++++++++++++
drivers/iommu/amd/nested.c | 8 +
include/linux/kho/abi/iommu.h | 70 ++++
include/linux/liveupdate.h | 4 +
init/main.c | 2 +
kernel/liveupdate/luo_core.c | 5 +-
11 files changed, 1039 insertions(+), 39 deletions(-)
create mode 100644 drivers/iommu/amd/liveupdate.c
base-commit: 1d195de0e8caf627e93c0a39af1e84bb19e3cc53
--
2.43.0
^ permalink raw reply [flat|nested] 8+ messages in thread
* [RFC PATCH 1/7] iommu/amd: defer device attach only on a kdump boot
2026-10-05 6:40 [RFC PATCH 0/7] iommu/amd: Implement live update state preservation Ankit Soni
@ 2026-10-05 6:40 ` Ankit Soni
2026-10-05 6:40 ` [RFC PATCH 2/7] liveupdate: parse the incoming handover tree before late_time_init() Ankit Soni
` (5 subsequent siblings)
6 siblings, 0 replies; 8+ messages in thread
From: Ankit Soni @ 2026-10-05 6:40 UTC (permalink / raw)
To: iommu, joro, will, jgg
Cc: suravee.suthikulpanit, vasant.hegde, robin.murphy,
joao.m.martins, alejandro.j.jimenez, pasha.tatashin, rppt,
pratyush, skhawaja, praan, baolu.lu, dwmw2, kevin.tian, dmatlack,
vipinsh, kexec, linux-kernel
find_dev_data() sets defer_attach whenever the IOMMU came up already
translating. Only a kdump boot can ever complete such an attach: the
core completes it from iommu_deferred_attach(), behind a static key that
is enabled only when is_kdump_kernel().
That was harmless while a pre-enabled unit surviving driver init implied
a kdump boot. A live-update handover breaks the implication, because a
unit carrying preserved devices is deliberately left translating. The
devices behind it that were not preserved had their DTEs blocked at
shutdown, so the skipped attach never replaces the blocked entry and
they come up unable to do DMA.
Restrict the deferral to kdump, which is the only case that can complete
it.
Signed-off-by: Ankit Soni <Ankit.Soni@amd.com>
---
drivers/iommu/amd/iommu.c | 7 ++++++-
1 file changed, 6 insertions(+), 1 deletion(-)
diff --git a/drivers/iommu/amd/iommu.c b/drivers/iommu/amd/iommu.c
index 4dc306a4b5c6..a83ce4521f7f 100644
--- a/drivers/iommu/amd/iommu.c
+++ b/drivers/iommu/amd/iommu.c
@@ -31,6 +31,7 @@
#include <linux/irqdomain.h>
#include <linux/percpu.h>
#include <linux/cc_platform.h>
+#include <linux/crash_dump.h>
#include <asm/irq_remapping.h>
#include <asm/io_apic.h>
#include <asm/apic.h>
@@ -500,7 +501,11 @@ static struct iommu_dev_data *find_dev_data(struct amd_iommu *iommu, u16 devid)
if (!dev_data)
return NULL;
- if (translation_pre_enabled(iommu))
+ /*
+ * Only a kdump boot can complete a deferred attach, so only a
+ * kdump boot may start one.
+ */
+ if (translation_pre_enabled(iommu) && is_kdump_kernel())
dev_data->defer_attach = true;
}
--
2.43.0
^ permalink raw reply [flat|nested] 8+ messages in thread
* [RFC PATCH 2/7] liveupdate: parse the incoming handover tree before late_time_init()
2026-10-05 6:40 [RFC PATCH 0/7] iommu/amd: Implement live update state preservation Ankit Soni
2026-10-05 6:40 ` [RFC PATCH 1/7] iommu/amd: defer device attach only on a kdump boot Ankit Soni
@ 2026-10-05 6:40 ` Ankit Soni
2026-10-05 6:40 ` [RFC PATCH 3/7] iommu/kho/abi: add AMD IOMMU live-update serialisation structs Ankit Soni
` (4 subsequent siblings)
6 siblings, 0 replies; 8+ messages in thread
From: Ankit Soni @ 2026-10-05 6:40 UTC (permalink / raw)
To: iommu, joro, will, jgg
Cc: suravee.suthikulpanit, vasant.hegde, robin.murphy,
joao.m.martins, alejandro.j.jimenez, pasha.tatashin, rppt,
pratyush, skhawaja, praan, baolu.lu, dwmw2, kevin.tian, dmatlack,
vipinsh, kexec, linux-kernel
LUO parses the handover tree from an early_initcall, which runs inside
rest_init() -- the last statement of start_kernel(). Anything earlier in
start_kernel() cannot see the incoming state, and cannot tell that it
cannot see it: luo_flb_retrieve_one() returns -ENODATA both before the
parse and when the previous kernel handed nothing over.
On x86 the IOMMU is such a consumer, because it is also the interrupt
remapping provider. amd_iommu_prepare() runs from late_time_init() via
enable_IR_x2apic() -> irq_remapping_prepare(), and reaches
init_iommu_one_late() long before any initcall.
Fix the ordering rather than teaching each consumer to cope. Call the
parse directly from start_kernel(), immediately before late_time_init().
Rename it to liveupdate_init_early() at the same time, since
liveupdate_early_init() was named after the early_initcall that no longer
exists.
Signed-off-by: Ankit Soni <Ankit.Soni@amd.com>
---
include/linux/liveupdate.h | 4 ++++
init/main.c | 2 ++
kernel/liveupdate/luo_core.c | 5 +----
3 files changed, 7 insertions(+), 4 deletions(-)
diff --git a/include/linux/liveupdate.h b/include/linux/liveupdate.h
index 6051abc0612c..07dbe8366f76 100644
--- a/include/linux/liveupdate.h
+++ b/include/linux/liveupdate.h
@@ -229,6 +229,8 @@ struct liveupdate_flb {
#ifdef CONFIG_LIVEUPDATE
+void liveupdate_init_early(void);
+
/* Return true if live update orchestrator is enabled */
bool liveupdate_enabled(void);
@@ -259,6 +261,8 @@ int liveupdate_get_token_outgoing(struct liveupdate_session *s,
#else /* CONFIG_LIVEUPDATE */
+static inline void liveupdate_init_early(void) { }
+
static inline bool liveupdate_enabled(void)
{
return false;
diff --git a/init/main.c b/init/main.c
index 2613d3f9b3ce..221c22489c04 100644
--- a/init/main.c
+++ b/init/main.c
@@ -108,6 +108,7 @@
#include <linux/time_namespace.h>
#include <linux/unaligned.h>
#include <linux/vdso_datastore.h>
+#include <linux/liveupdate.h>
#include <net/net_namespace.h>
#include <asm/io.h>
@@ -1145,6 +1146,7 @@ void start_kernel(void)
setup_per_cpu_pageset();
numa_policy_init();
acpi_early_init();
+ liveupdate_init_early();
if (late_time_init)
late_time_init();
sched_clock_init();
diff --git a/kernel/liveupdate/luo_core.c b/kernel/liveupdate/luo_core.c
index 91703715e993..54771d67d922 100644
--- a/kernel/liveupdate/luo_core.c
+++ b/kernel/liveupdate/luo_core.c
@@ -135,7 +135,7 @@ static int __init luo_early_startup(void)
return err;
}
-static int __init liveupdate_early_init(void)
+void __init liveupdate_init_early(void)
{
int err;
@@ -145,10 +145,7 @@ static int __init liveupdate_early_init(void)
luo_restore_fail("The incoming tree failed to initialize properly [%pe], disabling live update\n",
ERR_PTR(err));
}
-
- return err;
}
-early_initcall(liveupdate_early_init);
/* Called during boot to create outgoing LUO state */
static int __init luo_state_setup(void)
--
2.43.0
^ permalink raw reply [flat|nested] 8+ messages in thread
* [RFC PATCH 3/7] iommu/kho/abi: add AMD IOMMU live-update serialisation structs
2026-10-05 6:40 [RFC PATCH 0/7] iommu/amd: Implement live update state preservation Ankit Soni
2026-10-05 6:40 ` [RFC PATCH 1/7] iommu/amd: defer device attach only on a kdump boot Ankit Soni
2026-10-05 6:40 ` [RFC PATCH 2/7] liveupdate: parse the incoming handover tree before late_time_init() Ankit Soni
@ 2026-10-05 6:40 ` Ankit Soni
2026-10-05 6:40 ` [RFC PATCH 4/7] iommu/amd: preserve IOMMU and device state for live update Ankit Soni
` (3 subsequent siblings)
6 siblings, 0 replies; 8+ messages in thread
From: Ankit Soni @ 2026-10-05 6:40 UTC (permalink / raw)
To: iommu, joro, will, jgg
Cc: suravee.suthikulpanit, vasant.hegde, robin.murphy,
joao.m.martins, alejandro.j.jimenez, pasha.tatashin, rppt,
pratyush, skhawaja, praan, baolu.lu, dwmw2, kevin.tian, dmatlack,
vipinsh, kexec, linux-kernel
Describe the AMD IOMMU state that has to cross a live-update kexec: the
device table of a PCI segment, and the domain ID, page-table mode and
GCR3 tree of a preserved device.
Both structs fit inside the existing unions, so the layout and the
version are unchanged. Assert that, since growing either arm past the
Intel one would move the array stride.
Signed-off-by: Ankit Soni <Ankit.Soni@amd.com>
---
include/linux/kho/abi/iommu.h | 70 +++++++++++++++++++++++++++++++++++
1 file changed, 70 insertions(+)
diff --git a/include/linux/kho/abi/iommu.h b/include/linux/kho/abi/iommu.h
index 308cd83fd3e3..5cbc8b3589ed 100644
--- a/include/linux/kho/abi/iommu.h
+++ b/include/linux/kho/abi/iommu.h
@@ -8,6 +8,7 @@
#ifndef _LINUX_KHO_ABI_IOMMU_H
#define _LINUX_KHO_ABI_IOMMU_H
+#include <linux/bug.h>
#include <linux/mutex_types.h>
#include <linux/compiler.h>
#include <linux/types.h>
@@ -75,6 +76,7 @@
/**
* enum iommu_type_ser - Type of the IOMMU being preserved
* @IOMMU_INVALID: Invalid type of IOMMU
+ * @IOMMU_AMD: AMD-Vi, whose per-instance state is struct iommu_amd_ser
*
* IOMMU type is stored in the IOMMU HW state to differentiate between various
* IOMMU HWs.
@@ -82,6 +84,7 @@
enum iommu_type_ser {
IOMMU_INVALID,
IOMMU_INTEL,
+ IOMMU_AMD,
};
#define IOMMU_SER_FLAG_DELETED (1 << 0)
@@ -142,6 +145,29 @@ struct iommu_device_intel_ser {
u64 max_pasid;
} __packed;
+/*
+ * Page-table mode of the domain a preserved AMD device was attached to. These
+ * are wire values with their own numbering rather than the kernel's
+ * enum protection_domain_mode, so reordering that enum cannot silently change
+ * the handoff format.
+ */
+#define IOMMU_AMD_SER_PD_MODE_NONE 0
+#define IOMMU_AMD_SER_PD_MODE_V1 1
+#define IOMMU_AMD_SER_PD_MODE_V2 2
+
+/**
+ * struct iommu_device_amd_ser - AMD specific state of serialized device
+ * @gcr3_tbl_phys: Physical address of the device's GCR3 table root, or 0 if
+ * the device has no PASID/GCR3 table.
+ * @gcr3_glx: Number of GCR3 table levels (0, 1, or 2; see amd_iommu_max_glx_val)
+ * @pd_mode: Page-table format of the domain, one of IOMMU_AMD_SER_PD_MODE_*
+ */
+struct iommu_device_amd_ser {
+ u64 gcr3_tbl_phys;
+ u32 gcr3_glx;
+ u32 pd_mode;
+} __packed;
+
/**
* struct iommu_device_ser - Serialized state of a device
* @hdr: Common object header
@@ -150,6 +176,8 @@ struct iommu_device_intel_ser {
* @dma_owner_token: Token to identify the DMA owner of this device
* @domain_iommu_ser: Domain and IOMMU mapping
* @intel: Intel specific serialization data
+ * @amd: AMD-Vi per-device state, valid when the owning IOMMU record is
+ * of type IOMMU_AMD
*/
struct iommu_device_ser {
struct iommu_hdr_ser hdr;
@@ -159,9 +187,18 @@ struct iommu_device_ser {
struct iommu_dev_map_ser domain_iommu_ser;
union {
struct iommu_device_intel_ser intel;
+ struct iommu_device_amd_ser amd;
};
} __packed;
+/*
+ * The Intel arm is the larger one and so sets the size of the union, and with
+ * it the stride of the device array. Growing either arm changes that stride and
+ * needs a version bump.
+ */
+static_assert(sizeof(struct iommu_device_amd_ser) == 16);
+static_assert(sizeof(struct iommu_device_intel_ser) == 24);
+
/* There are maximum 256 buses, so maximum 512 context tables */
#define VTD_PRESERVED_BITMAP_LONGS DIV_ROUND_UP(512, BITS_PER_LONG_LONG)
@@ -181,12 +218,36 @@ struct iommu_intel_ser {
u64 context_tables_bitmap[VTD_PRESERVED_BITMAP_LONGS];
};
+/**
+ * struct iommu_amd_ser - Serialized state of an AMD IOMMU instance
+ * @restored: Whether IOMMU state is restored. Guards against double-restore.
+ * @mmio_phys: Physical address of the IOMMU MMIO register base.
+ * @dev_table_phys: Physical address of this IOMMU's PCI-segment device
+ * table (struct amd_iommu_pci_seg.dev_table).
+ * @dev_table_size: Size of the device table, in bytes.
+ * @pci_seg_id: PCI segment ID that owns the device table.
+ * @efr: Extended Feature Register bits (struct amd_iommu.features).
+ * Compared on restore; a mismatch is fatal.
+ * @efr2: Extended Feature Register 2 bits (struct amd_iommu.features2).
+ */
+struct iommu_amd_ser {
+ u8 restored;
+ u8 padding[7];
+ u64 mmio_phys;
+ u64 dev_table_phys;
+ u32 dev_table_size;
+ u32 pci_seg_id;
+ u64 efr;
+ u64 efr2;
+} __packed;
+
/**
* struct iommu_hw_ser - Serialized state of an IOMMU instance
* @hdr: Common object header
* @token: Unique token for the IOMMU
* @type: IOMMU type serialized state belongs to
* @intel: Intel specific serialization data
+ * @amd: AMD specific serialization data
*/
struct iommu_hw_ser {
struct iommu_hdr_ser hdr;
@@ -194,9 +255,18 @@ struct iommu_hw_ser {
u64 type;
union {
struct iommu_intel_ser intel;
+ struct iommu_amd_ser amd;
};
} __packed;
+/*
+ * The Intel arm is the larger one and so sets the size of the union, and with
+ * it the stride of the IOMMU array. Growing either arm changes that stride and
+ * needs a version bump.
+ */
+static_assert(sizeof(struct iommu_amd_ser) == 48);
+static_assert(sizeof(struct iommu_intel_ser) == 88);
+
/**
* struct iommu_array_hdr_ser - Header for an array of serialized objects
* @next_array_phys: Physical address of the next array of objects
--
2.43.0
^ permalink raw reply [flat|nested] 8+ messages in thread
* [RFC PATCH 4/7] iommu/amd: preserve IOMMU and device state for live update
2026-10-05 6:40 [RFC PATCH 0/7] iommu/amd: Implement live update state preservation Ankit Soni
` (2 preceding siblings ...)
2026-10-05 6:40 ` [RFC PATCH 3/7] iommu/kho/abi: add AMD IOMMU live-update serialisation structs Ankit Soni
@ 2026-10-05 6:40 ` Ankit Soni
2026-10-05 6:40 ` [RFC PATCH 5/7] iommu/amd: clear unpreserved DTEs and quiesce logs at live-update shutdown Ankit Soni
` (2 subsequent siblings)
6 siblings, 0 replies; 8+ messages in thread
From: Ankit Soni @ 2026-10-05 6:40 UTC (permalink / raw)
To: iommu, joro, will, jgg
Cc: suravee.suthikulpanit, vasant.hegde, robin.murphy,
joao.m.martins, alejandro.j.jimenez, pasha.tatashin, rppt,
pratyush, skhawaja, praan, baolu.lu, dwmw2, kevin.tian, dmatlack,
vipinsh, kexec, linux-kernel
Implement the .preserve and .preserve_device callbacks. Pin the PCI
segment's device table for KHO and record each device's domain ID,
page-table mode and GCR3 tree, so the next kernel can rebuild them.
The device table is shared by every IOMMU in a segment while the
callbacks run once per IOMMU, so the pin is reference counted.
Signed-off-by: Ankit Soni <Ankit.Soni@amd.com>
---
drivers/iommu/amd/Makefile | 1 +
drivers/iommu/amd/amd_iommu.h | 12 ++
drivers/iommu/amd/amd_iommu_types.h | 14 ++
drivers/iommu/amd/iommu.c | 7 +
drivers/iommu/amd/liveupdate.c | 281 ++++++++++++++++++++++++++++
5 files changed, 315 insertions(+)
create mode 100644 drivers/iommu/amd/liveupdate.c
diff --git a/drivers/iommu/amd/Makefile b/drivers/iommu/amd/Makefile
index 94b8ef2acb18..227bbe920c26 100644
--- a/drivers/iommu/amd/Makefile
+++ b/drivers/iommu/amd/Makefile
@@ -2,3 +2,4 @@
obj-y += iommu.o init.o quirks.o ppr.o pasid.o
obj-$(CONFIG_AMD_IOMMU_IOMMUFD) += iommufd.o nested.o
obj-$(CONFIG_AMD_IOMMU_DEBUGFS) += debugfs.o
+obj-$(CONFIG_IOMMU_LIVEUPDATE) += liveupdate.o
diff --git a/drivers/iommu/amd/amd_iommu.h b/drivers/iommu/amd/amd_iommu.h
index a2fe804b038b..5cf32e4898dc 100644
--- a/drivers/iommu/amd/amd_iommu.h
+++ b/drivers/iommu/amd/amd_iommu.h
@@ -225,4 +225,16 @@ amd_iommu_make_clear_dte(struct iommu_dev_data *dev_data, struct dev_table_entry
struct iommu_domain *
amd_iommu_alloc_domain_nested(struct iommufd_viommu *viommu, u32 flags,
const struct iommu_user_data *user_data);
+
+#ifdef CONFIG_IOMMU_LIVEUPDATE
+/* LIVE UPDATE (drivers/iommu/amd/liveupdate.c) */
+int amd_iommu_preserve(struct iommu_device *iommu_dev,
+ struct iommu_hw_ser *ser);
+void amd_iommu_unpreserve(struct iommu_device *iommu_dev,
+ struct iommu_hw_ser *ser);
+int amd_iommu_preserve_device(struct device *dev,
+ struct iommu_device_ser *device_ser);
+void amd_iommu_unpreserve_device(struct device *dev,
+ struct iommu_device_ser *device_ser);
+#endif /* CONFIG_IOMMU_LIVEUPDATE */
#endif /* AMD_IOMMU_H */
diff --git a/drivers/iommu/amd/amd_iommu_types.h b/drivers/iommu/amd/amd_iommu_types.h
index 3dbe20023456..fc9d98dfc6cd 100644
--- a/drivers/iommu/amd/amd_iommu_types.h
+++ b/drivers/iommu/amd/amd_iommu_types.h
@@ -596,6 +596,20 @@ struct amd_iommu_pci_seg {
*/
struct dev_table_entry *dev_table;
+#ifdef CONFIG_IOMMU_LIVEUPDATE
+ /*
+ * The device table is shared by every IOMMU in this PCI segment, but
+ * the live-update .preserve callback runs once per IOMMU, and each of
+ * those IOMMUs is preserved and unpreserved independently by the core.
+ * KHO preservation is not refcounted, so the references are counted
+ * here and the shared pages are pinned on the first and unpinned only
+ * on the last -- otherwise one IOMMU being unpreserved would drop the
+ * pin from underneath the devices still preserved behind every other
+ * IOMMU in the segment.
+ */
+ unsigned int dev_table_preserve_count;
+#endif
+
/*
* The rlookup iommu table is used to find the IOMMU which is
* responsible for a specific device. It is indexed by the PCI
diff --git a/drivers/iommu/amd/iommu.c b/drivers/iommu/amd/iommu.c
index a83ce4521f7f..72df97b99589 100644
--- a/drivers/iommu/amd/iommu.c
+++ b/drivers/iommu/amd/iommu.c
@@ -32,6 +32,7 @@
#include <linux/percpu.h>
#include <linux/cc_platform.h>
#include <linux/crash_dump.h>
+#include <linux/iommu-liveupdate.h>
#include <asm/irq_remapping.h>
#include <asm/io_apic.h>
#include <asm/apic.h>
@@ -3221,6 +3222,12 @@ const struct iommu_ops amd_iommu_ops = {
.page_response = amd_iommu_page_response,
.get_viommu_size = amd_iommufd_get_viommu_size,
.viommu_init = amd_iommufd_viommu_init,
+#ifdef CONFIG_IOMMU_LIVEUPDATE
+ .preserve_device = amd_iommu_preserve_device,
+ .unpreserve_device = amd_iommu_unpreserve_device,
+ .preserve = amd_iommu_preserve,
+ .unpreserve = amd_iommu_unpreserve,
+#endif
};
#ifdef CONFIG_IRQ_REMAP
diff --git a/drivers/iommu/amd/liveupdate.c b/drivers/iommu/amd/liveupdate.c
new file mode 100644
index 000000000000..096a23bb4e7b
--- /dev/null
+++ b/drivers/iommu/amd/liveupdate.c
@@ -0,0 +1,281 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/*
+ * AMD IOMMU (AMD-Vi) live update support.
+ *
+ * Copyright (C) 2026 Advanced Micro Devices, Inc.
+ * Author: Ankit Soni <Ankit.Soni@amd.com>
+ */
+
+#include <linux/iommu-liveupdate.h>
+#include <linux/pci.h>
+
+#include "amd_iommu.h"
+#include "../iommu-pages.h"
+
+/* Each 4K GCR3 table level holds 512 u64 entries. */
+#define GCR3_ENTRIES_PER_LEVEL 512
+
+/**
+ * amd_iommu_preserve - Preserve one AMD IOMMU instance for live update
+ * @iommu_dev: Core handle for the IOMMU whose state is being preserved
+ * @ser: Serialized IOMMU-instance record to fill in
+ *
+ * Pins the PCI-segment Device Table. Command, event, PPR and GA buffers are
+ * not preserved: they are drained or quiesced at shutdown and the next kernel
+ * allocates fresh ones, matching Intel phase 1.
+ *
+ * Return: 0 on success, negative errno on failure (any pages pinned by this
+ * call are released before returning).
+ */
+int amd_iommu_preserve(struct iommu_device *iommu_dev, struct iommu_hw_ser *ser)
+{
+ struct amd_iommu *iommu = container_of(iommu_dev, struct amd_iommu, iommu);
+ struct amd_iommu_pci_seg *pci_seg = iommu->pci_seg;
+ int ret;
+
+ if (!pci_seg->dev_table_preserve_count) {
+ ret = iommu_preserve_pages(pci_seg->dev_table);
+ if (ret)
+ return ret;
+ }
+ pci_seg->dev_table_preserve_count++;
+
+ ser->type = IOMMU_AMD;
+ ser->token = iommu->mmio_phys;
+ ser->amd.mmio_phys = iommu->mmio_phys;
+ ser->amd.dev_table_phys = __pa(pci_seg->dev_table);
+ ser->amd.dev_table_size = pci_seg->dev_table_size;
+ ser->amd.pci_seg_id = pci_seg->id;
+ ser->amd.efr = iommu->features;
+ ser->amd.efr2 = iommu->features2;
+
+ return 0;
+}
+
+/**
+ * amd_iommu_unpreserve - Release live-update state of one AMD IOMMU instance
+ * @iommu_dev: Core handle for the IOMMU whose state is being released
+ * @ser: Serialized IOMMU-instance record (unused; state is derived from @iommu)
+ *
+ * Reverse of amd_iommu_preserve(): once the last IOMMU of the segment has
+ * dropped its reference, unpins the shared Device Table.
+ */
+void amd_iommu_unpreserve(struct iommu_device *iommu_dev,
+ struct iommu_hw_ser *ser)
+{
+ struct amd_iommu *iommu = container_of(iommu_dev, struct amd_iommu, iommu);
+ struct amd_iommu_pci_seg *pci_seg = iommu->pci_seg;
+
+ if (WARN_ON(!pci_seg->dev_table_preserve_count))
+ return;
+
+ if (!--pci_seg->dev_table_preserve_count)
+ iommu_unpreserve_pages(pci_seg->dev_table);
+}
+
+/*
+ * Map the domain's page-table mode onto its handoff wire value. Done with an
+ * explicit switch rather than a cast so that renumbering
+ * enum protection_domain_mode cannot silently change the ABI.
+ */
+static u32 pd_mode_to_ser(enum protection_domain_mode pd_mode)
+{
+ switch (pd_mode) {
+ case PD_MODE_V1:
+ return IOMMU_AMD_SER_PD_MODE_V1;
+ case PD_MODE_V2:
+ return IOMMU_AMD_SER_PD_MODE_V2;
+ default:
+ return IOMMU_AMD_SER_PD_MODE_NONE;
+ }
+}
+
+static void unpreserve_gcr3_level(u64 *tbl, int level)
+{
+ int i;
+
+ if (level > 0) {
+ for (i = 0; i < GCR3_ENTRIES_PER_LEVEL; i++) {
+ u64 *child;
+
+ if (!(tbl[i] & GCR3_VALID))
+ continue;
+
+ child = iommu_phys_to_virt(tbl[i] & PAGE_MASK);
+ unpreserve_gcr3_level(child, level - 1);
+ }
+ }
+
+ iommu_unpreserve_pages(tbl);
+}
+
+static int preserve_gcr3_level(u64 *tbl, int level)
+{
+ u64 *child;
+ int i, ret;
+
+ ret = iommu_preserve_pages(tbl);
+ if (ret)
+ return ret;
+
+ if (level == 0)
+ return 0;
+
+ for (i = 0; i < GCR3_ENTRIES_PER_LEVEL; i++) {
+ if (!(tbl[i] & GCR3_VALID))
+ continue;
+
+ child = iommu_phys_to_virt(tbl[i] & PAGE_MASK);
+ ret = preserve_gcr3_level(child, level - 1);
+ if (ret)
+ goto err_unwind;
+ }
+
+ return 0;
+
+err_unwind:
+ while (--i >= 0) {
+ if (!(tbl[i] & GCR3_VALID))
+ continue;
+
+ child = iommu_phys_to_virt(tbl[i] & PAGE_MASK);
+ unpreserve_gcr3_level(child, level - 1);
+ }
+ iommu_unpreserve_pages(tbl);
+ return ret;
+}
+
+/*
+ * A nested attach programs the DTE from the guest's own descriptor instead of
+ * from dev_data: the DOMID is the viommu's host domain ID for that guest domain,
+ * the GCR3 pointer is the guest's, and dev_data->domain is never assigned, so it
+ * still refers to whatever was attached before -- the nest parent, or nothing at
+ * all (see set_dte_nested()). None of that is recoverable from the state
+ * amd_iommu_preserve_device() serializes, and the core cannot filter it out
+ * either: a device on a nested domain is preserved against the domain of its
+ * paging parent, which is a legitimately preserved domain, so every check up to
+ * this point passes (see find_hwpt_paging()).
+ *
+ * Guest translation that this kernel set up itself always has a host-allocated
+ * GCR3 table behind it, which is what tells the two apart.
+ */
+static bool dev_is_nested_attached(struct iommu_dev_data *dev_data)
+{
+ struct amd_iommu *iommu = get_amd_iommu_from_dev_data(dev_data);
+ struct dev_table_entry *dev_table = get_dev_table(iommu);
+
+ if (!dev_table)
+ return false;
+
+ return (READ_ONCE(dev_table[dev_data->devid].data[0]) & DTE_FLAG_GV) &&
+ !dev_data->gcr3_info.gcr3_tbl;
+}
+
+/**
+ * amd_iommu_preserve_device - Preserve per-device live-update state
+ * @dev: The device being preserved
+ * @device_ser: Serialized per-device record to fill in
+ *
+ * The device's DTE rides across the kexec inside the preserved Device Table and
+ * keeps being used by the hardware, so everything the DTE still points at
+ * afterwards must be KHO-pinned. For a PASID-capable device that is the GCR3
+ * directory tree; the domain's page tables are pinned by the core when it
+ * preserves the domain. Anything not pinned must instead be dropped from the
+ * DTE, which is what amd_iommu_clear_unpreserved_dtes() does at shutdown.
+ *
+ * Records the domain ID the DTE tags this device with, and the domain's
+ * page-table mode so the next kernel can confirm its own default agrees before
+ * adopting the preserved tables. Both describe the DTE only as long as it was
+ * programmed from dev_data, so a device whose DTE came from a guest descriptor
+ * instead is refused outright; see dev_is_nested_attached().
+ *
+ * The walk runs during the quiesced live-update window, where the device is
+ * owned by its userspace driver and no PASID is being attached or detached, so
+ * the GCR3 tree is stable and no iommu-group lock is taken.
+ *
+ * Return: 0 on success, negative errno otherwise.
+ */
+int amd_iommu_preserve_device(struct device *dev,
+ struct iommu_device_ser *device_ser)
+{
+ struct gcr3_tbl_info *gcr3_info;
+ struct iommu_dev_data *dev_data;
+ int ret;
+
+ if (!dev_is_pci(dev)) {
+ dev_err(dev, "cannot preserve non-PCI device\n");
+ return -EOPNOTSUPP;
+ }
+
+ dev_data = dev_iommu_priv_get(dev);
+ if (!dev_data)
+ return -EINVAL;
+
+ if (dev_is_nested_attached(dev_data)) {
+ dev_err(dev, "cannot preserve device attached to a nested domain\n");
+ return -EOPNOTSUPP;
+ }
+
+ if (!dev_data->domain)
+ return -EINVAL;
+
+ /* Page-table mode is a domain property, independent of PASID use. */
+ device_ser->amd.pd_mode = pd_mode_to_ser(dev_data->domain->pd_mode);
+
+ gcr3_info = &dev_data->gcr3_info;
+ if (!gcr3_info->gcr3_tbl) {
+ /*
+ * Non-PASID device: the DTE's DOMID field holds the plain
+ * protection-domain ID (see amd_iommu_set_dte_v1()), and there
+ * is no GCR3 tree to pin.
+ */
+ device_ser->domain_iommu_ser.attachment_id = dev_data->domain->id;
+ device_ser->amd.gcr3_tbl_phys = 0;
+ device_ser->amd.gcr3_glx = 0;
+ return 0;
+ }
+
+ /* PASID device: pin the whole GCR3 directory tree. */
+ ret = preserve_gcr3_level(gcr3_info->gcr3_tbl, gcr3_info->glx);
+ if (ret)
+ return ret;
+
+ /*
+ * For a GCR3 device the DTE's DOMID field holds gcr3_info->domid, not
+ * domain->id (see set_dte_gcr3_table()).
+ */
+ device_ser->domain_iommu_ser.attachment_id = gcr3_info->domid;
+ device_ser->amd.gcr3_tbl_phys = __pa(gcr3_info->gcr3_tbl);
+ device_ser->amd.gcr3_glx = gcr3_info->glx;
+
+ return 0;
+}
+
+/**
+ * amd_iommu_unpreserve_device - Release per-device live-update state
+ * @dev: The device whose state is being released
+ * @device_ser: Serialized per-device record (unused)
+ *
+ * Reverse of amd_iommu_preserve_device(): unpins the GCR3 tree of a
+ * PASID-capable device. The DTE itself lives in the shared Device Table
+ * released by amd_iommu_unpreserve().
+ */
+void amd_iommu_unpreserve_device(struct device *dev,
+ struct iommu_device_ser *device_ser)
+{
+ struct gcr3_tbl_info *gcr3_info;
+ struct iommu_dev_data *dev_data;
+
+ if (!dev_is_pci(dev))
+ return;
+
+ dev_data = dev_iommu_priv_get(dev);
+ if (!dev_data)
+ return;
+
+ gcr3_info = &dev_data->gcr3_info;
+ if (!gcr3_info->gcr3_tbl)
+ return;
+
+ unpreserve_gcr3_level(gcr3_info->gcr3_tbl, gcr3_info->glx);
+}
--
2.43.0
^ permalink raw reply [flat|nested] 8+ messages in thread
* [RFC PATCH 5/7] iommu/amd: clear unpreserved DTEs and quiesce logs at live-update shutdown
2026-10-05 6:40 [RFC PATCH 0/7] iommu/amd: Implement live update state preservation Ankit Soni
` (3 preceding siblings ...)
2026-10-05 6:40 ` [RFC PATCH 4/7] iommu/amd: preserve IOMMU and device state for live update Ankit Soni
@ 2026-10-05 6:40 ` Ankit Soni
2026-10-05 6:40 ` [RFC PATCH 6/7] iommu/amd: restore preserved state on a live-update boot Ankit Soni
2026-10-05 6:40 ` [RFC PATCH 7/7] iommu/amd: reattach preserved devices to their restored domains Ankit Soni
6 siblings, 0 replies; 8+ messages in thread
From: Ankit Soni @ 2026-10-05 6:40 UTC (permalink / raw)
To: iommu, joro, will, jgg
Cc: suravee.suthikulpanit, vasant.hegde, robin.murphy,
joao.m.martins, alejandro.j.jimenez, pasha.tatashin, rppt,
pratyush, skhawaja, praan, baolu.lu, dwmw2, kevin.tian, dmatlack,
vipinsh, kexec, linux-kernel
Keep translation enabled on the units carrying preserved devices, but
block every device that was not preserved, so the next kernel cannot
translate through page tables it did not inherit.
Then stop the command, event, PPR and GA engines and wait for each to
report itself idle. Those buffers are not preserved, so an engine still
running would write into pages the next kernel is free to reuse.
If an engine does not go idle within the timeout, panic rather than
continue. Completing the handover would leave a running DMA engine
writing into pages the next kernel owns, and that corruption is both
silent and impossible to attribute later. Failing the live update is the
lesser harm.
Signed-off-by: Ankit Soni <Ankit.Soni@amd.com>
---
drivers/iommu/amd/amd_iommu.h | 5 ++
drivers/iommu/amd/init.c | 88 +++++++++++++++++++++++++++++++++-
drivers/iommu/amd/liveupdate.c | 83 ++++++++++++++++++++++++++++++++
3 files changed, 175 insertions(+), 1 deletion(-)
diff --git a/drivers/iommu/amd/amd_iommu.h b/drivers/iommu/amd/amd_iommu.h
index 5cf32e4898dc..4402724bfd06 100644
--- a/drivers/iommu/amd/amd_iommu.h
+++ b/drivers/iommu/amd/amd_iommu.h
@@ -236,5 +236,10 @@ int amd_iommu_preserve_device(struct device *dev,
struct iommu_device_ser *device_ser);
void amd_iommu_unpreserve_device(struct device *dev,
struct iommu_device_ser *device_ser);
+void amd_iommu_clear_unpreserved_dtes(struct amd_iommu *iommu);
+#else
+static inline void amd_iommu_clear_unpreserved_dtes(struct amd_iommu *iommu)
+{
+}
#endif /* CONFIG_IOMMU_LIVEUPDATE */
#endif /* AMD_IOMMU_H */
diff --git a/drivers/iommu/amd/init.c b/drivers/iommu/amd/init.c
index 40726dfef273..3b46d46f7143 100644
--- a/drivers/iommu/amd/init.c
+++ b/drivers/iommu/amd/init.c
@@ -32,6 +32,7 @@
#include <asm/sev.h>
#include <linux/crash_dump.h>
+#include <linux/iommu-liveupdate.h>
#include "amd_iommu.h"
#include "../irq_remapping.h"
@@ -3045,6 +3046,91 @@ static void disable_iommus(void)
#endif
}
+/*
+ * Bound the wait for a log engine to report itself idle. An engine only has
+ * to finish a write it has already started, which takes microseconds, so this
+ * is ample. It is deliberately far shorter than MMIO_STATUS_TIMEOUT because
+ * this runs on the live-update shutdown path, where any stall is downtime.
+ */
+#define LU_LOG_QUIESCE_RETRIES 10000 /* x 10us = 100ms */
+
+/*
+ * Clearing a log's enable bit only requests a stop; the engine may still be
+ * completing a write. Each log reports its real state in a separate "running"
+ * status bit, so wait for that to clear before handing over.
+ */
+static void wait_log_stopped(struct amd_iommu *iommu, u32 run_mask,
+ const char *name)
+{
+ u32 status;
+ int i;
+
+ for (i = 0; i < LU_LOG_QUIESCE_RETRIES; ++i) {
+ status = readl(iommu->mmio_base + MMIO_STATUS_OFFSET);
+ if (!(status & run_mask))
+ return;
+ udelay(10);
+ }
+
+ /*
+ * The buffer is not preserved, so the next kernel is free to reuse
+ * these pages. An engine still running here would keep writing into
+ * them after the kexec, corrupting whatever the next kernel puts
+ * there. That corruption is silent and unattributable, so refuse the
+ * handover instead of completing one we cannot prove is safe.
+ */
+ panic("AMD-Vi: IOMMU:%d %s log still running at handover; refusing to hand over a running DMA engine\n",
+ iommu->index, name);
+}
+
+/*
+ * Stop the hardware writing event/PPR/GA logs and reading the command
+ * buffer, without turning translation off. Call after DTE cleanup: the
+ * cache flush still needs the command buffer. The next kernel allocates
+ * fresh buffers.
+ */
+static void amd_iommu_quiesce_logs(struct amd_iommu *iommu)
+{
+ /*
+ * The completion wait at the end of amd_iommu_clear_unpreserved_dtes()
+ * has already drained the command buffer, so there is nothing left for
+ * the hardware to read.
+ */
+ iommu_disable_command_buffer(iommu);
+
+ iommu_feature_disable(iommu, CONTROL_EVT_INT_EN);
+ iommu_disable_event_buffer(iommu);
+ wait_log_stopped(iommu, MMIO_STATUS_EVT_RUN_MASK, "event");
+
+ iommu_feature_disable(iommu, CONTROL_GAINT_EN);
+ iommu_feature_disable(iommu, CONTROL_GALOG_EN);
+ wait_log_stopped(iommu, MMIO_STATUS_GALOG_RUN_MASK, "GA");
+
+ iommu_feature_disable(iommu, CONTROL_PPRINT_EN);
+ iommu_feature_disable(iommu, CONTROL_PPRLOG_EN);
+ iommu_feature_disable(iommu, CONTROL_PPR_EN);
+ wait_log_stopped(iommu, MMIO_STATUS_PPR_RUN_MASK, "PPR");
+}
+
+static void amd_iommu_shutdown(void)
+{
+ struct amd_iommu *iommu;
+
+ for_each_iommu(iommu) {
+ if (iommu_preserved_state(&iommu->iommu)) {
+ amd_iommu_clear_unpreserved_dtes(iommu);
+ amd_iommu_quiesce_logs(iommu);
+ } else {
+ iommu_disable(iommu);
+ }
+ }
+
+#ifdef CONFIG_IRQ_REMAP
+ if (AMD_IOMMU_GUEST_IR_VAPIC(amd_iommu_guest_ir))
+ amd_iommu_irq_ops.capability &= ~(1 << IRQ_POSTING_CAP);
+#endif
+}
+
/*
* Suspend/Resume support
* disable suspend until real resume implemented
@@ -3500,7 +3586,7 @@ static int __init state_next(void)
break;
case IOMMU_ACPI_FINISHED:
early_enable_iommus();
- x86_platform.iommu_shutdown = disable_iommus;
+ x86_platform.iommu_shutdown = amd_iommu_shutdown;
init_state = IOMMU_ENABLED;
break;
case IOMMU_ENABLED:
diff --git a/drivers/iommu/amd/liveupdate.c b/drivers/iommu/amd/liveupdate.c
index 096a23bb4e7b..a3a9ebea5138 100644
--- a/drivers/iommu/amd/liveupdate.c
+++ b/drivers/iommu/amd/liveupdate.c
@@ -279,3 +279,86 @@ void amd_iommu_unpreserve_device(struct device *dev,
unpreserve_gcr3_level(gcr3_info->gcr3_tbl, gcr3_info->glx);
}
+
+/*
+ * Reset one non-preserved device's DTE to the blocked state during live-update
+ * shutdown. Every unpreserved device is reset so the next kernel starts from a
+ * clean slate for it and cannot translate through a domain whose page tables
+ * were not preserved.
+ */
+static int clear_unpreserved_dte(struct device *dev,
+ struct iommu_device *iommu_dev, void *arg)
+{
+ struct amd_iommu *iommu = container_of(iommu_dev, struct amd_iommu,
+ iommu);
+ struct dev_table_entry new = {};
+ struct iommu_dev_data *dev_data;
+
+ dev_data = dev_iommu_priv_get(dev);
+ if (!dev_data)
+ return 0;
+
+ if (dev_is_pci(dev) && dev_iommu_preserved_state(dev))
+ return 0;
+
+ amd_iommu_make_clear_dte(dev_data, &new);
+ amd_iommu_update_dte(iommu, dev_data, &new);
+
+ return 0;
+}
+
+static void clear_irq_dtes(struct amd_iommu *iommu)
+{
+ struct amd_iommu_pci_seg *pci_seg = iommu->pci_seg;
+ struct dev_table_entry *dev_table = get_dev_table(iommu);
+ u32 devid;
+ u64 dte2;
+
+ if (!amd_iommu_irq_remap)
+ return;
+
+ /*
+ * Interrupt remapping tables are not preserved. Clear the interrupt
+ * fields on every DTE so none still points at a table the next kernel
+ * is free to recycle. DTE_DATA2_INTR_MASK is everything in data[2]
+ * except the guest page-table level, which preserved devices still
+ * need. The GCR3 pointer lives in data[0] and data[1], so it is not
+ * affected.
+ *
+ * This must run after the clear_unpreserved_dte() pass, because
+ * write_dte_upper128() deliberately copies DTE_DATA2_INTR_MASK back
+ * from the old entry. Clearing a DTE therefore keeps its interrupt
+ * fields, and only this walk removes them.
+ *
+ * The walk is by raw devid, so there is no iommu_dev_data to take
+ * dte_lock on, and looking one up per entry would make this O(n^2).
+ * Going without the lock is safe only because amd_iommu_shutdown()
+ * runs from native_machine_shutdown(), after the other CPUs are
+ * stopped and interrupts are off, so nothing can race these writes.
+ */
+ for (devid = 0; devid <= pci_seg->last_bdf; devid++) {
+ dte2 = READ_ONCE(dev_table[devid].data[2]);
+ if (!(dte2 & DTE_IRQ_REMAP_ENABLE))
+ continue;
+
+ WRITE_ONCE(dev_table[devid].data[2],
+ dte2 & ~DTE_DATA2_INTR_MASK);
+ }
+}
+
+/**
+ * amd_iommu_clear_unpreserved_dtes - Quiesce non-preserved devices at shutdown
+ * @iommu: The IOMMU whose device table is being cleaned up
+ */
+void amd_iommu_clear_unpreserved_dtes(struct amd_iommu *iommu)
+{
+ struct iommu_dev_iter iter = {
+ .fn = clear_unpreserved_dte,
+ .iommu = &iommu->iommu,
+ };
+
+ iommu_for_each_dev(&iter);
+ clear_irq_dtes(iommu);
+
+ amd_iommu_flush_all_caches(iommu);
+}
--
2.43.0
^ permalink raw reply [flat|nested] 8+ messages in thread
* [RFC PATCH 6/7] iommu/amd: restore preserved state on a live-update boot
2026-10-05 6:40 [RFC PATCH 0/7] iommu/amd: Implement live update state preservation Ankit Soni
` (4 preceding siblings ...)
2026-10-05 6:40 ` [RFC PATCH 5/7] iommu/amd: clear unpreserved DTEs and quiesce logs at live-update shutdown Ankit Soni
@ 2026-10-05 6:40 ` Ankit Soni
2026-10-05 6:40 ` [RFC PATCH 7/7] iommu/amd: reattach preserved devices to their restored domains Ankit Soni
6 siblings, 0 replies; 8+ messages in thread
From: Ankit Soni @ 2026-10-05 6:40 UTC (permalink / raw)
To: iommu, joro, will, jgg
Cc: suravee.suthikulpanit, vasant.hegde, robin.murphy,
joao.m.martins, alejandro.j.jimenez, pasha.tatashin, rppt,
pratyush, skhawaja, praan, baolu.lu, dwmw2, kevin.tian, dmatlack,
vipinsh, kexec, linux-kernel
Adopt the preserved device table for the PCI segment instead of
allocating a fresh one, and reserve the domain IDs it already carries so
a domain allocated by this kernel cannot be handed an ID the hardware is
still tagging cache entries with.
Also stop clearing the enable bit on a unit that was handed over
translating. Doing so would drop its devices into untranslated
passthrough while their DMA is still in flight.
Signed-off-by: Ankit Soni <Ankit.Soni@amd.com>
---
drivers/iommu/amd/amd_iommu.h | 7 ++
drivers/iommu/amd/init.c | 165 ++++++++++++++++++++++++++-------
drivers/iommu/amd/liveupdate.c | 33 +++++++
3 files changed, 174 insertions(+), 31 deletions(-)
diff --git a/drivers/iommu/amd/amd_iommu.h b/drivers/iommu/amd/amd_iommu.h
index 4402724bfd06..ac6d7a17eb40 100644
--- a/drivers/iommu/amd/amd_iommu.h
+++ b/drivers/iommu/amd/amd_iommu.h
@@ -237,9 +237,16 @@ int amd_iommu_preserve_device(struct device *dev,
void amd_iommu_unpreserve_device(struct device *dev,
struct iommu_device_ser *device_ser);
void amd_iommu_clear_unpreserved_dtes(struct amd_iommu *iommu);
+void amd_iommu_restore_dev_table(struct amd_iommu *iommu,
+ struct iommu_hw_ser *ser);
#else
static inline void amd_iommu_clear_unpreserved_dtes(struct amd_iommu *iommu)
{
}
+
+static inline void amd_iommu_restore_dev_table(struct amd_iommu *iommu,
+ struct iommu_hw_ser *ser)
+{
+}
#endif /* CONFIG_IOMMU_LIVEUPDATE */
#endif /* AMD_IOMMU_H */
diff --git a/drivers/iommu/amd/init.c b/drivers/iommu/amd/init.c
index 3b46d46f7143..8aba3b715a02 100644
--- a/drivers/iommu/amd/init.c
+++ b/drivers/iommu/amd/init.c
@@ -33,6 +33,7 @@
#include <linux/crash_dump.h>
#include <linux/iommu-liveupdate.h>
+#include <linux/kexec_handover.h>
#include "amd_iommu.h"
#include "../irq_remapping.h"
@@ -1135,14 +1136,42 @@ static void set_dte_bit(struct dev_table_entry *dte, u8 bit)
dte->data[i] |= (1UL << _bit);
}
-static bool __reuse_device_table(struct amd_iommu *iommu)
+/*
+ * Reserve the domain IDs the previous kernel programmed into the adopted Device
+ * Table, so this kernel never hands out an ID the hardware still tags cache
+ * entries with. @pci_seg->old_dev_tbl_cpy must already point at that table.
+ */
+static bool reserve_dev_table_domain_ids(struct amd_iommu_pci_seg *pci_seg)
{
- struct amd_iommu_pci_seg *pci_seg = iommu->pci_seg;
struct dev_table_entry *old_dev_tbl_entry;
- u32 lo, hi, old_devtb_size, devid;
- phys_addr_t old_devtb_phys;
+ u32 devid;
u16 dom_id;
bool dte_v;
+
+ for (devid = 0; devid <= pci_seg->last_bdf; devid++) {
+ old_dev_tbl_entry = &pci_seg->old_dev_tbl_cpy[devid];
+ dte_v = FIELD_GET(DTE_FLAG_V, old_dev_tbl_entry->data[0]);
+ dom_id = FIELD_GET(DTE_DOMID_MASK, old_dev_tbl_entry->data[1]);
+
+ if (!dte_v || !dom_id)
+ continue;
+ /*
+ * ID reservation can fail with -ENOSPC when there
+ * are multiple devices present in the same domain,
+ * hence check only for -ENOMEM.
+ */
+ if (amd_iommu_pdom_id_reserve(dom_id, GFP_KERNEL) == -ENOMEM)
+ return false;
+ }
+
+ return true;
+}
+
+static bool __reuse_device_table(struct amd_iommu *iommu)
+{
+ struct amd_iommu_pci_seg *pci_seg = iommu->pci_seg;
+ u32 lo, hi, old_devtb_size;
+ phys_addr_t old_devtb_phys;
u64 entry;
/* Each IOMMU use separate device table with the same size */
@@ -1178,22 +1207,6 @@ static bool __reuse_device_table(struct amd_iommu *iommu)
return false;
}
- for (devid = 0; devid <= pci_seg->last_bdf; devid++) {
- old_dev_tbl_entry = &pci_seg->old_dev_tbl_cpy[devid];
- dte_v = FIELD_GET(DTE_FLAG_V, old_dev_tbl_entry->data[0]);
- dom_id = FIELD_GET(DTE_DOMID_MASK, old_dev_tbl_entry->data[1]);
-
- if (!dte_v || !dom_id)
- continue;
- /*
- * ID reservation can fail with -ENOSPC when there
- * are multiple devices present in the same domain,
- * hence check only for -ENOMEM.
- */
- if (amd_iommu_pdom_id_reserve(dom_id, GFP_KERNEL) == -ENOMEM)
- return false;
- }
-
return true;
}
@@ -1201,27 +1214,69 @@ static bool reuse_device_table(void)
{
struct amd_iommu *iommu;
struct amd_iommu_pci_seg *pci_seg;
-
- if (!amd_iommu_pre_enabled)
- return false;
-
- pr_warn("Translation is already enabled - trying to reuse translation structures\n");
+ bool reused = false;
/*
* All IOMMUs within PCI segment shares common device table.
* Hence reuse device table only once per PCI segment.
*/
for_each_pci_segment(pci_seg) {
+ struct amd_iommu *preserved_iommu = NULL;
+ struct iommu_hw_ser *ser = NULL;
+
+ /*
+ * The device table is shared by the whole segment and pinned
+ * by whichever unit preserved it, so any member holding
+ * preserved state can describe it for all of them.
+ */
+ for_each_iommu(iommu) {
+ if (pci_seg->id != iommu->pci_seg->id)
+ continue;
+
+ ser = iommu_get_preserved_data(iommu->mmio_phys,
+ IOMMU_AMD);
+ if (ser) {
+ preserved_iommu = iommu;
+ break;
+ }
+ }
+
+ if (preserved_iommu) {
+ amd_iommu_restore_dev_table(preserved_iommu, ser);
+ /*
+ * Fatal for the same reason a failed table adopt is;
+ * see amd_iommu_restore_dev_table().
+ */
+ if (!reserve_dev_table_domain_ids(pci_seg))
+ panic("AMD-Vi: IOMMU:%d cannot reserve the preserved domain IDs\n",
+ preserved_iommu->index);
+ reused = true;
+ continue;
+ }
+
+ /*
+ * Nothing was handed over for this segment. The kdump path can
+ * still remap the previous kernel's table, but only when every
+ * IOMMU came up with translation already enabled.
+ */
+ if (!amd_iommu_pre_enabled)
+ return false;
+
+ pr_warn("Translation is already enabled - trying to reuse translation structures\n");
+
for_each_iommu(iommu) {
if (pci_seg->id != iommu->pci_seg->id)
continue;
if (!__reuse_device_table(iommu))
return false;
+ if (!reserve_dev_table_domain_ids(pci_seg))
+ return false;
+ reused = true;
break;
}
}
- return true;
+ return reused;
}
struct dev_table_entry *amd_iommu_get_ivhd_dte_flags(u16 segid, u16 devid)
@@ -1948,6 +2003,7 @@ static int __init init_iommu_one(struct amd_iommu *iommu, struct ivhd_header *h,
static int __init init_iommu_one_late(struct amd_iommu *iommu)
{
+ struct iommu_hw_ser *ser;
int ret;
ret = alloc_iommu_buffers(iommu);
@@ -1957,7 +2013,25 @@ static int __init init_iommu_one_late(struct amd_iommu *iommu)
iommu->int_enabled = false;
init_translation_status(iommu);
- if (translation_pre_enabled(iommu) && !is_kdump_kernel()) {
+
+ /* The MMIO physical base is the handoff token; NULL means a cold boot. */
+ ser = iommu_get_preserved_data(iommu->mmio_phys, IOMMU_AMD);
+
+ if (translation_pre_enabled(iommu) && !is_kdump_kernel() && !ser) {
+ /*
+ * A handover boot that left this unit translating but produced
+ * no record for it is not a stale enable. Either the handover
+ * was never consumable (this kernel booted without
+ * liveupdate=1) or it was not visible yet. Clearing the enable
+ * bit drops every device behind this unit into untranslated
+ * passthrough while its DMA is still in flight, aimed at
+ * addresses that mean nothing in this kernel's physical map,
+ * so stop instead of doing it silently.
+ */
+ if (is_kho_boot())
+ panic("AMD-Vi: IOMMU:%d was handed over translating but no preserved state is available; refusing to disable it under live DMA\n",
+ iommu->index);
+
iommu_disable(iommu);
clear_translation_pre_enabled(iommu);
pr_warn("Translation was enabled for IOMMU:%d but we are not in kdump mode\n",
@@ -2944,6 +3018,18 @@ static void early_enable_iommus(void)
}
for_each_iommu(iommu) {
+ /*
+ * Units with no preserved devices come back with
+ * translation off. Programming their buffers without
+ * enabling them leaves amd_iommu_flush_all_caches()
+ * waiting on a completion idle hardware never posts,
+ * so bring them up the normal way.
+ */
+ if (!translation_pre_enabled(iommu)) {
+ early_enable_iommu(iommu);
+ continue;
+ }
+
iommu_disable_command_buffer(iommu);
iommu_disable_event_buffer(iommu);
iommu_disable_irtcachedis(iommu);
@@ -3033,12 +3119,25 @@ static void enable_iommus_vapic(void)
#endif
}
-static void disable_iommus(void)
+static bool iommu_was_handed_over(struct amd_iommu *iommu)
+{
+ struct iommu_hw_ser *ser;
+
+ ser = iommu_get_preserved_data(iommu->mmio_phys, IOMMU_AMD);
+
+ return ser;
+}
+
+static void __disable_iommus(bool keep_handed_over)
{
struct amd_iommu *iommu;
- for_each_iommu(iommu)
+ for_each_iommu(iommu) {
+ if (keep_handed_over && iommu_was_handed_over(iommu))
+ continue;
+
iommu_disable(iommu);
+ }
#ifdef CONFIG_IRQ_REMAP
if (AMD_IOMMU_GUEST_IR_VAPIC(amd_iommu_guest_ir))
@@ -3046,6 +3145,11 @@ static void disable_iommus(void)
#endif
}
+static void disable_iommus(void)
+{
+ __disable_iommus(false);
+}
+
/*
* Bound the wait for a log engine to report itself idle. An engine only has
* to finish a write it has already started, which takes microseconds, so this
@@ -3388,9 +3492,8 @@ static int __init early_amd_iommu_init(void)
amd_iommu_pgtable = PD_MODE_NONE;
}
- /* Disable any previously enabled IOMMUs */
if (!is_kdump_kernel() || amd_iommu_disabled)
- disable_iommus();
+ __disable_iommus(true);
if (amd_iommu_irq_remap)
amd_iommu_irq_remap = check_ioapic_information();
diff --git a/drivers/iommu/amd/liveupdate.c b/drivers/iommu/amd/liveupdate.c
index a3a9ebea5138..2e9a1de6b1ff 100644
--- a/drivers/iommu/amd/liveupdate.c
+++ b/drivers/iommu/amd/liveupdate.c
@@ -362,3 +362,36 @@ void amd_iommu_clear_unpreserved_dtes(struct amd_iommu *iommu)
amd_iommu_flush_all_caches(iommu);
}
+
+void amd_iommu_restore_dev_table(struct amd_iommu *iommu,
+ struct iommu_hw_ser *ser)
+{
+ struct amd_iommu_pci_seg *pci_seg = iommu->pci_seg;
+
+ if (ser->amd.dev_table_size != pci_seg->dev_table_size)
+ panic("AMD-Vi: IOMMU:%d preserved device table size mismatch (0x%x vs 0x%x)\n",
+ iommu->index, ser->amd.dev_table_size,
+ pci_seg->dev_table_size);
+
+ if (ser->amd.pci_seg_id != pci_seg->id)
+ panic("AMD-Vi: IOMMU:%d preserved PCI segment mismatch (%u vs %u)\n",
+ iommu->index, ser->amd.pci_seg_id, pci_seg->id);
+
+ if (ser->amd.efr != iommu->features ||
+ ser->amd.efr2 != iommu->features2)
+ panic("AMD-Vi: IOMMU:%d preserved feature registers mismatch (EFR 0x%llx/0x%llx vs 0x%llx/0x%llx)\n",
+ iommu->index, ser->amd.efr, ser->amd.efr2,
+ iommu->features, iommu->features2);
+
+ /*
+ * Reclaim the folio from KHO the first time this segment's table is
+ * seen, then adopt it. Later IOMMUs in the same segment reference the
+ * same physical table and must not restore it again.
+ */
+ if (!ser->amd.restored) {
+ iommu_restore_pages(ser->amd.dev_table_phys);
+ ser->amd.restored = 1;
+ }
+
+ pci_seg->old_dev_tbl_cpy = __va(ser->amd.dev_table_phys);
+}
--
2.43.0
^ permalink raw reply [flat|nested] 8+ messages in thread
* [RFC PATCH 7/7] iommu/amd: reattach preserved devices to their restored domains
2026-10-05 6:40 [RFC PATCH 0/7] iommu/amd: Implement live update state preservation Ankit Soni
` (5 preceding siblings ...)
2026-10-05 6:40 ` [RFC PATCH 6/7] iommu/amd: restore preserved state on a live-update boot Ankit Soni
@ 2026-10-05 6:40 ` Ankit Soni
6 siblings, 0 replies; 8+ messages in thread
From: Ankit Soni @ 2026-10-05 6:40 UTC (permalink / raw)
To: iommu, joro, will, jgg
Cc: suravee.suthikulpanit, vasant.hegde, robin.murphy,
joao.m.martins, alejandro.j.jimenez, pasha.tatashin, rppt,
pratyush, skhawaja, praan, baolu.lu, dwmw2, kevin.tian, dmatlack,
vipinsh, kexec, linux-kernel
A device that rode across the kexec is already translating through the
DTE this kernel adopted, so adopt its domain ID and GCR3 table instead
of programming the DTE again, and verify the adopted state against the
live DTE. A mismatch means the kernel and the hardware disagree about
how the device's translations are cached, so fail the attach and leave
the DTE alone.
Refuse every other attach for such a device, in all four attach ops,
because the blocking, identity, nested and freshly built paging domains
would each rewrite that DTE under a device that is still doing DMA. The
paging path compares the target against the domain recorded for this
device, so a second restored domain cannot be swapped in either, and
clone_alias() is skipped for the same reason.
On device removal the core skips the release domain for a preserved
device, so detach_device() never runs and dev_data->domain stays set.
Drop the software state and leave the DTE alone; the restored domain
outlives the device.
Signed-off-by: Ankit Soni <Ankit.Soni@amd.com>
---
drivers/iommu/amd/amd_iommu.h | 10 +++
drivers/iommu/amd/iommu.c | 130 ++++++++++++++++++++++++++++-
drivers/iommu/amd/liveupdate.c | 146 +++++++++++++++++++++++++++++++++
drivers/iommu/amd/nested.c | 8 ++
4 files changed, 292 insertions(+), 2 deletions(-)
diff --git a/drivers/iommu/amd/amd_iommu.h b/drivers/iommu/amd/amd_iommu.h
index ac6d7a17eb40..93fe65b28d53 100644
--- a/drivers/iommu/amd/amd_iommu.h
+++ b/drivers/iommu/amd/amd_iommu.h
@@ -239,6 +239,9 @@ void amd_iommu_unpreserve_device(struct device *dev,
void amd_iommu_clear_unpreserved_dtes(struct amd_iommu *iommu);
void amd_iommu_restore_dev_table(struct amd_iommu *iommu,
struct iommu_hw_ser *ser);
+int amd_iommu_reattach_device(struct device *dev,
+ struct protection_domain *domain,
+ struct iommu_device_ser *device_ser);
#else
static inline void amd_iommu_clear_unpreserved_dtes(struct amd_iommu *iommu)
{
@@ -248,5 +251,12 @@ static inline void amd_iommu_restore_dev_table(struct amd_iommu *iommu,
struct iommu_hw_ser *ser)
{
}
+
+static inline int amd_iommu_reattach_device(struct device *dev,
+ struct protection_domain *domain,
+ struct iommu_device_ser *device_ser)
+{
+ return -EOPNOTSUPP;
+}
#endif /* CONFIG_IOMMU_LIVEUPDATE */
#endif /* AMD_IOMMU_H */
diff --git a/drivers/iommu/amd/iommu.c b/drivers/iommu/amd/iommu.c
index 72df97b99589..0bdd93016814 100644
--- a/drivers/iommu/amd/iommu.c
+++ b/drivers/iommu/amd/iommu.c
@@ -444,6 +444,13 @@ static int clone_alias(struct pci_dev *pdev_origin, u16 alias, void *data)
ret = -EINVAL;
goto out;
}
+ /*
+ * A preserved device is still translating through the DTE this kernel
+ * adopted, so never clone another device's DTE over it.
+ */
+ if (alias_data->dev && dev_iommu_restored_state(alias_data->dev))
+ goto out;
+
update_dte256(iommu, alias_data, &new);
amd_iommu_set_rlookup_table(iommu, alias);
@@ -2376,6 +2383,55 @@ static void pdom_detach_iommu(struct amd_iommu *iommu,
spin_unlock_irqrestore(&pdom->lock, flags);
}
+/*
+ * True when @dom is the one domain this kernel restored for @dev. Attaching a
+ * preserved device to anything else has to be refused, including a different
+ * restored domain.
+ */
+static bool dev_restored_domain_matches(struct device *dev,
+ struct iommu_domain *dom)
+{
+ struct iommu_device_ser *device_ser = dev_iommu_restored_state(dev);
+ struct iommu_domain_ser *domain_ser;
+
+ if (!device_ser || !device_ser->domain_iommu_ser.domain_phys)
+ return false;
+
+ domain_ser = phys_to_virt(device_ser->domain_iommu_ser.domain_phys);
+
+ return domain_ser->restored_domain == dom;
+}
+
+static int reattach_verify_dte(struct amd_iommu *iommu,
+ struct iommu_dev_data *dev_data, u16 domid)
+{
+ struct dev_table_entry dte;
+
+ get_dte256(iommu, dev_data, &dte);
+
+ if (!(dte.data[0] & DTE_FLAG_V) ||
+ FIELD_GET(DTE_DOMID_MASK, dte.data[1]) != domid) {
+ dev_err(dev_data->dev,
+ "preserved DTE does not describe the adopted domain ID %u (DTE 0x%llx/0x%llx)\n",
+ domid, dte.data[0], dte.data[1]);
+ return -EINVAL;
+ }
+
+ /*
+ * If these two disagree, either the device caches translations
+ * nobody invalidates or the IOMMU waits for completions from a
+ * device that will not send them.
+ */
+ if (!!(dte.data[1] & DTE_FLAG_IOTLB) != !!dev_data->ats_enabled) {
+ dev_err(dev_data->dev,
+ "preserved DTE and restored ATS state disagree (DTE 0x%llx/0x%llx, ats_enabled %u)\n",
+ dte.data[0], dte.data[1], dev_data->ats_enabled);
+ return -EINVAL;
+ }
+
+ return 0;
+}
+
/*
* If a device is not yet associated with a domain, this function makes the
* device visible in the domain
@@ -2385,6 +2441,7 @@ static int attach_device(struct device *dev,
{
struct iommu_dev_data *dev_data = dev_iommu_priv_get(dev);
struct amd_iommu *iommu = get_amd_iommu_from_dev_data(dev_data);
+ struct iommu_device_ser *device_ser = NULL;
struct pci_dev *pdev;
unsigned long flags;
int ret = 0;
@@ -2401,8 +2458,17 @@ static int attach_device(struct device *dev,
if (ret)
goto out;
+ if (iommu_domain_restored_state(&domain->domain))
+ device_ser = dev_iommu_restored_state(dev);
+
/* Setup GCR3 table */
- if (pdom_is_sva_capable(domain)) {
+ if (device_ser) {
+ ret = amd_iommu_reattach_device(dev, domain, device_ser);
+ if (ret) {
+ pdom_detach_iommu(iommu, domain);
+ goto out;
+ }
+ } else if (pdom_is_sva_capable(domain)) {
ret = init_gcr3_table(dev_data, domain);
if (ret) {
pdom_detach_iommu(iommu, domain);
@@ -2432,12 +2498,33 @@ static int attach_device(struct device *dev,
spin_unlock_irqrestore(&domain->lock, flags);
/* Update device table */
- dev_update_dte(dev_data, true);
+ if (device_ser) {
+ ret = reattach_verify_dte(iommu, dev_data,
+ device_ser->domain_iommu_ser.attachment_id);
+ if (ret)
+ goto err_reattach;
+ } else {
+ dev_update_dte(dev_data, true);
+ }
out:
mutex_unlock(&dev_data->mutex);
return ret;
+
+ /*
+ * The DTE is the previous kernel's and the hardware is still walking
+ * it, so unwind the software state only and leave it alone.
+ */
+err_reattach:
+ spin_lock_irqsave(&domain->lock, flags);
+ list_del(&dev_data->list);
+ spin_unlock_irqrestore(&domain->lock, flags);
+ dev_data->domain = NULL;
+ pdom_detach_iommu(iommu, domain);
+ mutex_unlock(&dev_data->mutex);
+
+ return ret;
}
/*
@@ -2493,6 +2580,30 @@ static void detach_device(struct device *dev)
mutex_unlock(&dev_data->mutex);
}
+/*
+ * The core skips the release domain for a preserved device, so its DTE is still
+ * live here. Drop the software state only. The restored domain outlives the
+ * device and the hardware keeps walking the page tables we adopted.
+ */
+static void detach_restored_device(struct device *dev)
+{
+ struct iommu_dev_data *dev_data = dev_iommu_priv_get(dev);
+ struct amd_iommu *iommu = get_amd_iommu_from_dev_data(dev_data);
+ struct protection_domain *domain = dev_data->domain;
+ unsigned long flags;
+
+ mutex_lock(&dev_data->mutex);
+
+ spin_lock_irqsave(&domain->lock, flags);
+ list_del(&dev_data->list);
+ spin_unlock_irqrestore(&domain->lock, flags);
+
+ dev_data->domain = NULL;
+ pdom_detach_iommu(iommu, domain);
+
+ mutex_unlock(&dev_data->mutex);
+}
+
static struct iommu_device *amd_iommu_probe_device(struct device *dev)
{
struct iommu_device *iommu_dev;
@@ -2561,6 +2672,9 @@ static void amd_iommu_release_device(struct device *dev)
{
struct iommu_dev_data *dev_data = dev_iommu_priv_get(dev);
+ if (dev_iommu_restored_state(dev) && dev_data->domain)
+ detach_restored_device(dev);
+
WARN_ON(dev_data->domain);
/*
@@ -2935,6 +3049,9 @@ static int blocked_domain_attach_device(struct iommu_domain *domain,
{
struct iommu_dev_data *dev_data = dev_iommu_priv_get(dev);
+ if (dev_iommu_restored_state(dev))
+ return -EBUSY;
+
if (dev_data->domain)
detach_device(dev);
@@ -3018,6 +3135,15 @@ static int amd_iommu_attach_device(struct iommu_domain *dom, struct device *dev,
if (dom->dirty_ops && !amd_iommu_hd_support(iommu))
return -EINVAL;
+ /*
+ * A preserved device is still translating through the DTE this kernel
+ * adopted, so only the one domain restored for it may be attached.
+ * Refuse before the detach below, which would tear that DTE down.
+ */
+ if (dev_iommu_restored_state(dev) &&
+ !dev_restored_domain_matches(dev, dom))
+ return -EBUSY;
+
if (dev_data->domain)
detach_device(dev);
diff --git a/drivers/iommu/amd/liveupdate.c b/drivers/iommu/amd/liveupdate.c
index 2e9a1de6b1ff..3d09054be438 100644
--- a/drivers/iommu/amd/liveupdate.c
+++ b/drivers/iommu/amd/liveupdate.c
@@ -395,3 +395,149 @@ void amd_iommu_restore_dev_table(struct amd_iommu *iommu,
pci_seg->old_dev_tbl_cpy = __va(ser->amd.dev_table_phys);
}
+
+/* Inverse of pd_mode_to_ser(), for a value coming off the wire. */
+static enum protection_domain_mode ser_to_pd_mode(u32 mode)
+{
+ switch (mode) {
+ case IOMMU_AMD_SER_PD_MODE_V1:
+ return PD_MODE_V1;
+ case IOMMU_AMD_SER_PD_MODE_V2:
+ return PD_MODE_V2;
+ default:
+ return PD_MODE_NONE;
+ }
+}
+
+static void restore_gcr3_level(u64 *tbl, int level)
+{
+ int i;
+
+ iommu_restore_pages(__pa(tbl));
+
+ if (level == 0)
+ return;
+
+ for (i = 0; i < GCR3_ENTRIES_PER_LEVEL; i++) {
+ if (!(tbl[i] & GCR3_VALID))
+ continue;
+
+ restore_gcr3_level(iommu_phys_to_virt(tbl[i] & PAGE_MASK),
+ level - 1);
+ }
+}
+
+static u64 *gcr3_pasid0_entry(u64 *tbl, int level)
+{
+ while (level--) {
+ if (!(tbl[0] & GCR3_VALID))
+ return NULL;
+
+ tbl = iommu_phys_to_virt(tbl[0] & PAGE_MASK);
+ }
+
+ return &tbl[0];
+}
+
+/*
+ * Every preserved domain ID is reserved before the first fresh allocation (see
+ * reserve_dev_table_domain_ids()), so @attachment_id already matching means an
+ * earlier device of this domain swapped it rather than a collision.
+ */
+static int reattach_domain_id(struct device *dev,
+ struct protection_domain *domain,
+ u16 attachment_id)
+{
+ bool disagree = false;
+ unsigned long flags;
+ int fresh_id = -1;
+
+ spin_lock_irqsave(&domain->lock, flags);
+ if (domain->id != attachment_id) {
+ if (list_empty(&domain->dev_list)) {
+ fresh_id = domain->id;
+ domain->id = attachment_id;
+ } else {
+ disagree = true;
+ }
+ }
+ spin_unlock_irqrestore(&domain->lock, flags);
+
+ if (disagree) {
+ dev_err(dev, "preserved domain ID %u conflicts with %u already adopted for its domain\n",
+ attachment_id, domain->id);
+ return -EINVAL;
+ }
+
+ if (fresh_id >= 0)
+ amd_iommu_pdom_id_free(fresh_id);
+
+ return 0;
+}
+
+static int reattach_gcr3_table(struct device *dev,
+ struct protection_domain *domain,
+ struct iommu_device_ser *device_ser)
+{
+ struct iommu_dev_data *dev_data = dev_iommu_priv_get(dev);
+ struct gcr3_tbl_info *gcr3_info = &dev_data->gcr3_info;
+ struct pt_iommu_x86_64_hw_info pt_info;
+ u32 glx = device_ser->amd.gcr3_glx;
+ u64 *gcr3_tbl, *pte;
+
+ if (amd_iommu_max_glx_val < 0 || glx > (u32)amd_iommu_max_glx_val) {
+ dev_err(dev, "cannot adopt a %u-level preserved GCR3 tree, hardware supports %d\n",
+ glx, amd_iommu_max_glx_val);
+ return -EINVAL;
+ }
+
+ gcr3_tbl = phys_to_virt(device_ser->amd.gcr3_tbl_phys);
+ restore_gcr3_level(gcr3_tbl, glx);
+
+ gcr3_info->gcr3_tbl = gcr3_tbl;
+ gcr3_info->glx = glx;
+ gcr3_info->domid = device_ser->domain_iommu_ser.attachment_id;
+
+ if (domain->pd_mode != PD_MODE_V2)
+ return 0;
+
+ /* Double-check the hardware is walking the page tables we restored. */
+ pt_iommu_x86_64_hw_info(&domain->amdv2, &pt_info);
+ pte = gcr3_pasid0_entry(gcr3_tbl, gcr3_info->glx);
+ if (!pte || (__sme_clr(*pte) & PAGE_MASK) != (pt_info.gcr3_pt & PAGE_MASK)) {
+ dev_err(dev, "preserved GCR3[0] 0x%llx does not match restored v2 page-table root 0x%llx\n",
+ pte ? *pte : 0, pt_info.gcr3_pt);
+ return -EINVAL;
+ }
+
+ return 0;
+}
+
+/**
+ * amd_iommu_reattach_device - Adopt a device's preserved DTE state on LU boot
+ * @dev: Device being attached
+ * @domain: Domain this kernel rebuilt from the same preserved state
+ * @device_ser: Preserved per-device record handed over by the previous kernel
+ *
+ * Return: 0 on success, or a negative error code if the preserved state does
+ * not describe the domain this kernel restored.
+ */
+int amd_iommu_reattach_device(struct device *dev,
+ struct protection_domain *domain,
+ struct iommu_device_ser *device_ser)
+{
+ enum protection_domain_mode pd_mode;
+
+ pd_mode = ser_to_pd_mode(device_ser->amd.pd_mode);
+ if (pd_mode != domain->pd_mode) {
+ dev_err(dev, "preserved page-table mode %d does not match restored domain mode %d\n",
+ pd_mode, domain->pd_mode);
+ return -EINVAL;
+ }
+
+ if (!device_ser->amd.gcr3_tbl_phys)
+ return reattach_domain_id(dev, domain,
+ device_ser->domain_iommu_ser.attachment_id);
+
+ return reattach_gcr3_table(dev, domain, device_ser);
+}
diff --git a/drivers/iommu/amd/nested.c b/drivers/iommu/amd/nested.c
index 63b53b29e029..4b84b7b4b829 100644
--- a/drivers/iommu/amd/nested.c
+++ b/drivers/iommu/amd/nested.c
@@ -6,6 +6,7 @@
#define dev_fmt(fmt) "AMD-Vi: " fmt
#include <linux/iommu.h>
+#include <linux/iommu-liveupdate.h>
#include <linux/refcount.h>
#include <uapi/linux/iommufd.h>
@@ -248,6 +249,13 @@ static int nested_attach_device(struct iommu_domain *dom, struct device *dev,
if (WARN_ON(dev_data->pasid_enabled))
return -EINVAL;
+ /*
+ * A preserved device is still translating through the DTE this kernel
+ * adopted, and a nested domain is never the domain restored for it.
+ */
+ if (dev_iommu_restored_state(dev))
+ return -EBUSY;
+
mutex_lock(&dev_data->mutex);
set_dte_nested(iommu, dom, dev_data, &new);
--
2.43.0
^ permalink raw reply [flat|nested] 8+ messages in thread
end of thread, other threads:[~2026-10-05 6:44 UTC | newest]
Thread overview: 8+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-05 6:40 [RFC PATCH 0/7] iommu/amd: Implement live update state preservation Ankit Soni
2026-10-05 6:40 ` [RFC PATCH 1/7] iommu/amd: defer device attach only on a kdump boot Ankit Soni
2026-10-05 6:40 ` [RFC PATCH 2/7] liveupdate: parse the incoming handover tree before late_time_init() Ankit Soni
2026-10-05 6:40 ` [RFC PATCH 3/7] iommu/kho/abi: add AMD IOMMU live-update serialisation structs Ankit Soni
2026-10-05 6:40 ` [RFC PATCH 4/7] iommu/amd: preserve IOMMU and device state for live update Ankit Soni
2026-10-05 6:40 ` [RFC PATCH 5/7] iommu/amd: clear unpreserved DTEs and quiesce logs at live-update shutdown Ankit Soni
2026-10-05 6:40 ` [RFC PATCH 6/7] iommu/amd: restore preserved state on a live-update boot Ankit Soni
2026-10-05 6:40 ` [RFC PATCH 7/7] iommu/amd: reattach preserved devices to their restored domains Ankit Soni
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®