* [PATCH v11 01/11] iommu/arm-smmu-v3: Init the vmid_map ida before the stream table setup
2026-10-05 19:30 [PATCH v11 00/11] iommu/arm-smmu-v3: Adopt the crashed kernel's stream table for kdump Nicolin Chen
@ 2026-10-05 19:30 ` Nicolin Chen
2026-10-05 19:30 ` [PATCH v11 02/11] iommu/arm-smmu-v3: Make the ASID space per SMMU instance Nicolin Chen
` (10 subsequent siblings)
11 siblings, 0 replies; 13+ messages in thread
From: Nicolin Chen @ 2026-10-05 19:30 UTC (permalink / raw)
To: will
Cc: robin.murphy, joro, jgg, praan, kevin.tian, smostafa,
linux-arm-kernel, iommu, linux-kernel, jamien, kas,
Cristian Prundeanu, Breno Leitao
The vmid_map ida is initialized after the stream table setup, along with
the devres action that destroys it.
An upcoming change will reserve the crashed kernel's in-use VMIDs in this
ida, from a kdump kernel's stream table adoption that runs in place of the
regular setup and returns early.
Move the ida_init() and its devres registration to the entry of the whole
function, so that the ida is ready by the time the adoption path runs.
Reviewed-by: Jason Gunthorpe <jgg@nvidia.com>
Tested-by: Breno Leitao <leitao@kernel.org>
Tested-by: Cristian Prundeanu <cpru@amazon.com>
Assisted-by: LLM
Signed-off-by: Nicolin Chen <nicolinc@nvidia.com>
---
drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c | 14 +++++++-------
1 file changed, 7 insertions(+), 7 deletions(-)
diff --git a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
index 34e916ea339ff..792627e1c790e 100644
--- a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
+++ b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
@@ -4842,17 +4842,17 @@ static int arm_smmu_init_strtab(struct arm_smmu_device *smmu)
{
int ret;
+ ida_init(&smmu->vmid_map);
+ ret = devm_add_action_or_reset(smmu->dev, arm_smmu_destroy_vmid_map,
+ &smmu->vmid_map);
+ if (ret)
+ return ret;
+
if (smmu->features & ARM_SMMU_FEAT_2_LVL_STRTAB)
ret = arm_smmu_init_strtab_2lvl(smmu);
else
ret = arm_smmu_init_strtab_linear(smmu);
- if (ret)
- return ret;
-
- ida_init(&smmu->vmid_map);
-
- return devm_add_action_or_reset(smmu->dev, arm_smmu_destroy_vmid_map,
- &smmu->vmid_map);
+ return ret;
}
static int arm_smmu_init_structures(struct arm_smmu_device *smmu)
--
2.43.0
^ permalink raw reply [flat|nested] 13+ messages in thread* [PATCH v11 02/11] iommu/arm-smmu-v3: Make the ASID space per SMMU instance
2026-10-05 19:30 [PATCH v11 00/11] iommu/arm-smmu-v3: Adopt the crashed kernel's stream table for kdump Nicolin Chen
2026-10-05 19:30 ` [PATCH v11 01/11] iommu/arm-smmu-v3: Init the vmid_map ida before the stream table setup Nicolin Chen
@ 2026-10-05 19:30 ` Nicolin Chen
2026-10-05 19:30 ` [PATCH v11 03/11] iommu/arm-smmu-v3: Swap EVTQ_MSI_INDEX and GERROR_MSI_INDEX Nicolin Chen
` (9 subsequent siblings)
11 siblings, 0 replies; 13+ messages in thread
From: Nicolin Chen @ 2026-10-05 19:30 UTC (permalink / raw)
To: will
Cc: robin.murphy, joro, jgg, praan, kevin.tian, smostafa,
linux-arm-kernel, iommu, linux-kernel, jamien, kas,
Cristian Prundeanu, Breno Leitao
An ASID tags the TLB entries within one SMMU, so two instances can use the
same ASID without ever aliasing each other. Yet the driver allocates them
out of a single global xarray, which makes the instances share a space that
the hardware keeps apart, and lets one instance exhaust the IDs of another.
That global space is a leftover from the BTM support that shared ASIDs with
the CPU. Now ARM_SMMU_FEAT_BTM is never set, nothing looks a domain up by
its ASID, and a domain is pinned to one SMMU at the attach.
Give each SMMU its own asid_map, mirroring the per-SMMU vmid_map, and clean
it up with devres, so that both of the ID maps get the same lifetime as the
SMMU device structure that holds them. An upcoming change will need this,
to reserve the crashed kernel's in-use ASIDs in the new map during a kdump
kernel's stream table adoption.
This also eases the BTM work that Jason's working on:
https://lore.kernel.org/linux-iommu/20261004162219.GA4064@nvidia.com/
Note that arm_smmu_asid_lock stays global, as it serializes the STE and CD
updates against any ASID change rather than guarding the map itself, which
does its own locking. Giving each SMMU its own lock looks possible now, but
that would touch every attach path and belongs to a separate change.
Reviewed-by: Jason Gunthorpe <jgg@nvidia.com>
Tested-by: Breno Leitao <leitao@kernel.org>
Tested-by: Cristian Prundeanu <cpru@amazon.com>
Assisted-by: LLM
Signed-off-by: Nicolin Chen <nicolinc@nvidia.com>
---
drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.h | 2 +-
.../iommu/arm/arm-smmu-v3/arm-smmu-v3-sva.c | 6 +++---
drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c | 18 +++++++++++++++---
3 files changed, 19 insertions(+), 7 deletions(-)
diff --git a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.h b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.h
index bb048e08d613d..8f2c0b1cc4a8d 100644
--- a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.h
+++ b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.h
@@ -1003,6 +1003,7 @@ struct arm_smmu_device {
#define ARM_SMMU_MAX_VMIDS (1 << 16)
unsigned int vmid_bits;
struct ida vmid_map;
+ struct xarray asid_map;
unsigned int ssid_bits;
unsigned int sid_bits;
@@ -1180,7 +1181,6 @@ to_smmu_nested_domain(struct iommu_domain *dom)
return container_of(dom, struct arm_smmu_nested_domain, domain);
}
-extern struct xarray arm_smmu_asid_xa;
extern struct mutex arm_smmu_asid_lock;
struct arm_smmu_domain *arm_smmu_domain_alloc(void);
diff --git a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3-sva.c b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3-sva.c
index faeec656da9e7..489531e0027bd 100644
--- a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3-sva.c
+++ b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3-sva.c
@@ -311,7 +311,7 @@ static void arm_smmu_sva_domain_free(struct iommu_domain *domain)
* reused, and if there is a race then it just suffers harmless
* unnecessary invalidation.
*/
- xa_erase(&arm_smmu_asid_xa, smmu_domain->cd.asid);
+ xa_erase(&smmu_domain->smmu->asid_map, smmu_domain->cd.asid);
/*
* Actual free is defered to the SRCU callback
@@ -352,7 +352,7 @@ struct iommu_domain *arm_smmu_sva_domain_alloc(struct device *dev,
smmu_domain->stage = ARM_SMMU_DOMAIN_SVA;
smmu_domain->smmu = smmu;
- ret = xa_alloc(&arm_smmu_asid_xa, &asid, smmu_domain,
+ ret = xa_alloc(&smmu->asid_map, &asid, smmu_domain,
XA_LIMIT(1, (1 << smmu->asid_bits) - 1), GFP_KERNEL);
if (ret)
goto err_free;
@@ -366,7 +366,7 @@ struct iommu_domain *arm_smmu_sva_domain_alloc(struct device *dev,
return &smmu_domain->domain;
err_asid:
- xa_erase(&arm_smmu_asid_xa, smmu_domain->cd.asid);
+ xa_erase(&smmu_domain->smmu->asid_map, smmu_domain->cd.asid);
err_free:
arm_smmu_domain_free(smmu_domain);
return ERR_PTR(ret);
diff --git a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
index 792627e1c790e..aeace4335abfe 100644
--- a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
+++ b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
@@ -90,7 +90,6 @@ struct arm_smmu_option_prop {
const char *prop;
};
-DEFINE_XARRAY_ALLOC1(arm_smmu_asid_xa);
DEFINE_MUTEX(arm_smmu_asid_lock);
static struct arm_smmu_option_prop arm_smmu_options[] = {
@@ -3015,7 +3014,7 @@ static void arm_smmu_domain_free_paging(struct iommu_domain *domain)
if (smmu_domain->stage == ARM_SMMU_DOMAIN_S1) {
/* Prevent SVA from touching the CD while we're freeing it */
mutex_lock(&arm_smmu_asid_lock);
- xa_erase(&arm_smmu_asid_xa, smmu_domain->cd.asid);
+ xa_erase(&smmu->asid_map, smmu_domain->cd.asid);
mutex_unlock(&arm_smmu_asid_lock);
} else {
struct arm_smmu_s2_cfg *cfg = &smmu_domain->s2_cfg;
@@ -3035,7 +3034,7 @@ static int arm_smmu_domain_finalise_s1(struct arm_smmu_device *smmu,
/* Prevent SVA from modifying the ASID until it is written to the CD */
mutex_lock(&arm_smmu_asid_lock);
- ret = xa_alloc(&arm_smmu_asid_xa, &asid, smmu_domain,
+ ret = xa_alloc(&smmu->asid_map, &asid, smmu_domain,
XA_LIMIT(1, (1 << smmu->asid_bits) - 1), GFP_KERNEL);
cd->asid = (u16)asid;
mutex_unlock(&arm_smmu_asid_lock);
@@ -4741,6 +4740,13 @@ static void arm_smmu_destroy_vmid_map(void *data)
ida_destroy(ida);
}
+static void arm_smmu_destroy_asid_map(void *data)
+{
+ struct xarray *xa = data;
+
+ xa_destroy(xa);
+}
+
static int arm_smmu_init_queues(struct arm_smmu_device *smmu)
{
int ret;
@@ -4848,6 +4854,12 @@ static int arm_smmu_init_strtab(struct arm_smmu_device *smmu)
if (ret)
return ret;
+ xa_init_flags(&smmu->asid_map, XA_FLAGS_ALLOC1);
+ ret = devm_add_action_or_reset(smmu->dev, arm_smmu_destroy_asid_map,
+ &smmu->asid_map);
+ if (ret)
+ return ret;
+
if (smmu->features & ARM_SMMU_FEAT_2_LVL_STRTAB)
ret = arm_smmu_init_strtab_2lvl(smmu);
else
--
2.43.0
^ permalink raw reply [flat|nested] 13+ messages in thread* [PATCH v11 03/11] iommu/arm-smmu-v3: Swap EVTQ_MSI_INDEX and GERROR_MSI_INDEX
2026-10-05 19:30 [PATCH v11 00/11] iommu/arm-smmu-v3: Adopt the crashed kernel's stream table for kdump Nicolin Chen
2026-10-05 19:30 ` [PATCH v11 01/11] iommu/arm-smmu-v3: Init the vmid_map ida before the stream table setup Nicolin Chen
2026-10-05 19:30 ` [PATCH v11 02/11] iommu/arm-smmu-v3: Make the ASID space per SMMU instance Nicolin Chen
@ 2026-10-05 19:30 ` Nicolin Chen
2026-10-05 19:30 ` [PATCH v11 04/11] iommu/arm-smmu-v3: Add ARM_SMMU_FEAT_EVTQ for the event queue Nicolin Chen
` (8 subsequent siblings)
11 siblings, 0 replies; 13+ messages in thread
From: Nicolin Chen @ 2026-10-05 19:30 UTC (permalink / raw)
To: will
Cc: robin.murphy, joro, jgg, praan, kevin.tian, smostafa,
linux-arm-kernel, iommu, linux-kernel, jamien, kas,
Cristian Prundeanu, Breno Leitao
The nvec trick in arm_smmu_setup_msis() works well on PRIQ since its index
is the last index in enum arm_smmu_msi_index, but it does not work well on
the EVTQ as its index is 0.
A subsequent change will make EVTQ optional so it can be disabled in kdump
mode. This can skip its MSI vector allocation.
The index is driver-internal to select arm_smmu_msi_cfg[] register set and
the corresponding MSI slot. Nothing else.
Swap EVTQ_MSI_INDEX and GERROR_MSI_INDEX. Then, arm_smmu_setup_msis() will
simply decrease nvec when EVTQ is disabled.
No functional change intended.
Signed-off-by: Nicolin Chen <nicolinc@nvidia.com>
---
drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c | 16 ++++++++--------
1 file changed, 8 insertions(+), 8 deletions(-)
diff --git a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
index aeace4335abfe..6ff61565f7583 100644
--- a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
+++ b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
@@ -57,8 +57,8 @@ static struct iommu_dirty_ops arm_smmu_dirty_ops;
static DEFINE_STATIC_KEY_FALSE(arm_smmu_erratum_repeat_tlbi_cfgi_key);
enum arm_smmu_msi_index {
- EVTQ_MSI_INDEX,
GERROR_MSI_INDEX,
+ EVTQ_MSI_INDEX,
PRIQ_MSI_INDEX,
ARM_SMMU_MAX_MSIS,
};
@@ -68,16 +68,16 @@ static_assert(sizeof(struct arm_smmu_ste) == NUM_ENTRY_QWORDS * sizeof(u64));
static_assert(sizeof(struct arm_smmu_cd) == NUM_ENTRY_QWORDS * sizeof(u64));
static phys_addr_t arm_smmu_msi_cfg[ARM_SMMU_MAX_MSIS][3] = {
- [EVTQ_MSI_INDEX] = {
- ARM_SMMU_EVTQ_IRQ_CFG0,
- ARM_SMMU_EVTQ_IRQ_CFG1,
- ARM_SMMU_EVTQ_IRQ_CFG2,
- },
[GERROR_MSI_INDEX] = {
ARM_SMMU_GERROR_IRQ_CFG0,
ARM_SMMU_GERROR_IRQ_CFG1,
ARM_SMMU_GERROR_IRQ_CFG2,
},
+ [EVTQ_MSI_INDEX] = {
+ ARM_SMMU_EVTQ_IRQ_CFG0,
+ ARM_SMMU_EVTQ_IRQ_CFG1,
+ ARM_SMMU_EVTQ_IRQ_CFG2,
+ },
[PRIQ_MSI_INDEX] = {
ARM_SMMU_PRIQ_IRQ_CFG0,
ARM_SMMU_PRIQ_IRQ_CFG1,
@@ -4965,15 +4965,15 @@ static void arm_smmu_setup_msis(struct arm_smmu_device *smmu)
return;
}
- /* Allocate MSIs for evtq, gerror and priq. Ignore cmdq */
+ /* Allocate MSIs for gerror, evtq and priq. Ignore cmdq */
ret = platform_device_msi_init_and_alloc_irqs(dev, nvec, arm_smmu_write_msi_msg);
if (ret) {
dev_warn(dev, "failed to allocate MSIs - falling back to wired irqs\n");
return;
}
- smmu->evtq.q.irq = msi_get_virq(dev, EVTQ_MSI_INDEX);
smmu->gerr_irq = msi_get_virq(dev, GERROR_MSI_INDEX);
+ smmu->evtq.q.irq = msi_get_virq(dev, EVTQ_MSI_INDEX);
smmu->priq.q.irq = msi_get_virq(dev, PRIQ_MSI_INDEX);
/* Add callback to free MSIs on teardown */
--
2.43.0
^ permalink raw reply [flat|nested] 13+ messages in thread* [PATCH v11 04/11] iommu/arm-smmu-v3: Add ARM_SMMU_FEAT_EVTQ for the event queue
2026-10-05 19:30 [PATCH v11 00/11] iommu/arm-smmu-v3: Adopt the crashed kernel's stream table for kdump Nicolin Chen
` (2 preceding siblings ...)
2026-10-05 19:30 ` [PATCH v11 03/11] iommu/arm-smmu-v3: Swap EVTQ_MSI_INDEX and GERROR_MSI_INDEX Nicolin Chen
@ 2026-10-05 19:30 ` Nicolin Chen
2026-10-05 19:30 ` [PATCH v11 05/11] iommu/arm-smmu-v3: Disable the EVTQ and the PRIQ in a kdump kernel Nicolin Chen
` (7 subsequent siblings)
11 siblings, 0 replies; 13+ messages in thread
From: Nicolin Chen @ 2026-10-05 19:30 UTC (permalink / raw)
To: will
Cc: robin.murphy, joro, jgg, praan, kevin.tian, smostafa,
linux-arm-kernel, iommu, linux-kernel, jamien, kas,
Cristian Prundeanu, Breno Leitao
The driver programs and enables the event queue unconditionally, while the
PRI queue has an ARM_SMMU_FEAT_PRI gating each of its touch points. Yet a
kdump kernel wants to leave both of the queues alone, which would take an
is_kdump_kernel() test at every one of those places.
Add an ARM_SMMU_FEAT_EVTQ that the probe always sets, as the event queue is
architecturally mandatory, and gate the queue's allocation, its interrupt,
its MSI vector, and its CR0 and IRQ_CTRL enables on it. A subsequent change
will turn the queue off in a single place.
No functional change intended.
Suggested-by: Jason Gunthorpe <jgg@nvidia.com>
Suggested-by: Robin Murphy <robin.murphy@arm.com>
Acked-by: Kiryl Shutsemau (Meta) <kas@kernel.org>
Reviewed-by: Jason Gunthorpe <jgg@nvidia.com>
Tested-by: Breno Leitao <leitao@kernel.org>
Tested-by: Cristian Prundeanu <cpru@amazon.com>
Assisted-by: LLM
Signed-off-by: Nicolin Chen <nicolinc@nvidia.com>
---
drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.h | 1 +
drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c | 89 ++++++++++++++-------
2 files changed, 61 insertions(+), 29 deletions(-)
diff --git a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.h b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.h
index 8f2c0b1cc4a8d..559bfa8cbc5e0 100644
--- a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.h
+++ b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.h
@@ -972,6 +972,7 @@ struct arm_smmu_device {
#define ARM_SMMU_FEAT_BBML2 (1 << 24)
#define ARM_SMMU_FEAT_HAFT (1 << 25)
#define ARM_SMMU_FEAT_DS (1 << 26)
+#define ARM_SMMU_FEAT_EVTQ (1 << 27)
u32 features;
#define ARM_SMMU_OPT_SKIP_PREFETCH (1 << 0)
diff --git a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
index 6ff61565f7583..756825e004774 100644
--- a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
+++ b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
@@ -2381,7 +2381,8 @@ static irqreturn_t arm_smmu_combined_irq_thread(int irq, void *dev)
{
struct arm_smmu_device *smmu = dev;
- arm_smmu_evtq_thread(irq, dev);
+ if (smmu->features & ARM_SMMU_FEAT_EVTQ)
+ arm_smmu_evtq_thread(irq, dev);
if (smmu->features & ARM_SMMU_FEAT_PRI)
arm_smmu_priq_thread(irq, dev);
@@ -2390,7 +2391,12 @@ static irqreturn_t arm_smmu_combined_irq_thread(int irq, void *dev)
static irqreturn_t arm_smmu_combined_irq_handler(int irq, void *dev)
{
- arm_smmu_gerror_handler(irq, dev);
+ struct arm_smmu_device *smmu = dev;
+ irqreturn_t ret = arm_smmu_gerror_handler(irq, dev);
+
+ /* Without either queue, the thread would have nothing to drain */
+ if (!(smmu->features & (ARM_SMMU_FEAT_EVTQ | ARM_SMMU_FEAT_PRI)))
+ return ret;
return IRQ_WAKE_THREAD;
}
@@ -4763,11 +4769,14 @@ static int arm_smmu_init_queues(struct arm_smmu_device *smmu)
return ret;
/* evtq */
- ret = arm_smmu_init_one_queue(smmu, &smmu->evtq.q, smmu->page1,
- ARM_SMMU_EVTQ_PROD, ARM_SMMU_EVTQ_CONS,
- EVTQ_ENT_DWORDS, "evtq");
- if (ret)
- return ret;
+ if (smmu->features & ARM_SMMU_FEAT_EVTQ) {
+ ret = arm_smmu_init_one_queue(smmu, &smmu->evtq.q, smmu->page1,
+ ARM_SMMU_EVTQ_PROD,
+ ARM_SMMU_EVTQ_CONS,
+ EVTQ_ENT_DWORDS, "evtq");
+ if (ret)
+ return ret;
+ }
if ((smmu->features & ARM_SMMU_FEAT_SVA) &&
(smmu->features & ARM_SMMU_FEAT_STALLS)) {
@@ -4950,7 +4959,15 @@ static void arm_smmu_setup_msis(struct arm_smmu_device *smmu)
/* Clear the MSI address regs */
writeq_relaxed(0, smmu->base + ARM_SMMU_GERROR_IRQ_CFG0);
- writeq_relaxed(0, smmu->base + ARM_SMMU_EVTQ_IRQ_CFG0);
+
+ /* PRIQ's vector needs EVTQ's to be kept */
+ WARN_ON_ONCE(!(smmu->features & ARM_SMMU_FEAT_EVTQ) &&
+ (smmu->features & ARM_SMMU_FEAT_PRI));
+
+ if (smmu->features & ARM_SMMU_FEAT_EVTQ)
+ writeq_relaxed(0, smmu->base + ARM_SMMU_EVTQ_IRQ_CFG0);
+ else
+ nvec--;
if (smmu->features & ARM_SMMU_FEAT_PRI)
writeq_relaxed(0, smmu->base + ARM_SMMU_PRIQ_IRQ_CFG0);
@@ -4987,16 +5004,20 @@ static void arm_smmu_setup_unique_irqs(struct arm_smmu_device *smmu)
arm_smmu_setup_msis(smmu);
/* Request interrupt lines */
- irq = smmu->evtq.q.irq;
- if (irq) {
- ret = devm_request_threaded_irq(smmu->dev, irq, NULL,
- arm_smmu_evtq_thread,
- IRQF_ONESHOT,
- "arm-smmu-v3-evtq", smmu);
- if (ret < 0)
- dev_warn(smmu->dev, "failed to enable evtq irq\n");
- } else {
- dev_warn(smmu->dev, "no evtq irq - events will not be reported!\n");
+ if (smmu->features & ARM_SMMU_FEAT_EVTQ) {
+ irq = smmu->evtq.q.irq;
+ if (irq) {
+ ret = devm_request_threaded_irq(smmu->dev, irq, NULL,
+ arm_smmu_evtq_thread,
+ IRQF_ONESHOT,
+ "arm-smmu-v3-evtq", smmu);
+ if (ret < 0)
+ dev_warn(smmu->dev,
+ "failed to enable evtq irq\n");
+ } else {
+ dev_warn(smmu->dev,
+ "no evtq irq - events will not be reported!\n");
+ }
}
irq = smmu->gerr_irq;
@@ -5029,7 +5050,7 @@ static void arm_smmu_setup_unique_irqs(struct arm_smmu_device *smmu)
static int arm_smmu_setup_irqs(struct arm_smmu_device *smmu)
{
int ret, irq;
- u32 irqen_flags = IRQ_CTRL_EVTQ_IRQEN | IRQ_CTRL_GERROR_IRQEN;
+ u32 irqen_flags = IRQ_CTRL_GERROR_IRQEN;
/* Disable IRQs first */
ret = arm_smmu_write_reg_sync(smmu, 0, ARM_SMMU_IRQ_CTRL,
@@ -5055,6 +5076,8 @@ static int arm_smmu_setup_irqs(struct arm_smmu_device *smmu)
} else
arm_smmu_setup_unique_irqs(smmu);
+ if (smmu->features & ARM_SMMU_FEAT_EVTQ)
+ irqen_flags |= IRQ_CTRL_EVTQ_IRQEN;
if (smmu->features & ARM_SMMU_FEAT_PRI)
irqen_flags |= IRQ_CTRL_PRIQ_IRQEN;
@@ -5173,16 +5196,21 @@ static int arm_smmu_device_reset(struct arm_smmu_device *smmu)
smmu, arm_smmu_make_cmd_op(CMDQ_OP_TLBI_NSNH_ALL));
/* Event queue */
- writeq_relaxed(smmu->evtq.q.q_base, smmu->base + ARM_SMMU_EVTQ_BASE);
- writel_relaxed(smmu->evtq.q.llq.prod, smmu->page1 + ARM_SMMU_EVTQ_PROD);
- writel_relaxed(smmu->evtq.q.llq.cons, smmu->page1 + ARM_SMMU_EVTQ_CONS);
-
- enables |= CR0_EVTQEN;
- ret = arm_smmu_write_reg_sync(smmu, enables, ARM_SMMU_CR0,
- ARM_SMMU_CR0ACK);
- if (ret) {
- dev_err(smmu->dev, "failed to enable event queue\n");
- return ret;
+ if (smmu->features & ARM_SMMU_FEAT_EVTQ) {
+ writeq_relaxed(smmu->evtq.q.q_base,
+ smmu->base + ARM_SMMU_EVTQ_BASE);
+ writel_relaxed(smmu->evtq.q.llq.prod,
+ smmu->page1 + ARM_SMMU_EVTQ_PROD);
+ writel_relaxed(smmu->evtq.q.llq.cons,
+ smmu->page1 + ARM_SMMU_EVTQ_CONS);
+
+ enables |= CR0_EVTQEN;
+ ret = arm_smmu_write_reg_sync(smmu, enables, ARM_SMMU_CR0,
+ ARM_SMMU_CR0ACK);
+ if (ret) {
+ dev_err(smmu->dev, "failed to enable event queue\n");
+ return ret;
+ }
}
/* PRI queue */
@@ -5330,6 +5358,9 @@ static int arm_smmu_device_hw_probe(struct arm_smmu_device *smmu)
u32 reg;
bool coherent = smmu->features & ARM_SMMU_FEAT_COHERENCY;
+ /* The event queue is architecturally mandatory, unlike the PRI queue */
+ smmu->features |= ARM_SMMU_FEAT_EVTQ;
+
/* IDR0 */
reg = readl_relaxed(smmu->base + ARM_SMMU_IDR0);
--
2.43.0
^ permalink raw reply [flat|nested] 13+ messages in thread* [PATCH v11 05/11] iommu/arm-smmu-v3: Disable the EVTQ and the PRIQ in a kdump kernel
2026-10-05 19:30 [PATCH v11 00/11] iommu/arm-smmu-v3: Adopt the crashed kernel's stream table for kdump Nicolin Chen
` (3 preceding siblings ...)
2026-10-05 19:30 ` [PATCH v11 04/11] iommu/arm-smmu-v3: Add ARM_SMMU_FEAT_EVTQ for the event queue Nicolin Chen
@ 2026-10-05 19:30 ` Nicolin Chen
2026-10-05 19:30 ` [PATCH v11 06/11] iommu/arm-smmu-v3: Add ARM_SMMU_OPT_KDUMP_ADOPT for " Nicolin Chen
` (6 subsequent siblings)
11 siblings, 0 replies; 13+ messages in thread
From: Nicolin Chen @ 2026-10-05 19:30 UTC (permalink / raw)
To: will
Cc: robin.murphy, joro, jgg, praan, kevin.tian, smostafa,
linux-arm-kernel, iommu, linux-kernel, jamien, kas,
Cristian Prundeanu, Breno Leitao
A kdump kernel cannot use either queue. The crashed kernel's CDs and page
tables might be corrupted, so events would spam the EVTQ, and there is no
way to serve the page requests that would arrive at the PRIQ.
The reset routine still enables both of the queues and then masks the two
enable bits back out, having already programmed the queue bases and taken
the interrupts of both.
Clear ARM_SMMU_FEAT_EVTQ and ARM_SMMU_FEAT_PRI in the probe instead, so as
to entirely skip the queue initialization and handling in a kdump kernel.
Suggested-by: Will Deacon <will@kernel.org>
Link: https://lore.kernel.org/all/amiBagGKn-Aym1DK@willie-the-truck/
Acked-by: Kiryl Shutsemau (Meta) <kas@kernel.org>
Reviewed-by: Jason Gunthorpe <jgg@nvidia.com>
Tested-by: Breno Leitao <leitao@kernel.org>
Tested-by: Cristian Prundeanu <cpru@amazon.com>
Assisted-by: LLM
Signed-off-by: Nicolin Chen <nicolinc@nvidia.com>
---
drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c | 12 +++++++++---
1 file changed, 9 insertions(+), 3 deletions(-)
diff --git a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
index 756825e004774..42cd30c79211e 100644
--- a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
+++ b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
@@ -5247,9 +5247,6 @@ static int arm_smmu_device_reset(struct arm_smmu_device *smmu)
return ret;
}
- if (is_kdump_kernel())
- enables &= ~(CR0_EVTQEN | CR0_PRIQEN);
-
/* Enable the SMMU interface */
enables |= CR0_SMMUEN;
ret = arm_smmu_write_reg_sync(smmu, enables, ARM_SMMU_CR0,
@@ -5576,6 +5573,15 @@ static int arm_smmu_device_hw_probe(struct arm_smmu_device *smmu)
if (arm_smmu_sva_supported(smmu))
smmu->features |= ARM_SMMU_FEAT_SVA;
+ /*
+ * A kdump kernel wants neither queue: the crashed kernel's CDs and page
+ * tables might be corrupted, spamming events, and page requests cannot
+ * be served. A disabled queue discards new records without raising any
+ * global error.
+ */
+ if (is_kdump_kernel())
+ smmu->features &= ~(ARM_SMMU_FEAT_EVTQ | ARM_SMMU_FEAT_PRI);
+
dev_info(smmu->dev, "oas %lu-bit (features 0x%08x)\n",
smmu->oas, smmu->features);
return 0;
--
2.43.0
^ permalink raw reply [flat|nested] 13+ messages in thread* [PATCH v11 06/11] iommu/arm-smmu-v3: Add ARM_SMMU_OPT_KDUMP_ADOPT for kdump kernel
2026-10-05 19:30 [PATCH v11 00/11] iommu/arm-smmu-v3: Adopt the crashed kernel's stream table for kdump Nicolin Chen
` (4 preceding siblings ...)
2026-10-05 19:30 ` [PATCH v11 05/11] iommu/arm-smmu-v3: Disable the EVTQ and the PRIQ in a kdump kernel Nicolin Chen
@ 2026-10-05 19:30 ` Nicolin Chen
2026-10-05 19:30 ` [PATCH v11 07/11] iommu/arm-smmu-v3-kexec: Reserve crashed kernel's ASIDs and VMIDs Nicolin Chen
` (5 subsequent siblings)
11 siblings, 0 replies; 13+ messages in thread
From: Nicolin Chen @ 2026-10-05 19:30 UTC (permalink / raw)
To: will
Cc: robin.murphy, joro, jgg, praan, kevin.tian, smostafa,
linux-arm-kernel, iommu, linux-kernel, jamien, kas,
Cristian Prundeanu, Breno Leitao
When transitioning to a kdump kernel, the primary kernel might have crashed
while endpoint devices were actively bus-mastering DMA. Currently, the SMMU
driver aggressively resets the hardware during probe by clearing CR0_SMMUEN
and setting the Global Bypass Attribute (GBPA) to ABORT.
In a kdump scenario, this aggressive reset is highly destructive:
a) If GBPA is set to ABORT, in-flight DMA will be aborted, generating fatal
PCIe AER or SErrors that may panic the kdump kernel
b) If GBPA is set to BYPASS, in-flight DMA targeting some IOVAs will bypass
the SMMU and corrupt the physical memory at those 1:1 mapped IOVAs.
To safely absorb in-flight DMAs, a kdump kernel will have to leave SMMUEN=1
intact and avoid modifying STRTAB_BASE, allowing HW to continue translating
in-flight DMAs reusing the crashed kernel's page tables until the endpoint
device drivers probe and quiesce their respective hardware.
However, the ARM SMMUv3 architecture specification states that updating the
SMMU_STRTAB_BASE register while SMMUEN == 1 is UNPREDICTABLE or ignored.
This leaves a kdump kernel no choice but to adopt the stream table from the
crashed kernel.
Introduce ARM_SMMU_OPT_KDUMP_ADOPT and adoption functions that memremap the
stream tables extracted from STRTAB_BASE and STRTAB_BASE_CFG.
Note that the adoption of the crashed kernel's stream table follows certain
strict rules, since the old stream table might be compromised. Thus, apply
some basic validations against the values read from the registers. If tests
fail, it means the stream table cannot be trusted, so toss it entirely. To
avoid OOM due to a potentially corrupted stream table, the memremap for l2
tables is done lazily on the kdump kernel's demand.
The new option will be set in a following change, once the device reset and
the RMR setup are reworked not to overwrite the adopted stream table, and
the crashed kernel's in-use ASIDs and VMIDs are reserved.
Keep all kdump functions and kexec helpers in a new arm-smmu-v3-kexec file,
which will be shared with the liveupdate series in the future.
Suggested-by: Jason Gunthorpe <jgg@nvidia.com>
Suggested-by: Pranjal Shrivastava <praan@google.com>
Reviewed-by: Jason Gunthorpe <jgg@nvidia.com>
Tested-by: Breno Leitao <leitao@kernel.org>
Tested-by: Cristian Prundeanu <cpru@amazon.com>
Assisted-by: LLM
Signed-off-by: Nicolin Chen <nicolinc@nvidia.com>
---
drivers/iommu/arm/Kconfig | 3 +
drivers/iommu/arm/arm-smmu-v3/Makefile | 1 +
drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.h | 21 +
.../iommu/arm/arm-smmu-v3/arm-smmu-v3-kexec.c | 398 ++++++++++++++++++
drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c | 21 +-
5 files changed, 441 insertions(+), 3 deletions(-)
create mode 100644 drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3-kexec.c
diff --git a/drivers/iommu/arm/Kconfig b/drivers/iommu/arm/Kconfig
index 5fac08b89deea..9a7bebe16f395 100644
--- a/drivers/iommu/arm/Kconfig
+++ b/drivers/iommu/arm/Kconfig
@@ -109,6 +109,9 @@ config ARM_SMMU_V3_IOMMUFD
Say Y here if you are doing development and testing on this feature.
+config ARM_SMMU_V3_KEXEC
+ def_bool CRASH_DUMP
+
config ARM_SMMU_V3_KUNIT_TEST
tristate "KUnit tests for arm-smmu-v3 driver" if !KUNIT_ALL_TESTS
depends on KUNIT
diff --git a/drivers/iommu/arm/arm-smmu-v3/Makefile b/drivers/iommu/arm/arm-smmu-v3/Makefile
index 493a659cc66bb..89a27de8dd2f0 100644
--- a/drivers/iommu/arm/arm-smmu-v3/Makefile
+++ b/drivers/iommu/arm/arm-smmu-v3/Makefile
@@ -3,6 +3,7 @@ obj-$(CONFIG_ARM_SMMU_V3) += arm_smmu_v3.o
arm_smmu_v3-y := arm-smmu-v3.o
arm_smmu_v3-$(CONFIG_ARM_SMMU_V3_IOMMUFD) += arm-smmu-v3-iommufd.o
arm_smmu_v3-$(CONFIG_ARM_SMMU_V3_SVA) += arm-smmu-v3-sva.o
+arm_smmu_v3-$(CONFIG_ARM_SMMU_V3_KEXEC) += arm-smmu-v3-kexec.o
arm_smmu_v3-$(CONFIG_TEGRA241_CMDQV) += tegra241-cmdqv.o
obj-$(CONFIG_ARM_SMMU_V3_KUNIT_TEST) += arm-smmu-v3-test.o
diff --git a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.h b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.h
index 559bfa8cbc5e0..a1f947e7d9fe4 100644
--- a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.h
+++ b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.h
@@ -986,6 +986,7 @@ struct arm_smmu_device {
* invalidated CONT
*/
#define ARM_SMMU_OPT_FULL_CONT_RANGE_INV (1 << 6)
+#define ARM_SMMU_OPT_KDUMP_ADOPT (1 << 7)
u32 options;
struct arm_smmu_cmdq cmdq;
@@ -1312,6 +1313,26 @@ tegra241_cmdqv_probe(struct arm_smmu_device *smmu)
}
#endif /* CONFIG_TEGRA241_CMDQV */
+#ifdef CONFIG_CRASH_DUMP
+int arm_smmu_kdump_adopt_strtab(struct arm_smmu_device *smmu);
+int arm_smmu_kdump_adopt_deferred_l2_strtab(struct arm_smmu_device *smmu,
+ u32 sid, phys_addr_t base, u32 span,
+ struct arm_smmu_strtab_l2 **l2table);
+#else /* CONFIG_CRASH_DUMP */
+static inline int arm_smmu_kdump_adopt_strtab(struct arm_smmu_device *smmu)
+{
+ return -EOPNOTSUPP;
+}
+
+static inline int
+arm_smmu_kdump_adopt_deferred_l2_strtab(struct arm_smmu_device *smmu, u32 sid,
+ phys_addr_t base, u32 span,
+ struct arm_smmu_strtab_l2 **l2table)
+{
+ return -EOPNOTSUPP;
+}
+#endif /* CONFIG_CRASH_DUMP */
+
struct arm_vsmmu {
struct iommufd_viommu core;
struct arm_smmu_device *smmu;
diff --git a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3-kexec.c b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3-kexec.c
new file mode 100644
index 0000000000000..6bbc369df86bb
--- /dev/null
+++ b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3-kexec.c
@@ -0,0 +1,398 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2026, NVIDIA CORPORATION & AFFILIATES
+ */
+
+#define dev_fmt(fmt) "kexec: " fmt
+
+#include <linux/io.h>
+#include <linux/slab.h>
+
+#include "arm-smmu-v3.h"
+
+/*
+ * Common helpers for a kexec'd kernel to parse, validate, and walk through the
+ * previous kernel's SMMU table structures, shared by the kdump adoption and a
+ * future live-update restoration.
+ *
+ * The common helpers are read-only against the previous kernel's structures: a
+ * table that is not yet mapped by this kernel gets a transient memremap during
+ * a walk, followed by an immediate memunmap. They never allocate memory or take
+ * ownership of the previous kernel's tables; the callers make those decisions.
+ */
+
+/**
+ * arm_smmu_kexec_parse_strtab_2lvl() - Validate a 2-level stream table
+ * @smmu: SMMU device of this kernel
+ * @cfg_reg: STRTAB_BASE_CFG register value set by the previous kernel
+ * @base: stream table base address extracted from the STRTAB_BASE register
+ * @num_l1_ents: pointer to return the number of L1 entries
+ *
+ * Validate the 2-level stream table geometry in @cfg_reg and @base's alignment
+ * against this kernel's hardware limits.
+ *
+ * Return: 0 on success with @num_l1_ents set, or -EINVAL on a bad geometry
+ */
+static int arm_smmu_kexec_parse_strtab_2lvl(struct arm_smmu_device *smmu,
+ u32 cfg_reg, phys_addr_t base,
+ u32 *num_l1_ents)
+{
+ u32 log2size = FIELD_GET(STRTAB_BASE_CFG_LOG2SIZE, cfg_reg);
+ u32 split = FIELD_GET(STRTAB_BASE_CFG_SPLIT, cfg_reg);
+ u32 num_ents;
+ size_t size;
+
+ if (log2size < split || log2size > smmu->sid_bits) {
+ dev_err(smmu->dev, "log2size %u out of range [%u, %u]\n",
+ log2size, split, smmu->sid_bits);
+ return -EINVAL;
+ }
+ if (split != STRTAB_SPLIT) {
+ dev_err(smmu->dev,
+ "unsupported STRTAB_SPLIT %u (expected %u)\n", split,
+ STRTAB_SPLIT);
+ return -EINVAL;
+ }
+
+ /*
+ * Bound the entry count before the shift, as a log2size wider than what
+ * this kernel itself supports would overflow it.
+ */
+ if (log2size - split > ilog2(STRTAB_MAX_L1_ENTRIES)) {
+ dev_err(smmu->dev, "l1 entries 2^%u exceeds max %u\n",
+ log2size - split, STRTAB_MAX_L1_ENTRIES);
+ return -EINVAL;
+ }
+
+ num_ents = 1U << (log2size - split);
+
+ size = num_ents * sizeof(struct arm_smmu_strtab_l1);
+ /*
+ * According to spec (6.3.24), HW aligns the base down to the L1 table
+ * size, i.e. min 64 bytes, so an unaligned base would make this kernel
+ * read another table.
+ */
+ if (!IS_ALIGNED(base, size)) {
+ dev_err(smmu->dev, "unaligned l1 stream table base %pa\n",
+ &base);
+ return -EINVAL;
+ }
+
+ *num_l1_ents = num_ents;
+ return 0;
+}
+
+/**
+ * arm_smmu_kexec_parse_strtab_linear() - Validate a linear stream table
+ * @smmu: SMMU device of this kernel
+ * @cfg_reg: STRTAB_BASE_CFG register value set by the previous kernel
+ * @base: stream table base address extracted from the STRTAB_BASE register
+ * @num_ents: pointer to return the number of STEs
+ *
+ * Validate the linear stream table geometry in @cfg_reg and @base's alignment
+ * against this kernel's own limits.
+ *
+ * Return: 0 on success with @num_ents set, or -EINVAL on a bad geometry
+ */
+static int arm_smmu_kexec_parse_strtab_linear(struct arm_smmu_device *smmu,
+ u32 cfg_reg, phys_addr_t base,
+ u32 *num_ents)
+{
+ u32 log2size = FIELD_GET(STRTAB_BASE_CFG_LOG2SIZE, cfg_reg);
+ unsigned int max_log2size = smmu->sid_bits;
+ size_t size;
+
+ /* Cap the size at what this kernel itself would have allocated */
+ if (smmu->features & ARM_SMMU_FEAT_2_LVL_STRTAB)
+ max_log2size = min_t(
+ unsigned int, max_log2size,
+ ilog2(STRTAB_MAX_L1_ENTRIES * STRTAB_NUM_L2_STES));
+
+ /* num_ents is limited to a u32, so cap log2size at 31 */
+ max_log2size = min(max_log2size, 31U);
+ if (log2size > max_log2size) {
+ dev_err(smmu->dev, "unsupported log2size %u (> %u)\n", log2size,
+ max_log2size);
+ return -EINVAL;
+ }
+
+ size = (1U << log2size) * sizeof(struct arm_smmu_ste);
+ /*
+ * According to spec (6.3.24), HW aligns the base down to the table size
+ * and ignores the low bits, so an unaligned base would make this kernel
+ * read a different table.
+ */
+ if (!IS_ALIGNED(base, size)) {
+ dev_err(smmu->dev, "unaligned stream table base %pa\n", &base);
+ return -EINVAL;
+ }
+
+ *num_ents = 1U << log2size;
+ return 0;
+}
+
+/**
+ * arm_smmu_kexec_check_strtab_l1_desc() - Check one stream table L1 descriptor
+ * @smmu: SMMU device of this kernel
+ * @l1_desc: L1 descriptor value from the previous kernel's stream table
+ * @idx: index of the L1 descriptor, for diagnostics
+ * @l2_base: pointer to return the L2 table's physical address
+ *
+ * Return: 1 if the descriptor is unused, 0 if it is valid with @l2_base set, or
+ * -EINVAL if it is malformed
+ */
+static int arm_smmu_kexec_check_strtab_l1_desc(struct arm_smmu_device *smmu,
+ u64 l1_desc, u32 idx,
+ phys_addr_t *l2_base)
+{
+ phys_addr_t base = l1_desc & STRTAB_L1_DESC_L2PTR_MASK;
+ u32 span = FIELD_GET(STRTAB_L1_DESC_SPAN, l1_desc);
+
+ /* L1STD.L2Ptr is invalid */
+ if (!span)
+ return 1;
+
+ if (span != STRTAB_SPLIT + 1) {
+ dev_err(smmu->dev, "L1[%u] unsupported span %u (vs %u)\n", idx,
+ span, STRTAB_SPLIT + 1);
+ return -EINVAL;
+ }
+
+ /*
+ * A valid descriptor never carries a null pointer. Also, HW aligns the
+ * pointer down to the L2 table size, so an unaligned pointer would make
+ * this kernel read a different table.
+ */
+ if (!base || !IS_ALIGNED(base, sizeof(struct arm_smmu_strtab_l2))) {
+ dev_err(smmu->dev, "L1[%u] bad l2 table base %pa\n", idx,
+ &base);
+ return -EINVAL;
+ }
+
+ *l2_base = base;
+ return 0;
+}
+
+#ifdef CONFIG_CRASH_DUMP
+/*
+ * Helper functions of the kdump stream table adoption for ARM SMMUv3
+ *
+ * When the crashed kernel left the SMMU enabled with in-flight DMAs, the kdump
+ * kernel adopts the crashed kernel's stream tables, instead of doing a regular
+ * reset, to keep in-flight DMAs translating until the endpoint device drivers
+ * re-probe and quiesce their devices.
+ *
+ * Note:
+ * - Adoption only starts on an SMMU that the crashed kernel left enabled, as a
+ * disabled SMMU (CR0_SMMUEN=0) could hold meaningless register values.
+ * - Values read from the crashed kernel's registers get structural validation
+ * only (format, size, span, alignment, and ID range); the physical addresses
+ * are not vetted, as the kdump kernel has no record of which pages held the
+ * tables.
+ * - A structural inconsistency at adoption time tosses the entire adoption and
+ * makes the SMMU fall back to a full reset blocking in-flight DMAs.
+ * - L2 stream tables are adopted lazily at master-inserting time, to bound the
+ * peak memory use against a corrupted L1 table; any lazy L2 adoption failure
+ * rejects that device alone, as its blast radius is bounded to the bus.
+ * - Only a coherent SMMU (ARM_SMMU_FEAT_COHERENCY) is supported, as the stream
+ * table adoption is done by memremap with MEMREMAP_WB, which is verified on
+ * the real hardware. Callers of these functions are responsible for gating
+ * ARM_SMMU_FEAT_COHERENCY once during the probe.
+ */
+
+int arm_smmu_kdump_adopt_deferred_l2_strtab(struct arm_smmu_device *smmu,
+ u32 sid, phys_addr_t base, u32 span,
+ struct arm_smmu_strtab_l2 **l2table)
+{
+ struct arm_smmu_strtab_l2 *table;
+ size_t size;
+
+ /*
+ * Retest the span in case the L1 descriptor has been overwritten since
+ * the adopt. Reject this master's insert; panic or SMMU-disable would
+ * either lose the vmcore or cascade aborts. Do not try to fix it, as it
+ * would break all other SIDs in the same bus (PCI case). The corruption
+ * blast radius is already bounded to that bus range.
+ */
+ if (span != STRTAB_SPLIT + 1) {
+ dev_err(smmu->dev,
+ "L1[%u] span %u changed since adopt (was %u)\n",
+ arm_smmu_strtab_l1_idx(sid), span, STRTAB_SPLIT + 1);
+ return -EINVAL;
+ }
+
+ size = (1UL << (span - 1)) * sizeof(struct arm_smmu_ste);
+
+ /* Same live-corruption check as the span; reject an overwritten base */
+ if (!base || !IS_ALIGNED(base, size)) {
+ dev_err(smmu->dev, "L1[%u] bad l2 table base %pa\n",
+ arm_smmu_strtab_l1_idx(sid), &base);
+ return -EINVAL;
+ }
+
+ /*
+ * This L2 table is mapped lazily per master; devres frees it at unbind,
+ * as with the dmam_alloc_coherent() used for a fresh L2.
+ */
+ table = devm_memremap(smmu->dev, base, size, MEMREMAP_WB);
+ if (IS_ERR(table)) {
+ dev_err(smmu->dev,
+ "failed to adopt l2 stream table for SID %u\n", sid);
+ return PTR_ERR(table);
+ }
+
+ *l2table = table;
+ return 0;
+}
+
+static int arm_smmu_kdump_adopt_strtab_2lvl(struct arm_smmu_device *smmu,
+ u32 cfg_reg, phys_addr_t base)
+{
+ struct arm_smmu_strtab_cfg *cfg = &smmu->strtab_cfg;
+ u32 num_l1_ents;
+ size_t size;
+ int ret, i;
+
+ ret = arm_smmu_kexec_parse_strtab_2lvl(smmu, cfg_reg, base,
+ &num_l1_ents);
+ if (ret)
+ return ret;
+
+ cfg->l2.num_l1_ents = num_l1_ents;
+
+ size = num_l1_ents * sizeof(struct arm_smmu_strtab_l1);
+ cfg->l2.l1tab = memremap(base, size, MEMREMAP_WB);
+ if (!cfg->l2.l1tab)
+ return -ENOMEM;
+
+ cfg->l2.l2ptrs =
+ kcalloc(num_l1_ents, sizeof(*cfg->l2.l2ptrs), GFP_KERNEL);
+ if (!cfg->l2.l2ptrs)
+ return -ENOMEM;
+
+ for (i = 0; i < num_l1_ents; i++) {
+ u64 l2ptr = le64_to_cpu(cfg->l2.l1tab[i].l2ptr);
+ phys_addr_t l2_base;
+
+ ret = arm_smmu_kexec_check_strtab_l1_desc(smmu, l2ptr, i,
+ &l2_base);
+ if (ret < 0)
+ return ret;
+
+ /*
+ * If the crashed kernel's l1 descriptors are deeply corrupted,
+ * blindly memremapping every l2 table here could lead to OOM.
+ *
+ * Defer the l2 memremap to arm_smmu_init_l2_strtab(), so peak
+ * memory is bounded by the kdump kernel's actual demand.
+ */
+ }
+
+ return 0;
+}
+
+static int arm_smmu_kdump_adopt_strtab_linear(struct arm_smmu_device *smmu,
+ u32 cfg_reg, phys_addr_t base)
+{
+ struct arm_smmu_strtab_cfg *cfg = &smmu->strtab_cfg;
+ u32 num_ents;
+ size_t size;
+ int ret;
+
+ ret = arm_smmu_kexec_parse_strtab_linear(smmu, cfg_reg, base,
+ &num_ents);
+ if (ret)
+ return ret;
+
+ /*
+ * We might end up with a num_ents != sid_bits, which is fine, since the
+ * ARM_SMMU_OPT_KDUMP_ADOPT case bypasses arm_smmu_write_strtab().
+ */
+ cfg->linear.num_ents = num_ents;
+
+ size = num_ents * sizeof(struct arm_smmu_ste);
+ cfg->linear.table = memremap(base, size, MEMREMAP_WB);
+ if (!cfg->linear.table)
+ return -ENOMEM;
+ return 0;
+}
+
+static void arm_smmu_kdump_adopt_cleanup(void *data)
+{
+ struct arm_smmu_device *smmu = data;
+ struct arm_smmu_strtab_cfg *cfg = &smmu->strtab_cfg;
+
+ if (smmu->features & ARM_SMMU_FEAT_2_LVL_STRTAB) {
+ kfree(cfg->l2.l2ptrs);
+ if (cfg->l2.l1tab)
+ memunmap(cfg->l2.l1tab);
+ } else {
+ if (cfg->linear.table)
+ memunmap(cfg->linear.table);
+ }
+}
+
+int arm_smmu_kdump_adopt_strtab(struct arm_smmu_device *smmu)
+{
+ u32 cfg_reg = readl_relaxed(smmu->base + ARM_SMMU_STRTAB_BASE_CFG);
+ u64 base_reg = readq_relaxed(smmu->base + ARM_SMMU_STRTAB_BASE);
+ bool was_2lvl = smmu->features & ARM_SMMU_FEAT_2_LVL_STRTAB;
+ phys_addr_t base = base_reg & STRTAB_BASE_ADDR_MASK;
+ u32 fmt = FIELD_GET(STRTAB_BASE_CFG_FMT, cfg_reg);
+ int ret;
+
+ dev_dbg(smmu->dev, "adopting crashed kernel's stream table\n");
+
+ if (fmt == STRTAB_BASE_CFG_FMT_2LVL) {
+ /*
+ * Both kernels run on the same hardware, so it's impossible for
+ * kdump kernel to see the support for linear stream table only.
+ */
+ if (WARN_ON(!(smmu->features & ARM_SMMU_FEAT_2_LVL_STRTAB)))
+ ret = -EINVAL;
+ else
+ ret = arm_smmu_kdump_adopt_strtab_2lvl(smmu, cfg_reg,
+ base);
+ } else if (fmt == STRTAB_BASE_CFG_FMT_LINEAR) {
+ /*
+ * The kdump kernel need not match the crashed kernel. An older
+ * crashed kernel that predates two-level stream table support
+ * may have used a linear table on 2-level-capable hardware, so
+ * enforce the same format here to match the adopted table.
+ */
+ ret = arm_smmu_kdump_adopt_strtab_linear(smmu, cfg_reg, base);
+ if (!ret)
+ smmu->features &= ~ARM_SMMU_FEAT_2_LVL_STRTAB;
+ } else {
+ dev_err(smmu->dev, "invalid STRTAB format %u\n", fmt);
+ ret = -EINVAL;
+ }
+
+ if (ret) {
+ arm_smmu_kdump_adopt_cleanup(smmu);
+ goto err;
+ }
+
+ ret = devm_add_action_or_reset(smmu->dev, arm_smmu_kdump_adopt_cleanup,
+ smmu);
+ /* devm_add_action_or_reset ran the cleanup upon failure */
+ if (ret) {
+ dev_warn(smmu->dev, "failed to set up cleanup action\n");
+ goto err;
+ }
+
+ return 0;
+
+err:
+ dev_warn(smmu->dev, "falling back to full reset\n");
+ /*
+ * Undo the linear adoption's clearing of FEAT_2_LVL_STRTAB so that the
+ * full-reset fallback uses the hardware-supported format.
+ */
+ if (was_2lvl)
+ smmu->features |= ARM_SMMU_FEAT_2_LVL_STRTAB;
+ memset(&smmu->strtab_cfg, 0, sizeof(smmu->strtab_cfg));
+ smmu->options &= ~ARM_SMMU_OPT_KDUMP_ADOPT;
+ return ret;
+}
+#endif /* CONFIG_CRASH_DUMP */
diff --git a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
index 42cd30c79211e..db80b980d1dd9 100644
--- a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
+++ b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
@@ -2034,11 +2034,23 @@ static int arm_smmu_init_l2_strtab(struct arm_smmu_device *smmu, u32 sid)
dma_addr_t l2ptr_dma;
struct arm_smmu_strtab_cfg *cfg = &smmu->strtab_cfg;
struct arm_smmu_strtab_l2 **l2table;
+ u32 l1_idx = arm_smmu_strtab_l1_idx(sid);
- l2table = &cfg->l2.l2ptrs[arm_smmu_strtab_l1_idx(sid)];
+ l2table = &cfg->l2.l2ptrs[l1_idx];
if (*l2table)
return 0;
+ /* Deferred adoption of the crashed kernel's L2 table */
+ if (smmu->options & ARM_SMMU_OPT_KDUMP_ADOPT) {
+ u64 l2ptr = le64_to_cpu(cfg->l2.l1tab[l1_idx].l2ptr);
+ phys_addr_t base = l2ptr & STRTAB_L1_DESC_L2PTR_MASK;
+ u32 span = FIELD_GET(STRTAB_L1_DESC_SPAN, l2ptr);
+
+ if (span)
+ return arm_smmu_kdump_adopt_deferred_l2_strtab(
+ smmu, sid, base, span, l2table);
+ }
+
*l2table = dmam_alloc_coherent(smmu->dev, sizeof(**l2table),
&l2ptr_dma, GFP_KERNEL);
if (!*l2table) {
@@ -2050,8 +2062,7 @@ static int arm_smmu_init_l2_strtab(struct arm_smmu_device *smmu, u32 sid)
arm_smmu_init_initial_stes((*l2table)->stes,
ARRAY_SIZE((*l2table)->stes));
- arm_smmu_write_strtab_l1_desc(&cfg->l2.l1tab[arm_smmu_strtab_l1_idx(sid)],
- l2ptr_dma);
+ arm_smmu_write_strtab_l1_desc(&cfg->l2.l1tab[l1_idx], l2ptr_dma);
return 0;
}
@@ -4869,6 +4880,10 @@ static int arm_smmu_init_strtab(struct arm_smmu_device *smmu)
if (ret)
return ret;
+ if ((smmu->options & ARM_SMMU_OPT_KDUMP_ADOPT) &&
+ !arm_smmu_kdump_adopt_strtab(smmu))
+ return 0;
+
if (smmu->features & ARM_SMMU_FEAT_2_LVL_STRTAB)
ret = arm_smmu_init_strtab_2lvl(smmu);
else
--
2.43.0
^ permalink raw reply [flat|nested] 13+ messages in thread* [PATCH v11 07/11] iommu/arm-smmu-v3-kexec: Reserve crashed kernel's ASIDs and VMIDs
2026-10-05 19:30 [PATCH v11 00/11] iommu/arm-smmu-v3: Adopt the crashed kernel's stream table for kdump Nicolin Chen
` (5 preceding siblings ...)
2026-10-05 19:30 ` [PATCH v11 06/11] iommu/arm-smmu-v3: Add ARM_SMMU_OPT_KDUMP_ADOPT for " Nicolin Chen
@ 2026-10-05 19:30 ` Nicolin Chen
2026-10-05 19:30 ` [PATCH v11 08/11] iommu/arm-smmu-v3-kexec: Implement is_attach_deferred() Nicolin Chen
` (4 subsequent siblings)
11 siblings, 0 replies; 13+ messages in thread
From: Nicolin Chen @ 2026-10-05 19:30 UTC (permalink / raw)
To: will
Cc: robin.murphy, joro, jgg, praan, kevin.tian, smostafa,
linux-arm-kernel, iommu, linux-kernel, jamien, kas,
Cristian Prundeanu, Breno Leitao
The adopted stream table keeps translating in-flight DMA, so the SMMU keeps
caching TLB entries tagged with the crashed kernel's ASIDs and VMIDs. If
this kernel handed one of those IDs to its own domain, the new domain's DMA
could hit the crashed kernel's cached translations, e.g. a stale entry left
behind by an invalidation that the crash cut short.
Scan the adopted stream table at adoption time, reserving every ID in use
via arm_smmu_kexec_scan_and_resv_ids(), and roll all of them back through
arm_smmu_kexec_unresv_ids() should the scan fail. These two kexec helpers
will be shared with the liveupdate code.
The scan memremaps each table transiently, since the IDs must be reserved
within the SMMU probe, long before any master re-probes to claim its table.
The scan walks untrusted tables, yet every loop is strictly index-bounded:
the iteration counts derive from the log2size and s1cdmax fields, which are
validated against this kernel's own sid_bits and ssid_bits, so a corrupted
table cannot extend the walk.
A nested STE's guest-owned CD table is left alone, since its ASIDs live in
a space of their own under that STE's VMID. Reserving that VMID covers all
of them, so there is no reason for the scan to walk the (VMID, ASID) pairs
behind it. The only ASIDs needing a reservation are those in the space that
this kernel uses for its own domains.
Note that, on an E2H/VHE host, the kernel's stage-1 domains are tagged by
the EL2 ASID, and the TLBI_EL2_* commands take no VMID. So isolating this
kernel by a reserved VMID alone would not work. Reserving the ASIDs covers
both the E2H and the NSEL1 cases.
Reservations are never released: a kdump kernel reboots after it saves the
vmcore, and the full-reset fallback flushes the entire TLB, which turns any
stale reservation into a merely unused ID.
If the scan finds any inconsistent structure, toss the entire adoption and
fall back to the full reset.
Suggested-by: Jason Gunthorpe <jgg@nvidia.com>
Reviewed-by: Jason Gunthorpe <jgg@nvidia.com>
Tested-by: Breno Leitao <leitao@kernel.org>
Tested-by: Cristian Prundeanu <cpru@amazon.com>
Assisted-by: LLM
Signed-off-by: Nicolin Chen <nicolinc@nvidia.com>
---
.../iommu/arm/arm-smmu-v3/arm-smmu-v3-kexec.c | 304 +++++++++++++++++-
drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c | 1 +
2 files changed, 304 insertions(+), 1 deletion(-)
diff --git a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3-kexec.c b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3-kexec.c
index 6bbc369df86bb..22d997480db1c 100644
--- a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3-kexec.c
+++ b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3-kexec.c
@@ -173,6 +173,298 @@ static int arm_smmu_kexec_check_strtab_l1_desc(struct arm_smmu_device *smmu,
return 0;
}
+/**
+ * arm_smmu_kexec_check_ste_cdtab() - Decode the CD table geometry of an STE
+ * @smmu: SMMU device of this kernel
+ * @ste0: first 64 bits of the previous kernel's S1 STE
+ * @cdtab: pointer to return the CD table's physical address
+ * @s1fmt: pointer to return the CD table format
+ * @max_contexts: pointer to return the number of CDs
+ *
+ * A linear CD table on the 2-level capable hardware is accepted, as a previous
+ * kernel might have used one, like the linear stream table.
+ *
+ * Note that the spec requires a CD table to be aligned to its own size, so an
+ * unaligned @cdtab gets rejected here: HW may then zero the low bits or fetch
+ * any CD in the table, leaving the live ASIDs unknowable to this scan.
+ *
+ * Return: 0 on success with the three outputs set, or -EINVAL on a bad geometry
+ */
+static int arm_smmu_kexec_check_ste_cdtab(struct arm_smmu_device *smmu,
+ u64 ste0, phys_addr_t *cdtab,
+ u32 *s1fmt, u32 *max_contexts)
+{
+ phys_addr_t base = ste0 & STRTAB_STE_0_S1CTXPTR_MASK;
+ u32 s1cdmax = FIELD_GET(STRTAB_STE_0_S1CDMAX, ste0);
+ u32 fmt = FIELD_GET(STRTAB_STE_0_S1FMT, ste0);
+ size_t size;
+
+ if (!base || s1cdmax > smmu->ssid_bits)
+ return -EINVAL;
+
+ if (fmt != STRTAB_STE_0_S1FMT_LINEAR &&
+ fmt != STRTAB_STE_0_S1FMT_64K_L2)
+ return -EINVAL;
+
+ /* Both kernels run on the same HW, so a genuine STE never has this */
+ if (fmt == STRTAB_STE_0_S1FMT_64K_L2 &&
+ !(smmu->features & ARM_SMMU_FEAT_2_LVL_CDTAB))
+ return -EINVAL;
+
+ if (fmt == STRTAB_STE_0_S1FMT_LINEAR)
+ size = (1UL << s1cdmax) * sizeof(struct arm_smmu_cd);
+ else
+ size = DIV_ROUND_UP(1UL << s1cdmax, CTXDESC_L2_ENTRIES) *
+ sizeof(struct arm_smmu_cdtab_l1);
+
+ /*
+ * An unaligned base is CONSTRAINED UNPREDICTABLE: HW may zero the low
+ * bits or fetch any CD in the table, so live ASIDs become unknowable.
+ */
+ if (!IS_ALIGNED(base, size))
+ return -EINVAL;
+
+ *cdtab = base;
+ *s1fmt = fmt;
+ *max_contexts = 1U << s1cdmax;
+ return 0;
+}
+
+static int arm_smmu_kexec_resv_asid(struct arm_smmu_device *smmu, u32 asid)
+{
+ /* A valid CD never has ASID 0; both kernels share the same HW limit */
+ if (!asid || asid >= 1UL << smmu->asid_bits)
+ return -EINVAL;
+
+ guard(mutex)(&arm_smmu_asid_lock);
+
+ /*
+ * The scan runs before this SMMU registers with the IOMMU core, so no
+ * domain of its own holds an ASID yet, while xa_reserve() does nothing
+ * if the entry is there, covering a domain's ASID that many CDs share.
+ */
+ return xa_reserve(&smmu->asid_map, asid, GFP_KERNEL);
+}
+
+static int arm_smmu_kexec_resv_vmid(struct arm_smmu_device *smmu, u32 vmid)
+{
+ int ret;
+
+ /* A translating STE never has VMID 0, which is reserved for bypass */
+ if (!vmid || vmid >= 1UL << smmu->vmid_bits)
+ return -EINVAL;
+
+ ret = ida_alloc_range(&smmu->vmid_map, vmid, vmid, GFP_KERNEL);
+ if (ret < 0 && ret != -ENOSPC) /* -ENOSPC means already reserved */
+ return ret;
+ return 0;
+}
+
+static int arm_smmu_kexec_resv_cd_asids(struct arm_smmu_device *smmu,
+ struct arm_smmu_cd *cds, u32 num_cds)
+{
+ int ret = 0;
+ u32 i;
+
+ for (i = 0; i < num_cds; i++) {
+ u64 val = le64_to_cpu(cds[i].data[0]);
+ u32 asid = FIELD_GET(CTXDESC_CD_0_ASID, val);
+
+ if (!(val & CTXDESC_CD_0_V))
+ continue;
+ ret = arm_smmu_kexec_resv_asid(smmu, asid);
+ if (ret)
+ break;
+ }
+ return ret;
+}
+
+/*
+ * Reserve the ASIDs of all the valid CDs of an S1 STE in the previous kernel's
+ * CD tables. The CD tables are transiently memremapped for the scan.
+ */
+static int arm_smmu_kexec_resv_s1_asids(struct arm_smmu_device *smmu, u64 ste0)
+{
+ struct arm_smmu_cdtab_l1 *l1tab;
+ u32 num_l1_ents, num_cds, i;
+ u32 max_contexts, s1fmt;
+ phys_addr_t cdtab;
+ int ret;
+
+ ret = arm_smmu_kexec_check_ste_cdtab(smmu, ste0, &cdtab, &s1fmt,
+ &max_contexts);
+ if (ret)
+ return ret;
+
+ if (s1fmt == STRTAB_STE_0_S1FMT_LINEAR) {
+ struct arm_smmu_cd *cds;
+
+ cds = memremap(cdtab, max_contexts * sizeof(*cds), MEMREMAP_WB);
+ if (!cds)
+ return -ENOMEM;
+ ret = arm_smmu_kexec_resv_cd_asids(smmu, cds, max_contexts);
+ memunmap(cds);
+ return ret;
+ }
+
+ num_l1_ents = DIV_ROUND_UP(max_contexts, CTXDESC_L2_ENTRIES);
+ l1tab = memremap(cdtab, num_l1_ents * sizeof(*l1tab), MEMREMAP_WB);
+ if (!l1tab)
+ return -ENOMEM;
+
+ /* max_contexts being under a full leaf makes the only leaf partial */
+ num_cds = min_t(u32, max_contexts, CTXDESC_L2_ENTRIES);
+
+ /* Aliased L2 tables cannot extend the walk; they only repeat a scan */
+ for (i = 0; i < num_l1_ents; i++) {
+ u64 l1_desc = le64_to_cpu(l1tab[i].l2ptr);
+ phys_addr_t l2_base = l1_desc & CTXDESC_L1_DESC_L2PTR_MASK;
+ struct arm_smmu_cdtab_l2 *l2;
+
+ if (!(l1_desc & CTXDESC_L1_DESC_V))
+ continue;
+
+ /*
+ * A valid descriptor never carries a null pointer. Also, an L2
+ * table is always 64KB-aligned, so an unaligned pointer would
+ * make this kernel read a different table.
+ */
+ if (!l2_base || !IS_ALIGNED(l2_base, sizeof(*l2))) {
+ ret = -EINVAL;
+ break;
+ }
+
+ l2 = memremap(l2_base, num_cds * sizeof(*l2->cds), MEMREMAP_WB);
+ if (!l2) {
+ ret = -ENOMEM;
+ break;
+ }
+ ret = arm_smmu_kexec_resv_cd_asids(smmu, l2->cds, num_cds);
+ memunmap(l2);
+ if (ret)
+ break;
+ }
+ memunmap(l1tab);
+ return ret;
+}
+
+static int arm_smmu_kexec_resv_ste_ids(struct arm_smmu_device *smmu,
+ struct arm_smmu_ste *ste)
+{
+ u32 vmid = FIELD_GET(STRTAB_STE_2_S2VMID, le64_to_cpu(ste->data[2]));
+ u64 ste0 = le64_to_cpu(ste->data[0]);
+
+ if (!(ste0 & STRTAB_STE_0_V))
+ return 0;
+
+ switch (FIELD_GET(STRTAB_STE_0_CFG, ste0)) {
+ case STRTAB_STE_0_CFG_ABORT:
+ case STRTAB_STE_0_CFG_BYPASS:
+ return 0;
+ case STRTAB_STE_0_CFG_S1_TRANS:
+ return arm_smmu_kexec_resv_s1_asids(smmu, ste0);
+ case STRTAB_STE_0_CFG_NESTED:
+ /*
+ * A guest-owned CD table is in the IPA space, unreachable. Its
+ * ASIDs are only tagged with the S2VMID reserved below, so they
+ * cannot alias this kernel's VMID-0 or EL2 S1 domains.
+ */
+ fallthrough;
+ case STRTAB_STE_0_CFG_S2_TRANS:
+ return arm_smmu_kexec_resv_vmid(smmu, vmid);
+ default:
+ return -EINVAL;
+ }
+}
+
+/**
+ * arm_smmu_kexec_scan_and_resv_ids() - Reserve a stream table's in-use IDs
+ * @smmu: SMMU device of this kernel, with an adopted or restored strtab_cfg
+ *
+ * Scan the stream table set up in the strtab_cfg and every CD table behind an
+ * S1 STE, reserving all of the in-use ASIDs and VMIDs. A failing scan rolls
+ * back through arm_smmu_kexec_unresv_ids().
+ *
+ * Note that the scan selects the linear or 2-level walk per this kernel's own
+ * ARM_SMMU_FEAT_2_LVL_STRTAB, so the caller must have matched the feature bit
+ * to the format of the adopted stream table in the strtab_cfg.
+ *
+ * Return: 0 on success, -EINVAL on any malformed table entry, or -ENOMEM on a
+ * memory shortage
+ */
+static int arm_smmu_kexec_scan_and_resv_ids(struct arm_smmu_device *smmu)
+{
+ struct arm_smmu_strtab_cfg *cfg = &smmu->strtab_cfg;
+ int ret = 0;
+ u32 i, j;
+
+ if (!(smmu->features & ARM_SMMU_FEAT_2_LVL_STRTAB)) {
+ for (i = 0; i < cfg->linear.num_ents; i++) {
+ ret = arm_smmu_kexec_resv_ste_ids(
+ smmu, &cfg->linear.table[i]);
+ if (ret)
+ return ret;
+ }
+ return 0;
+ }
+
+ /* Aliased L2 tables cannot extend the scan; they only repeat a scan */
+ for (i = 0; i < cfg->l2.num_l1_ents; i++) {
+ u64 l1_desc = le64_to_cpu(cfg->l2.l1tab[i].l2ptr);
+ struct arm_smmu_strtab_l2 *l2;
+ phys_addr_t base;
+
+ ret = arm_smmu_kexec_check_strtab_l1_desc(smmu, l1_desc, i,
+ &base);
+ if (ret == 1)
+ continue;
+ if (ret)
+ return ret;
+
+ /*
+ * This kernel will map the previous kernel's L2 tables lazily
+ * or not at all. Here, take a transient view for this scan.
+ */
+ l2 = memremap(base, sizeof(*l2), MEMREMAP_WB);
+ if (!l2)
+ return -ENOMEM;
+ for (j = 0; j < ARRAY_SIZE(l2->stes); j++) {
+ ret = arm_smmu_kexec_resv_ste_ids(smmu, &l2->stes[j]);
+ if (ret)
+ break;
+ }
+ memunmap(l2);
+ if (ret)
+ return ret;
+ }
+ return 0;
+}
+
+/**
+ * arm_smmu_kexec_unresv_ids() - Release the IDs that a failing scan reserved
+ * @smmu: SMMU device of this kernel that failed its reservation scan
+ *
+ * Undo the reservations of a failing arm_smmu_kexec_scan_and_resv_ids() call,
+ * for a caller that falls back to a full reset.
+ *
+ * That reset flushes the whole TLB, so the previous kernel's IDs no longer need
+ * any protection. A scan that fails halfway would otherwise keep a good share
+ * of an 8-bit ASID or VMID space reserved for nothing.
+ */
+static void arm_smmu_kexec_unresv_ids(struct arm_smmu_device *smmu)
+{
+ /*
+ * Emptying both maps releases exactly this scan's IDs, as no domain of
+ * this SMMU can hold one until it registers with the IOMMU core, later
+ * in the probe. Both stay initialized and usable for the full reset.
+ */
+ mutex_lock(&arm_smmu_asid_lock);
+ xa_destroy(&smmu->asid_map);
+ mutex_unlock(&arm_smmu_asid_lock);
+
+ ida_destroy(&smmu->vmid_map);
+}
+
#ifdef CONFIG_CRASH_DUMP
/*
* Helper functions of the kdump stream table adoption for ARM SMMUv3
@@ -373,16 +665,26 @@ int arm_smmu_kdump_adopt_strtab(struct arm_smmu_device *smmu)
goto err;
}
+ ret = arm_smmu_kexec_scan_and_resv_ids(smmu);
+ if (ret) {
+ dev_warn(smmu->dev, "failed to reserve in-use ASIDs/VMIDs\n");
+ arm_smmu_kdump_adopt_cleanup(smmu);
+ goto err_unresv;
+ }
+
ret = devm_add_action_or_reset(smmu->dev, arm_smmu_kdump_adopt_cleanup,
smmu);
/* devm_add_action_or_reset ran the cleanup upon failure */
if (ret) {
dev_warn(smmu->dev, "failed to set up cleanup action\n");
- goto err;
+ goto err_unresv;
}
return 0;
+err_unresv:
+ /* The full reset will flush the entire TLB, so release everything */
+ arm_smmu_kexec_unresv_ids(smmu);
err:
dev_warn(smmu->dev, "falling back to full reset\n");
/*
diff --git a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
index db80b980d1dd9..3ec49c1e16aa7 100644
--- a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
+++ b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
@@ -4868,6 +4868,7 @@ static int arm_smmu_init_strtab(struct arm_smmu_device *smmu)
{
int ret;
+ /* Init both first, as a kdump adoption reserves in-use ASIDs/VMIDs */
ida_init(&smmu->vmid_map);
ret = devm_add_action_or_reset(smmu->dev, arm_smmu_destroy_vmid_map,
&smmu->vmid_map);
--
2.43.0
^ permalink raw reply [flat|nested] 13+ messages in thread* [PATCH v11 08/11] iommu/arm-smmu-v3-kexec: Implement is_attach_deferred()
2026-10-05 19:30 [PATCH v11 00/11] iommu/arm-smmu-v3: Adopt the crashed kernel's stream table for kdump Nicolin Chen
` (6 preceding siblings ...)
2026-10-05 19:30 ` [PATCH v11 07/11] iommu/arm-smmu-v3-kexec: Reserve crashed kernel's ASIDs and VMIDs Nicolin Chen
@ 2026-10-05 19:30 ` Nicolin Chen
2026-10-05 19:30 ` [PATCH v11 09/11] iommu/arm-smmu-v3: Retain CR0_SMMUEN during kdump device reset Nicolin Chen
` (3 subsequent siblings)
11 siblings, 0 replies; 13+ messages in thread
From: Nicolin Chen @ 2026-10-05 19:30 UTC (permalink / raw)
To: will
Cc: robin.murphy, joro, jgg, praan, kevin.tian, smostafa,
linux-arm-kernel, iommu, linux-kernel, jamien, kas,
Cristian Prundeanu, Breno Leitao
Though the kdump kernel adopts the crashed kernel's stream table, the iommu
core will still try to attach each probed device to a default domain, which
overwrites the adopted STE and breaks in-flight DMA from that device.
Implement an is_attach_deferred() callback to prevent this. For each device
that has STE.V=1 and STE.Cfg!=Abort in the adopted table, defer the default
domain attachment, until the device driver explicitly requests it.
Also, move arm_smmu_get_step_for_sid() to the header for the kdump function
to use.
Reviewed-by: Kevin Tian <kevin.tian@intel.com>
Reviewed-by: Jason Gunthorpe <jgg@nvidia.com>
Reviewed-by: Pranjal Shrivastava <praan@google.com>
Tested-by: Breno Leitao <leitao@kernel.org>
Tested-by: Cristian Prundeanu <cpru@amazon.com>
Signed-off-by: Nicolin Chen <nicolinc@nvidia.com>
---
drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.h | 23 +++++++++++++++++
.../iommu/arm/arm-smmu-v3/arm-smmu-v3-kexec.c | 19 ++++++++++++++
drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c | 25 ++++++++-----------
3 files changed, 52 insertions(+), 15 deletions(-)
diff --git a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.h b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.h
index a1f947e7d9fe4..49604550e3b9f 100644
--- a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.h
+++ b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.h
@@ -1280,6 +1280,22 @@ int __arm_smmu_cmdq_issue_cmdlist(struct arm_smmu_device *smmu,
struct arm_smmu_cmdq *cmdq,
struct arm_smmu_cmd *cmds, int n,
bool sync);
+
+static inline struct arm_smmu_ste *
+arm_smmu_get_step_for_sid(struct arm_smmu_device *smmu, u32 sid)
+{
+ struct arm_smmu_strtab_cfg *cfg = &smmu->strtab_cfg;
+
+ if (smmu->features & ARM_SMMU_FEAT_2_LVL_STRTAB) {
+ /* Two-level walk */
+ return &cfg->l2.l2ptrs[arm_smmu_strtab_l1_idx(sid)]
+ ->stes[arm_smmu_strtab_l2_idx(sid)];
+ } else {
+ /* Simple linear lookup */
+ return &cfg->linear.table[sid];
+ }
+}
+
int arm_smmu_cmdq_issue_cmdlist(struct arm_smmu_device *smmu,
struct arm_smmu_cmdq *cmdq,
struct arm_smmu_cmd *cmds, int n,
@@ -1318,6 +1334,7 @@ int arm_smmu_kdump_adopt_strtab(struct arm_smmu_device *smmu);
int arm_smmu_kdump_adopt_deferred_l2_strtab(struct arm_smmu_device *smmu,
u32 sid, phys_addr_t base, u32 span,
struct arm_smmu_strtab_l2 **l2table);
+bool arm_smmu_kdump_is_attach_deferred(struct arm_smmu_master *master);
#else /* CONFIG_CRASH_DUMP */
static inline int arm_smmu_kdump_adopt_strtab(struct arm_smmu_device *smmu)
{
@@ -1331,6 +1348,12 @@ arm_smmu_kdump_adopt_deferred_l2_strtab(struct arm_smmu_device *smmu, u32 sid,
{
return -EOPNOTSUPP;
}
+
+static inline bool
+arm_smmu_kdump_is_attach_deferred(struct arm_smmu_master *master)
+{
+ return false;
+}
#endif /* CONFIG_CRASH_DUMP */
struct arm_vsmmu {
diff --git a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3-kexec.c b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3-kexec.c
index 22d997480db1c..8a76b8cb4ef0a 100644
--- a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3-kexec.c
+++ b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3-kexec.c
@@ -697,4 +697,23 @@ int arm_smmu_kdump_adopt_strtab(struct arm_smmu_device *smmu)
smmu->options &= ~ARM_SMMU_OPT_KDUMP_ADOPT;
return ret;
}
+
+bool arm_smmu_kdump_is_attach_deferred(struct arm_smmu_master *master)
+{
+ struct arm_smmu_device *smmu = master->smmu;
+ int i;
+
+ for (i = 0; i < master->num_streams; i++) {
+ struct arm_smmu_ste *ste =
+ arm_smmu_get_step_for_sid(smmu, master->streams[i].id);
+ u64 ent0 = le64_to_cpu(ste->data[0]);
+
+ /* Defer only when there might be in-flight DMAs */
+ if ((ent0 & STRTAB_STE_0_V) &&
+ FIELD_GET(STRTAB_STE_0_CFG, ent0) != STRTAB_STE_0_CFG_ABORT)
+ return true;
+ }
+
+ return false;
+}
#endif /* CONFIG_CRASH_DUMP */
diff --git a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
index 3ec49c1e16aa7..d301f313ba1aa 100644
--- a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
+++ b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
@@ -3142,21 +3142,6 @@ static int arm_smmu_domain_finalise(struct arm_smmu_domain *smmu_domain,
return 0;
}
-static struct arm_smmu_ste *
-arm_smmu_get_step_for_sid(struct arm_smmu_device *smmu, u32 sid)
-{
- struct arm_smmu_strtab_cfg *cfg = &smmu->strtab_cfg;
-
- if (smmu->features & ARM_SMMU_FEAT_2_LVL_STRTAB) {
- /* Two-level walk */
- return &cfg->l2.l2ptrs[arm_smmu_strtab_l1_idx(sid)]
- ->stes[arm_smmu_strtab_l2_idx(sid)];
- } else {
- /* Simple linear lookup */
- return &cfg->linear.table[sid];
- }
-}
-
void arm_smmu_install_ste_for_dev(struct arm_smmu_master *master,
const struct arm_smmu_ste *target)
{
@@ -4598,6 +4583,15 @@ static int arm_smmu_def_domain_type(struct device *dev)
return 0;
}
+static bool arm_smmu_is_attach_deferred(struct device *dev)
+{
+ struct arm_smmu_master *master = dev_iommu_priv_get(dev);
+
+ if (master->smmu->options & ARM_SMMU_OPT_KDUMP_ADOPT)
+ return arm_smmu_kdump_is_attach_deferred(master);
+ return false;
+}
+
static const struct iommu_ops arm_smmu_ops = {
.identity_domain = &arm_smmu_identity_domain,
.blocked_domain = &arm_smmu_blocked_domain,
@@ -4606,6 +4600,7 @@ static const struct iommu_ops arm_smmu_ops = {
.hw_info = arm_smmu_hw_info,
.domain_alloc_sva = arm_smmu_sva_domain_alloc,
.domain_alloc_paging_flags = arm_smmu_domain_alloc_paging_flags,
+ .is_attach_deferred = arm_smmu_is_attach_deferred,
.probe_device = arm_smmu_probe_device,
.release_device = arm_smmu_release_device,
.device_group = arm_smmu_device_group,
--
2.43.0
^ permalink raw reply [flat|nested] 13+ messages in thread* [PATCH v11 09/11] iommu/arm-smmu-v3: Retain CR0_SMMUEN during kdump device reset
2026-10-05 19:30 [PATCH v11 00/11] iommu/arm-smmu-v3: Adopt the crashed kernel's stream table for kdump Nicolin Chen
` (7 preceding siblings ...)
2026-10-05 19:30 ` [PATCH v11 08/11] iommu/arm-smmu-v3-kexec: Implement is_attach_deferred() Nicolin Chen
@ 2026-10-05 19:30 ` Nicolin Chen
2026-10-05 19:30 ` [PATCH v11 10/11] iommu/arm-smmu-v3: Skip RMR bypass for kdump adoption Nicolin Chen
` (2 subsequent siblings)
11 siblings, 0 replies; 13+ messages in thread
From: Nicolin Chen @ 2026-10-05 19:30 UTC (permalink / raw)
To: will
Cc: robin.murphy, joro, jgg, praan, kevin.tian, smostafa,
linux-arm-kernel, iommu, linux-kernel, jamien, kas,
Cristian Prundeanu, Breno Leitao
When ARM_SMMU_OPT_KDUMP_ADOPT is detected, do not disable SMMUEN and skip
the CR1/CR2/STRTAB_BASE update sequence in arm_smmu_device_reset(). Those
register writes are all CONSTRAINED UNPREDICTABLE while CR0_SMMUEN==1, so
leaving them intact lets in-flight DMAs continue to be translated by the
adopted stream table.
Initialize 'enables' to 0, so it can carry the retained CR0 fields in the
kdump case, clearing only the queue enable bits. Then, preserve them when
enabling the command queue.
The retained CR0 keeps the crashed kernel's ATSCHK too, which selects the
fast (0) or the safe (1) mode for the ATS translated traffic. Switching to
the safe mode would start checking the in-flight traffic against STE.EATS
fields that the crashed kernel never set up for a check, aborting the very
DMAs being carried.
Clear latched gerror bits if necessary.
Reviewed-by: Kevin Tian <kevin.tian@intel.com>
Reviewed-by: Pranjal Shrivastava <praan@google.com>
Tested-by: Breno Leitao <leitao@kernel.org>
Tested-by: Cristian Prundeanu <cpru@amazon.com>
Signed-off-by: Nicolin Chen <nicolinc@nvidia.com>
---
drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c | 60 +++++++++++++++++++--
1 file changed, 56 insertions(+), 4 deletions(-)
diff --git a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
index d301f313ba1aa..5542d136c1d77 100644
--- a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
+++ b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
@@ -5148,10 +5148,28 @@ static void arm_smmu_write_strtab(struct arm_smmu_device *smmu)
static int arm_smmu_device_reset(struct arm_smmu_device *smmu)
{
int ret;
- u32 reg, enables;
+ u32 reg, enables = 0;
- /* Clear CR0 and sync (disables SMMU and queue processing) */
reg = readl_relaxed(smmu->base + ARM_SMMU_CR0);
+
+ /*
+ * In a kdump case (set when CR0_SMMUEN=1 and !GERROR_SFM_ERR), retain
+ * all the live CR0 fields, e.g. CR0_SMMUEN to avoid aborting in-flight
+ * DMA and CR0_ATSCHK to carry on the ATS-check policy, while clearing
+ * only the queue enable bits for this kernel to take over the queues.
+ *
+ * According to spec, updating STRTAB_BASE/CR1/CR2 when CR0_SMMUEN=1 is
+ * CONSTRAINED UNPREDICTABLE. So, skip those register updates and rely
+ * on the adopted stream table from the crashed kernel.
+ */
+ if (smmu->options & ARM_SMMU_OPT_KDUMP_ADOPT) {
+ dev_info(smmu->dev,
+ "kdump: retaining SMMUEN for in-flight DMA\n");
+ enables = reg & ~(CR0_CMDQEN | CR0_EVTQEN | CR0_PRIQEN);
+ goto reset_queues;
+ }
+
+ /* Clear CR0 and sync (disables SMMU and queue processing) */
if (reg & CR0_SMMUEN) {
dev_warn(smmu->dev, "SMMU currently enabled! Resetting...\n");
arm_smmu_update_gbpa(smmu, GBPA_ABORT, 0);
@@ -5181,12 +5199,41 @@ static int arm_smmu_device_reset(struct arm_smmu_device *smmu)
/* Stream table */
arm_smmu_write_strtab(smmu);
+reset_queues:
+ if (smmu->options & ARM_SMMU_OPT_KDUMP_ADOPT) {
+ /*
+ * Disable queues since arm_smmu_device_disable() was skipped.
+ * CR0 fields are independent per spec, so the queue enable bits
+ * can be cleared while retaining SMMUEN=1.
+ */
+ ret = arm_smmu_write_reg_sync(smmu, enables, ARM_SMMU_CR0,
+ ARM_SMMU_CR0ACK);
+ if (ret) {
+ dev_err(smmu->dev, "failed to disable queues\n");
+ return ret;
+ }
+ }
+
+ /*
+ * GERROR bits are latched. Read after queue disabling so that unhandled
+ * errors would be visible. Ack everything prior to re-enabling the CMDQ
+ * as a stale CMDQ_ERR would halt the CMDQ and new command will timeout.
+ * Acking SFM_ERR is defined too, although it would not exit the SFM.
+ */
+ if (is_kdump_kernel()) {
+ u32 gerror = readl_relaxed(smmu->base + ARM_SMMU_GERROR);
+ u32 gerrorn = readl_relaxed(smmu->base + ARM_SMMU_GERRORN);
+
+ if ((gerror ^ gerrorn) & GERROR_ERR_MASK)
+ writel(gerror, smmu->base + ARM_SMMU_GERRORN);
+ }
+
/* Command queue */
writeq_relaxed(smmu->cmdq.q.q_base, smmu->base + ARM_SMMU_CMDQ_BASE);
writel_relaxed(smmu->cmdq.q.llq.prod, smmu->base + ARM_SMMU_CMDQ_PROD);
writel_relaxed(smmu->cmdq.q.llq.cons, smmu->base + ARM_SMMU_CMDQ_CONS);
- enables = CR0_CMDQEN;
+ enables |= CR0_CMDQEN;
ret = arm_smmu_write_reg_sync(smmu, enables, ARM_SMMU_CR0,
ARM_SMMU_CR0ACK);
if (ret) {
@@ -5242,7 +5289,12 @@ static int arm_smmu_device_reset(struct arm_smmu_device *smmu)
}
}
- if (smmu->features & ARM_SMMU_FEAT_ATS) {
+ /*
+ * In a kdump adopt case, retain the crashed kernel's ATS-check policy
+ * captured above rather than forcing it on.
+ */
+ if (!(smmu->options & ARM_SMMU_OPT_KDUMP_ADOPT) &&
+ (smmu->features & ARM_SMMU_FEAT_ATS)) {
enables |= CR0_ATSCHK;
ret = arm_smmu_write_reg_sync(smmu, enables, ARM_SMMU_CR0,
ARM_SMMU_CR0ACK);
--
2.43.0
^ permalink raw reply [flat|nested] 13+ messages in thread* [PATCH v11 10/11] iommu/arm-smmu-v3: Skip RMR bypass for kdump adoption
2026-10-05 19:30 [PATCH v11 00/11] iommu/arm-smmu-v3: Adopt the crashed kernel's stream table for kdump Nicolin Chen
` (8 preceding siblings ...)
2026-10-05 19:30 ` [PATCH v11 09/11] iommu/arm-smmu-v3: Retain CR0_SMMUEN during kdump device reset Nicolin Chen
@ 2026-10-05 19:30 ` Nicolin Chen
2026-10-05 19:30 ` [PATCH v11 11/11] iommu/arm-smmu-v3: Detect ARM_SMMU_OPT_KDUMP_ADOPT in probe() Nicolin Chen
2026-10-05 20:48 ` [PATCH v11 00/11] iommu/arm-smmu-v3: Adopt the crashed kernel's stream table for kdump Nicolin Chen
11 siblings, 0 replies; 13+ messages in thread
From: Nicolin Chen @ 2026-10-05 19:30 UTC (permalink / raw)
To: will
Cc: robin.murphy, joro, jgg, praan, kevin.tian, smostafa,
linux-arm-kernel, iommu, linux-kernel, jamien, kas,
Cristian Prundeanu, Breno Leitao
RMR bypass STEs are installed during SMMUv3 probe for StreamIDs listed by
IORT RMR nodes. A normal boot switches the driver to a fresh stream table
whose initial STEs abort, so those RMR SIDs need bypass entries before it
becomes live. This preserves firmware/guest-owned traffic, including vSMMU
guest MSI cases built around RMR-described SIDs.
ARM_SMMU_OPT_KDUMP_ADOPT is the opposite case: the driver keeps SMMUEN set
and adopts the crashed kernel's stream table, so RMR SIDs already have the
only translation state known to be safe for active in-flight DMA. Replacing
an adopted STE with bypass can turn translated DMA into physical DMA, then
point it at the wrong memory.
arm_smmu_make_bypass_ste() also rewrites the STE in place after clearing it
first. While the table is live, a concurrent hardware STE fetch can observe
V=0 or mixed old/new state.
Leaving the adopted STE unmodified keeps the kdump kernel using the crashed
kernel's translation. That gives the endpoint driver a chance to probe and
quiesce the device.
If the old STE was already abort or invalid, installing bypass would create
new DMA permission; leaving it alone is a safer failure mode. Later domain
setup still gets the RMR direct mappings through the reserved-region path.
Reviewed-by: Pranjal Shrivastava <praan@google.com>
Reviewed-by: Jason Gunthorpe <jgg@nvidia.com>
Tested-by: Breno Leitao <leitao@kernel.org>
Tested-by: Cristian Prundeanu <cpru@amazon.com>
Assisted-by: LLM
Signed-off-by: Nicolin Chen <nicolinc@nvidia.com>
---
drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c | 13 +++++++++----
1 file changed, 9 insertions(+), 4 deletions(-)
diff --git a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
index 5542d136c1d77..78a24bca44202 100644
--- a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
+++ b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
@@ -5816,6 +5816,14 @@ static void arm_smmu_rmr_install_bypass_ste(struct arm_smmu_device *smmu)
struct list_head rmr_list;
struct iommu_resv_region *e;
+ /*
+ * Kdump adoption keeps the crashed kernel's table live. Rewriting the
+ * adopted STE here could expose an in-flight fetch to a transient V=0
+ * entry, or change Cfg=translate to Cfg=bypass. Must skip here.
+ */
+ if (smmu->options & ARM_SMMU_OPT_KDUMP_ADOPT)
+ return;
+
INIT_LIST_HEAD(&rmr_list);
iort_get_rmr_sids(dev_fwnode(smmu->dev), &rmr_list);
@@ -5832,10 +5840,7 @@ static void arm_smmu_rmr_install_bypass_ste(struct arm_smmu_device *smmu)
continue;
}
- /*
- * STE table is not programmed to HW, see
- * arm_smmu_initial_bypass_stes()
- */
+ /* The fresh stream table is not yet live. */
arm_smmu_make_bypass_ste(smmu,
arm_smmu_get_step_for_sid(smmu, rmr->sids[i]));
}
--
2.43.0
^ permalink raw reply [flat|nested] 13+ messages in thread* [PATCH v11 11/11] iommu/arm-smmu-v3: Detect ARM_SMMU_OPT_KDUMP_ADOPT in probe()
2026-10-05 19:30 [PATCH v11 00/11] iommu/arm-smmu-v3: Adopt the crashed kernel's stream table for kdump Nicolin Chen
` (9 preceding siblings ...)
2026-10-05 19:30 ` [PATCH v11 10/11] iommu/arm-smmu-v3: Skip RMR bypass for kdump adoption Nicolin Chen
@ 2026-10-05 19:30 ` Nicolin Chen
2026-10-05 20:48 ` [PATCH v11 00/11] iommu/arm-smmu-v3: Adopt the crashed kernel's stream table for kdump Nicolin Chen
11 siblings, 0 replies; 13+ messages in thread
From: Nicolin Chen @ 2026-10-05 19:30 UTC (permalink / raw)
To: will
Cc: robin.murphy, joro, jgg, praan, kevin.tian, smostafa,
linux-arm-kernel, iommu, linux-kernel, jamien, kas,
Cristian Prundeanu, Breno Leitao
arm_smmu_device_hw_probe() runs before arm_smmu_init_structures(), so it's
natural to decide whether the kdump kernel must adopt the crashed kernel's
stream table.
Given that memremap is used to adopt the old stream table, set this option
only on a coherent SMMU.
And make sure SMMU isn't in Service Failure Mode.
Reviewed-by: Kevin Tian <kevin.tian@intel.com>
Reviewed-by: Jason Gunthorpe <jgg@nvidia.com>
Reviewed-by: Pranjal Shrivastava <praan@google.com>
Tested-by: Breno Leitao <leitao@kernel.org>
Tested-by: Cristian Prundeanu <cpru@amazon.com>
Signed-off-by: Nicolin Chen <nicolinc@nvidia.com>
---
drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.h | 5 ++++
.../iommu/arm/arm-smmu-v3/arm-smmu-v3-kexec.c | 27 +++++++++++++++++++
drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c | 4 +++
3 files changed, 36 insertions(+)
diff --git a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.h b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.h
index 49604550e3b9f..228ccdf9c6f5d 100644
--- a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.h
+++ b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.h
@@ -1335,6 +1335,7 @@ int arm_smmu_kdump_adopt_deferred_l2_strtab(struct arm_smmu_device *smmu,
u32 sid, phys_addr_t base, u32 span,
struct arm_smmu_strtab_l2 **l2table);
bool arm_smmu_kdump_is_attach_deferred(struct arm_smmu_master *master);
+void arm_smmu_device_kdump_probe(struct arm_smmu_device *smmu);
#else /* CONFIG_CRASH_DUMP */
static inline int arm_smmu_kdump_adopt_strtab(struct arm_smmu_device *smmu)
{
@@ -1354,6 +1355,10 @@ arm_smmu_kdump_is_attach_deferred(struct arm_smmu_master *master)
{
return false;
}
+
+static inline void arm_smmu_device_kdump_probe(struct arm_smmu_device *smmu)
+{
+}
#endif /* CONFIG_CRASH_DUMP */
struct arm_vsmmu {
diff --git a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3-kexec.c b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3-kexec.c
index 8a76b8cb4ef0a..f664de3697e0e 100644
--- a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3-kexec.c
+++ b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3-kexec.c
@@ -716,4 +716,31 @@ bool arm_smmu_kdump_is_attach_deferred(struct arm_smmu_master *master)
return false;
}
+
+void arm_smmu_device_kdump_probe(struct arm_smmu_device *smmu)
+{
+ u32 gerror, gerrorn, active;
+
+ /* No adoption if SMMU is disabled (i.e., there is no in-flight DMA) */
+ if (!(readl_relaxed(smmu->base + ARM_SMMU_CR0) & CR0_SMMUEN))
+ return;
+
+ /* For now, only support a coherent SMMU that works with MEMREMAP_WB */
+ if (!(smmu->features & ARM_SMMU_FEAT_COHERENCY)) {
+ dev_warn(smmu->dev,
+ "non-coherent SMMU unsupported; reset to block all DMAs\n");
+ return;
+ }
+
+ gerror = readl_relaxed(smmu->base + ARM_SMMU_GERROR);
+ gerrorn = readl_relaxed(smmu->base + ARM_SMMU_GERRORN);
+ active = gerror ^ gerrorn;
+ if (active & GERROR_SFM_ERR) {
+ dev_warn(smmu->dev,
+ "SMMU in Service Failure Mode, must reset\n");
+ return;
+ }
+
+ smmu->options |= ARM_SMMU_OPT_KDUMP_ADOPT;
+}
#endif /* CONFIG_CRASH_DUMP */
diff --git a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
index 78a24bca44202..3ada71fe7239f 100644
--- a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
+++ b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
@@ -5647,6 +5647,10 @@ static int arm_smmu_device_hw_probe(struct arm_smmu_device *smmu)
dev_info(smmu->dev, "oas %lu-bit (features 0x%08x)\n",
smmu->oas, smmu->features);
+
+ if (is_kdump_kernel())
+ arm_smmu_device_kdump_probe(smmu);
+
return 0;
}
--
2.43.0
^ permalink raw reply [flat|nested] 13+ messages in thread* Re: [PATCH v11 00/11] iommu/arm-smmu-v3: Adopt the crashed kernel's stream table for kdump
2026-10-05 19:30 [PATCH v11 00/11] iommu/arm-smmu-v3: Adopt the crashed kernel's stream table for kdump Nicolin Chen
` (10 preceding siblings ...)
2026-10-05 19:30 ` [PATCH v11 11/11] iommu/arm-smmu-v3: Detect ARM_SMMU_OPT_KDUMP_ADOPT in probe() Nicolin Chen
@ 2026-10-05 20:48 ` Nicolin Chen
11 siblings, 0 replies; 13+ messages in thread
From: Nicolin Chen @ 2026-10-05 20:48 UTC (permalink / raw)
To: will
Cc: robin.murphy, joro, jgg, praan, kevin.tian, smostafa,
linux-arm-kernel, iommu, linux-kernel, jamien, kas,
Cristian Prundeanu, Breno Leitao
On Mon, Oct 05, 2026 at 12:30:43PM -0700, Nicolin Chen wrote:
> This is on Github:
> https://github.com/nicolinc/iommufd/commits/smmuv3_kdump-v11
>
> Changelog
> v11
> * Rebase on arm/smmu/updates branch
> * Add Tested-by from Breno and Cristian
> * Add Reviewed/Acked-by from Jason and Kiryl
> * Update commit messages and Assisted-by tags
> * New prep patch: swap EVTQ/GERROR MSI indexes
> * Skip EVTQ's CFG0 write and vector allocation
> * Merge arm-smmu-v3-kdump code into arm-smmu-v3-kexec
> * Fold patches accordingly as static functions require callers
Will,
I have checked the new Sashiko finding against this v11.
To approach its claimed 35-trillion bound, every L1 descriptor,
STE and CD L1 descriptor must pass the validations, which takes
a deliberately built aliasing table but is not likely reachable
in practice: with a corrupted memory, one of the entries would
more likely either skip the iteration (V=0) or fail the test.
I tend to ignore that. But if you think it is worth addressing
and have some good idea about it, I don't mind a respin.
Thanks
Nicolin
^ permalink raw reply [flat|nested] 13+ messages in thread