From: Lu Baolu <baolu.lu@linux.intel.com>
To: Joerg Roedel <joro@8bytes.org>
Cc: Guanghui Feng <guanghuifeng@linux.alibaba.com>,
Zhenzhong Duan <zhenzhong.duan@intel.com>,
iommu@lists.linux.dev, linux-kernel@vger.kernel.org
Subject: [PATCH 9/9] iommu/vt-d: Fix IQE handling to cover all descriptors in submission range
Date: Mon, 28 Sep 2026 11:27:22 +0800 [thread overview]
Message-ID: <20260928032722.2868623-10-baolu.lu@linux.intel.com> (raw)
In-Reply-To: <20260928032722.2868623-1-baolu.lu@linux.intel.com>
From: Guanghui Feng <guanghuifeng@linux.alibaba.com>
When an Invalidation Queue Error (IQE) occurs, hardware halts fetching
and IQH points at the faulting descriptor. The previous code only
checked whether IQH matched the first descriptor index of the current
submission, missing faults on any other descriptor within the batch.
Expand the check to cover the entire submission range [index, wait_index],
accounting for circular wrap-around.
Furthermore, after detecting IQE, the old recovery only replaced the
single faulting slot with a copy of the wait descriptor and immediately
returned -EINVAL. This left two problems:
a) Hardware resumed fetching and could hit another invalid descriptor
in the same abandoned batch, raising a second IQE that no submitter
would claim — permanently deadlocking the queue.
b) The caller reclaimed all batch slots (QI_FREE) while hardware might
still be asynchronously processing descriptors from that batch,
allowing concurrent overwrite and descriptor corruption.
Fix both by introducing qi_drain_remaining_descs(): upon IQE detection,
overwrite the stranded slots in [IQH, wait_index) with fenced no-op wait
descriptors, resubmit the wait descriptor at wait_index, clear IQE, and
spin until hardware signals QI_DONE (with DMAR_OPERATION_TIMEOUT). This
guarantees hardware has fully drained the batch before the caller
reclaims slots.
Fixes: 8a1d82462540 ("iommu/vt-d: Multiple descriptors per qi_submit_sync()")
Signed-off-by: Guanghui Feng <guanghuifeng@linux.alibaba.com>
Signed-off-by: Lu Baolu <baolu.lu@linux.intel.com>
---
drivers/iommu/intel/dmar.c | 70 ++++++++++++++++++++++++++++++++------
1 file changed, 60 insertions(+), 10 deletions(-)
diff --git a/drivers/iommu/intel/dmar.c b/drivers/iommu/intel/dmar.c
index ba675b08cd20..6377ba97a9ea 100644
--- a/drivers/iommu/intel/dmar.c
+++ b/drivers/iommu/intel/dmar.c
@@ -1344,6 +1344,49 @@ static void qi_dump_fault(struct intel_iommu *iommu, u32 fault)
(unsigned long long)desc->qw1);
}
+static int qi_drain_remaining_descs(struct intel_iommu *iommu, int head,
+ int wait_index)
+{
+ struct q_inval *qi = iommu->qi;
+ int shift = qi_shift(iommu);
+ struct qi_desc wait_desc = {};
+ struct qi_desc nop_desc = {};
+ cycles_t start;
+
+ nop_desc.qw0 = QI_IWD_FENCE | QI_IWD_TYPE;
+
+ wait_desc.qw0 = QI_IWD_STATUS_DATA(QI_DONE) | QI_IWD_PRQ_DRAIN |
+ QI_IWD_STATUS_WRITE | QI_IWD_FENCE | QI_IWD_TYPE;
+ wait_desc.qw1 = virt_to_phys(&qi->desc_status[wait_index]);
+
+ while (head != wait_index) {
+ memcpy(qi->desc + (head << shift), &nop_desc, 1 << shift);
+ head = (head + 1) % QI_LENGTH;
+ }
+
+ WRITE_ONCE(qi->desc_status[wait_index], QI_IN_USE);
+ memcpy(qi->desc + (wait_index << shift), &wait_desc, 1 << shift);
+
+ /*
+ * Order the descriptor rewrites before the writel() that restarts
+ * the fetch engine.
+ */
+ wmb();
+
+ writel(DMA_FSTS_IQE, iommu->reg + DMAR_FSTS_REG);
+
+ start = get_cycles();
+ while (READ_ONCE(qi->desc_status[wait_index]) != QI_DONE) {
+ if (DMAR_OPERATION_TIMEOUT < (get_cycles() - start)) {
+ pr_err("Timeout draining invalidation queue after IQE\n");
+ return -ETIMEDOUT;
+ }
+ cpu_relax();
+ }
+
+ return 0;
+}
+
static int qi_check_fault(struct intel_iommu *iommu, int index, int wait_index)
{
u32 fault;
@@ -1366,18 +1409,25 @@ static int qi_check_fault(struct intel_iommu *iommu, int index, int wait_index)
* is cleared.
*/
if (fault & DMA_FSTS_IQE) {
+ int head_idx, ret;
+
head = readl(iommu->reg + DMAR_IQH_REG);
- if ((head >> shift) == index) {
- struct qi_desc *desc = qi->desc + head;
+ head_idx = (head >> shift) % QI_LENGTH;
+
+ /*
+ * The faulting descriptor can be anywhere within the current
+ * submission's range [index, wait_index]. Since the queue is
+ * circular, this submission may wrap around QI_LENGTH
+ * (index > wait_index in that case), so check both the
+ * non-wrapped and wrapped cases of the range.
+ */
+ if (index <= wait_index ?
+ (head_idx >= index && head_idx <= wait_index) :
+ (head_idx >= index || head_idx <= wait_index)) {
+ ret = qi_drain_remaining_descs(iommu, head_idx, wait_index);
+ if (ret)
+ return ret;
- /*
- * desc->qw2 and desc->qw3 are either reserved or
- * used by software as private data. We won't print
- * out these two qw's for security consideration.
- */
- memcpy(desc, qi->desc + (wait_index << shift),
- 1 << shift);
- writel(DMA_FSTS_IQE, iommu->reg + DMAR_FSTS_REG);
pr_info("Invalidation Queue Error (IQE) cleared\n");
return -EINVAL;
}
--
2.43.0
next prev parent reply other threads:[~2026-09-28 3:39 UTC|newest]
Thread overview: 11+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-28 3:27 [PATCH 0/9][PULL REQUEST] Intel IOMMU updates for v7.4 Lu Baolu
2026-09-28 3:27 ` [PATCH 1/9] iommu/vt-d: Fix page table level calculation in compute_vasz_lg2_ss() Lu Baolu
2026-09-28 3:27 ` [PATCH 2/9] iommu/vt-d: Avoid out-of-range shift in qi_desc_dev_iotlb_pasid() Lu Baolu
2026-09-28 3:27 ` [PATCH 3/9] iommu/vt-d: Do not ignore context table copy failures Lu Baolu
2026-09-28 3:27 ` [PATCH 4/9] iommu/vt-d: Handle DID reservation errors when copying context tables Lu Baolu
2026-09-28 3:27 ` [PATCH 5/9] iommu/vt-d: Reserve scalable-mode DIDs from PASID entries during copy Lu Baolu
2026-09-28 3:27 ` [PATCH 6/9] iommu/vt-d: Use old domain parameter when attaching the blocking domain Lu Baolu
2026-09-28 3:27 ` [PATCH 7/9] iommu/vt-d: Fix iopf refcount leak in nested attach Lu Baolu
2026-09-28 3:27 ` [PATCH 8/9] iommu/vt-d: Drop old iopf ref only after attach succeeds Lu Baolu
2026-09-28 3:27 ` Lu Baolu [this message]
2026-09-28 8:18 ` [PATCH 0/9][PULL REQUEST] Intel IOMMU updates for v7.4 Joerg Roedel
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260928032722.2868623-10-baolu.lu@linux.intel.com \
--to=baolu.lu@linux.intel.com \
--cc=guanghuifeng@linux.alibaba.com \
--cc=iommu@lists.linux.dev \
--cc=joro@8bytes.org \
--cc=linux-kernel@vger.kernel.org \
--cc=zhenzhong.duan@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®