From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mgamail.intel.com (mgamail.intel.com [192.198.163.9]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 600FE356760 for ; Mon, 28 Sep 2026 03:39:30 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=192.198.163.9 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790566772; cv=none; b=BW+ufRUSUEb1+8ef66mCGdKpxwC75/WREimfRRq1FK5Dmm4UZVnRv9/EjCsfedoOT+6HiEL+Ht5cZV713fcp22U/FHzcpaLHCYcQZcc89fjsuZwr0/lCVq0d4065uPJuqUqkih7hHSHijO9A5nNYVexb+OAp05+0jRa+O+KgfQc= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790566772; c=relaxed/simple; bh=kxs2ChAqdr9Vuh0TMe9O9iozN0Jgi4CBJP2zi/0AXXo=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=ZZgi4ezdUyKjWpI3+FQ78NGLdEM0obfuTg9G0bo4jBSjTxGUVRnI9QKpGBQKMn9ugbHOa8Io1MTqq8OgZUHCI8d0NchjDudFsTiYp+dU6vIG9UInRaFawU+EUhK+Loz62alD5c1tUrzAk2aWkAkVjV/ELhF2WrS8KweEiYj5zXA= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com; spf=pass smtp.mailfrom=linux.intel.com; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b=Wg2Jd4ck; arc=none smtp.client-ip=192.198.163.9 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b="Wg2Jd4ck" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1790566770; x=1822102770; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=kxs2ChAqdr9Vuh0TMe9O9iozN0Jgi4CBJP2zi/0AXXo=; b=Wg2Jd4ck46n6Dt2BGX7jLo/NQCsDp5hUQ2cb/SXmdkjdP7Dv7WzBTf6x mhN1cABAQqm+2AOVjZGlBleb0z4agjvIZZxrybvW/xM5XDYkYrXYlSQAH G2ke8r0XuHH3UycYFTV0rALFady3efAtNnNrWnm5NqqIg2+AZ/EFzK0EL wozhGjJ1TgaoBf0YT48MtX3yc7NCqFksq10I+ysAwfKNuoWc2GgL8LTZt IH0gnjXycyDkTGzipQzvNCrzG4Sug1qhYVfqQAAVtBvg19/rdW5LlQpK8 PES0k1t3zISC97cPbebotvnIHXzU2zZ1W3dQ4tySAB+Lg0iwpS2eWbCYu w==; X-CSE-ConnectionGUID: i7sFxcShQgujXJNEtN2ADw== X-CSE-MsgGUID: NyO73dPISVC0sAx1HJE3Dw== X-IronPort-AV: E=McAfee;i="6800,10657,11918"; a="101917233" X-IronPort-AV: E=Sophos;i="6.27,127,1787036400"; d="scan'208";a="101917233" Received: from fmviesa005.fm.intel.com ([10.60.135.145]) by fmvoesa103.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 27 Sep 2026 20:39:30 -0700 X-CSE-ConnectionGUID: k4IcTJoxRLKGv6gIHP+/zw== X-CSE-MsgGUID: fglChjFPSKGpQVO2jrvXlA== X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="6.27,127,1787036400"; d="scan'208";a="283004119" Received: from allen-box.sh.intel.com ([10.239.48.101]) by fmviesa005.fm.intel.com with ESMTP; 27 Sep 2026 20:39:28 -0700 From: Lu Baolu To: Joerg Roedel Cc: Guanghui Feng , Zhenzhong Duan , iommu@lists.linux.dev, linux-kernel@vger.kernel.org Subject: [PATCH 9/9] iommu/vt-d: Fix IQE handling to cover all descriptors in submission range Date: Mon, 28 Sep 2026 11:27:22 +0800 Message-ID: <20260928032722.2868623-10-baolu.lu@linux.intel.com> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20260928032722.2868623-1-baolu.lu@linux.intel.com> References: <20260928032722.2868623-1-baolu.lu@linux.intel.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit From: Guanghui Feng When an Invalidation Queue Error (IQE) occurs, hardware halts fetching and IQH points at the faulting descriptor. The previous code only checked whether IQH matched the first descriptor index of the current submission, missing faults on any other descriptor within the batch. Expand the check to cover the entire submission range [index, wait_index], accounting for circular wrap-around. Furthermore, after detecting IQE, the old recovery only replaced the single faulting slot with a copy of the wait descriptor and immediately returned -EINVAL. This left two problems: a) Hardware resumed fetching and could hit another invalid descriptor in the same abandoned batch, raising a second IQE that no submitter would claim — permanently deadlocking the queue. b) The caller reclaimed all batch slots (QI_FREE) while hardware might still be asynchronously processing descriptors from that batch, allowing concurrent overwrite and descriptor corruption. Fix both by introducing qi_drain_remaining_descs(): upon IQE detection, overwrite the stranded slots in [IQH, wait_index) with fenced no-op wait descriptors, resubmit the wait descriptor at wait_index, clear IQE, and spin until hardware signals QI_DONE (with DMAR_OPERATION_TIMEOUT). This guarantees hardware has fully drained the batch before the caller reclaims slots. Fixes: 8a1d82462540 ("iommu/vt-d: Multiple descriptors per qi_submit_sync()") Signed-off-by: Guanghui Feng Signed-off-by: Lu Baolu --- drivers/iommu/intel/dmar.c | 70 ++++++++++++++++++++++++++++++++------ 1 file changed, 60 insertions(+), 10 deletions(-) diff --git a/drivers/iommu/intel/dmar.c b/drivers/iommu/intel/dmar.c index ba675b08cd20..6377ba97a9ea 100644 --- a/drivers/iommu/intel/dmar.c +++ b/drivers/iommu/intel/dmar.c @@ -1344,6 +1344,49 @@ static void qi_dump_fault(struct intel_iommu *iommu, u32 fault) (unsigned long long)desc->qw1); } +static int qi_drain_remaining_descs(struct intel_iommu *iommu, int head, + int wait_index) +{ + struct q_inval *qi = iommu->qi; + int shift = qi_shift(iommu); + struct qi_desc wait_desc = {}; + struct qi_desc nop_desc = {}; + cycles_t start; + + nop_desc.qw0 = QI_IWD_FENCE | QI_IWD_TYPE; + + wait_desc.qw0 = QI_IWD_STATUS_DATA(QI_DONE) | QI_IWD_PRQ_DRAIN | + QI_IWD_STATUS_WRITE | QI_IWD_FENCE | QI_IWD_TYPE; + wait_desc.qw1 = virt_to_phys(&qi->desc_status[wait_index]); + + while (head != wait_index) { + memcpy(qi->desc + (head << shift), &nop_desc, 1 << shift); + head = (head + 1) % QI_LENGTH; + } + + WRITE_ONCE(qi->desc_status[wait_index], QI_IN_USE); + memcpy(qi->desc + (wait_index << shift), &wait_desc, 1 << shift); + + /* + * Order the descriptor rewrites before the writel() that restarts + * the fetch engine. + */ + wmb(); + + writel(DMA_FSTS_IQE, iommu->reg + DMAR_FSTS_REG); + + start = get_cycles(); + while (READ_ONCE(qi->desc_status[wait_index]) != QI_DONE) { + if (DMAR_OPERATION_TIMEOUT < (get_cycles() - start)) { + pr_err("Timeout draining invalidation queue after IQE\n"); + return -ETIMEDOUT; + } + cpu_relax(); + } + + return 0; +} + static int qi_check_fault(struct intel_iommu *iommu, int index, int wait_index) { u32 fault; @@ -1366,18 +1409,25 @@ static int qi_check_fault(struct intel_iommu *iommu, int index, int wait_index) * is cleared. */ if (fault & DMA_FSTS_IQE) { + int head_idx, ret; + head = readl(iommu->reg + DMAR_IQH_REG); - if ((head >> shift) == index) { - struct qi_desc *desc = qi->desc + head; + head_idx = (head >> shift) % QI_LENGTH; + + /* + * The faulting descriptor can be anywhere within the current + * submission's range [index, wait_index]. Since the queue is + * circular, this submission may wrap around QI_LENGTH + * (index > wait_index in that case), so check both the + * non-wrapped and wrapped cases of the range. + */ + if (index <= wait_index ? + (head_idx >= index && head_idx <= wait_index) : + (head_idx >= index || head_idx <= wait_index)) { + ret = qi_drain_remaining_descs(iommu, head_idx, wait_index); + if (ret) + return ret; - /* - * desc->qw2 and desc->qw3 are either reserved or - * used by software as private data. We won't print - * out these two qw's for security consideration. - */ - memcpy(desc, qi->desc + (wait_index << shift), - 1 << shift); - writel(DMA_FSTS_IQE, iommu->reg + DMAR_FSTS_REG); pr_info("Invalidation Queue Error (IQE) cleared\n"); return -EINVAL; } -- 2.43.0