mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Lu Baolu <baolu.lu@linux.intel.com>
To: Joerg Roedel <joro@8bytes.org>
Cc: Guanghui Feng <guanghuifeng@linux.alibaba.com>,
	Zhenzhong Duan <zhenzhong.duan@intel.com>,
	iommu@lists.linux.dev, linux-kernel@vger.kernel.org
Subject: [PATCH 1/9] iommu/vt-d: Fix page table level calculation in compute_vasz_lg2_ss()
Date: Mon, 28 Sep 2026 11:27:14 +0800	[thread overview]
Message-ID: <20260928032722.2868623-2-baolu.lu@linux.intel.com> (raw)
In-Reply-To: <20260928032722.2868623-1-baolu.lu@linux.intel.com>

From: Zhenzhong Duan <zhenzhong.duan@intel.com>

compute_vasz_lg2_ss() finds the optimal Second-Stage page table level by
intersecting the maximum guest address width (mgaw) with the hardware's
SAGAW capability register.

The VT-d spec maps the SAGAW bit field positions as:
  - Bit 1: 39-bit AGAW (3-level page table, top_level = 2)
  - Bit 2: 48-bit AGAW (4-level page table, top_level = 3)
  - Bit 3: 57-bit AGAW (5-level page table, top_level = 4)

The fallback paths use bit shifts that are one position too large,
causing ffs() to select a deeper page table level than the mgaw window
requires:

  - mgaw > 39: "3 + ffs(sagaw >> 3)" evaluates to top_level = 4 (5-level)
    instead of top_level = 3 (4-level) when hardware supports both
    48-bit (Bit 2) and 57-bit (Bit 3) AGAW.
  - mgaw > 30: "2 + ffs(sagaw >> 2)" evaluates to top_level = 3 (4-level)
    instead of top_level = 2 (3-level) when hardware supports both
    39-bit (Bit 1) and 48-bit (Bit 2) AGAW.

In both cases the selected level is still one that the hardware advertises
in its SAGAW capability, so IOVA translation remains functionally correct.
However, an unnecessarily deep page table may be selected, adding an extra
level of page walk overhead and reducing TLB and cache efficiency without
providing any increase in addressable IOVA space beyond what the mgaw
window already caps.

Fix by decreasing the shift offset by one in each fallback case, ensuring
ffs() targets the correct SAGAW bit position and selects the smallest
page table level that fully covers the mgaw range:

  - mgaw > 39: "2 + ffs(sagaw >> 2)" correctly yields top_level = 3
  - mgaw > 30: "1 + ffs(sagaw >> 1)" correctly yields top_level = 2

Fixes: d856f9d27885 ("iommupt/vtd: Allow VT-d to have a larger table top than the vasz requires")
Signed-off-by: Zhenzhong Duan <zhenzhong.duan@intel.com>
Reviewed-by: Jason Gunthorpe <jgg@nvidia.com>
Signed-off-by: Lu Baolu <baolu.lu@linux.intel.com>
---
 drivers/iommu/intel/iommu.c | 4 ++--
 1 file changed, 2 insertions(+), 2 deletions(-)

diff --git a/drivers/iommu/intel/iommu.c b/drivers/iommu/intel/iommu.c
index 2e3b3ab216f8..05f351833d0b 100644
--- a/drivers/iommu/intel/iommu.c
+++ b/drivers/iommu/intel/iommu.c
@@ -2911,10 +2911,10 @@ static unsigned int compute_vasz_lg2_ss(struct intel_iommu *iommu,
 		*top_level = 4;
 		return min(57, mgaw);
 	} else if (mgaw > 39 && sagaw >= BIT(2)) {
-		*top_level = 3 + ffs(sagaw >> 3);
+		*top_level = 2 + ffs(sagaw >> 2);
 		return min(48, mgaw);
 	} else if (mgaw > 30 && sagaw >= BIT(1)) {
-		*top_level = 2 + ffs(sagaw >> 2);
+		*top_level = 1 + ffs(sagaw >> 1);
 		return min(39, mgaw);
 	}
 	return 0;
-- 
2.43.0


  reply	other threads:[~2026-09-28  3:39 UTC|newest]

Thread overview: 11+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-28  3:27 [PATCH 0/9][PULL REQUEST] Intel IOMMU updates for v7.4 Lu Baolu
2026-09-28  3:27 ` Lu Baolu [this message]
2026-09-28  3:27 ` [PATCH 2/9] iommu/vt-d: Avoid out-of-range shift in qi_desc_dev_iotlb_pasid() Lu Baolu
2026-09-28  3:27 ` [PATCH 3/9] iommu/vt-d: Do not ignore context table copy failures Lu Baolu
2026-09-28  3:27 ` [PATCH 4/9] iommu/vt-d: Handle DID reservation errors when copying context tables Lu Baolu
2026-09-28  3:27 ` [PATCH 5/9] iommu/vt-d: Reserve scalable-mode DIDs from PASID entries during copy Lu Baolu
2026-09-28  3:27 ` [PATCH 6/9] iommu/vt-d: Use old domain parameter when attaching the blocking domain Lu Baolu
2026-09-28  3:27 ` [PATCH 7/9] iommu/vt-d: Fix iopf refcount leak in nested attach Lu Baolu
2026-09-28  3:27 ` [PATCH 8/9] iommu/vt-d: Drop old iopf ref only after attach succeeds Lu Baolu
2026-09-28  3:27 ` [PATCH 9/9] iommu/vt-d: Fix IQE handling to cover all descriptors in submission range Lu Baolu
2026-09-28  8:18 ` [PATCH 0/9][PULL REQUEST] Intel IOMMU updates for v7.4 Joerg Roedel

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260928032722.2868623-2-baolu.lu@linux.intel.com \
    --to=baolu.lu@linux.intel.com \
    --cc=guanghuifeng@linux.alibaba.com \
    --cc=iommu@lists.linux.dev \
    --cc=joro@8bytes.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=zhenzhong.duan@intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®