* [PATCH V2 0/4] Hyper-V: root VM iommu kernel only driver
@ 2026-09-29 22:36 Mukesh R
2026-09-29 22:36 ` [PATCH V2 1/4] mshv: Add gfp_flags parameter to hv_call_deposit_pages() Mukesh R
` (3 more replies)
0 siblings, 4 replies; 5+ messages in thread
From: Mukesh R @ 2026-09-29 22:36 UTC (permalink / raw)
To: linux-hyperv, linux-kernel, iommu, linux-arch
Cc: jgg, jacob.pan, kys, haiyangz, wei.liu, decui, tglx, mingo, bp,
dave.hansen, x86, hpa, joro, will, robin.murphy, arnd, mrathor
This short series introduces hyperv root VM iommu kernel only driver. This
driver acts as liason between the hypervisor and iommu layer. Details are
both in the actual commit and in the comments in the file implementing it.
Thanks,
-Mukesh
V0: based on upstream hyperv-next commit be0cfab740e5
V1:
- change order of lookup tree deletions to appease all AI agents,
carving out hv_iommu_hyp_unmap_pages() from hv_iommu_unmap_pages().
- change unique_id to ida subsystem management so we never ever run
out of domain ids.
- change hv_max_iova_width to hv_iommu_max_iova and generate proper
mask.
- some other minor changes.
V2:
- Add patch (0001) to allow gfp_flags to be passed to deposit pages
function.
- In last main patch:
o remove hv_iommu_setup_skip to skip devices being managed
o remove hv_special_domain()
o Add comments around pgsize_bitmap and IOMMU capabilities.
o Add to lookup tree before mapping iova in the map callback
o Unmap iova before removing from the tree in the unmap callback
Mukesh R (4):
mshv: Add gfp_flags parameter to hv_call_deposit_pages()
PCI: hv: Export hv_build_devid_type_pci() and change return type
mshv: Import data structs around device domains from hyperv headers
x86/hyperv: Implement Hyper-V virtual IOMMU
arch/x86/hyperv/irqdomain.c | 9 +-
arch/x86/include/asm/mshyperv.h | 6 +
arch/x86/kernel/pci-dma.c | 2 +
drivers/hv/hv_proc.c | 14 +-
drivers/hv/mshv_root_hv_call.c | 6 +-
drivers/iommu/Kconfig | 1 +
drivers/iommu/hyperv/Kconfig | 15 +
drivers/iommu/hyperv/Makefile | 1 +
drivers/iommu/hyperv/hv-iommu-root.c | 611 +++++++++++++++++++++++++++
include/asm-generic/mshyperv.h | 9 +-
include/hyperv/hvgdk_mini.h | 9 +
include/hyperv/hvhdk_mini.h | 90 ++++
include/linux/hyperv.h | 6 +
13 files changed, 764 insertions(+), 15 deletions(-)
create mode 100644 drivers/iommu/hyperv/Kconfig
create mode 100644 drivers/iommu/hyperv/hv-iommu-root.c
--
2.51.2.vfs.0.1
^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH V2 1/4] mshv: Add gfp_flags parameter to hv_call_deposit_pages()
2026-09-29 22:36 [PATCH V2 0/4] Hyper-V: root VM iommu kernel only driver Mukesh R
@ 2026-09-29 22:36 ` Mukesh R
2026-09-29 22:36 ` [PATCH V2 2/4] PCI: hv: Export hv_build_devid_type_pci() and change return type Mukesh R
` (2 subsequent siblings)
3 siblings, 0 replies; 5+ messages in thread
From: Mukesh R @ 2026-09-29 22:36 UTC (permalink / raw)
To: linux-hyperv, linux-kernel, iommu, linux-arch
Cc: jgg, jacob.pan, kys, haiyangz, wei.liu, decui, tglx, mingo, bp,
dave.hansen, x86, hpa, joro, will, robin.murphy, arnd, mrathor
Currently GFP_KERNEL is hard coded in hv_call_deposit_pages(). Add
gfp_flags parameter so it can be called in atomic context from
map_pages callback in struct iommu_domain_ops and other places in
future.
Signed-off-by: Mukesh R <mrathor@linux.microsoft.com>
---
drivers/hv/hv_proc.c | 14 +++++++-------
drivers/hv/mshv_root_hv_call.c | 6 ++++--
include/asm-generic/mshyperv.h | 6 ++++--
3 files changed, 15 insertions(+), 11 deletions(-)
diff --git a/drivers/hv/hv_proc.c b/drivers/hv/hv_proc.c
index 57b2c64197cb..ce4276d37d45 100644
--- a/drivers/hv/hv_proc.c
+++ b/drivers/hv/hv_proc.c
@@ -15,8 +15,8 @@
*/
#define HV_DEPOSIT_MAX (HV_HYP_PAGE_SIZE / sizeof(u64) - 1)
-/* Deposits exact number of pages. Must be called with interrupts enabled. */
-int hv_call_deposit_pages(int node, u64 partition_id, u32 num_pages)
+int hv_call_deposit_pages(int node, u64 partition_id, u32 num_pages,
+ gfp_t gfp_flags)
{
struct page **pages, *page;
int *counts;
@@ -35,12 +35,12 @@ int hv_call_deposit_pages(int node, u64 partition_id, u32 num_pages)
return 0;
/* One buffer for page pointers and counts */
- page = alloc_page(GFP_KERNEL);
+ page = alloc_page(gfp_flags);
if (!page)
return -ENOMEM;
pages = page_address(page);
- counts = kzalloc_objs(int, HV_DEPOSIT_MAX);
+ counts = kzalloc_objs(int, HV_DEPOSIT_MAX, gfp_flags);
if (!counts) {
free_page((unsigned long)pages);
return -ENOMEM;
@@ -54,7 +54,7 @@ int hv_call_deposit_pages(int node, u64 partition_id, u32 num_pages)
order = 31 - __builtin_clz(num_pages);
while (1) {
- pages[i] = alloc_pages_node(node, GFP_KERNEL, order);
+ pages[i] = alloc_pages_node(node, gfp_flags, order);
if (pages[i])
break;
if (!order) {
@@ -137,7 +137,7 @@ int hv_deposit_memory_node(int node, u64 partition_id,
hv_status_err(hv_status, "Unexpected!\n");
return -ENOMEM;
}
- return hv_call_deposit_pages(node, partition_id, num_pages);
+ return hv_call_deposit_pages(node, partition_id, num_pages, GFP_KERNEL);
}
EXPORT_SYMBOL_GPL(hv_deposit_memory_node);
@@ -206,7 +206,7 @@ int hv_call_create_vp(int node, u64 partition_id, u32 vp_index, u32 flags)
/* Root VPs don't seem to need pages deposited */
if (partition_id != hv_current_partition_id) {
/* The value 90 is empirically determined. It may change. */
- ret = hv_call_deposit_pages(node, partition_id, 90);
+ ret = hv_call_deposit_pages(node, partition_id, 90, GFP_KERNEL);
if (ret)
return ret;
}
diff --git a/drivers/hv/mshv_root_hv_call.c b/drivers/hv/mshv_root_hv_call.c
index cb55d4d4be2e..bd975239f3bc 100644
--- a/drivers/hv/mshv_root_hv_call.c
+++ b/drivers/hv/mshv_root_hv_call.c
@@ -141,7 +141,8 @@ int hv_call_initialize_partition(u64 partition_id)
input.partition_id = partition_id;
ret = hv_call_deposit_pages(NUMA_NO_NODE, partition_id,
- HV_INIT_PARTITION_DEPOSIT_PAGES);
+ HV_INIT_PARTITION_DEPOSIT_PAGES,
+ GFP_KERNEL);
if (ret)
return ret;
@@ -249,7 +250,8 @@ static int hv_do_map_gpa_hcall(u64 partition_id, u64 gfn, u64 page_struct_count,
if (hv_result_needs_memory(status)) {
ret = hv_call_deposit_pages(NUMA_NO_NODE, partition_id,
- HV_MAP_GPA_DEPOSIT_PAGES);
+ HV_MAP_GPA_DEPOSIT_PAGES,
+ GFP_KERNEL);
if (ret)
break;
diff --git a/include/asm-generic/mshyperv.h b/include/asm-generic/mshyperv.h
index bf601d67cecb..af65dd3e3725 100644
--- a/include/asm-generic/mshyperv.h
+++ b/include/asm-generic/mshyperv.h
@@ -345,7 +345,8 @@ static inline bool hv_parent_partition(void)
bool hv_result_needs_memory(u64 status);
int hv_deposit_memory_node(int node, u64 partition_id, u64 status);
-int hv_call_deposit_pages(int node, u64 partition_id, u32 num_pages);
+int hv_call_deposit_pages(int node, u64 partition_id, u32 num_pages,
+ gfp_t gfp_flags);
int hv_call_add_logical_proc(int node, u32 lp_index, u32 acpi_id);
int hv_call_notify_all_processors_started(void);
bool hv_lp_exists(u32 lp_index);
@@ -360,7 +361,8 @@ static inline int hv_deposit_memory_node(int node, u64 partition_id, u64 status)
{
return -EOPNOTSUPP;
}
-static inline int hv_call_deposit_pages(int node, u64 partition_id, u32 num_pages)
+static inline int hv_call_deposit_pages(int node, u64 partition_id,
+ u32 num_pages, gfp_t gfp_flags)
{
return -EOPNOTSUPP;
}
--
2.51.2.vfs.0.1
^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH V2 2/4] PCI: hv: Export hv_build_devid_type_pci() and change return type
2026-09-29 22:36 [PATCH V2 0/4] Hyper-V: root VM iommu kernel only driver Mukesh R
2026-09-29 22:36 ` [PATCH V2 1/4] mshv: Add gfp_flags parameter to hv_call_deposit_pages() Mukesh R
@ 2026-09-29 22:36 ` Mukesh R
2026-09-29 22:36 ` [PATCH V2 3/4] mshv: Import data structs around device domains from hyperv headers Mukesh R
2026-09-29 22:36 ` [PATCH V2 4/4] x86/hyperv: Implement root partition IOMMU kernel only driver Mukesh R
3 siblings, 0 replies; 5+ messages in thread
From: Mukesh R @ 2026-09-29 22:36 UTC (permalink / raw)
To: linux-hyperv, linux-kernel, iommu, linux-arch
Cc: jgg, jacob.pan, kys, haiyangz, wei.liu, decui, tglx, mingo, bp,
dave.hansen, x86, hpa, joro, will, robin.murphy, arnd, mrathor
On Hyper-V, most hypercalls related to device domain management like
attach device to a device domain, or to map/unmap regions, interrupts,
etc need a device ID as a parameter. So, make hv_build_devid_type_pci()
public and change return type to u64 to enforce it's size.
Signed-off-by: Mukesh R <mrathor@linux.microsoft.com>
Reviewed-by: Souradeep Chakrabarti <schakrabarti@linux.microsoft.com>
---
arch/x86/hyperv/irqdomain.c | 9 +++++----
arch/x86/include/asm/mshyperv.h | 6 ++++++
2 files changed, 11 insertions(+), 4 deletions(-)
diff --git a/arch/x86/hyperv/irqdomain.c b/arch/x86/hyperv/irqdomain.c
index b3ad50a874dc..8780573a4332 100644
--- a/arch/x86/hyperv/irqdomain.c
+++ b/arch/x86/hyperv/irqdomain.c
@@ -112,7 +112,7 @@ static int get_rid_cb(struct pci_dev *pdev, u16 alias, void *data)
return 0;
}
-static union hv_device_id hv_build_devid_type_pci(struct pci_dev *pdev)
+u64 hv_build_devid_type_pci(struct pci_dev *pdev)
{
int pos;
union hv_device_id hv_devid;
@@ -172,8 +172,9 @@ static union hv_device_id hv_build_devid_type_pci(struct pci_dev *pdev)
}
out:
- return hv_devid;
+ return hv_devid.as_uint64;
}
+EXPORT_SYMBOL_GPL(hv_build_devid_type_pci);
/*
* hv_map_msi_interrupt() - Map the MSI IRQ in the hypervisor.
@@ -196,7 +197,7 @@ int hv_map_msi_interrupt(struct irq_data *data,
msidesc = irq_data_get_msi_desc(data);
pdev = msi_desc_to_pci_dev(msidesc);
- hv_devid = hv_build_devid_type_pci(pdev);
+ hv_devid.as_uint64 = hv_build_devid_type_pci(pdev);
cpu = cpumask_first(irq_data_get_effective_affinity_mask(data));
return hv_map_interrupt(hv_devid, false, cpu, cfg->vector,
@@ -271,7 +272,7 @@ static int hv_unmap_msi_interrupt(struct pci_dev *pdev,
{
union hv_device_id hv_devid;
- hv_devid = hv_build_devid_type_pci(pdev);
+ hv_devid.as_uint64 = hv_build_devid_type_pci(pdev);
return hv_unmap_interrupt(hv_devid.as_uint64, irq_entry);
}
diff --git a/arch/x86/include/asm/mshyperv.h b/arch/x86/include/asm/mshyperv.h
index f64393e853ee..8ebbd1cb7c8c 100644
--- a/arch/x86/include/asm/mshyperv.h
+++ b/arch/x86/include/asm/mshyperv.h
@@ -248,6 +248,12 @@ void hv_crash_asm_end(void);
static inline void hv_root_crash_init(void) {}
#endif /* CONFIG_MSHV_ROOT && CONFIG_CRASH_DUMP */
+#ifdef CONFIG_PCI_MSI
+u64 hv_build_devid_type_pci(struct pci_dev *pdev);
+#else
+static inline u64 hv_build_devid_type_pci(struct pci_dev *pdev) { return 0; }
+#endif
+
#else /* CONFIG_HYPERV */
static inline void hyperv_init(void) {}
static inline void hyperv_setup_mmu_ops(void) {}
--
2.51.2.vfs.0.1
^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH V2 3/4] mshv: Import data structs around device domains from hyperv headers
2026-09-29 22:36 [PATCH V2 0/4] Hyper-V: root VM iommu kernel only driver Mukesh R
2026-09-29 22:36 ` [PATCH V2 1/4] mshv: Add gfp_flags parameter to hv_call_deposit_pages() Mukesh R
2026-09-29 22:36 ` [PATCH V2 2/4] PCI: hv: Export hv_build_devid_type_pci() and change return type Mukesh R
@ 2026-09-29 22:36 ` Mukesh R
2026-09-29 22:36 ` [PATCH V2 4/4] x86/hyperv: Implement root partition IOMMU kernel only driver Mukesh R
3 siblings, 0 replies; 5+ messages in thread
From: Mukesh R @ 2026-09-29 22:36 UTC (permalink / raw)
To: linux-hyperv, linux-kernel, iommu, linux-arch
Cc: jgg, jacob.pan, kys, haiyangz, wei.liu, decui, tglx, mingo, bp,
dave.hansen, x86, hpa, joro, will, robin.murphy, arnd, mrathor
Copy/import from Hyper-V public headers, definitions and declarations that
are related to creating iommu domains in the hypervisor, attaching devices
to them, doing the reverse, etc.
Signed-off-by: Mukesh R <mrathor@linux.microsoft.com>
---
include/hyperv/hvgdk_mini.h | 9 ++++
include/hyperv/hvhdk_mini.h | 87 +++++++++++++++++++++++++++++++++++++
2 files changed, 96 insertions(+)
diff --git a/include/hyperv/hvgdk_mini.h b/include/hyperv/hvgdk_mini.h
index 6a4e8b9d570f..3b74f3c0229d 100644
--- a/include/hyperv/hvgdk_mini.h
+++ b/include/hyperv/hvgdk_mini.h
@@ -326,6 +326,8 @@ union hv_hypervisor_version_info {
/* stimer Direct Mode is available */
#define HV_STIMER_DIRECT_MODE_AVAILABLE BIT(19)
+#define HV_DEVICE_DOMAIN_AVAILABLE BIT(24)
+
/*
* Implementation recommendations. Indicates which behaviors the hypervisor
* recommends the OS implement for optimal performance.
@@ -486,9 +488,15 @@ union hv_vp_assist_msr_contents { /* HV_REGISTER_VP_ASSIST_PAGE */
#define HVCALL_GET_VP_INDEX_FROM_APIC_ID 0x009a
#define HVCALL_FLUSH_GUEST_PHYSICAL_ADDRESS_SPACE 0x00af
#define HVCALL_FLUSH_GUEST_PHYSICAL_ADDRESS_LIST 0x00b0
+#define HVCALL_CREATE_DEVICE_DOMAIN 0x00b1
+#define HVCALL_ATTACH_DEVICE_DOMAIN 0x00b2
+#define HVCALL_MAP_DEVICE_GPA_PAGES 0x00b3
+#define HVCALL_UNMAP_DEVICE_GPA_PAGES 0x00b4
#define HVCALL_SIGNAL_EVENT_DIRECT 0x00c0
#define HVCALL_POST_MESSAGE_DIRECT 0x00c1
#define HVCALL_DISPATCH_VP 0x00c2
+#define HVCALL_DETACH_DEVICE_DOMAIN 0x00c4
+#define HVCALL_DELETE_DEVICE_DOMAIN 0x00c5
#define HVCALL_GET_GPA_PAGES_ACCESS_STATES 0x00c9
#define HVCALL_ACQUIRE_SPARSE_SPA_PAGE_HOST_ACCESS 0x00d7
#define HVCALL_RELEASE_SPARSE_SPA_PAGE_HOST_ACCESS 0x00d8
@@ -502,6 +510,7 @@ union hv_vp_assist_msr_contents { /* HV_REGISTER_VP_ASSIST_PAGE */
#define HVCALL_MMIO_READ 0x0106
#define HVCALL_MMIO_WRITE 0x0107
#define HVCALL_DISABLE_HYP_EX 0x010f
+#define HVCALL_GET_IOMMU_CAPABILITIES 0x0125
#define HVCALL_MAP_STATS_PAGE2 0x0131
/* HV_HYPERCALL_INPUT */
diff --git a/include/hyperv/hvhdk_mini.h b/include/hyperv/hvhdk_mini.h
index 035ba20870f7..3dbfb338bcba 100644
--- a/include/hyperv/hvhdk_mini.h
+++ b/include/hyperv/hvhdk_mini.h
@@ -548,4 +548,91 @@ union hv_device_id { /* HV_DEVICE_ID */
} acpi;
} __packed;
+#define HV_DEVICE_DOMAIN_TYPE_S2 0 /* HV_DEVICE_DOMAIN_ID_TYPE_S2 */
+#define HV_DEVICE_DOMAIN_TYPE_S1 1 /* HV_DEVICE_DOMAIN_ID_TYPE_S1 */
+
+#define HV_DEVICE_DOMAIN_ID_S2_DEFAULT 0
+#define HV_DEVICE_DOMAIN_ID_S2_NULL 0xFFFFFFFFULL
+
+union hv_device_domain_id {
+ u64 as_uint64;
+ struct {
+ u32 type : 4;
+ u32 reserved : 28;
+ u32 id;
+ };
+} __packed;
+
+struct hv_input_device_domain { /* HV_INPUT_DEVICE_DOMAIN */
+ u64 partition_id;
+ union hv_input_vtl owner_vtl;
+ u8 padding[7];
+ union hv_device_domain_id domain_id;
+} __packed;
+
+union hv_create_device_domain_flags { /* HV_CREATE_DEVICE_DOMAIN_FLAGS */
+ u32 as_uint32;
+ struct {
+ u32 forward_progress_required : 1;
+ u32 inherit_owning_vtl : 1;
+ u32 reserved : 30;
+ } __packed;
+} __packed;
+
+struct hv_input_create_device_domain { /* HV_INPUT_CREATE_DEVICE_DOMAIN */
+ struct hv_input_device_domain device_domain;
+ union hv_create_device_domain_flags create_device_domain_flags;
+} __packed;
+
+struct hv_input_delete_device_domain { /* HV_INPUT_DELETE_DEVICE_DOMAIN */
+ struct hv_input_device_domain device_domain;
+} __packed;
+
+struct hv_input_attach_device_domain { /* HV_INPUT_ATTACH_DEVICE_DOMAIN */
+ struct hv_input_device_domain device_domain;
+ union hv_device_id device_id;
+} __packed;
+
+struct hv_input_detach_device_domain { /* HV_INPUT_DETACH_DEVICE_DOMAIN */
+ u64 partition_id;
+ union hv_device_id device_id;
+} __packed;
+
+/* HV_INPUT_GET_IOMMU_CAPABILITIES */
+struct hv_input_get_iommu_capabilities {
+ u64 partition_id;
+ u64 reserved;
+} __packed;
+
+#define HV_IOMMU_CAP_PRESENT BIT_ULL(0)
+#define HV_IOMMU_CAP_S2 BIT_ULL(1)
+#define HV_IOMMU_CAP_S1 BIT_ULL(2)
+#define HV_IOMMU_CAP_S1_5LVL BIT_ULL(3)
+#define HV_IOMMU_CAP_PASID BIT_ULL(4)
+#define HV_IOMMU_CAP_ATS BIT_ULL(5)
+#define HV_IOMMU_CAP_PRI BIT_ULL(6)
+
+struct hv_output_get_iommu_capabilities { /* HV_OUTPUT_GET_IOMMU_CAPABILITIES */
+ u32 size;
+ u16 reserved;
+ u8 max_iova_width;
+ u8 max_pasid_width;
+ u64 iommu_cap; /* HV_IOMMU_CAP_* above */
+ u64 pgsize_bitmap;
+} __packed;
+
+struct hv_input_map_device_gpa_pages { /* HV_INPUT_MAP_DEVICE_GPA_PAGES */
+ struct hv_input_device_domain device_domain;
+ union hv_input_vtl target_vtl;
+ u8 padding[3];
+ u32 map_flags;
+ u64 target_device_va_base;
+ u64 gpa_page_list[];
+} __packed;
+
+struct hv_input_unmap_device_gpa_pages { /* HV_INPUT_UNMAP_DEVICE_GPA_PAGES */
+ struct hv_input_device_domain device_domain;
+ u64 target_device_va_base;
+} __packed;
+
#endif /* _HV_HVHDK_MINI_H */
--
2.51.2.vfs.0.1
^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH V2 4/4] x86/hyperv: Implement root partition IOMMU kernel only driver
2026-09-29 22:36 [PATCH V2 0/4] Hyper-V: root VM iommu kernel only driver Mukesh R
` (2 preceding siblings ...)
2026-09-29 22:36 ` [PATCH V2 3/4] mshv: Import data structs around device domains from hyperv headers Mukesh R
@ 2026-09-29 22:36 ` Mukesh R
3 siblings, 0 replies; 5+ messages in thread
From: Mukesh R @ 2026-09-29 22:36 UTC (permalink / raw)
To: linux-hyperv, linux-kernel, iommu, linux-arch
Cc: jgg, jacob.pan, kys, haiyangz, wei.liu, decui, tglx, mingo, bp,
dave.hansen, x86, hpa, joro, will, robin.murphy, arnd, mrathor
Add a new file to implement a kernel only virtual IOMMU that works
with Microsoft Hyper-V hypervisor (aka MSHV) on privileged VMs aka
root VMs. The hypervisor claims the IOMMU upon boot, and this driver
communicates with it for creating and deleting paging domains, attaching
of devices, mapping and unmapping of pages, etc. During boot, hypervisor
automatically creates identity and blocked domains, so there is no need
to do hypercalls to create them. This is a kernel only driver and only
supported on baremetal root (and not L1VH root) without any guest passthru
support. Support for guest device passthru will be added incrementally.
Signed-off-by: Mukesh R <mrathor@linux.microsoft.com>
---
arch/x86/kernel/pci-dma.c | 2 +
drivers/iommu/Kconfig | 1 +
drivers/iommu/hyperv/Kconfig | 15 +
drivers/iommu/hyperv/Makefile | 1 +
drivers/iommu/hyperv/hv-iommu-root.c | 608 +++++++++++++++++++++++++++
include/asm-generic/mshyperv.h | 3 +
include/linux/hyperv.h | 6 +
7 files changed, 636 insertions(+)
create mode 100644 drivers/iommu/hyperv/Kconfig
create mode 100644 drivers/iommu/hyperv/hv-iommu-root.c
diff --git a/arch/x86/kernel/pci-dma.c b/arch/x86/kernel/pci-dma.c
index 6267363e0189..37fed8c7a8c2 100644
--- a/arch/x86/kernel/pci-dma.c
+++ b/arch/x86/kernel/pci-dma.c
@@ -8,6 +8,7 @@
#include <linux/gfp.h>
#include <linux/pci.h>
#include <linux/amd-iommu.h>
+#include <linux/hyperv.h>
#include <asm/proto.h>
#include <asm/dma.h>
@@ -103,6 +104,7 @@ void __init pci_iommu_alloc(void)
}
pci_swiotlb_detect();
gart_iommu_hole_init();
+ hv_iommu_detect();
amd_iommu_detect();
detect_intel_iommu();
swiotlb_init(x86_swiotlb_enable, x86_swiotlb_flags);
diff --git a/drivers/iommu/Kconfig b/drivers/iommu/Kconfig
index 6e07bd69467a..3e410e0f3e1d 100644
--- a/drivers/iommu/Kconfig
+++ b/drivers/iommu/Kconfig
@@ -197,6 +197,7 @@ source "drivers/iommu/arm/Kconfig"
source "drivers/iommu/intel/Kconfig"
source "drivers/iommu/iommufd/Kconfig"
source "drivers/iommu/riscv/Kconfig"
+source "drivers/iommu/hyperv/Kconfig"
config IRQ_REMAP
bool "Support for Interrupt Remapping"
diff --git a/drivers/iommu/hyperv/Kconfig b/drivers/iommu/hyperv/Kconfig
new file mode 100644
index 000000000000..b9be8de4d901
--- /dev/null
+++ b/drivers/iommu/hyperv/Kconfig
@@ -0,0 +1,15 @@
+# SPDX-License-Identifier: GPL-2.0-only
+# Hyper-V IOMMU support
+
+config HYPERV_ROOT_IOMMU
+ bool "Hyper-V IOMMU Device in root partition"
+ depends on HYPERV && X86 && PCI_MSI && MSHV_ROOT
+ select IOMMU_API
+ default HYPERV
+ help
+ This enables Hyper-V pseudo IOMMU device. When running as privileged
+ VM aka root on Microsoft Hyper-V hypervisor, this must be enabled
+ for doing any PCI passthru of devices to guest VMs. This applies to
+ both PFs and VFs. When enabling this, it is best to disable amd/intel
+ iommus via: intel_iommu=off amd_iommu=off as the hypervisor really
+ owns the iommu.
diff --git a/drivers/iommu/hyperv/Makefile b/drivers/iommu/hyperv/Makefile
index 6ef0ef97f3dd..c7e7d0dac2a4 100644
--- a/drivers/iommu/hyperv/Makefile
+++ b/drivers/iommu/hyperv/Makefile
@@ -1,2 +1,3 @@
# SPDX-License-Identifier: GPL-2.0
obj-$(CONFIG_IRQ_REMAP) += hv-irq-remap-x86.o
+obj-$(CONFIG_HYPERV_ROOT_IOMMU) += hv-iommu-root.o
diff --git a/drivers/iommu/hyperv/hv-iommu-root.c b/drivers/iommu/hyperv/hv-iommu-root.c
new file mode 100644
index 000000000000..2502db5624de
--- /dev/null
+++ b/drivers/iommu/hyperv/hv-iommu-root.c
@@ -0,0 +1,608 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Hyper-V root vIOMMU driver.
+ * Copyright (C) 2026, Microsoft, Inc.
+ */
+#include <linux/pci.h>
+#include <linux/dma-map-ops.h>
+#include <linux/interval_tree.h>
+#include <linux/hyperv.h>
+#include <asm/iommu.h>
+#include <asm/mshyperv.h>
+#include "../dma-iommu.h"
+
+static dma_addr_t hv_iommu_max_iova;
+static struct iommu_domain_ops hv_paging_domain_ops;
+
+/* IOMMU device that we export to the world. MSHV supports max of one */
+static struct iommu_device hv_virt_iommu;
+
+struct hv_iommu_domain {
+ struct iommu_domain iommu_dom;
+ u32 domid_num; /* as opposed to domain_id.type */
+ spinlock_t mappings_lock; /* protects mappings_tree */
+ struct rb_root_cached mappings_tree; /* iova to pa lookup tree */
+};
+
+#define to_hv_domain(d) container_of(d, struct hv_iommu_domain, iommu_dom)
+
+struct hv_iommu_mapping {
+ phys_addr_t paddr;
+ struct interval_tree_node iova;
+ u32 flags;
+};
+
+/*
+ * By default, during boot the hypervisor creates one Stage 2 (S2) default
+ * domain. It has two types:
+ * S2 default: access to entire root partition memory. This for the host
+ * maps to IOMMU_DOMAIN_IDENTITY in the iommu subsystem, and
+ * is called HV_DEVICE_DOMAIN_ID_S2_DEFAULT in the hypervisor.
+ * S2 NULL: Blocks everything, ie, IOMMU_DOMAIN_BLOCKED in linux.
+ */
+
+/*
+ * Create dummy domains to correspond to hypervisor prebuilt default identity
+ * and null domains (dummy because we do not make hypercalls to create them).
+ */
+static struct hv_iommu_domain hv_def_identity_dom;
+static struct hv_iommu_domain hv_def_blocked_dom;
+
+/*
+ * The hypervisor owns the IOMMU. As such it will choose the best page size
+ * and whenever possible use larger page sizes automatically. The map gpa
+ * hypercall to map the iova however currently uses only 4K pfns as input.
+ * So we specify 4K as the IOMMU PAGE_SIZE.
+ */
+#define HV_IOMMU_PGSIZES SZ_4K
+
+static DEFINE_IDA(hv_iommu_domids); /* unique numeric ids for new domains */
+
+static bool hv_iommu_capable(struct device *dev, enum iommu_cap cap)
+{
+ switch (cap) {
+ case IOMMU_CAP_CACHE_COHERENCY:
+ return true;
+ default:
+ return false;
+ }
+}
+
+static int hv_iommu_add_tree_mapping(struct hv_iommu_domain *hvdom, ulong iova,
+ phys_addr_t paddr, size_t size, u32 flags)
+{
+ ulong irqflags;
+ struct hv_iommu_mapping *mapping;
+
+ mapping = kzalloc_obj(struct hv_iommu_mapping, GFP_ATOMIC);
+ if (!mapping)
+ return -ENOMEM;
+
+ mapping->paddr = paddr;
+ mapping->iova.start = iova;
+ mapping->iova.last = iova + size - 1;
+ mapping->flags = flags;
+
+ spin_lock_irqsave(&hvdom->mappings_lock, irqflags);
+ interval_tree_insert(&mapping->iova, &hvdom->mappings_tree);
+ spin_unlock_irqrestore(&hvdom->mappings_lock, irqflags);
+
+ return 0;
+}
+
+/* If size == 0, then last = ULONG_MAX. With iova 0, will remove everything */
+static size_t hv_iommu_del_tree_mappings(struct hv_iommu_domain *hvdom,
+ ulong iova, size_t size)
+{
+ ulong flags;
+ size_t unmapped = 0;
+ ulong last = iova + size - 1;
+ struct hv_iommu_mapping *mapping = NULL;
+ struct interval_tree_node *node, *next;
+
+ spin_lock_irqsave(&hvdom->mappings_lock, flags);
+ next = interval_tree_iter_first(&hvdom->mappings_tree, iova, last);
+ while (next) {
+ node = next;
+ mapping = container_of(node, struct hv_iommu_mapping, iova);
+ next = interval_tree_iter_next(node, iova, last);
+
+ /* Splitting of a mapping is not supported at present */
+ if (mapping->iova.start < iova)
+ break;
+
+ unmapped += mapping->iova.last - mapping->iova.start + 1;
+
+ interval_tree_remove(node, &hvdom->mappings_tree);
+ kfree(mapping);
+ }
+ spin_unlock_irqrestore(&hvdom->mappings_lock, flags);
+
+ return unmapped;
+}
+
+/* Create a new device domain in the hypervisor */
+static int hv_iommu_create_hyp_devdom(struct hv_iommu_domain *hvdom)
+{
+ u64 status;
+ struct hv_input_device_domain *ddp;
+ struct hv_input_create_device_domain *input;
+ ulong flags;
+
+ local_irq_save(flags);
+ input = *this_cpu_ptr(hyperv_pcpu_input_arg);
+ memset(input, 0, sizeof(*input));
+
+ ddp = &input->device_domain;
+ ddp->partition_id = HV_PARTITION_ID_SELF;
+ ddp->domain_id.type = HV_DEVICE_DOMAIN_TYPE_S2;
+ ddp->domain_id.id = hvdom->domid_num;
+
+ input->create_device_domain_flags.forward_progress_required = 1;
+ input->create_device_domain_flags.inherit_owning_vtl = 0;
+
+ status = hv_do_hypercall(HVCALL_CREATE_DEVICE_DOMAIN, input, NULL);
+
+ local_irq_restore(flags);
+
+ if (!hv_result_success(status))
+ hv_status_err(status, "\n");
+
+ return hv_result_to_errno(status);
+}
+
+static struct iommu_domain *hv_iommu_domain_alloc_paging(struct device *dev)
+{
+ struct hv_iommu_domain *hvdom;
+ int rc;
+
+ if (!dev_is_pci(dev))
+ return NULL;
+
+ hvdom = kzalloc_obj(struct hv_iommu_domain);
+ if (hvdom == NULL)
+ return ERR_PTR(-ENOMEM);
+
+ spin_lock_init(&hvdom->mappings_lock);
+ hvdom->mappings_tree = RB_ROOT_CACHED;
+
+ rc = ida_alloc_range(&hv_iommu_domids,
+ HV_DEVICE_DOMAIN_ID_S2_DEFAULT + 1,
+ HV_DEVICE_DOMAIN_ID_S2_NULL - 1, GFP_KERNEL);
+ if (rc < 0)
+ goto out_err;
+
+ hvdom->domid_num = rc;
+ hvdom->iommu_dom.pgsize_bitmap = HV_IOMMU_PGSIZES;
+ hvdom->iommu_dom.geometry.aperture_start = 0;
+ hvdom->iommu_dom.geometry.aperture_end = hv_iommu_max_iova;
+ hvdom->iommu_dom.geometry.force_aperture = true;
+ hvdom->iommu_dom.ops = &hv_paging_domain_ops;
+
+ rc = hv_iommu_create_hyp_devdom(hvdom);
+ if (rc)
+ goto out_err1;
+
+ return &hvdom->iommu_dom;
+
+out_err1:
+ ida_free(&hv_iommu_domids, hvdom->domid_num);
+out_err:
+ kfree(hvdom);
+ return ERR_PTR(rc);
+}
+
+static void hv_iommu_domain_free(struct iommu_domain *immdom)
+{
+ ulong flags;
+ u64 status;
+ struct hv_input_delete_device_domain *input;
+ struct hv_input_device_domain *ddp;
+ struct hv_iommu_domain *hvdom = to_hv_domain(immdom);
+
+ /* Cleanup any remaining. 0 for size results in ULONG_MAX as the last */
+ hv_iommu_del_tree_mappings(hvdom, 0, 0);
+
+ local_irq_save(flags);
+ input = *this_cpu_ptr(hyperv_pcpu_input_arg);
+ ddp = &input->device_domain;
+ memset(input, 0, sizeof(*input));
+
+ ddp->partition_id = HV_PARTITION_ID_SELF;
+ ddp->domain_id.type = HV_DEVICE_DOMAIN_TYPE_S2;
+ ddp->domain_id.id = hvdom->domid_num;
+
+ status = hv_do_hypercall(HVCALL_DELETE_DEVICE_DOMAIN, input,
+ NULL);
+ local_irq_restore(flags);
+
+ if (!hv_result_success(status))
+ hv_status_err(status, "\n");
+
+ ida_free(&hv_iommu_domids, hvdom->domid_num);
+ kfree(hvdom);
+}
+
+/*
+ * Attach a device to the default domain, or the null domain, or to a domain
+ * previously created in the hypervisor.
+ */
+static int hv_iommu_att_dev2dom(struct hv_iommu_domain *hvdom,
+ struct pci_dev *pdev)
+{
+ ulong flags;
+ u64 status;
+ struct hv_input_attach_device_domain *input;
+
+ local_irq_save(flags);
+ input = *this_cpu_ptr(hyperv_pcpu_input_arg);
+ memset(input, 0, sizeof(*input));
+
+ /* For null domain, hvdom->domid_num == HV_DEVICE_DOMAIN_ID_S2_NULL */
+ input->device_domain.partition_id = HV_PARTITION_ID_SELF;
+ input->device_domain.domain_id.type = HV_DEVICE_DOMAIN_TYPE_S2;
+ input->device_domain.domain_id.id = hvdom->domid_num;
+
+ input->device_id.as_uint64 = hv_build_devid_type_pci(pdev);
+
+ status = hv_do_hypercall(HVCALL_ATTACH_DEVICE_DOMAIN, input, NULL);
+ local_irq_restore(flags);
+
+ if (!hv_result_success(status))
+ hv_status_err(status, "\n");
+
+ return hv_result_to_errno(status);
+}
+
+static int hv_iommu_attach_dev(struct iommu_domain *immdom, struct device *dev,
+ struct iommu_domain *old)
+{
+ struct pci_dev *pdev;
+ int rc;
+ struct hv_iommu_domain *hvdom = to_hv_domain(immdom);
+
+ pdev = to_pci_dev(dev);
+
+ rc = hv_iommu_att_dev2dom(hvdom, pdev);
+ if (rc)
+ WARN(1, "Failed to attach pdev:%s\n", pci_name(pdev));
+
+ return rc;
+}
+
+static u64 hv_iommu_unmap_batch(u32 domid_num, ulong iova, u16 count)
+{
+ ulong flags;
+ struct hv_input_unmap_device_gpa_pages *input;
+ u64 status;
+
+ local_irq_save(flags);
+ input = *this_cpu_ptr(hyperv_pcpu_input_arg);
+ memset(input, 0, sizeof(*input));
+
+ input->device_domain.partition_id = HV_PARTITION_ID_SELF;
+ input->device_domain.domain_id.type = HV_DEVICE_DOMAIN_TYPE_S2;
+ input->device_domain.domain_id.id = domid_num;
+ input->target_device_va_base = iova;
+
+ status = hv_do_rep_hypercall(HVCALL_UNMAP_DEVICE_GPA_PAGES, count,
+ 0, input, NULL);
+ local_irq_restore(flags);
+
+ if (!hv_result_success(status))
+ hv_status_err(status, "iova:0x%lx count:0x%x\n", iova, count);
+
+ return status;
+}
+
+static size_t hv_iommu_hyp_unmap_pages(struct iommu_domain *immdom, ulong iova,
+ size_t size)
+{
+ u64 status;
+ struct hv_iommu_domain *hvdom = to_hv_domain(immdom);
+ size_t tot_done = 0;
+ ulong npages = size >> HV_HYP_PAGE_SHIFT;
+
+ while (npages) {
+ int done, count = min(npages, HV_REP_COUNT_MAX);
+
+ status = hv_iommu_unmap_batch(hvdom->domid_num, iova, count);
+
+ done = hv_repcomp(status);
+ tot_done += done;
+ npages -= done;
+ iova += done << HV_HYP_PAGE_SHIFT;
+
+ if (!hv_result_success(status))
+ break;
+ }
+
+ return tot_done << HV_HYP_PAGE_SHIFT;
+}
+
+static size_t hv_iommu_unmap_pages(struct iommu_domain *immdom, ulong iova,
+ size_t pgsize, size_t pgcount,
+ struct iommu_iotlb_gather *gather)
+{
+ struct hv_iommu_domain *hvdom = to_hv_domain(immdom);
+ size_t unmap_sz, size = pgsize * pgcount;
+
+ /* For valid input, the hypervisor guarantees unmap will succeed */
+ unmap_sz = hv_iommu_hyp_unmap_pages(immdom, iova, size);
+ if (unmap_sz != size)
+ WARN(1, "Failed to unmap device gpa(%lx/%lx)\n", unmap_sz,
+ size);
+
+ size = hv_iommu_del_tree_mappings(hvdom, iova, unmap_sz);
+ if (size != unmap_sz)
+ pr_err("%s: could not delete tree mappings (%lx:%lx/%lx)\n",
+ __func__, iova, unmap_sz, size);
+
+ return size;
+}
+
+/* Return: must return exact status from the hypercall without changes */
+static u64 hv_iommu_map_pgs(struct hv_iommu_domain *hvdom, ulong iova,
+ phys_addr_t paddr, ulong npages, u32 map_flags)
+{
+ u64 status;
+ int i;
+ struct hv_input_map_device_gpa_pages *input;
+ unsigned long flags, pfn;
+
+ local_irq_save(flags);
+ input = *this_cpu_ptr(hyperv_pcpu_input_arg);
+ memset(input, 0, sizeof(*input));
+
+ input->device_domain.partition_id = HV_PARTITION_ID_SELF;
+ input->device_domain.domain_id.type = HV_DEVICE_DOMAIN_TYPE_S2;
+ input->device_domain.domain_id.id = hvdom->domid_num;
+ input->map_flags = map_flags;
+ input->target_device_va_base = iova;
+
+ pfn = paddr >> HV_HYP_PAGE_SHIFT;
+ for (i = 0; i < npages; i++, pfn++)
+ input->gpa_page_list[i] = pfn;
+
+ status = hv_do_rep_hypercall(HVCALL_MAP_DEVICE_GPA_PAGES, npages, 0,
+ input, NULL);
+ local_irq_restore(flags);
+
+ if (!hv_result_success(status))
+ hv_status_err(status, "npgs:%lx paddr:%llx iova:%lx mf:0x%x\n",
+ npages, paddr, iova, map_flags);
+ return status;
+}
+
+#define HV_MAP_DEVICE_GPA_BATCH_SIZE \
+ ((HV_HYP_PAGE_SIZE - sizeof(struct hv_input_map_device_gpa_pages)) \
+ / sizeof(u64))
+
+/* NB: cond_resched() is in vfio_iommu_map() */
+static int hv_iommu_map_pages(struct iommu_domain *immdom, ulong iova,
+ phys_addr_t paddr, size_t pgsize, size_t pgcount,
+ int prot, gfp_t gfp, size_t *mapped)
+{
+ u32 map_flags;
+ int ret;
+ u64 status = 0;
+ ulong npages, sav_iova = iova, done = 0;
+ struct hv_iommu_domain *hvdom = to_hv_domain(immdom);
+ size_t size = pgsize * pgcount;
+
+ map_flags = HV_MAP_GPA_READABLE; /* required */
+ map_flags |= prot & IOMMU_WRITE ? HV_MAP_GPA_WRITABLE : 0;
+
+ ret = hv_iommu_add_tree_mapping(hvdom, iova, paddr, size, map_flags);
+ if (ret)
+ return ret;
+
+ npages = size >> HV_HYP_PAGE_SHIFT;
+ while (done < npages) {
+ ulong completed, remain = npages - done;
+
+ remain = min(remain, HV_MAP_DEVICE_GPA_BATCH_SIZE);
+
+ status = hv_iommu_map_pgs(hvdom, iova, paddr, remain,
+ map_flags);
+
+ completed = hv_repcomp(status);
+ done = done + completed;
+ iova = iova + (completed << HV_HYP_PAGE_SHIFT);
+ paddr = paddr + (completed << HV_HYP_PAGE_SHIFT);
+
+ if (hv_result_needs_memory(status)) {
+ ret = hv_call_deposit_pages(NUMA_NO_NODE,
+ hv_current_partition_id,
+ 4, gfp);
+ if (ret)
+ break;
+ continue;
+ }
+ if (!hv_result_success(status))
+ break;
+ }
+
+ if (hv_result_success(status)) {
+ if (mapped)
+ *mapped = size;
+ } else {
+ size_t done_size = done << HV_HYP_PAGE_SHIFT;
+
+ hv_status_err(status, "pgs:%lx/%lx iova:%lx\n", done, npages,
+ sav_iova);
+
+ hv_iommu_del_tree_mappings(hvdom, sav_iova, size);
+
+ /* Hypervisor guarantees unmap will succeed here */
+ hv_iommu_hyp_unmap_pages(immdom, sav_iova, done_size);
+
+ if (mapped)
+ *mapped = 0;
+ }
+
+ return hv_result_to_errno(status);
+}
+
+static phys_addr_t hv_iommu_iova_to_phys(struct iommu_domain *immdom,
+ dma_addr_t iova)
+{
+ unsigned long flags;
+ struct hv_iommu_mapping *mapping;
+ struct interval_tree_node *node;
+ u64 paddr = 0;
+ struct hv_iommu_domain *hvdom = to_hv_domain(immdom);
+
+ spin_lock_irqsave(&hvdom->mappings_lock, flags);
+ node = interval_tree_iter_first(&hvdom->mappings_tree, iova, iova);
+ if (node) {
+ mapping = container_of(node, struct hv_iommu_mapping, iova);
+ paddr = mapping->paddr + (iova - mapping->iova.start);
+ }
+ spin_unlock_irqrestore(&hvdom->mappings_lock, flags);
+
+ return paddr;
+}
+
+static struct iommu_device *hv_iommu_probe_device(struct device *dev)
+{
+ if (!dev_is_pci(dev))
+ return ERR_PTR(-ENODEV);
+
+ return &hv_virt_iommu;
+}
+
+static void hv_iommu_get_resv_regions(struct device *dev,
+ struct list_head *head)
+{
+ struct iommu_resv_region *reg;
+
+ /* reserve the entire LAPIC region */
+ reg = iommu_alloc_resv_region(0xfee00000, SZ_1M, 0, IOMMU_RESV_MSI,
+ GFP_KERNEL);
+ if (reg)
+ list_add_tail(®->list, head);
+}
+
+static struct iommu_domain_ops hv_paging_domain_ops = {
+ .attach_dev = hv_iommu_attach_dev,
+ .map_pages = hv_iommu_map_pages,
+ .unmap_pages = hv_iommu_unmap_pages,
+ .iova_to_phys = hv_iommu_iova_to_phys,
+ .free = hv_iommu_domain_free,
+};
+
+static struct iommu_ops hv_iommu_ops = {
+ .capable = hv_iommu_capable,
+ .domain_alloc_paging = hv_iommu_domain_alloc_paging,
+ .probe_device = hv_iommu_probe_device,
+ .device_group = pci_device_group,
+ .get_resv_regions = hv_iommu_get_resv_regions,
+ .owner = THIS_MODULE,
+ .identity_domain = &hv_def_identity_dom.iommu_dom,
+ .blocked_domain = &hv_def_blocked_dom.iommu_dom,
+ .release_domain = &hv_def_blocked_dom.iommu_dom,
+};
+
+static const struct iommu_domain_ops hv_special_domain_ops = {
+ .attach_dev = hv_iommu_attach_dev,
+};
+
+static void __init hv_initialize_special_domains(void)
+{
+ hv_def_identity_dom.iommu_dom.type = IOMMU_DOMAIN_IDENTITY;
+ hv_def_identity_dom.iommu_dom.ops = &hv_special_domain_ops;
+ hv_def_identity_dom.iommu_dom.owner = &hv_iommu_ops;
+ hv_def_identity_dom.domid_num = HV_DEVICE_DOMAIN_ID_S2_DEFAULT; /* 0 */
+
+ hv_def_blocked_dom.iommu_dom.type = IOMMU_DOMAIN_BLOCKED;
+ hv_def_blocked_dom.iommu_dom.ops = &hv_special_domain_ops;
+ hv_def_blocked_dom.iommu_dom.owner = &hv_iommu_ops;
+ hv_def_blocked_dom.domid_num = HV_DEVICE_DOMAIN_ID_S2_NULL; /* INTMAX */
+}
+
+/*
+ * In general, the hypervisor does not support the capabilities hypercall for
+ * root partitions, an exception was made for max_iova_width. For IOMMU cap,
+ * hv_iommu_detect() below checks for HV_DEVICE_DOMAIN_AVAILABLE which
+ * supersedes HV_IOMMU_CAP_PRESENT. As for the pgsize_bitmap cap, see comment
+ * above HV_IOMMU_PGSIZES.
+ */
+static int hv_iommu_get_caps(struct hv_output_get_iommu_capabilities *caps)
+{
+ u64 status;
+ unsigned long flags;
+ struct hv_input_get_iommu_capabilities *input;
+ struct hv_output_get_iommu_capabilities *output;
+
+ local_irq_save(flags);
+
+ input = *this_cpu_ptr(hyperv_pcpu_input_arg);
+ output = *this_cpu_ptr(hyperv_pcpu_output_arg);
+ memset(input, 0, sizeof(*input));
+ input->partition_id = HV_PARTITION_ID_SELF;
+ status = hv_do_hypercall(HVCALL_GET_IOMMU_CAPABILITIES, input, output);
+ *caps = *output;
+
+ local_irq_restore(flags);
+
+ if (!hv_result_success(status))
+ hv_status_err(status, "\n");
+
+ return hv_result_to_errno(status);
+}
+
+static int __init hv_iommu_init(void)
+{
+ int rc;
+ struct iommu_device *iommup = &hv_virt_iommu;
+ struct hv_output_get_iommu_capabilities caps;
+
+ if (!hv_is_hyperv_initialized())
+ return -ENODEV;
+
+ rc = hv_iommu_get_caps(&caps);
+ if (rc)
+ return rc;
+
+ hv_iommu_max_iova = DMA_BIT_MASK(caps.max_iova_width);
+
+ /* This must come before iommu_device_register() because the latter
+ * calls into the hooks.
+ */
+ hv_initialize_special_domains();
+
+ rc = iommu_device_sysfs_add(iommup, NULL, NULL, "%s", "hyperv-iommu");
+ if (rc) {
+ pr_err("Hyper-V: iommu_device_sysfs_add failed: %d\n", rc);
+ return rc;
+ }
+
+ rc = iommu_device_register(iommup, &hv_iommu_ops, NULL);
+ if (rc) {
+ pr_err("Hyper-V: iommu_device_register failed: %d\n", rc);
+ goto err_sysfs_remove;
+ }
+
+ pr_info("Hyper-V IOMMU initialized\n");
+
+ return 0;
+
+err_sysfs_remove:
+ iommu_device_sysfs_remove(iommup);
+ return rc;
+}
+
+void __init hv_iommu_detect(void)
+{
+ if (no_iommu || iommu_detected || hv_l1vh_partition())
+ return;
+
+ if (!(ms_hyperv.misc_features & HV_DEVICE_DOMAIN_AVAILABLE))
+ return;
+
+ iommu_detected = 1;
+ x86_init.iommu.iommu_init = hv_iommu_init;
+
+ pci_request_acs();
+}
diff --git a/include/asm-generic/mshyperv.h b/include/asm-generic/mshyperv.h
index af65dd3e3725..a35e30b87815 100644
--- a/include/asm-generic/mshyperv.h
+++ b/include/asm-generic/mshyperv.h
@@ -28,6 +28,9 @@
#define VTPM_BASE_ADDRESS 0xfed40000
+#define HV_REP_COUNT_MAX \
+ (HV_HYPERCALL_REP_COMP_MASK >> HV_HYPERCALL_REP_COMP_OFFSET)
+
enum hv_partition_type {
HV_PARTITION_TYPE_GUEST,
HV_PARTITION_TYPE_ROOT,
diff --git a/include/linux/hyperv.h b/include/linux/hyperv.h
index 9e109d91aa14..01a69f88cfa7 100644
--- a/include/linux/hyperv.h
+++ b/include/linux/hyperv.h
@@ -1783,4 +1783,10 @@ static inline unsigned long virt_to_hvpfn(void *addr)
#define HVPFN_DOWN(x) ((x) >> HV_HYP_PAGE_SHIFT)
#define page_to_hvpfn(page) (page_to_pfn(page) * NR_HV_HYP_PAGES_IN_PAGE)
+#ifdef CONFIG_HYPERV_ROOT_IOMMU
+void __init hv_iommu_detect(void);
+#else
+static inline void hv_iommu_detect(void) { }
+#endif /* CONFIG_HYPERV_ROOT_IOMMU */
+
#endif /* _HYPERV_H */
--
2.51.2.vfs.0.1
^ permalink raw reply [flat|nested] 5+ messages in thread
end of thread, other threads:[~2026-09-29 22:36 UTC | newest]
Thread overview: 5+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-29 22:36 [PATCH V2 0/4] Hyper-V: root VM iommu kernel only driver Mukesh R
2026-09-29 22:36 ` [PATCH V2 1/4] mshv: Add gfp_flags parameter to hv_call_deposit_pages() Mukesh R
2026-09-29 22:36 ` [PATCH V2 2/4] PCI: hv: Export hv_build_devid_type_pci() and change return type Mukesh R
2026-09-29 22:36 ` [PATCH V2 3/4] mshv: Import data structs around device domains from hyperv headers Mukesh R
2026-09-29 22:36 ` [PATCH V2 4/4] x86/hyperv: Implement root partition IOMMU kernel only driver Mukesh R
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®