mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH 0/2] PCI/P2PDMA: Allow P2PDMA in VMware and Hyper-V guests on Intel hosts
@ 2026-10-09 15:40 Pavel Popov
  2026-10-09 15:40 ` [PATCH 1/2] PCI/P2PDMA: Allow P2PDMA in VMware " Pavel Popov
                   ` (2 more replies)
  0 siblings, 3 replies; 11+ messages in thread
From: Pavel Popov @ 2026-10-09 15:40 UTC (permalink / raw)
  To: Bjorn Helgaas, Logan Gunthorpe
  Cc: pavel.e.popov, Jim Chow, Radu Rugina, Alexey Makhalov, Wei Liu,
	Michael Kelley, Lukas Wunner, Leon Romanovsky, Nathan Ciobanu,
	bcm-kernel-feedback-list, virtualization, linux-hyperv,
	linux-pci, linux-kernel

VMware and Hyper-V expose passthrough devices to guest VMs without an
explicit host bridge device. Consequently, host_bridge_whitelist()
cannot automatically identify the underlying host bridge structure,
causing Peer-to-Peer DMA (P2PDMA) requests to be rejected unless the
device IDs are explicitly listed in pci_p2pdma_whitelist[]. To enable
P2PDMA support in virtualized environments on Intel hosts, this series
introduces hypervisor_supports_p2pdma().

Although alternative P2PDMA mechanisms are in development [1], their
adoption timeline in hypervisors remains uncertain. This solution
addresses the immediate need to support both new and existing
hypervisor deployments.

The patches are organized by hypervisor for clarity. While the primary
focus is VMware, a corresponding Hyper-V patch addresses the same
pattern and is included for maintainer consideration.

[1] https://lore.kernel.org/r/20260812-hmat-p2p-v1-0-75ac41380585@nvidia.com

Pavel Popov (2):
  PCI/P2PDMA: Allow P2PDMA in VMware guests on Intel hosts
  PCI/P2PDMA: Allow P2PDMA in Hyper-V guests on Intel hosts

 drivers/pci/p2pdma.c | 21 +++++++++++++++++++++
 1 file changed, 21 insertions(+)


base-commit: 6c377d19d4a5116d9bec5203aa3c6c11523e7898
-- 
2.53.0


^ permalink raw reply	[flat|nested] 11+ messages in thread

* [PATCH 1/2] PCI/P2PDMA: Allow P2PDMA in VMware guests on Intel hosts
  2026-10-09 15:40 [PATCH 0/2] PCI/P2PDMA: Allow P2PDMA in VMware and Hyper-V guests on Intel hosts Pavel Popov
@ 2026-10-09 15:40 ` Pavel Popov
  2026-10-09 16:38   ` Logan Gunthorpe
  2026-10-09 15:40 ` [PATCH 2/2] PCI/P2PDMA: Allow P2PDMA in Hyper-V " Pavel Popov
  2026-10-09 16:58 ` [PATCH 0/2] PCI/P2PDMA: Allow P2PDMA in VMware and " Bjorn Helgaas
  2 siblings, 1 reply; 11+ messages in thread
From: Pavel Popov @ 2026-10-09 15:40 UTC (permalink / raw)
  To: Bjorn Helgaas, Logan Gunthorpe
  Cc: pavel.e.popov, Jim Chow, Radu Rugina, Alexey Makhalov,
	bcm-kernel-feedback-list, linux-pci, linux-kernel

Peer-to-Peer DMA (P2PDMA) between two devices requires either a shared
upstream bridge or a host bridge matching pci_p2pdma_whitelist[]. In
VMware guests, passthrough devices appear directly on virtual root
buses without host bridge devices. Consequently, P2PDMA checks
evaluate endpoint Device IDs against the host bridge whitelist,
causing P2PDMA requests to fail regardless of host support.

This affects discrete devices such as GPUs, NICs and NVMe drives.

The P2PDMA host bridge check does not work in a VMware guest. Bypass it on
Intel hosts under VMware, starting with Skylake-SP, the first Xeon
Scalable generation. Xeon Scalable introduced UPI and is whitelisted for
P2PDMA without restrictions. AMD Zen and newer are already covered by
cpu_supports_p2pdma(), so the bypass is implemented for Intel hosts only.

Tested with the Intel Level Zero P2P benchmark (ze_peer) on Arc Pro B60
cards in a Linux guest on VMware ESXi 9.0 (Emerald Rapids host).

Reviewed-by: Jim Chow <jim.chow@broadcom.com>
Signed-off-by: Pavel Popov <pavel.e.popov@intel.com>
---
 drivers/pci/p2pdma.c | 20 ++++++++++++++++++++
 1 file changed, 20 insertions(+)

diff --git a/drivers/pci/p2pdma.c b/drivers/pci/p2pdma.c
index 9334eb314663..d96f44d0738a 100644
--- a/drivers/pci/p2pdma.c
+++ b/drivers/pci/p2pdma.c
@@ -21,6 +21,12 @@
 #include <linux/seq_buf.h>
 #include <linux/xarray.h>
 
+#ifdef CONFIG_X86
+#include <asm/cpu_device_id.h>
+#include <asm/hypervisor.h>
+#include <asm/intel-family.h>
+#endif
+
 struct pci_p2pdma {
 	struct gen_pool *pool;
 	bool p2pmem_published;
@@ -532,6 +538,19 @@ static bool cpu_supports_p2pdma(void)
 	return false;
 }
 
+static bool hypervisor_supports_p2pdma(void)
+{
+#ifdef CONFIG_X86
+	/* Hypervisor hides host topology, allow p2pdma on Skylake and newer */
+	if (boot_cpu_data.x86_vendor == X86_VENDOR_INTEL &&
+	    boot_cpu_data.x86_vfm >= INTEL_SKYLAKE_X &&
+	    hypervisor_is_type(X86_HYPER_VMWARE))
+		return true;
+#endif
+
+	return false;
+}
+
 static const struct pci_p2pdma_whitelist_entry {
 	unsigned short vendor;
 	int device;
@@ -781,6 +800,7 @@ calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client,
 
 map_through_host_bridge:
 	if (!cpu_supports_p2pdma() &&
+	    !hypervisor_supports_p2pdma() &&
 	    !host_bridge_whitelist(provider, client, acs_redirects)) {
 		if (verbose)
 			pci_warn(client, "cannot be used for peer-to-peer DMA as the client and provider (%s) do not share an upstream bridge or whitelisted host bridge\n",
-- 
2.53.0


^ permalink raw reply	[flat|nested] 11+ messages in thread

* [PATCH 2/2] PCI/P2PDMA: Allow P2PDMA in Hyper-V guests on Intel hosts
  2026-10-09 15:40 [PATCH 0/2] PCI/P2PDMA: Allow P2PDMA in VMware and Hyper-V guests on Intel hosts Pavel Popov
  2026-10-09 15:40 ` [PATCH 1/2] PCI/P2PDMA: Allow P2PDMA in VMware " Pavel Popov
@ 2026-10-09 15:40 ` Pavel Popov
  2026-10-09 16:58 ` [PATCH 0/2] PCI/P2PDMA: Allow P2PDMA in VMware and " Bjorn Helgaas
  2 siblings, 0 replies; 11+ messages in thread
From: Pavel Popov @ 2026-10-09 15:40 UTC (permalink / raw)
  To: Bjorn Helgaas, Logan Gunthorpe
  Cc: pavel.e.popov, Wei Liu, Michael Kelley, Jim Chow, linux-hyperv,
	linux-pci, linux-kernel

Peer-to-Peer DMA (P2PDMA) between two devices requires either a shared
upstream bridge or a host bridge matching pci_p2pdma_whitelist[]. In
Hyper-V guests, each passthrough device appears in a PCI domain of its
own without a host bridge device. Consequently, P2PDMA checks evaluate
endpoint Device IDs against the host bridge whitelist, causing P2PDMA
requests to fail regardless of host support.

This affects discrete devices such as GPUs, NICs and NVMe drives.

The P2PDMA host bridge check does not work in a Hyper-V guest. Bypass it
on Intel hosts under Hyper-V, starting with Skylake-SP, the first Xeon
Scalable generation. Xeon Scalable introduced UPI and is whitelisted for
P2PDMA without restrictions. AMD Zen and newer are already covered by
cpu_supports_p2pdma(), so the bypass is implemented for Intel hosts only.

Tested with the Intel Level Zero P2P benchmark (ze_peer) on Arc Pro B60
cards in a Linux guest on Windows Server 2025 (Emerald Rapids host).

Reviewed-by: Jim Chow <jim.chow@broadcom.com>
Signed-off-by: Pavel Popov <pavel.e.popov@intel.com>
---
 drivers/pci/p2pdma.c | 3 ++-
 1 file changed, 2 insertions(+), 1 deletion(-)

diff --git a/drivers/pci/p2pdma.c b/drivers/pci/p2pdma.c
index d96f44d0738a..76fadfbc1814 100644
--- a/drivers/pci/p2pdma.c
+++ b/drivers/pci/p2pdma.c
@@ -544,7 +544,8 @@ static bool hypervisor_supports_p2pdma(void)
 	/* Hypervisor hides host topology, allow p2pdma on Skylake and newer */
 	if (boot_cpu_data.x86_vendor == X86_VENDOR_INTEL &&
 	    boot_cpu_data.x86_vfm >= INTEL_SKYLAKE_X &&
-	    hypervisor_is_type(X86_HYPER_VMWARE))
+	    (hypervisor_is_type(X86_HYPER_VMWARE) ||
+	     hypervisor_is_type(X86_HYPER_MS_HYPERV)))
 		return true;
 #endif
 
-- 
2.53.0


^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH 1/2] PCI/P2PDMA: Allow P2PDMA in VMware guests on Intel hosts
  2026-10-09 15:40 ` [PATCH 1/2] PCI/P2PDMA: Allow P2PDMA in VMware " Pavel Popov
@ 2026-10-09 16:38   ` Logan Gunthorpe
  2026-10-09 16:49     ` Popov, Pavel E
  0 siblings, 1 reply; 11+ messages in thread
From: Logan Gunthorpe @ 2026-10-09 16:38 UTC (permalink / raw)
  To: Pavel Popov, Bjorn Helgaas
  Cc: Jim Chow, Radu Rugina, Alexey Makhalov, bcm-kernel-feedback-list,
	linux-pci, linux-kernel



On 2026-10-09 9:40 a.m., Pavel Popov wrote:
> Peer-to-Peer DMA (P2PDMA) between two devices requires either a shared
> upstream bridge or a host bridge matching pci_p2pdma_whitelist[]. In
> VMware guests, passthrough devices appear directly on virtual root
> buses without host bridge devices. Consequently, P2PDMA checks
> evaluate endpoint Device IDs against the host bridge whitelist,
> causing P2PDMA requests to fail regardless of host support.
> 
> This affects discrete devices such as GPUs, NICs and NVMe drives.
> 
> The P2PDMA host bridge check does not work in a VMware guest. Bypass it on
> Intel hosts under VMware, starting with Skylake-SP, the first Xeon
> Scalable generation. Xeon Scalable introduced UPI and is whitelisted for
> P2PDMA without restrictions. AMD Zen and newer are already covered by
> cpu_supports_p2pdma(), so the bypass is implemented for Intel hosts only.
> 
> Tested with the Intel Level Zero P2P benchmark (ze_peer) on Arc Pro B60
> cards in a Linux guest on VMware ESXi 9.0 (Emerald Rapids host).
> 
> Reviewed-by: Jim Chow <jim.chow@broadcom.com>
> Signed-off-by: Pavel Popov <pavel.e.popov@intel.com>
> ---
>  drivers/pci/p2pdma.c | 20 ++++++++++++++++++++
>  1 file changed, 20 insertions(+)
> 
> diff --git a/drivers/pci/p2pdma.c b/drivers/pci/p2pdma.c
> index 9334eb314663..d96f44d0738a 100644
> --- a/drivers/pci/p2pdma.c
> +++ b/drivers/pci/p2pdma.c
> @@ -21,6 +21,12 @@
>  #include <linux/seq_buf.h>
>  #include <linux/xarray.h>
>  
> +#ifdef CONFIG_X86
> +#include <asm/cpu_device_id.h>
> +#include <asm/hypervisor.h>
> +#include <asm/intel-family.h>
> +#endif
> +
>  struct pci_p2pdma {
>  	struct gen_pool *pool;
>  	bool p2pmem_published;
> @@ -532,6 +538,19 @@ static bool cpu_supports_p2pdma(void)
>  	return false;
>  }
>  
> +static bool hypervisor_supports_p2pdma(void)
> +{
> +#ifdef CONFIG_X86
> +	/* Hypervisor hides host topology, allow p2pdma on Skylake and newer */
> +	if (boot_cpu_data.x86_vendor == X86_VENDOR_INTEL &&
> +	    boot_cpu_data.x86_vfm >= INTEL_SKYLAKE_X &&
> +	    hypervisor_is_type(X86_HYPER_VMWARE))
> +		return true;
> +#endif
> +
> +	return false;
> +}

The changes seem fine, but one minor nit from me. Can we avoid putting
the #ifdefs inside the function?

Something like:

#ifdef CONFIG_X86
static bool hypervisor_supports_p2pdma(void)
{
...
}
#else /* !CONFIG_X86 */
static bool hypervisor_supports_p2pdma(void)
{
	return false;
}
#endif /* CONFIG_X86 */


Other than that, the core concept seems fine to me:

Reviewed-by: Logan Gunthorpe <logang@deltatee.com>

Thanks!

Logan

^ permalink raw reply	[flat|nested] 11+ messages in thread

* RE: [PATCH 1/2] PCI/P2PDMA: Allow P2PDMA in VMware guests on Intel hosts
  2026-10-09 16:38   ` Logan Gunthorpe
@ 2026-10-09 16:49     ` Popov, Pavel E
  0 siblings, 0 replies; 11+ messages in thread
From: Popov, Pavel E @ 2026-10-09 16:49 UTC (permalink / raw)
  To: Logan Gunthorpe, Bjorn Helgaas
  Cc: Jim Chow, Radu Rugina, Alexey Makhalov, bcm-kernel-feedback-list,
	linux-pci, linux-kernel

> The changes seem fine, but one minor nit from me. Can we avoid putting
> the #ifdefs inside the function?

Sure, will do in v2. Thanks for the review.

^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH 0/2] PCI/P2PDMA: Allow P2PDMA in VMware and Hyper-V guests on Intel hosts
  2026-10-09 15:40 [PATCH 0/2] PCI/P2PDMA: Allow P2PDMA in VMware and Hyper-V guests on Intel hosts Pavel Popov
  2026-10-09 15:40 ` [PATCH 1/2] PCI/P2PDMA: Allow P2PDMA in VMware " Pavel Popov
  2026-10-09 15:40 ` [PATCH 2/2] PCI/P2PDMA: Allow P2PDMA in Hyper-V " Pavel Popov
@ 2026-10-09 16:58 ` Bjorn Helgaas
  2026-10-09 18:08   ` Leon Romanovsky
  2 siblings, 1 reply; 11+ messages in thread
From: Bjorn Helgaas @ 2026-10-09 16:58 UTC (permalink / raw)
  To: Pavel Popov
  Cc: Bjorn Helgaas, Logan Gunthorpe, Jim Chow, Radu Rugina,
	Alexey Makhalov, Wei Liu, Michael Kelley, Lukas Wunner,
	Leon Romanovsky, Nathan Ciobanu, bcm-kernel-feedback-list,
	virtualization, linux-hyperv, linux-pci, linux-kernel

On Fri, Oct 09, 2026 at 08:40:32AM -0700, Pavel Popov wrote:
> VMware and Hyper-V expose passthrough devices to guest VMs without an
> explicit host bridge device. Consequently, host_bridge_whitelist()
> cannot automatically identify the underlying host bridge structure,
> causing Peer-to-Peer DMA (P2PDMA) requests to be rejected unless the
> device IDs are explicitly listed in pci_p2pdma_whitelist[]. To enable
> P2PDMA support in virtualized environments on Intel hosts, this series
> introduces hypervisor_supports_p2pdma().
> 
> Although alternative P2PDMA mechanisms are in development [1], their
> adoption timeline in hypervisors remains uncertain. This solution
> addresses the immediate need to support both new and existing
> hypervisor deployments.
> 
> The patches are organized by hypervisor for clarity. While the primary
> focus is VMware, a corresponding Hyper-V patch addresses the same
> pattern and is included for maintainer consideration.
> 
> [1] https://lore.kernel.org/r/20260812-hmat-p2p-v1-0-75ac41380585@nvidia.com
> 
> Pavel Popov (2):
>   PCI/P2PDMA: Allow P2PDMA in VMware guests on Intel hosts
>   PCI/P2PDMA: Allow P2PDMA in Hyper-V guests on Intel hosts
> 
>  drivers/pci/p2pdma.c | 21 +++++++++++++++++++++
>  1 file changed, 21 insertions(+)

Applied with the #ifdef tweak Logan suggested to pci/p2pdma for v7.4,
thanks!

FWIW, it's better if responses like Jim's Reviewed-by actually appear
on the mailing list.  I don't see anything in lore, so I assume it
must have been private review.

^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH 0/2] PCI/P2PDMA: Allow P2PDMA in VMware and Hyper-V guests on Intel hosts
  2026-10-09 16:58 ` [PATCH 0/2] PCI/P2PDMA: Allow P2PDMA in VMware and " Bjorn Helgaas
@ 2026-10-09 18:08   ` Leon Romanovsky
  2026-10-09 18:21     ` Popov, Pavel E
  2026-10-09 19:12     ` Bjorn Helgaas
  0 siblings, 2 replies; 11+ messages in thread
From: Leon Romanovsky @ 2026-10-09 18:08 UTC (permalink / raw)
  To: Bjorn Helgaas
  Cc: Pavel Popov, Bjorn Helgaas, Logan Gunthorpe, Jim Chow,
	Radu Rugina, Alexey Makhalov, Wei Liu, Michael Kelley,
	Lukas Wunner, Nathan Ciobanu, bcm-kernel-feedback-list,
	virtualization, linux-hyperv, linux-pci, linux-kernel

On Fri, Oct 09, 2026 at 11:58:33AM -0500, Bjorn Helgaas wrote:
> On Fri, Oct 09, 2026 at 08:40:32AM -0700, Pavel Popov wrote:
> > VMware and Hyper-V expose passthrough devices to guest VMs without an
> > explicit host bridge device. Consequently, host_bridge_whitelist()
> > cannot automatically identify the underlying host bridge structure,
> > causing Peer-to-Peer DMA (P2PDMA) requests to be rejected unless the
> > device IDs are explicitly listed in pci_p2pdma_whitelist[]. To enable
> > P2PDMA support in virtualized environments on Intel hosts, this series
> > introduces hypervisor_supports_p2pdma().
> > 
> > Although alternative P2PDMA mechanisms are in development [1], their
> > adoption timeline in hypervisors remains uncertain. This solution
> > addresses the immediate need to support both new and existing
> > hypervisor deployments.
> > 
> > The patches are organized by hypervisor for clarity. While the primary
> > focus is VMware, a corresponding Hyper-V patch addresses the same
> > pattern and is included for maintainer consideration.
> > 
> > [1] https://lore.kernel.org/r/20260812-hmat-p2p-v1-0-75ac41380585@nvidia.com
> > 
> > Pavel Popov (2):
> >   PCI/P2PDMA: Allow P2PDMA in VMware guests on Intel hosts
> >   PCI/P2PDMA: Allow P2PDMA in Hyper-V guests on Intel hosts
> > 
> >  drivers/pci/p2pdma.c | 21 +++++++++++++++++++++
> >  1 file changed, 21 insertions(+)
> 
> Applied with the #ifdef tweak Logan suggested to pci/p2pdma for v7.4,
> thanks!

Bjorn,

I disagree with this decision for several reasons:

1. Note where `hypervisor_supports_p2pdma()` is called: before
   `host_bridge_whitelist()`. This means the code ignores the hypervisor
   topology. The claim that the VM has a virtual bridge, causing
   `pci_p2pdma_whitelist()` to fail, describes exactly how P2P is expected
   to work today.

2. We are not developing an alternative solution. This is the right
   solution, and there is broad agreement that it is the only reliable
   way to enable P2P in VMs, for ALL emulation software stacks.

3. This problem has existed and been known for at least the past eight
   years, since Logan upstreamed P2P support. The claim "we need it now
   and ASAP" is not valid at all.

4. The proposed hack does not solve the P2P-in-VM problem; it only makes it work
   in some random cases. An HMAT-based solution will still be needed, even on systems
   that use this hack.

Let's get the HMAT solution merged sooner rather than later. To make that happen,
we need a coordinated effort across the industry, not hacks to work around the
problem.

> 
> FWIW, it's better if responses like Jim's Reviewed-by actually appear
> on the mailing list.  I don't see anything in lore, so I assume it
> must have been private review.

Unfortunately, this patch series was not given enough time for review.

Thanks

> 

^ permalink raw reply	[flat|nested] 11+ messages in thread

* RE: [PATCH 0/2] PCI/P2PDMA: Allow P2PDMA in VMware and Hyper-V guests on Intel hosts
  2026-10-09 18:08   ` Leon Romanovsky
@ 2026-10-09 18:21     ` Popov, Pavel E
  2026-10-09 18:49       ` Leon Romanovsky
  2026-10-09 19:12     ` Bjorn Helgaas
  1 sibling, 1 reply; 11+ messages in thread
From: Popov, Pavel E @ 2026-10-09 18:21 UTC (permalink / raw)
  To: Leon Romanovsky, Bjorn Helgaas, Jim Chow
  Cc: Bjorn Helgaas, Logan Gunthorpe, Radu Rugina, Alexey Makhalov,
	Wei Liu, Michael Kelley, Lukas Wunner, Nathan Ciobanu,
	bcm-kernel-feedback-list, virtualization, linux-hyperv,
	linux-pci, linux-kernel

> 1. Note where `hypervisor_supports_p2pdma()` is called: before
>    `host_bridge_whitelist()`. This means the code ignores the hypervisor
>    topology. The claim that the VM has a virtual bridge, causing
>    `pci_p2pdma_whitelist()` to fail, describes exactly how P2P is expected
>    to work today.

It does not ignore the topology. The check is reached only after the
common-upstream-bridge walk and the ACS checks have failed, at the point
where the host bridge allowlist is consulted. It skips that lookup on
the matched hypervisors, the same way cpu_supports_p2pdma() already
skips it for AMD.

The lookup cannot work there because VMware and Hyper-V expose no host
bridge device at all on the root buses carrying passthrough devices, so
pci_host_bridge_dev() returns the passthrough endpoint itself. That is
described in the commit messages.


> 2. We are not developing an alternative solution. This is the right
>    solution, and there is broad agreement that it is the only reliable
>    way to enable P2P in VMs, for ALL emulation software stacks.

Agreed, and the cover letter says so. Our understanding is that HMAT
becomes the primary source of P2PDMA information and the allowlist
remains as a fallback; the hypervisor check is part of that fallback,
not a competing path. When the HMAT lookup lands it takes precedence in
the same decision path, and this check only matters where no HMAT is
exposed. The adoption timeline of HMAT in hypervisors is uncertain, and
guests on existing hypervisor releases will not get it at all, so the
fallback is needed for both new and existing deployments.


> 3. This problem has existed and been known for at least the past eight
>    years, since Logan upstreamed P2P support. The claim "we need it now
>    and ASAP" is not valid at all.

The problem is old, the demand is not. Passing several accelerators
through to one guest and expecting them to talk to each other is a
recent requirement, and today P2PDMA does not work at all for
passthrough devices in VMware or Hyper-V guests on Intel hosts. The
series does not claim urgency beyond that: a working path exists now,
HMAT does not yet, and the two do not conflict.


> 4. The proposed hack does not solve the P2P-in-VM problem; it only makes it work
>    in some random cases. An HMAT-based solution will still be needed, even on systems
>    that use this hack.

It makes it work on Xeon Scalable hosts under VMware and Hyper-V, which
is where passthrough accelerators are deployed today. Nothing more is
claimed. I agree HMAT is still needed on top for latency, bandwidth and
ordering attributes.

^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH 0/2] PCI/P2PDMA: Allow P2PDMA in VMware and Hyper-V guests on Intel hosts
  2026-10-09 18:21     ` Popov, Pavel E
@ 2026-10-09 18:49       ` Leon Romanovsky
  2026-10-09 19:13         ` Popov, Pavel E
  0 siblings, 1 reply; 11+ messages in thread
From: Leon Romanovsky @ 2026-10-09 18:49 UTC (permalink / raw)
  To: Popov, Pavel E
  Cc: Bjorn Helgaas, Jim Chow, Bjorn Helgaas, Logan Gunthorpe,
	Radu Rugina, Alexey Makhalov, Wei Liu, Michael Kelley,
	Lukas Wunner, Nathan Ciobanu, bcm-kernel-feedback-list,
	virtualization, linux-hyperv, linux-pci, linux-kernel

On Fri, Oct 09, 2026 at 06:21:49PM +0000, Popov, Pavel E wrote:
> > 1. Note where `hypervisor_supports_p2pdma()` is called: before
> >    `host_bridge_whitelist()`. This means the code ignores the hypervisor
> >    topology. The claim that the VM has a virtual bridge, causing
> >    `pci_p2pdma_whitelist()` to fail, describes exactly how P2P is expected
> >    to work today.
> 
> It does not ignore the topology. The check is reached only after the
> common-upstream-bridge walk and the ACS checks have failed, at the point
> where the host bridge allowlist is consulted.

Since the request comes from a VM, we always take the "skip" path,
regardless of the topology.

<...>

> 
> > 2. We are not developing an alternative solution. This is the right
> >    solution, and there is broad agreement that it is the only reliable
> >    way to enable P2P in VMs, for ALL emulation software stacks.
> 
> Agreed, and the cover letter says so. Our understanding is that HMAT
> becomes the primary source of P2PDMA information and the allowlist
> remains as a fallback; the hypervisor check is part of that fallback,
> not a competing path. When the HMAT lookup lands it takes precedence in
> the same decision path, and this check only matters where no HMAT is
> exposed. The adoption timeline of HMAT in hypervisors is uncertain, and
> guests on existing hypervisor releases will not get it at all, so the
> fallback is needed for both new and existing deployments.

I won't worry for hypervisors, they have enough brilliant developers to
take the upstream code to their codebase.

> 
> 
> > 3. This problem has existed and been known for at least the past eight
> >    years, since Logan upstreamed P2P support. The claim "we need it now
> >    and ASAP" is not valid at all.
> 
> The problem is old, the demand is not. Passing several accelerators
> through to one guest and expecting them to talk to each other is a
> recent requirement

This is not correct, at least for mlx5 devices. The need was already recognized
in 2020 in commit 90da7dc8206a ("RDMA/mlx5: Support dma-buf based userspace
memory region").

> and today P2PDMA does not work at all for
> passthrough devices in VMware or Hyper-V guests on Intel hosts. The
> series does not claim urgency beyond that: a working path exists now,
> HMAT does not yet, and the two do not conflict.
> 
> 
> > 4. The proposed hack does not solve the P2P-in-VM problem; it only makes it work
> >    in some random cases. An HMAT-based solution will still be needed, even on systems
> >    that use this hack.
> 
> It makes it work on Xeon Scalable hosts under VMware and Hyper-V, which
> is where passthrough accelerators are deployed today.

This is a very Intel-centric view. The vast majority of systems probably run on ARM
and use QEMU.

> Nothing more is claimed. I agree HMAT is still needed on top for latency, bandwidth and
> ordering attributes.
> 

The primary goal of HMAT is to eliminate the need for whitelists and to
simulate a P2P route for the VM that accounts for the hypervisor topology.
Latency and bandwidth are secondary considerations.

Thanks

^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH 0/2] PCI/P2PDMA: Allow P2PDMA in VMware and Hyper-V guests on Intel hosts
  2026-10-09 18:08   ` Leon Romanovsky
  2026-10-09 18:21     ` Popov, Pavel E
@ 2026-10-09 19:12     ` Bjorn Helgaas
  1 sibling, 0 replies; 11+ messages in thread
From: Bjorn Helgaas @ 2026-10-09 19:12 UTC (permalink / raw)
  To: Leon Romanovsky
  Cc: Pavel Popov, Bjorn Helgaas, Logan Gunthorpe, Jim Chow,
	Radu Rugina, Alexey Makhalov, Wei Liu, Michael Kelley,
	Lukas Wunner, Nathan Ciobanu, bcm-kernel-feedback-list,
	virtualization, linux-hyperv, linux-pci, linux-kernel

On Fri, Oct 09, 2026 at 09:08:47PM +0300, Leon Romanovsky wrote:
> On Fri, Oct 09, 2026 at 11:58:33AM -0500, Bjorn Helgaas wrote:
> > On Fri, Oct 09, 2026 at 08:40:32AM -0700, Pavel Popov wrote:
> > > VMware and Hyper-V expose passthrough devices to guest VMs without an
> > > explicit host bridge device. Consequently, host_bridge_whitelist()
> > > cannot automatically identify the underlying host bridge structure,
> > > causing Peer-to-Peer DMA (P2PDMA) requests to be rejected unless the
> > > device IDs are explicitly listed in pci_p2pdma_whitelist[]. To enable
> > > P2PDMA support in virtualized environments on Intel hosts, this series
> > > introduces hypervisor_supports_p2pdma().
> > > 
> > > Although alternative P2PDMA mechanisms are in development [1], their
> > > adoption timeline in hypervisors remains uncertain. This solution
> > > addresses the immediate need to support both new and existing
> > > hypervisor deployments.
> > > 
> > > The patches are organized by hypervisor for clarity. While the primary
> > > focus is VMware, a corresponding Hyper-V patch addresses the same
> > > pattern and is included for maintainer consideration.
> > > 
> > > [1] https://lore.kernel.org/r/20260812-hmat-p2p-v1-0-75ac41380585@nvidia.com
> > > 
> > > Pavel Popov (2):
> > >   PCI/P2PDMA: Allow P2PDMA in VMware guests on Intel hosts
> > >   PCI/P2PDMA: Allow P2PDMA in Hyper-V guests on Intel hosts
> > > 
> > >  drivers/pci/p2pdma.c | 21 +++++++++++++++++++++
> > >  1 file changed, 21 insertions(+)
> > 
> > Applied with the #ifdef tweak Logan suggested to pci/p2pdma for v7.4,
> > thanks!
> 
> Bjorn,
> 
> I disagree with this decision for several reasons:
> 
> 1. Note where `hypervisor_supports_p2pdma()` is called: before
>    `host_bridge_whitelist()`. This means the code ignores the hypervisor
>    topology. The claim that the VM has a virtual bridge, causing
>    `pci_p2pdma_whitelist()` to fail, describes exactly how P2P is expected
>    to work today.
> 
> 2. We are not developing an alternative solution. This is the right
>    solution, and there is broad agreement that it is the only reliable
>    way to enable P2P in VMs, for ALL emulation software stacks.
> 
> 3. This problem has existed and been known for at least the past eight
>    years, since Logan upstreamed P2P support. The claim "we need it now
>    and ASAP" is not valid at all.
> 
> 4. The proposed hack does not solve the P2P-in-VM problem; it only makes it work
>    in some random cases. An HMAT-based solution will still be needed, even on systems
>    that use this hack.
> 
> Let's get the HMAT solution merged sooner rather than later. To make that happen,
> we need a coordinated effort across the industry, not hacks to work around the
> problem.

OK, I'll drop this while we discuss it.

^ permalink raw reply	[flat|nested] 11+ messages in thread

* RE: [PATCH 0/2] PCI/P2PDMA: Allow P2PDMA in VMware and Hyper-V guests on Intel hosts
  2026-10-09 18:49       ` Leon Romanovsky
@ 2026-10-09 19:13         ` Popov, Pavel E
  0 siblings, 0 replies; 11+ messages in thread
From: Popov, Pavel E @ 2026-10-09 19:13 UTC (permalink / raw)
  To: Leon Romanovsky, Bjorn Helgaas, Jim Chow, Bjorn Helgaas, Logan Gunthorpe
  Cc: Radu Rugina, Alexey Makhalov, Wei Liu, Michael Kelley,
	Lukas Wunner, Nathan Ciobanu, bcm-kernel-feedback-list,
	virtualization, linux-hyperv, linux-pci, linux-kernel

Leon,

Understood on all points, and I agree HMAT is the proper fix. We discussed
it with Lukas Wunner as well. Exposing HMAT is on the hypervisor side; Jim
Chow from VMware can speak to their plans. Even once HMAT lands in Linux,
it only helps guests after the hypervisors add support for it.

The key point for us is that today there is no path at all for
passthrough devices on VMware or Hyper-V, and that is what this patch
addresses. We reviewed it with VMware and aligned on this approach for the
guest in the meantime.

Is there anything you would recommend to make the series acceptable as an
interim? Bjorn, Logan, your view would help here as well.

^ permalink raw reply	[flat|nested] 11+ messages in thread

end of thread, other threads:[~2026-10-09 19:13 UTC | newest]

Thread overview: 11+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-09 15:40 [PATCH 0/2] PCI/P2PDMA: Allow P2PDMA in VMware and Hyper-V guests on Intel hosts Pavel Popov
2026-10-09 15:40 ` [PATCH 1/2] PCI/P2PDMA: Allow P2PDMA in VMware " Pavel Popov
2026-10-09 16:38   ` Logan Gunthorpe
2026-10-09 16:49     ` Popov, Pavel E
2026-10-09 15:40 ` [PATCH 2/2] PCI/P2PDMA: Allow P2PDMA in Hyper-V " Pavel Popov
2026-10-09 16:58 ` [PATCH 0/2] PCI/P2PDMA: Allow P2PDMA in VMware and " Bjorn Helgaas
2026-10-09 18:08   ` Leon Romanovsky
2026-10-09 18:21     ` Popov, Pavel E
2026-10-09 18:49       ` Leon Romanovsky
2026-10-09 19:13         ` Popov, Pavel E
2026-10-09 19:12     ` Bjorn Helgaas

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®