mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Bjorn Helgaas <helgaas@kernel.org>
To: Scott Lee <dsix123@gmail.com>
Cc: Bjorn Helgaas <bhelgaas@google.com>,
	Logan Gunthorpe <logang@deltatee.com>,
	linux-pci@vger.kernel.org, linux-kernel@vger.kernel.org
Subject: Re: [PATCH] PCI/P2PDMA: Add Intel Haswell client host bridge to the whitelist
Date: Mon, 5 Oct 2026 10:00:30 -0500	[thread overview]
Message-ID: <20261005150030.GA566065@bhelgaas> (raw)
In-Reply-To: <20261005132651.18368-1-dsix123@gmail.com>

On Mon, Oct 05, 2026 at 09:26:51PM +0800, Scott Lee wrote:
> P2P DMA between two devices that do not share an upstream bridge is only
> permitted when the traffic goes through a host bridge that is listed in
> pci_p2pdma_whitelist[] with REQ_SAME_HOST_BRIDGE.
> 
> The Haswell client host bridge (8086:0c08, e.g. H97/H87/B85) is not in that
> list, and cpu_supports_p2pdma() returns false on Intel, so
> calc_map_type_and_dist() refuses P2P for any pair of devices behind
> different root ports of the same host bridge, even though that root complex
> handles the traffic correctly.
> 
> Measured on an ASUS H97-PRO with a Xeon E3-1231 v3 (Haswell, 8086:0c08 at
> 00:00.0) and two GPUs on separate root ports of that host bridge
> (00:01.0 and 00:1c.4):
> 
>   - pci_p2pdma_distance() < 0 before the change; peer access granted after
>     (amdgpu stops reporting "PCIe P2P access ... is not supported by the
>     chipset" and hipDeviceCanAccessPeer() returns true in both directions)
>   - 10 KB cross-device copy: 62-75 us vs 130-131 us host-staged (~2x)
>   - 512 MB cross-device copy verified byte-wise against a known buffer:
>     no corruption
>   - tensor-parallel LLM inference (>25 GB model across both GPUs):
>     +53% tokens/s
> 
> REQ_SAME_HOST_BRIDGE is the correct flag here: the two devices share the
> upstream bridge (two root ports of one host bridge) rather than sitting
> behind a PCIe switch.
> 
> This entry only permits P2P DMA to be used on this platform, it does not
> force it: the existing checks on ACS redirect, on the map type and on
> pci_p2pdma_distance() all still apply. Tested on one board only, so the
> usual caveat applies - broad testing on other Haswell boards would be
> needed before this can be considered generally safe.
> 
> Signed-off-by: Scott Lee <dsix123@gmail.com>

Applied to pci/p2pdma for v7.4, thanks!

> ---
> # --- evidence from the machine this was tested on (stripped by git am) ---
> #
> # Hardware: ASUS H97-PRO (BIOS 2906), Xeon E3-1231 v3 (Haswell), 32 GB RAM,
> #           2x AMD Radeon RX 9060 XT (16 GB each), no PCIe switch
> #
> # lspci -nn:
> #   00:00.0 Host bridge: Intel 4th Gen Core Processor DRAM Controller [8086:0c08]
> #   00:01.0 PCI bridge: CPU PEG port; GPU0 is 0000:03:00.0 behind it
> #   00:1c.4 PCI bridge: PCH root port; GPU1 is 0000:09:00.0 behind it
> #
> # unpatched kernel (7.0.0-34-generic, distro stock):
> #   $ sudo dmesg | grep 'not supported by the chipset'
> #   amdgpu 0000:09:00.0: PCIe P2P access from peer device 0000:03:00.0 is not
> #   supported by the chipset
> #
> # patched kernel (7.0.0-99-generic, same source tree + only this hunk):
> #   $ sudo dmesg | grep 'not supported by the chipset'      (no output)
> #   $ python3 p2p_verify.py
> #     can_access_peer: True / True
> #     peer 512 MiB copy, byte-wise compared against a known buffer: OK
> #   $ python3 p2p_latency.py     (10 KiB, HIP events)
> #     peer 0->1: 62.1 us    peer 1->0: 74.5 us
> #     host 0->1: 130.0 us   host 1->0: 131.0 us
> #   $ llama-bench, 27B model split across both GPUs, tensor vs layer split:
> #     25.05 +/- 0.91 tok/s vs 16.37 tok/s (+53%), greedy output identical
> #
> # Note the second GPU sits behind the PCH (Gen2 x4) - the setup still works
> # correctly despite the asymmetric and slow link, which is what makes it
> # interesting for the whitelist argument.
> #
>  drivers/pci/p2pdma.c | 2 ++
>  1 file changed, 2 insertions(+)
> 
> diff --git a/drivers/pci/p2pdma.c b/drivers/pci/p2pdma.c
> index 9334eb3..4f58564 100644
> --- a/drivers/pci/p2pdma.c
> +++ b/drivers/pci/p2pdma.c
> @@ -545,6 +545,8 @@ static const struct pci_p2pdma_whitelist_entry {
>  	/* Intel Xeon E7 v3/Xeon E5 v3/Core i7 */
>  	{PCI_VENDOR_ID_INTEL,	0x2f00, REQ_SAME_HOST_BRIDGE},
>  	{PCI_VENDOR_ID_INTEL,	0x2f01, REQ_SAME_HOST_BRIDGE},
> +	/* Intel Haswell (client) */
> +	{PCI_VENDOR_ID_INTEL,	0x0c08, REQ_SAME_HOST_BRIDGE},
>  	/* Intel Skylake-E */
>  	{PCI_VENDOR_ID_INTEL,	0x2030, 0},
>  	{PCI_VENDOR_ID_INTEL,	0x2031, 0},
> -- 
> 2.43.0
> 

      reply	other threads:[~2026-10-05 15:00 UTC|newest]

Thread overview: 2+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-05 13:26 Scott Lee
2026-10-05 15:00 ` Bjorn Helgaas [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261005150030.GA566065@bhelgaas \
    --to=helgaas@kernel.org \
    --cc=bhelgaas@google.com \
    --cc=dsix123@gmail.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-pci@vger.kernel.org \
    --cc=logang@deltatee.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®