mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: "Christian König" <christian.koenig@amd.com>
To: Leon Romanovsky <leon@kernel.org>
Cc: "Bjorn Helgaas" <bhelgaas@google.com>,
	"Logan Gunthorpe" <logang@deltatee.com>,
	"Jason Gunthorpe" <jgg@ziepe.ca>,
	"Joerg Roedel (AMD)" <joro@8bytes.org>,
	"Will Deacon" <will@kernel.org>,
	"Robin Murphy" <robin.murphy@arm.com>,
	"Thomas Hellström" <thomas.hellstrom@linux.intel.com>,
	linux-pci@vger.kernel.org, linux-kernel@vger.kernel.org,
	linux-doc@vger.kernel.org, iommu@lists.linux.dev,
	"Tushar Dave" <tdave@nvidia.com>,
	linux-media@vger.kernel.org, dri-devel@lists.freedesktop.org,
	linaro-mm-sig@lists.linaro.org, linux-rdma@vger.kernel.org,
	kvm@vger.kernel.org, "Chaitanya Kulkarni" <kch@nvidia.com>,
	"Greg Kroah-Hartman" <gregkh@linuxfoundation.org>,
	"Jens Axboe" <axboe@kernel.dk>,
	"Alex Williamson" <alex@shazbot.org>,
	"Ankit Agrawal" <ankita@nvidia.com>,
	"Jonathan Corbet" <corbet@lwn.net>,
	"Shuah Khan" <skhan@linuxfoundation.org>,
	"Randy Dunlap" <rdunlap@infradead.org>,
	"Sumit Semwal" <sumit.semwal@linaro.org>
Subject: Re: [PATCH v8 16/23] dma-buf: Let importers ask how peer-to-peer traffic is routed
Date: Thu, 1 Oct 2026 09:06:54 +0200	[thread overview]
Message-ID: <f5c5319b-e8fd-4a54-b7ea-6bb64a4ff7e2@amd.com> (raw)
In-Reply-To: <20261001061752.GJ3401365@unreal>

On 10/1/26 08:17, Leon Romanovsky wrote:
> On Wed, Sep 30, 2026 at 04:52:20PM +0200, Christian König wrote:
>> On 9/30/26 16:32, Leon Romanovsky wrote:
>>> On Wed, Sep 30, 2026 at 01:52:39PM +0200, Christian König wrote:
>> ...
>>>>> The importer needs a way to obtain device information from the exporter so
>>>>> that it can configure itself correctly.
>>>>
>>>> That won't work with DMA-buf then, the exporter is completely opaque to the importer and that is for really good reasons.
>>>>
>>>> Why in the world does the importer needs to know the information from the exporter before the mapping is created?
>>>>
>>>> It is the exporter who decides how data is accessed by the importer and not the other way around.
>>>
>>> There are several reasons:
>>>
>>> 1. This is how DMA-BUF MRs are built in RDMA. In mlx5, they rely on the ODP
>>>    mechanism, which requires an MKEY to be created first. See commit
>>>    90da7dc8206a (“RDMA/mlx5: Support dma-buf based userspace memory region”).
>>
>> I know that patch, but I absolutely don't understand where the problem is.
>>
>>> 2. P2P routing is a property of devices, not memory. It is known and remains
>>>    stable.
>>
>> No, absolutely not. You are making completely incorrect assumptions how DMA-buf works.
> 
> Maybe, but I understand how the kernel, PCI, and P2P DMA work, and dma-buf
> seems to operate in a parallel reality that is not aligned with any of them.

Yes, but that behavior predates the P2P DMA work by over a decade.

You can't ignore how existing drivers work just because you want PCI P2P to work this way.

>> Again: What is exported and from which path is up to the exporter and only determined when you create a DMA-buf mapping!
> 
> Let's discuss the PCI P2P case, where `p2pdma_provider` is present. In this
> flow, the exporter has a PCI device, and the importer also has one when it
> attaches.

No, they don't.

> The decision about how to construct the addresses is made at that
> point. Note that both the exporter and importer are PCI P2P devices, so the
> importer should make the request.

Again, no.

The decision where to place things is done when the first mapping is created and that is documented and full intentional behavior since the very first merged version in 2011.

>> In most cases you don't even have a single device which is the DMA-buf exporter.
> 
> How is it possible for exporter with p2pdma_provider?
> 
>>
>> Only while creating the mapping the underlying access path is finally determined. It can be that P2P is used, it can be that internal connections are used, it can be that the data is moved to system memory....
> 
> But we are talking about in-tree exporters: RDMA, VFIO e.t.c
> 
>>
>> That's why we have the distinction between attaching and importer and the importer creating a mapping.
>>
>>> 3. See the VFIO TPH ST discussion, where the requirement to obtain the
>>>    exporter’s P2P information in the importer was raised again.
>>
>> As far as I can see there isn't any. The TPH/ST information are just passed through from the exporter to the importer.
>>
>> We could define a bit better what needs to come first the mapping or the TPH/ST query but for their use case that is actually irrelevant.
> 
> They need an access to p2pdma_provider too.
> https://lore.kernel.org/linux-rdma/CAH3zFs2rhg6-b-pVDps1V8LsM4Dn48t4J4LFVyKOZtT4axdY=Q@mail.gmail.com/

No they don't. They just need the routing information the same way as you do it here.

And Zhiping reply actually makes it 100% clear were the misunderstanding is here:

> Two properties are worth stating explicitly:

>  - The metadata and its callback are per-dmabuf, while routing is per
> attachment, so this deliberately uses a conservative all-or-nothing
> gate across the current attachments. It may withhold TPH from a direct
> importer when another attachment is not BUS_ADDR, but it cannot return
> a tag while any current attachment has a non-direct route.

Yes that is fully correct.

> An importer
> attaching later does not change an existing importer's route, and
> future queries reevaluate the current attachment list.

No, that is incorrect! An importer attaching later eventually *does* change the routing!

The final routing is only determined on the first mapping call.

In other word the semantics is like this:

attachmentA = dma_buf_attach(dma_buf, importerA);
...
attachmentB = dma_buf_attach(dma_buf, importerB);
...
...
dma_resv_lock(dma_buf->resv, NULL);
...
pci_info = dma_buf_get_tph_st_and_routing(attachmentA);
mappingA = dma_buf_map_attachment(attachmentA, ...);
...
dma_resv_add_fence(dma_buf->resv, async_signal_preventing_unmap);
dma_resv_unlock(dma_buf->resv);

So that PCI info including TPH, ST and routing is only valid as long as you hold the lock of the DMA-buf.

When another dma_buf_attach() call comes after the mapping is already created the invalidate mappings callback is called and eventually the buffer is moved to a completely different location.

So on the next mapping call you can get different TPH, ST and routing information and depending on the internal topology of the exporter eventually a different PCI device you can do P2P with.

What can be is that not all importers support the invalidate mappings callback and/or dma_buf-pin() is called, but then exporters usually don't allow PCI P2P in the first place or at least only under restrictive limits (like for example cgroups).

And yes that semantics has been the one of DMA-buf for 15 years now and yes we can't change any of that on existing exporters since that is uAPI.

I hope that I finally made it clear how things work here. I'm kind of running out of ideas how to explain that.

Regards,
Christian.

> 
> Thanks
> 
>>
>> Regards,
>> Christian.
>>
>>>
>>> Thanks
>>


  reply	other threads:[~2026-10-01  7:07 UTC|newest]

Thread overview: 48+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-28 11:19 [PATCH v8 00/23] PCI/P2PDMA: Route peer-to-peer DMA by TLP class Leon Romanovsky
2026-09-28 11:19 ` [PATCH v8 01/23] PCI/P2PDMA: Document the TLP attribute assumptions Leon Romanovsky
2026-09-28 11:19 ` [PATCH v8 02/23] PCI/P2PDMA: Derive routing from directional ACS controls Leon Romanovsky
2026-09-28 11:19 ` [PATCH v8 03/23] PCI: Reject unreadable ACS controls in isolation checks Leon Romanovsky
2026-09-28 11:19 ` [PATCH v8 04/23] PCI/P2PDMA: Evaluate ACS controls at the path divergence Leon Romanovsky
2026-09-28 11:19 ` [PATCH v8 05/23] PCI/P2PDMA: Document directional ACS routing Leon Romanovsky
2026-09-28 11:19 ` [PATCH v8 06/23] PCI/P2PDMA: Collect the path's ACS controls before deciding Leon Romanovsky
2026-09-28 11:19 ` [PATCH v8 07/23] PCI/P2PDMA: Answer routing per TLP class Leon Romanovsky
2026-09-28 11:19 ` [PATCH v8 08/23] PCI/P2PDMA: Route Relaxed Ordering Completions directly Leon Romanovsky
2026-09-28 11:19 ` [PATCH v8 09/23] PCI/P2PDMA: Reject Translated Requests blocked by Translation Blocking Leon Romanovsky
2026-09-28 11:19 ` [PATCH v8 10/23] PCI/P2PDMA: Route Translated Requests under Direct Translated P2P Leon Romanovsky
2026-09-28 11:19 ` [PATCH v8 11/23] PCI/P2PDMA: Log detailed ACS routing diagnostics Leon Romanovsky
2026-09-28 11:19 ` [PATCH v8 12/23] PCI/P2PDMA: Add KUnit tests for the ACS routing decisions Leon Romanovsky
2026-09-28 11:19 ` [PATCH v8 13/23] PCI/P2PDMA: Test the ACS P2P routing walk Leon Romanovsky
2026-09-28 11:19 ` [PATCH v8 14/23] PCI: Add KUnit coverage for ACS isolation checks Leon Romanovsky
2026-09-28 11:19 ` [PATCH v8 15/23] PCI/P2PDMA: Document TLP-class routing Leon Romanovsky
2026-09-28 11:19 ` [PATCH v8 16/23] dma-buf: Let importers ask how peer-to-peer traffic is routed Leon Romanovsky
2026-09-29  9:34   ` Christian König
2026-09-29 13:23     ` Leon Romanovsky
2026-09-29 13:39       ` Christian König
2026-09-29 17:57         ` Leon Romanovsky
2026-09-30  7:00           ` Christian König
2026-09-30  8:18             ` Leon Romanovsky
2026-09-30  8:34               ` Christian König
2026-09-30 11:43                 ` Leon Romanovsky
2026-09-30 11:52                   ` Christian König
2026-09-30 14:32                     ` Leon Romanovsky
2026-09-30 14:52                       ` Christian König
2026-10-01  6:17                         ` Leon Romanovsky
2026-10-01  7:06                           ` Christian König [this message]
2026-09-30 14:57                       ` Thomas Hellström
2026-09-30 15:03                         ` Christian König
2026-09-30 16:52                         ` Jason Gunthorpe
2026-10-02  8:19                           ` Alistair Popple
2026-10-01  6:37                         ` Leon Romanovsky
2026-10-01  7:49                           ` Christian König
2026-10-01  8:27                             ` Leon Romanovsky
2026-10-01  9:07                               ` Christian König
2026-10-01  8:11                           ` Thomas Hellström
2026-09-28 11:19 ` [PATCH v8 17/23] vfio/pci: Hand out the P2PDMA provider behind a dma-buf Leon Romanovsky
2026-09-28 11:19 ` [PATCH v8 18/23] RDMA/uverbs: " Leon Romanovsky
2026-09-28 11:19 ` [PATCH v8 19/23] RDMA/mlx5: Ask P2PDMA whether ATS takes a direct peer-to-peer route Leon Romanovsky
2026-09-28 11:19 ` [PATCH v8 20/23] PCI/P2PDMA: Let a client declare that it selects ATS per mapping Leon Romanovsky
2026-09-28 12:44   ` Thomas Hellström
2026-09-28 11:19 ` [PATCH v8 21/23] RDMA/mlx5: Declare to P2PDMA that ATS is selected per memory key Leon Romanovsky
2026-09-28 11:19 ` [PATCH v8 22/23] PCI/P2PDMA: Evaluate the ATS path for clients with ATS enabled Leon Romanovsky
2026-09-28 12:43   ` Thomas Hellström
2026-09-28 11:19 ` [PATCH v8 23/23] PCI/P2PDMA: Test the routing of " Leon Romanovsky

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=f5c5319b-e8fd-4a54-b7ea-6bb64a4ff7e2@amd.com \
    --to=christian.koenig@amd.com \
    --cc=alex@shazbot.org \
    --cc=ankita@nvidia.com \
    --cc=axboe@kernel.dk \
    --cc=bhelgaas@google.com \
    --cc=corbet@lwn.net \
    --cc=dri-devel@lists.freedesktop.org \
    --cc=gregkh@linuxfoundation.org \
    --cc=iommu@lists.linux.dev \
    --cc=jgg@ziepe.ca \
    --cc=joro@8bytes.org \
    --cc=kch@nvidia.com \
    --cc=kvm@vger.kernel.org \
    --cc=leon@kernel.org \
    --cc=linaro-mm-sig@lists.linaro.org \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-media@vger.kernel.org \
    --cc=linux-pci@vger.kernel.org \
    --cc=linux-rdma@vger.kernel.org \
    --cc=logang@deltatee.com \
    --cc=rdunlap@infradead.org \
    --cc=robin.murphy@arm.com \
    --cc=skhan@linuxfoundation.org \
    --cc=sumit.semwal@linaro.org \
    --cc=tdave@nvidia.com \
    --cc=thomas.hellstrom@linux.intel.com \
    --cc=will@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®