mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Jason Gunthorpe <jgg@ziepe.ca>
To: "Christian König" <christian.koenig@amd.com>
Cc: "Thomas Hellström" <thomas.hellstrom@linux.intel.com>,
	"Christoph Hellwig" <hch@lst.de>,
	"Leon Romanovsky" <leon@kernel.org>,
	"Bjorn Helgaas" <bhelgaas@google.com>,
	"Logan Gunthorpe" <logang@deltatee.com>,
	"Chaitanya Kulkarni" <kch@nvidia.com>,
	"Greg Kroah-Hartman" <gregkh@linuxfoundation.org>,
	"Jens Axboe" <axboe@kernel.dk>,
	"Alex Williamson" <alex@shazbot.org>,
	"Ankit Agrawal" <ankita@nvidia.com>,
	"Jonathan Corbet" <corbet@lwn.net>,
	"Shuah Khan" <skhan@linuxfoundation.org>,
	"Joerg Roedel (AMD)" <joro@8bytes.org>,
	"Will Deacon" <will@kernel.org>,
	"Robin Murphy" <robin.murphy@arm.com>,
	"Randy Dunlap" <rdunlap@infradead.org>,
	"Sumit Semwal" <sumit.semwal@linaro.org>,
	linux-pci@vger.kernel.org, linux-kernel@vger.kernel.org,
	linux-doc@vger.kernel.org, iommu@lists.linux.dev,
	"Tushar Dave" <tdave@nvidia.com>,
	linux-media@vger.kernel.org, dri-devel@lists.freedesktop.org,
	linaro-mm-sig@lists.linaro.org, linux-rdma@vger.kernel.org,
	kvm@vger.kernel.org
Subject: Re: [PATCH v6 18/18] RDMA/mlx5: Ask P2PDMA whether ATS takes a direct peer-to-peer route
Date: Mon, 21 Sep 2026 10:08:58 -0300	[thread overview]
Message-ID: <20260921130858.GO11599@ziepe.ca> (raw)
In-Reply-To: <9656f2f9-2e39-4007-b065-3b763b53c6de@amd.com>

On Mon, Sep 21, 2026 at 08:43:49AM +0200, Christian König wrote:
> On 9/18/26 19:05, Jason Gunthorpe wrote:
> > On Fri, Sep 18, 2026 at 03:42:28PM +0200, Thomas Hellström wrote:
> >>
> >> 1) Xe attachment check if pci_p2pdma_distance() returns OK for the
> >> path. Then Xe always sets up dma-addresses using dma_map_resource(). 
> > 
> > Open coding pci_p2pdma_distance() in drivers is a hack. Using
> > dma_map_resource() like this was never "allowed".
> > 
> > We've fixed things so these hacks are not needed, the drivers need to
> > move over to things like dma_buf_phys_vec_to_sgt() and the hmm helpers
> > to use the DMA API correctly.
> 
> That is a completely broken approach as well since it limits the
> exported resources to addresses the CPU can reach.

Yes, of course it does. The DMA API only works on phy_addr_t. If you
have a struct p2pdma_provider * then you have a phys_addr_t for it.

If the exporter knows it is working with a MMIO mapping on a PCI
device it gets to acquire a p2pdma_provider and use the helper.

None of this is suposed to solve your "resources the CPU cannot reach"
problem, that has nothing to do with DMA API or P2P. It is not broken
just because it doesn't solve every problem.

> Christoph Hellwig is right that drivers should never use that stuff
> directly, not even through that dma_buf_phys_vec_to_sgt() function.

Hellwig's point was that the subsystem needs a mapping helper that
goes from the subsystem address representation to the HW
representation and hides these details from the drivers.

Look at what he built in nvme around biovec.

The dma_buf_phys_vec_to_sgt() is the dmabuf version of the same idea,
the subystem provides the mapping helper. We can try to do better, but
better is not making the exporters touch the mapping algorithm.

Ultimately I want to see something in lib/ handle this with a
non-scatterlist datastructure, but there is a huge gulf between where
dmabuf is now and it being able to work with a non-scatterlist
datastructure.

Look, it is easy to complain you don't like how it looks, but this
stuff is hard there are lots of competing concerns, if you have a
better idea now is a good time to present it. Maybe if you look
closely you will appreciate how much work has gone into even getting
things this far.

> >> It seems to me that a pci-device settable flag "ATS always enabled"
> >> should be enough to fix both issues?
> > 
> > It should be be per-mapping to support the NIC workflow that isn't a
> > global operation.
> 
> I just realized what you guys are doing and I'm not sure if the
> Linux PCI subsystem should support such hacks at all.

I don't know how to respond to that. It is spec complaint, it is
shipping in enormous volumes, of course Linux needs to support the HW
that exists.

> Basically from the point of view of the TA the NIC has ATS enabled
> all the time, but has a per request option to use translated
> addresses directly without previously translating and caching them
> using ATS, correct?

Yep. Very few devices in the world can use ATS for every single
operation. Many have this split operating model. Some even only use
ATS for DMA flows that are faultable.

Mostly the OS can't tell what the device is doing and doesn't
care. The P2P routing is the main issue and we have hacked around it
in our systems till now. Leon is trying to fix it. Thomas needs it
fixed too for Xe. So what's the issue here?

DMABUF needs to learn how to do interconnect specific behaviors. PCI
is an interconnect, it has lots and lots of crazy rules. An
importer/exporter that chooses to use PCI for their DMA should have a
way to exchange PCI specific information.

So should UALink and all the other zoo of options we have now. It
cannot be completely generic and meet everyones needs.

Can we focus on that instead of arguing if the PCI craziness should
exist or not?

> If yes than that is extremely questionable behavior, I'm not sure if
> that is covered by the PCIe spec.

Spec doesn't say anything about when a device has to translated vs
untranslated. The ATS flags only say translated is allowed to be used.

> ATS is meant to be an optimization which moves the TLB from the root
> complex (TA) into the devices at the cost of TLB invalidation
> complexity. But what you do here is abusing that functionality as
> far as I can see.

ATS is for alot more than that, and there is no abuse here.

Jason

  parent reply	other threads:[~2026-09-21 13:09 UTC|newest]

Thread overview: 48+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-14 11:22 [PATCH v6 00/18] PCI/P2PDMA: Route peer-to-peer DMA by TLP class Leon Romanovsky
2026-09-14 11:22 ` [PATCH v6 01/18] PCI/P2PDMA: Document pdev->p2pdma lifetime rules Leon Romanovsky
2026-09-18 18:25   ` Logan Gunthorpe
2026-09-19  4:39     ` Leon Romanovsky
2026-09-14 11:22 ` [PATCH v6 02/18] PCI/P2PDMA: Document the TLP attribute assumptions Leon Romanovsky
2026-09-14 11:22 ` [PATCH v6 03/18] PCI/P2PDMA: Derive routing from directional ACS controls Leon Romanovsky
2026-09-18 19:39   ` Logan Gunthorpe
2026-09-20 11:34     ` Leon Romanovsky
2026-09-14 11:22 ` [PATCH v6 04/18] PCI: Reject unreadable ACS controls in isolation checks Leon Romanovsky
2026-09-18 19:59   ` Logan Gunthorpe
2026-09-14 11:22 ` [PATCH v6 05/18] PCI/P2PDMA: Evaluate ACS controls at the path divergence Leon Romanovsky
2026-09-18 20:07   ` Logan Gunthorpe
2026-09-14 11:22 ` [PATCH v6 06/18] PCI/P2PDMA: Document directional ACS routing Leon Romanovsky
2026-09-18 20:53   ` Logan Gunthorpe
2026-09-14 11:22 ` [PATCH v6 07/18] PCI/P2PDMA: Collect the path's ACS controls before deciding Leon Romanovsky
2026-09-18 21:11   ` Logan Gunthorpe
2026-09-14 11:22 ` [PATCH v6 08/18] PCI/P2PDMA: Answer routing per TLP class Leon Romanovsky
2026-09-18 21:55   ` Logan Gunthorpe
2026-09-14 11:22 ` [PATCH v6 09/18] PCI/P2PDMA: Route Relaxed Ordering Completions directly Leon Romanovsky
2026-09-18 22:07   ` Logan Gunthorpe
2026-09-14 11:22 ` [PATCH v6 10/18] PCI/P2PDMA: Reject Translated Requests blocked by Translation Blocking Leon Romanovsky
2026-09-14 11:22 ` [PATCH v6 11/18] PCI/P2PDMA: Route Translated Requests under Direct Translated P2P Leon Romanovsky
2026-09-14 11:22 ` [PATCH v6 12/18] PCI/P2PDMA: Log detailed ACS routing diagnostics Leon Romanovsky
2026-09-14 11:22 ` [PATCH v6 13/18] PCI/P2PDMA: Add KUnit tests for the ACS routing decisions Leon Romanovsky
2026-09-14 11:22 ` [PATCH v6 14/18] PCI/P2PDMA: Test the ACS P2P routing walk Leon Romanovsky
2026-09-14 11:22 ` [PATCH v6 15/18] PCI: Add KUnit coverage for ACS isolation checks Leon Romanovsky
2026-09-14 11:22 ` [PATCH v6 16/18] PCI/P2PDMA: Document TLP-class routing Leon Romanovsky
2026-09-14 11:22 ` [PATCH v6 17/18] dma-buf: Let importers ask how peer-to-peer traffic is routed Leon Romanovsky
2026-09-14 14:27   ` Christian König
2026-09-14 22:10     ` Jason Gunthorpe
2026-09-15  6:28       ` Leon Romanovsky
2026-09-15  6:23     ` Leon Romanovsky
2026-09-14 11:22 ` [PATCH v6 18/18] RDMA/mlx5: Ask P2PDMA whether ATS takes a direct peer-to-peer route Leon Romanovsky
2026-09-17 13:56   ` Thomas Hellström
2026-09-17 14:00     ` Christian König
2026-09-17 14:28       ` Jason Gunthorpe
2026-09-17 14:02     ` Jason Gunthorpe
2026-09-17 14:16       ` Thomas Hellström
2026-09-17 15:15         ` Jason Gunthorpe
2026-09-18 12:15     ` Leon Romanovsky
2026-09-18 13:42       ` Thomas Hellström
2026-09-18 17:05         ` Jason Gunthorpe
2026-09-19 11:25           ` Leon Romanovsky
2026-09-21  6:43           ` Christian König
2026-09-21  6:57             ` Thomas Hellström
2026-09-21 13:08             ` Jason Gunthorpe [this message]
     [not found]           ` <f9ced8cab6eb929a5f65d789fe6922458c6f3396.camel@linux.intel.com>
     [not found]             ` <20260921131055.GP11599@ziepe.ca>
2026-09-21 13:27               ` Thomas Hellström
2026-09-21 13:25       ` Jason Gunthorpe

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260921130858.GO11599@ziepe.ca \
    --to=jgg@ziepe.ca \
    --cc=alex@shazbot.org \
    --cc=ankita@nvidia.com \
    --cc=axboe@kernel.dk \
    --cc=bhelgaas@google.com \
    --cc=christian.koenig@amd.com \
    --cc=corbet@lwn.net \
    --cc=dri-devel@lists.freedesktop.org \
    --cc=gregkh@linuxfoundation.org \
    --cc=hch@lst.de \
    --cc=iommu@lists.linux.dev \
    --cc=joro@8bytes.org \
    --cc=kch@nvidia.com \
    --cc=kvm@vger.kernel.org \
    --cc=leon@kernel.org \
    --cc=linaro-mm-sig@lists.linaro.org \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-media@vger.kernel.org \
    --cc=linux-pci@vger.kernel.org \
    --cc=linux-rdma@vger.kernel.org \
    --cc=logang@deltatee.com \
    --cc=rdunlap@infradead.org \
    --cc=robin.murphy@arm.com \
    --cc=skhan@linuxfoundation.org \
    --cc=sumit.semwal@linaro.org \
    --cc=tdave@nvidia.com \
    --cc=thomas.hellstrom@linux.intel.com \
    --cc=will@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®