From: "Christian König" <christian.koenig@amd.com>
To: Kelly Devilliv <kelly.devilliv@outlook.com>,
Robin Murphy <robin.murphy@arm.com>,
"joro@8bytes.org" <joro@8bytes.org>,
"will@kernel.org" <will@kernel.org>
Cc: "iommu@lists.linux.dev" <iommu@lists.linux.dev>,
"linux-kernel@vger.kernel.org" <linux-kernel@vger.kernel.org>
Subject: Re: dma_map_resource() has a bad performance in pcie peer to peer transactions when iommu enabled in Linux
Date: Mon, 25 Sep 2023 19:58:29 +0200 [thread overview]
Message-ID: <39dff6ed-2fc1-0462-0317-9d1e4c1c718e@amd.com> (raw)
In-Reply-To: <TY2PR0101MB31361E2EE3391EBFAB78014B84FCA@TY2PR0101MB3136.apcprd01.prod.exchangelabs.com>
Am 25.09.23 um 16:17 schrieb Kelly Devilliv:
>> On 2023-09-25 04:59, Kelly Devilliv wrote:
>>> Dear all,
>>>
>>> I am working on an ARM-V8 server with two gpu cards on it. Recently, I need
>> to test pcie peer to peer communication between the two gpu cards, but the
>> throughput is only 4GB/s.
>>> After I explored the gpu's kernel mode driver, I found it was using the
>> dma_map_resource() API to map the peer device's MMIO space. The arm
>> iommu driver then will hardcode a 'IOMMU_MMIO' prot in the later dma map:
>>> static dma_addr_t iommu_dma_map_resource(struct device *dev,
>> phys_addr_t phys,
>>> size_t size, enum dma_data_direction
>> dir, unsigned long attrs)
>>> {
>>> return __iommu_dma_map(dev, phys, size,
>>> dma_info_to_prot(dir, false,
>> attrs) | IOMMU_MMIO,
>>> dma_get_mask(dev));
>>> }
>>>
>>> And that will finally set the 'ARM_LPAE_PTE_MEMATTR_DEV' attribute in PTE,
>> which may have a negative impact on the performance of the pcie peer to peer
>> transactions.
>>> /*
>>> * Note that this logic is structured to accommodate Mali LPAE
>>> * having stage-1-like attributes but stage-2-like permissions.
>>> */
>>> if (data->iop.fmt == ARM_64_LPAE_S2 ||
>>> data->iop.fmt == ARM_32_LPAE_S2) {
>>> if (prot & IOMMU_MMIO)
>>> pte |= ARM_LPAE_PTE_MEMATTR_DEV;
>>> else if (prot & IOMMU_CACHE)
>>> pte |= ARM_LPAE_PTE_MEMATTR_OIWB;
>>> else
>>> pte |= ARM_LPAE_PTE_MEMATTR_NC;
>>> } else {
>>> if (prot & IOMMU_MMIO)
>>> pte |= (ARM_LPAE_MAIR_ATTR_IDX_DEV
>>> << ARM_LPAE_PTE_ATTRINDX_SHIFT);
>>> else if (prot & IOMMU_CACHE)
>>> pte |= (ARM_LPAE_MAIR_ATTR_IDX_CACHE
>>> << ARM_LPAE_PTE_ATTRINDX_SHIFT);
>>> }
>>>
>>> I tried to remove the 'IOMMU_MMIO' prot in the dma_map_resource() API
>> and re-compile the linux kernel, the throughput then can be up to 28GB/s.
>>> Is there an elegant way to solve this issue without modifying the linux kernel?
>> e.g., a substitution of dma_map_resource() API?
>>
>> Not really. Other use-cases for dma_map_resource() include DMA offload
>> engines accessing FIFO registers, where allowing reordering, write-gathering,
>> etc. would be a terrible idea. Thus it needs to assume a "safe" MMIO memory
>> type, which on Arm means Device-nGnRE.
>>
>> However, the "proper" PCI peer-to-peer support under CONFIG_PCI_P2PDMA
>> ended up moving away from the dma_map_resource() approach anyway, and
>> allows this kind of device memory to be treated more like regular memory (via
>> ZONE_DEVICE) rather than arbitrary MMIO resources, so your best bet would
>> be to get the GPU driver converted over to using that.
> Thanks Robin.
> So your suggestion is we'd better work out a new implementation just as what it
> does under CONFIG_PCI_P2PDMA instead of just using the dma_map_resource()
> API?
>
> I have explored the GPU drivers from AMD, Nvidia and habanalabs, e.g., and found
> they all using the dma_map_resource() API to map the peer device's bar address.
> If so, is it possible to be a common performance issue in PCI peer-to-peer scenario?
That's not an issue, but expected behavior.
When you enable IOMMU every transaction needs to go through the root
complex for address translation and you completely lose the performance
benefit of PCIe P2P.
This is a hardware limitation and not really related to
dma_map_resource() in any way.
Regards,
Christian.
>
>> Thanks,
>> Robin.
>>
>>> Thank you!
>>>
>>> Platform info:
>>> Linux kernel version: 5.10
>>> PCIE GEN4 x16
>>>
>>> Sincerely,
>>> Kelly
>>>
next prev parent reply other threads:[~2023-09-25 17:58 UTC|newest]
Thread overview: 10+ messages / expand[flat|nested] mbox.gz Atom feed top
2023-09-25 14:17 Kelly Devilliv
2023-09-25 17:58 ` Christian König [this message]
2023-09-26 4:33 ` 答复: " Kelly Devilliv
2023-09-26 5:32 ` Christian König
-- strict thread matches above, loose matches on Subject: below --
2023-09-26 15:30 Kelly Devilliv
2023-10-02 6:30 ` Christoph Hellwig
2023-10-02 20:08 ` Robin Murphy
2023-10-04 6:26 ` Christian König
2023-09-25 3:59 Kelly Devilliv
2023-09-25 11:15 ` Robin Murphy
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=39dff6ed-2fc1-0462-0317-9d1e4c1c718e@amd.com \
--to=christian.koenig@amd.com \
--cc=iommu@lists.linux.dev \
--cc=joro@8bytes.org \
--cc=kelly.devilliv@outlook.com \
--cc=linux-kernel@vger.kernel.org \
--cc=robin.murphy@arm.com \
--cc=will@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®