From: xiaqinxin <xiaqinxin@huawei.com>
To: Barry Song <21cnbao@gmail.com>
Cc: yangyicong <yangyicong@huawei.com>, "hch@lst.de" <hch@lst.de>,
"iommu@lists.linux.dev" <iommu@lists.linux.dev>,
Jonathan Cameron <jonathan.cameron@huawei.com>,
"Zengtao (B)" <prime.zeng@hisilicon.com>,
"fanghao (A)" <fanghao11@huawei.com>,
"linux-kernel@vger.kernel.org" <linux-kernel@vger.kernel.org>
Subject: 回复: [PATCH 2/3] dma-mapping: benchmark: add support for dma_map_sg
Date: Fri, 21 Feb 2025 03:16:50 +0000 [thread overview]
Message-ID: <43618c9167654c68945ec5e7d9bf69d5@huawei.com> (raw)
In-Reply-To: <CAGsJ_4yDBT4rJyG4-Ow4T3xLq8VujBjG+-uxjnWUm_vW1nzT_A@mail.gmail.com>
-----邮件原件-----
发件人: Barry Song <21cnbao@gmail.com>
发送时间: 2025年2月18日 4:59
收件人: xiaqinxin <xiaqinxin@huawei.com>
抄送: chenxiang66@hisilicon.com; yangyicong <yangyicong@huawei.com>; hch@lst.de; iommu@lists.linux.dev; Jonathan Cameron <jonathan.cameron@huawei.com>; Zengtao (B) <prime.zeng@hisilicon.com>; fanghao (A) <fanghao11@huawei.com>; linux-kernel@vger.kernel.org
主题: Re: [PATCH 2/3] dma-mapping: benchmark: add support for dma_map_sg
On Wed, Feb 12, 2025 at 3:27 PM Qinxin Xia <xiaqinxin@huawei.com> wrote:
>
> Support for dma scatter-gather mapping and is intended for testing
> mapping performance. It achieves by introducing the dma_sg_map_param
> structure and related functions, which enable the implementation of
> scatter-gather mapping preparation, mapping, and unmapping operations.
> Additionally, the dma_map_benchmark_ops array is updated to include
> operations for scatter-gather mapping. This commit aims to provide a
> wider range of mapping performance test to cater to different scenarios.
This benchmark is mainly designed to debug contention issues, such as IOMMU TLB flushes or IOMMU driver bottlenecks. I don't fully understand how SG or single will impact the evaluation of the IOMMU driver, making it unclear if the added complexity is justified.
Can you add some explanation on why single mode is not sufficient for profiling and improving IOMMU drivers?
Hello Barry ! 😊
Currently, the HiSilicon accelerator service uses the dma_map_sg interface. We want to evaluate the performance of the entire DMA map process. (including not only the iommu, but also the map framework). In addition, for scatterlist, __iommu_map is executed for each nent. This increases the complexity and time overhead of mapping. The effect of this fragmentation is not obvious in dma_map_single, which only handles a single contiguous block of memory.
>
> Signed-off-by: Qinxin Xia <xiaqinxin@huawei.com>
> ---
> include/linux/map_benchmark.h | 1 +
> kernel/dma/map_benchmark.c | 102 ++++++++++++++++++++++++++++++++++
> 2 files changed, 103 insertions(+)
>
> diff --git a/include/linux/map_benchmark.h
> b/include/linux/map_benchmark.h index 054db02a03a7..a9c1a104ba4f
> 100644
> --- a/include/linux/map_benchmark.h
> +++ b/include/linux/map_benchmark.h
> @@ -17,6 +17,7 @@
>
> enum {
> DMA_MAP_SINGLE_MODE,
> + DMA_MAP_SG_MODE,
> DMA_MAP_MODE_MAX
> };
>
> diff --git a/kernel/dma/map_benchmark.c b/kernel/dma/map_benchmark.c
> index d8ec0ce058d8..b5828eeb3db7 100644
> --- a/kernel/dma/map_benchmark.c
> +++ b/kernel/dma/map_benchmark.c
> @@ -17,6 +17,7 @@
> #include <linux/module.h>
> #include <linux/pci.h>
> #include <linux/platform_device.h>
> +#include <linux/scatterlist.h>
> #include <linux/slab.h>
> #include <linux/timekeeping.h>
>
> @@ -111,8 +112,109 @@ static struct map_benchmark_ops dma_single_map_benchmark_ops = {
> .do_unmap = dma_single_map_benchmark_do_unmap,
> };
>
> +struct dma_sg_map_param {
> + struct sg_table sgt;
> + struct device *dev;
> + void **buf;
> + u32 npages;
> + u32 dma_dir;
> +};
> +
> +static void *dma_sg_map_benchmark_prepare(struct map_benchmark_data
> +*map) {
> + struct scatterlist *sg;
> + int i = 0;
> +
> + struct dma_sg_map_param *mparam __free(kfree) = kzalloc(sizeof(*mparam), GFP_KERNEL);
> + if (!mparam)
> + return NULL;
> +
> + mparam->npages = map->bparam.granule;
> + mparam->dma_dir = map->bparam.dma_dir;
> + mparam->dev = map->dev;
> + mparam->buf = kmalloc_array(mparam->npages, sizeof(*mparam->buf),
> + GFP_KERNEL);
> + if (!mparam->buf)
> + goto err1;
> +
> + if (sg_alloc_table(&mparam->sgt, mparam->npages, GFP_KERNEL))
> + goto err2;
> +
> + for_each_sgtable_sg(&mparam->sgt, sg, i) {
> + mparam->buf[i] = (void *)__get_free_page(GFP_KERNEL);
> + if (!mparam->buf[i])
> + goto err3;
> +
> + if (mparam->dma_dir != DMA_FROM_DEVICE)
> + memset(mparam->buf[i], 0x66, PAGE_SIZE);
> +
> + sg_set_buf(sg, mparam->buf[i], PAGE_SIZE);
> + }
> +
> + return_ptr(mparam);
> +
> +err3:
> + while (i-- > 0)
> + free_page((unsigned long)mparam->buf[i]);
> +
> + pr_err("dma_map_sg failed get free page on %s\n", dev_name(mparam->dev));
> + sg_free_table(&mparam->sgt);
> +err2:
> + pr_err("dma_map_sg failed alloc sg table on %s\n", dev_name(mparam->dev));
> + kfree(mparam->buf);
> +err1:
> + pr_err("dma_map_sg failed alloc mparam buf on %s\n", dev_name(mparam->dev));
> + return NULL;
> +}
> +
> +static void dma_sg_map_benchmark_unprepare(void *arg) {
> + struct dma_sg_map_param *mparam = arg;
> + int i;
> +
> + for (i = 0; i < mparam->npages; i++)
> + free_page((unsigned long)mparam->buf[i]);
> +
> + sg_free_table(&mparam->sgt);
> +
> + kfree(mparam->buf);
> + kfree(mparam);
> +}
> +
> +static int dma_sg_map_benchmark_do_map(void *arg) {
> + struct dma_sg_map_param *mparam = arg;
> +
> + int sg_mapped = dma_map_sg(mparam->dev, mparam->sgt.sgl,
> + mparam->npages, mparam->dma_dir);
> + if (!sg_mapped) {
> + pr_err("dma_map_sg failed on %s\n", dev_name(mparam->dev));
> + return -ENOMEM;
> + }
> +
> + return 0;
> +}
> +
> +static int dma_sg_map_benchmark_do_unmap(void *arg) {
> + struct dma_sg_map_param *mparam = arg;
> +
> + dma_unmap_sg(mparam->dev, mparam->sgt.sgl, mparam->npages,
> + mparam->dma_dir);
> +
> + return 0;
> +}
> +
> +static struct map_benchmark_ops dma_sg_map_benchmark_ops = {
> + .prepare = dma_sg_map_benchmark_prepare,
> + .unprepare = dma_sg_map_benchmark_unprepare,
> + .do_map = dma_sg_map_benchmark_do_map,
> + .do_unmap = dma_sg_map_benchmark_do_unmap, };
> +
> static struct map_benchmark_ops *dma_map_benchmark_ops[DMA_MAP_MODE_MAX] = {
> [DMA_MAP_SINGLE_MODE] = &dma_single_map_benchmark_ops,
> + [DMA_MAP_SG_MODE] = &dma_sg_map_benchmark_ops,
> };
>
> static int map_benchmark_thread(void *data)
> --
> 2.33.0
>
Thanks
Barry
next prev parent reply other threads:[~2025-02-21 3:17 UTC|newest]
Thread overview: 15+ messages / expand[flat|nested] mbox.gz Atom feed top
2025-02-12 2:27 [PATCH 0/3] dma mapping " Qinxin Xia
2025-02-12 2:27 ` [PATCH 1/3] dma mapping benchmark: modify the framework to adapt to more map modes Qinxin Xia
2025-04-07 5:28 ` Barry Song
2025-04-08 9:42 ` Qinxin Xia
2025-02-12 2:27 ` [PATCH 2/3] dma-mapping: benchmark: add support for dma_map_sg Qinxin Xia
2025-02-17 20:59 ` Barry Song
2025-02-21 3:16 ` xiaqinxin [this message]
2025-02-22 6:36 ` Barry Song
2025-03-04 13:49 ` Qinxin Xia
2025-03-04 13:56 ` Qinxin Xia
2025-03-06 9:28 ` Barry Song
2025-04-01 12:46 ` Qinxin Xia
2025-04-07 5:50 ` Barry Song
2025-04-08 9:53 ` Qinxin Xia
2025-02-12 2:27 ` [PATCH 3/3] dma mapping benchmark:add " Qinxin Xia
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=43618c9167654c68945ec5e7d9bf69d5@huawei.com \
--to=xiaqinxin@huawei.com \
--cc=21cnbao@gmail.com \
--cc=fanghao11@huawei.com \
--cc=hch@lst.de \
--cc=iommu@lists.linux.dev \
--cc=jonathan.cameron@huawei.com \
--cc=linux-kernel@vger.kernel.org \
--cc=prime.zeng@hisilicon.com \
--cc=yangyicong@huawei.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®