* [PATCH v7 0/3] dma-mapping: benchmark: Add support for dma_map_sg
@ 2026-02-25 9:37 Qinxin Xia
2026-02-25 9:37 ` [PATCH v7 1/3] dma-mapping: benchmark: modify the framework to adapt to more map modes Qinxin Xia
` (4 more replies)
0 siblings, 5 replies; 8+ messages in thread
From: Qinxin Xia @ 2026-02-25 9:37 UTC (permalink / raw)
To: 21cnbao, jonathan.cameron, wangzhou1, xiaqinxin
Cc: iommu, prime.zeng, fanghao11, linux-kernel, linuxarm
Modify the framework to adapt to more map modes, add benchmark
support for dma_map_sg, and add support sg map mode in ioctl.
The result:
[root@localhost]# ./dma_map_benchmark -m 1 -g 8 -t 8 -s 30 -d 2
dma mapping benchmark(SG_MODE): threads:8 seconds:30 node:-1 dir:FROM_DEVICE granule/sg_nents: 8
average map latency(us):1.4 standard deviation:0.3
average unmap latency(us):1.3 standard deviation:0.3
[root@localhost]# ./dma_map_benchmark -m 0 -g 8 -t 8 -s 30 -d 2
dma mapping benchmark(SINGLE_MODE): threads:8 seconds:30 node:-1 dir:FROM_DEVICE granule/sg_nents: 8
average map latency(us):1.0 standard deviation:0.3
average unmap latency(us):1.3 standard deviation:0.5
---
Changes since V6:
- Address the comments from Barry, update the comment of the granule.
- Link: https://lore.kernel.org/lkml/20260112093436.3456315-1-xiaqinxin@huawei.com/
Changes since V5:
- Address the comments from Barry, the incorrect and unnecessary 'prepare_data' judgment
is deleted and renamed 'prepare_data()' to 'initialize_data()'. The mode and
other parameters in the output print are placed in the same print. In addition, the incorrect
map_mode judgment in ioctrl and the wrong path in sg:prepare() does not release 'params' is fixed.
- Link: https://lore.kernel.org/all/20251222153246.2220659-1-xiaqinxin@huawei.com/
Changes since V4:
- Address the comments from Barry and Jonathan, irrelevant patches are deleted and
the original cache processing logic is restored with add prepare_data().
- Link: https://lore.kernel.org/all/20250614143454.2927363-3-xiaqinxin@huawei.com/
Changes since V3:
- Address the comments from Barry, change mode to a more specific namespace.
- Link: https://lore.kernel.org/all/20250509020238.3378396-1-xiaqinxin@huawei.com/
Changes since V2:
- Address the comments from Barry and ALOK, some commit information and function
input parameter names are modified to make them more accurate.
- Link: https://lore.kernel.org/all/20250506030100.394376-1-xiaqinxin@huawei.com/
Changes since V1:
- Address the comments from Barry, added some comments and changed the unmap type to void.
- Link: https://lore.kernel.org/lkml/20250212022718.1995504-1-xiaqinxin@huawei.com/
Qinxin Xia (3):
dma-mapping: benchmark: modify the framework to adapt to more map
modes
dma-mapping: benchmark: add support for dma_map_sg
tools/dma: Add dma_map_sg support
include/uapi/linux/map_benchmark.h | 13 +-
kernel/dma/map_benchmark.c | 246 ++++++++++++++++++++++++++---
tools/dma/dma_map_benchmark.c | 23 ++-
3 files changed, 254 insertions(+), 28 deletions(-)
--
2.33.0
^ permalink raw reply [flat|nested] 8+ messages in thread
* [PATCH v7 1/3] dma-mapping: benchmark: modify the framework to adapt to more map modes
2026-02-25 9:37 [PATCH v7 0/3] dma-mapping: benchmark: Add support for dma_map_sg Qinxin Xia
@ 2026-02-25 9:37 ` Qinxin Xia
2026-02-25 9:37 ` [PATCH v7 2/3] dma-mapping: benchmark: add support for dma_map_sg Qinxin Xia
` (3 subsequent siblings)
4 siblings, 0 replies; 8+ messages in thread
From: Qinxin Xia @ 2026-02-25 9:37 UTC (permalink / raw)
To: 21cnbao, jonathan.cameron, wangzhou1, xiaqinxin
Cc: iommu, prime.zeng, fanghao11, linux-kernel, linuxarm, Barry Song
This patch adjusts the DMA map benchmark framework to make the DMA
map benchmark framework more flexible and adaptable to other mapping
modes in the future. By abstracting the framework into five interfaces:
prepare, unprepare, initialize_data, do_map, and do_unmap.
The new map schema can be introduced more easily
without major modifications to the existing code structure.
Reviewed-by: Barry Song <baohua@kernel.org>
Signed-off-by: Qinxin Xia <xiaqinxin@huawei.com>
---
include/uapi/linux/map_benchmark.h | 8 +-
kernel/dma/map_benchmark.c | 131 ++++++++++++++++++++++++-----
2 files changed, 115 insertions(+), 24 deletions(-)
diff --git a/include/uapi/linux/map_benchmark.h b/include/uapi/linux/map_benchmark.h
index c2d91088a40d..e076748f2120 100644
--- a/include/uapi/linux/map_benchmark.h
+++ b/include/uapi/linux/map_benchmark.h
@@ -17,6 +17,11 @@
#define DMA_MAP_TO_DEVICE 1
#define DMA_MAP_FROM_DEVICE 2
+enum {
+ DMA_MAP_BENCH_SINGLE_MODE,
+ DMA_MAP_BENCH_MODE_MAX
+};
+
struct map_benchmark {
__u64 avg_map_100ns; /* average map latency in 100ns */
__u64 map_stddev; /* standard deviation of map latency */
@@ -29,7 +34,8 @@ struct map_benchmark {
__u32 dma_dir; /* DMA data direction */
__u32 dma_trans_ns; /* time for DMA transmission in ns */
__u32 granule; /* how many PAGE_SIZE will do map/unmap once a time */
- __u8 expansion[76]; /* For future use */
+ __u8 map_mode; /* the mode of dma map */
+ __u8 expansion[75]; /* For future use */
};
#endif /* _UAPI_DMA_BENCHMARK_H */
diff --git a/kernel/dma/map_benchmark.c b/kernel/dma/map_benchmark.c
index 0f33b3ea7daf..312e7060c7a9 100644
--- a/kernel/dma/map_benchmark.c
+++ b/kernel/dma/map_benchmark.c
@@ -5,6 +5,7 @@
#define pr_fmt(fmt) KBUILD_MODNAME ": " fmt
+#include <linux/cleanup.h>
#include <linux/debugfs.h>
#include <linux/delay.h>
#include <linux/device.h>
@@ -31,17 +32,105 @@ struct map_benchmark_data {
atomic64_t loops;
};
+struct map_benchmark_ops {
+ void *(*prepare)(struct map_benchmark_data *map);
+ void (*unprepare)(void *mparam);
+ void (*initialize_data)(void *mparam);
+ int (*do_map)(void *mparam);
+ void (*do_unmap)(void *mparam);
+};
+
+struct dma_single_map_param {
+ struct device *dev;
+ dma_addr_t addr;
+ void *xbuf;
+ u32 npages;
+ u32 dma_dir;
+};
+
+static void *dma_single_map_benchmark_prepare(struct map_benchmark_data *map)
+{
+ struct dma_single_map_param *params __free(kfree) = kzalloc(sizeof(*params),
+ GFP_KERNEL);
+ if (!params)
+ return NULL;
+
+ params->npages = map->bparam.granule;
+ params->dma_dir = map->bparam.dma_dir;
+ params->dev = map->dev;
+ params->xbuf = alloc_pages_exact(params->npages * PAGE_SIZE, GFP_KERNEL);
+ if (!params->xbuf)
+ return NULL;
+
+ return_ptr(params);
+}
+
+static void dma_single_map_benchmark_unprepare(void *mparam)
+{
+ struct dma_single_map_param *params = mparam;
+
+ free_pages_exact(params->xbuf, params->npages * PAGE_SIZE);
+ kfree(params);
+}
+
+static void dma_single_map_benchmark_initialize_data(void *mparam)
+{
+ struct dma_single_map_param *params = mparam;
+
+ /*
+ * for a non-coherent device, if we don't stain them in the
+ * cache, this will give an underestimate of the real-world
+ * overhead of BIDIRECTIONAL or TO_DEVICE mappings;
+ * 66 means everything goes well! 66 is lucky.
+ */
+ if (params->dma_dir != DMA_FROM_DEVICE)
+ memset(params->xbuf, 0x66, params->npages * PAGE_SIZE);
+}
+
+static int dma_single_map_benchmark_do_map(void *mparam)
+{
+ struct dma_single_map_param *params = mparam;
+
+ params->addr = dma_map_single(params->dev, params->xbuf,
+ params->npages * PAGE_SIZE, params->dma_dir);
+ if (unlikely(dma_mapping_error(params->dev, params->addr))) {
+ pr_err("dma_map_single failed on %s\n", dev_name(params->dev));
+ return -ENOMEM;
+ }
+
+ return 0;
+}
+
+static void dma_single_map_benchmark_do_unmap(void *mparam)
+{
+ struct dma_single_map_param *params = mparam;
+
+ dma_unmap_single(params->dev, params->addr,
+ params->npages * PAGE_SIZE, params->dma_dir);
+}
+
+static struct map_benchmark_ops dma_single_map_benchmark_ops = {
+ .prepare = dma_single_map_benchmark_prepare,
+ .unprepare = dma_single_map_benchmark_unprepare,
+ .initialize_data = dma_single_map_benchmark_initialize_data,
+ .do_map = dma_single_map_benchmark_do_map,
+ .do_unmap = dma_single_map_benchmark_do_unmap,
+};
+
+static struct map_benchmark_ops *dma_map_benchmark_ops[DMA_MAP_BENCH_MODE_MAX] = {
+ [DMA_MAP_BENCH_SINGLE_MODE] = &dma_single_map_benchmark_ops,
+};
+
static int map_benchmark_thread(void *data)
{
- void *buf;
- dma_addr_t dma_addr;
struct map_benchmark_data *map = data;
- int npages = map->bparam.granule;
- u64 size = npages * PAGE_SIZE;
+ __u8 map_mode = map->bparam.map_mode;
int ret = 0;
- buf = alloc_pages_exact(size, GFP_KERNEL);
- if (!buf)
+ struct map_benchmark_ops *mb_ops = dma_map_benchmark_ops[map_mode];
+ void *mparam = mb_ops->prepare(map);
+
+ if (!mparam)
return -ENOMEM;
while (!kthread_should_stop()) {
@@ -49,23 +138,12 @@ static int map_benchmark_thread(void *data)
ktime_t map_stime, map_etime, unmap_stime, unmap_etime;
ktime_t map_delta, unmap_delta;
- /*
- * for a non-coherent device, if we don't stain them in the
- * cache, this will give an underestimate of the real-world
- * overhead of BIDIRECTIONAL or TO_DEVICE mappings;
- * 66 means evertything goes well! 66 is lucky.
- */
- if (map->dir != DMA_FROM_DEVICE)
- memset(buf, 0x66, size);
-
+ mb_ops->initialize_data(mparam);
map_stime = ktime_get();
- dma_addr = dma_map_single(map->dev, buf, size, map->dir);
- if (unlikely(dma_mapping_error(map->dev, dma_addr))) {
- pr_err("dma_map_single failed on %s\n",
- dev_name(map->dev));
- ret = -ENOMEM;
+ ret = mb_ops->do_map(mparam);
+ if (ret)
goto out;
- }
+
map_etime = ktime_get();
map_delta = ktime_sub(map_etime, map_stime);
@@ -73,7 +151,8 @@ static int map_benchmark_thread(void *data)
ndelay(map->bparam.dma_trans_ns);
unmap_stime = ktime_get();
- dma_unmap_single(map->dev, dma_addr, size, map->dir);
+ mb_ops->do_unmap(mparam);
+
unmap_etime = ktime_get();
unmap_delta = ktime_sub(unmap_etime, unmap_stime);
@@ -108,7 +187,7 @@ static int map_benchmark_thread(void *data)
}
out:
- free_pages_exact(buf, size);
+ mb_ops->unprepare(mparam);
return ret;
}
@@ -209,6 +288,12 @@ static long map_benchmark_ioctl(struct file *file, unsigned int cmd,
switch (cmd) {
case DMA_MAP_BENCHMARK:
+ if (map->bparam.map_mode < 0 ||
+ map->bparam.map_mode >= DMA_MAP_BENCH_MODE_MAX) {
+ pr_err("invalid map mode\n");
+ return -EINVAL;
+ }
+
if (map->bparam.threads == 0 ||
map->bparam.threads > DMA_MAP_MAX_THREADS) {
pr_err("invalid thread number\n");
--
2.33.0
^ permalink raw reply [flat|nested] 8+ messages in thread
* [PATCH v7 2/3] dma-mapping: benchmark: add support for dma_map_sg
2026-02-25 9:37 [PATCH v7 0/3] dma-mapping: benchmark: Add support for dma_map_sg Qinxin Xia
2026-02-25 9:37 ` [PATCH v7 1/3] dma-mapping: benchmark: modify the framework to adapt to more map modes Qinxin Xia
@ 2026-02-25 9:37 ` Qinxin Xia
2026-02-25 9:38 ` [PATCH v7 3/3] tools/dma: Add dma_map_sg support Qinxin Xia
` (2 subsequent siblings)
4 siblings, 0 replies; 8+ messages in thread
From: Qinxin Xia @ 2026-02-25 9:37 UTC (permalink / raw)
To: 21cnbao, jonathan.cameron, wangzhou1, xiaqinxin
Cc: iommu, prime.zeng, fanghao11, linux-kernel, linuxarm, Barry Song
Support for dma scatter-gather mapping and is intended for testing
mapping performance. It achieves by introducing the dma_sg_map_param
structure and related functions, which enable the implementation of
scatter-gather mapping preparation, mapping, and unmapping operations.
Additionally, the dma_map_benchmark_ops array is updated to include
operations for scatter-gather mapping. This commit aims to provide
a wider range of mapping performance test to cater to different scenarios.
Reviewed-by: Barry Song <baohua@kernel.org>
Signed-off-by: Qinxin Xia <xiaqinxin@huawei.com>
---
include/uapi/linux/map_benchmark.h | 5 +-
kernel/dma/map_benchmark.c | 115 +++++++++++++++++++++++++++++
2 files changed, 119 insertions(+), 1 deletion(-)
diff --git a/include/uapi/linux/map_benchmark.h b/include/uapi/linux/map_benchmark.h
index e076748f2120..4b17829a9f17 100644
--- a/include/uapi/linux/map_benchmark.h
+++ b/include/uapi/linux/map_benchmark.h
@@ -19,6 +19,7 @@
enum {
DMA_MAP_BENCH_SINGLE_MODE,
+ DMA_MAP_BENCH_SG_MODE,
DMA_MAP_BENCH_MODE_MAX
};
@@ -33,7 +34,9 @@ struct map_benchmark {
__u32 dma_bits; /* DMA addressing capability */
__u32 dma_dir; /* DMA data direction */
__u32 dma_trans_ns; /* time for DMA transmission in ns */
- __u32 granule; /* how many PAGE_SIZE will do map/unmap once a time */
+ __u32 granule; /* - SINGLE_MODE: number of pages mapped/unmapped per operation
+ * - SG_MODE: number of scatterlist entries (each maps one page)
+ */
__u8 map_mode; /* the mode of dma map */
__u8 expansion[75]; /* For future use */
};
diff --git a/kernel/dma/map_benchmark.c b/kernel/dma/map_benchmark.c
index 312e7060c7a9..e1689fdd3878 100644
--- a/kernel/dma/map_benchmark.c
+++ b/kernel/dma/map_benchmark.c
@@ -16,6 +16,7 @@
#include <linux/module.h>
#include <linux/pci.h>
#include <linux/platform_device.h>
+#include <linux/scatterlist.h>
#include <linux/slab.h>
#include <linux/timekeeping.h>
#include <uapi/linux/map_benchmark.h>
@@ -117,8 +118,122 @@ static struct map_benchmark_ops dma_single_map_benchmark_ops = {
.do_unmap = dma_single_map_benchmark_do_unmap,
};
+struct dma_sg_map_param {
+ struct sg_table sgt;
+ struct device *dev;
+ void **buf;
+ u32 npages;
+ u32 dma_dir;
+};
+
+static void *dma_sg_map_benchmark_prepare(struct map_benchmark_data *map)
+{
+ struct scatterlist *sg;
+ int i;
+
+ struct dma_sg_map_param *params = kzalloc(sizeof(*params), GFP_KERNEL);
+
+ if (!params)
+ return NULL;
+ /*
+ * Set the number of scatterlist entries based on the granule.
+ * In SG mode, 'granule' represents the number of scatterlist entries.
+ * Each scatterlist entry corresponds to a single page.
+ */
+ params->npages = map->bparam.granule;
+ params->dma_dir = map->bparam.dma_dir;
+ params->dev = map->dev;
+ params->buf = kmalloc_array(params->npages, sizeof(*params->buf),
+ GFP_KERNEL);
+ if (!params->buf)
+ goto out;
+
+ if (sg_alloc_table(¶ms->sgt, params->npages, GFP_KERNEL))
+ goto free_buf;
+
+ for_each_sgtable_sg(¶ms->sgt, sg, i) {
+ params->buf[i] = (void *)__get_free_page(GFP_KERNEL);
+ if (!params->buf[i])
+ goto free_page;
+
+ sg_set_buf(sg, params->buf[i], PAGE_SIZE);
+ }
+
+ return params;
+
+free_page:
+ while (i-- > 0)
+ free_page((unsigned long)params->buf[i]);
+
+ sg_free_table(¶ms->sgt);
+free_buf:
+ kfree(params->buf);
+out:
+ kfree(params);
+ return NULL;
+}
+
+static void dma_sg_map_benchmark_unprepare(void *mparam)
+{
+ struct dma_sg_map_param *params = mparam;
+ int i;
+
+ for (i = 0; i < params->npages; i++)
+ free_page((unsigned long)params->buf[i]);
+
+ sg_free_table(¶ms->sgt);
+
+ kfree(params->buf);
+ kfree(params);
+}
+
+static void dma_sg_map_benchmark_initialize_data(void *mparam)
+{
+ struct dma_sg_map_param *params = mparam;
+ struct scatterlist *sg;
+ int i = 0;
+
+ if (params->dma_dir == DMA_FROM_DEVICE)
+ return;
+
+ for_each_sgtable_sg(¶ms->sgt, sg, i)
+ memset(params->buf[i], 0x66, PAGE_SIZE);
+}
+
+static int dma_sg_map_benchmark_do_map(void *mparam)
+{
+ struct dma_sg_map_param *params = mparam;
+ int ret = 0;
+
+ int sg_mapped = dma_map_sg(params->dev, params->sgt.sgl,
+ params->npages, params->dma_dir);
+ if (!sg_mapped) {
+ pr_err("dma_map_sg failed on %s\n", dev_name(params->dev));
+ ret = -ENOMEM;
+ }
+
+ return ret;
+}
+
+static void dma_sg_map_benchmark_do_unmap(void *mparam)
+{
+ struct dma_sg_map_param *params = mparam;
+
+ dma_unmap_sg(params->dev, params->sgt.sgl, params->npages,
+ params->dma_dir);
+}
+
+static struct map_benchmark_ops dma_sg_map_benchmark_ops = {
+ .prepare = dma_sg_map_benchmark_prepare,
+ .unprepare = dma_sg_map_benchmark_unprepare,
+ .initialize_data = dma_sg_map_benchmark_initialize_data,
+ .do_map = dma_sg_map_benchmark_do_map,
+ .do_unmap = dma_sg_map_benchmark_do_unmap,
+};
+
static struct map_benchmark_ops *dma_map_benchmark_ops[DMA_MAP_BENCH_MODE_MAX] = {
[DMA_MAP_BENCH_SINGLE_MODE] = &dma_single_map_benchmark_ops,
+ [DMA_MAP_BENCH_SG_MODE] = &dma_sg_map_benchmark_ops,
};
static int map_benchmark_thread(void *data)
--
2.33.0
^ permalink raw reply [flat|nested] 8+ messages in thread
* [PATCH v7 3/3] tools/dma: Add dma_map_sg support
2026-02-25 9:37 [PATCH v7 0/3] dma-mapping: benchmark: Add support for dma_map_sg Qinxin Xia
2026-02-25 9:37 ` [PATCH v7 1/3] dma-mapping: benchmark: modify the framework to adapt to more map modes Qinxin Xia
2026-02-25 9:37 ` [PATCH v7 2/3] dma-mapping: benchmark: add support for dma_map_sg Qinxin Xia
@ 2026-02-25 9:38 ` Qinxin Xia
2026-02-25 10:42 ` [PATCH v7 0/3] dma-mapping: benchmark: Add support for dma_map_sg Barry Song
[not found] ` <CGME20260306085057eucas1p176a39f28514d3c7481f5813d597086ca@eucas1p1.samsung.com>
4 siblings, 0 replies; 8+ messages in thread
From: Qinxin Xia @ 2026-02-25 9:38 UTC (permalink / raw)
To: 21cnbao, jonathan.cameron, wangzhou1, xiaqinxin
Cc: iommu, prime.zeng, fanghao11, linux-kernel, linuxarm, Barry Song
Support for dma_map_sg, add option '-m' to distinguish mode.
i) Users can set option '-m' to select mode:
DMA_MAP_BENCH_SINGLE_MODE=0, DMA_MAP_BENCH_SG_MODE:=1
(The mode is also show in the test result).
ii) Users can set option '-g' to set sg_nents
(total count of entries in scatterlist)
the maximum number is 1024. Each of sg buf size is PAGE_SIZE.
e.g
[root@localhost]# ./dma_map_benchmark -m 1 -g 8 -t 8 -s 30 -d 2
dma mapping mode: DMA_MAP_BENCH_SG_MODE
dma mapping benchmark: threads:8 seconds:30 node:-1
dir:FROM_DEVICE granule/sg_nents: 8
average map latency(us):1.4 standard deviation:0.3
average unmap latency(us):1.3 standard deviation:0.3
[root@localhost]# ./dma_map_benchmark -m 0 -g 8 -t 8 -s 30 -d 2
dma mapping mode: DMA_MAP_BENCH_SINGLE_MODE
dma mapping benchmark: threads:8 seconds:30 node:-1
dir:FROM_DEVICE granule/sg_nents: 8
average map latency(us):1.0 standard deviation:0.3
average unmap latency(us):1.3 standard deviation:0.5
Reviewed-by: Barry Song <baohua@kernel.org>
Signed-off-by: Qinxin Xia <xiaqinxin@huawei.com>
---
tools/dma/dma_map_benchmark.c | 23 ++++++++++++++++++++---
1 file changed, 20 insertions(+), 3 deletions(-)
diff --git a/tools/dma/dma_map_benchmark.c b/tools/dma/dma_map_benchmark.c
index dd0ed528e6df..eab0ac611a23 100644
--- a/tools/dma/dma_map_benchmark.c
+++ b/tools/dma/dma_map_benchmark.c
@@ -20,12 +20,19 @@ static char *directions[] = {
"FROM_DEVICE",
};
+static char *mode[] = {
+ "SINGLE_MODE",
+ "SG_MODE",
+};
+
int main(int argc, char **argv)
{
struct map_benchmark map;
int fd, opt;
/* default single thread, run 20 seconds on NUMA_NO_NODE */
int threads = 1, seconds = 20, node = -1;
+ /* default single map mode */
+ int map_mode = DMA_MAP_BENCH_SINGLE_MODE;
/* default dma mask 32bit, bidirectional DMA */
int bits = 32, xdelay = 0, dir = DMA_MAP_BIDIRECTIONAL;
/* default granule 1 PAGESIZE */
@@ -33,7 +40,7 @@ int main(int argc, char **argv)
int cmd = DMA_MAP_BENCHMARK;
- while ((opt = getopt(argc, argv, "t:s:n:b:d:x:g:")) != -1) {
+ while ((opt = getopt(argc, argv, "t:s:n:b:d:x:g:m:")) != -1) {
switch (opt) {
case 't':
threads = atoi(optarg);
@@ -56,11 +63,20 @@ int main(int argc, char **argv)
case 'g':
granule = atoi(optarg);
break;
+ case 'm':
+ map_mode = atoi(optarg);
+ break;
default:
return -1;
}
}
+ if (map_mode < 0 || map_mode >= DMA_MAP_BENCH_MODE_MAX) {
+ fprintf(stderr, "invalid map mode, SINGLE_MODE:%d, SG_MODE: %d\n",
+ DMA_MAP_BENCH_SINGLE_MODE, DMA_MAP_BENCH_SG_MODE);
+ exit(1);
+ }
+
if (threads <= 0 || threads > DMA_MAP_MAX_THREADS) {
fprintf(stderr, "invalid number of threads, must be in 1-%d\n",
DMA_MAP_MAX_THREADS);
@@ -110,14 +126,15 @@ int main(int argc, char **argv)
map.dma_dir = dir;
map.dma_trans_ns = xdelay;
map.granule = granule;
+ map.map_mode = map_mode;
if (ioctl(fd, cmd, &map)) {
perror("ioctl");
exit(1);
}
- printf("dma mapping benchmark: threads:%d seconds:%d node:%d dir:%s granule: %d\n",
- threads, seconds, node, directions[dir], granule);
+ printf("dma mapping benchmark(%s): threads:%d seconds:%d node:%d dir:%s granule:%d\n",
+ mode[map_mode], threads, seconds, node, directions[dir], granule);
printf("average map latency(us):%.1f standard deviation:%.1f\n",
map.avg_map_100ns/10.0, map.map_stddev/10.0);
printf("average unmap latency(us):%.1f standard deviation:%.1f\n",
--
2.33.0
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH v7 0/3] dma-mapping: benchmark: Add support for dma_map_sg
2026-02-25 9:37 [PATCH v7 0/3] dma-mapping: benchmark: Add support for dma_map_sg Qinxin Xia
` (2 preceding siblings ...)
2026-02-25 9:38 ` [PATCH v7 3/3] tools/dma: Add dma_map_sg support Qinxin Xia
@ 2026-02-25 10:42 ` Barry Song
2026-03-06 7:54 ` Marek Szyprowski
[not found] ` <CGME20260306085057eucas1p176a39f28514d3c7481f5813d597086ca@eucas1p1.samsung.com>
4 siblings, 1 reply; 8+ messages in thread
From: Barry Song @ 2026-02-25 10:42 UTC (permalink / raw)
To: Qinxin Xia, Marek Szyprowski
Cc: jonathan.cameron, wangzhou1, iommu, prime.zeng, fanghao11,
linux-kernel, linuxarm
On Wed, Feb 25, 2026 at 5:38 PM Qinxin Xia <xiaqinxin@huawei.com> wrote:
>
> Modify the framework to adapt to more map modes, add benchmark
> support for dma_map_sg, and add support sg map mode in ioctl.
>
> The result:
> [root@localhost]# ./dma_map_benchmark -m 1 -g 8 -t 8 -s 30 -d 2
> dma mapping benchmark(SG_MODE): threads:8 seconds:30 node:-1 dir:FROM_DEVICE granule/sg_nents: 8
> average map latency(us):1.4 standard deviation:0.3
> average unmap latency(us):1.3 standard deviation:0.3
> [root@localhost]# ./dma_map_benchmark -m 0 -g 8 -t 8 -s 30 -d 2
> dma mapping benchmark(SINGLE_MODE): threads:8 seconds:30 node:-1 dir:FROM_DEVICE granule/sg_nents: 8
> average map latency(us):1.0 standard deviation:0.3
> average unmap latency(us):1.3 standard deviation:0.5
>
+ Marek,
Hi Marek,
Would you be willing to pick up Qinxin's series into the dma-mapping
tree? I think that would be helpful, at least for my
"dma-mapping: arm64: support batched cache sync" series[1] and for
potential further optimizations to dma_map_sg().
[1] https://lore.kernel.org/lkml/20251226225254.46197-1-21cnbao@gmail.com/
Thanks
Barry
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH v7 0/3] dma-mapping: benchmark: Add support for dma_map_sg
2026-02-25 10:42 ` [PATCH v7 0/3] dma-mapping: benchmark: Add support for dma_map_sg Barry Song
@ 2026-03-06 7:54 ` Marek Szyprowski
0 siblings, 0 replies; 8+ messages in thread
From: Marek Szyprowski @ 2026-03-06 7:54 UTC (permalink / raw)
To: Barry Song, Qinxin Xia
Cc: jonathan.cameron, wangzhou1, iommu, prime.zeng, fanghao11,
linux-kernel, linuxarm
On 25.02.2026 11:42, Barry Song wrote:
> On Wed, Feb 25, 2026 at 5:38 PM Qinxin Xia <xiaqinxin@huawei.com> wrote:
>> Modify the framework to adapt to more map modes, add benchmark
>> support for dma_map_sg, and add support sg map mode in ioctl.
>>
>> The result:
>> [root@localhost]# ./dma_map_benchmark -m 1 -g 8 -t 8 -s 30 -d 2
>> dma mapping benchmark(SG_MODE): threads:8 seconds:30 node:-1 dir:FROM_DEVICE granule/sg_nents: 8
>> average map latency(us):1.4 standard deviation:0.3
>> average unmap latency(us):1.3 standard deviation:0.3
>> [root@localhost]# ./dma_map_benchmark -m 0 -g 8 -t 8 -s 30 -d 2
>> dma mapping benchmark(SINGLE_MODE): threads:8 seconds:30 node:-1 dir:FROM_DEVICE granule/sg_nents: 8
>> average map latency(us):1.0 standard deviation:0.3
>> average unmap latency(us):1.3 standard deviation:0.5
>>
> + Marek,
>
> Hi Marek,
>
> Would you be willing to pick up Qinxin's series into the dma-mapping
> tree? I think that would be helpful, at least for my
> "dma-mapping: arm64: support batched cache sync" series[1] and for
> potential further optimizations to dma_map_sg().
>
> [1] https://lore.kernel.org/lkml/20251226225254.46197-1-21cnbao@gmail.com/
I would be easier for me to track that work if I were on CC:, but I will
take it.
Best regards
--
Marek Szyprowski, PhD
Samsung R&D Institute Poland
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH v7 0/3] dma-mapping: benchmark: Add support for dma_map_sg
[not found] ` <CGME20260306085057eucas1p176a39f28514d3c7481f5813d597086ca@eucas1p1.samsung.com>
@ 2026-03-06 8:50 ` Marek Szyprowski
2026-03-07 7:46 ` Qinxin Xia
0 siblings, 1 reply; 8+ messages in thread
From: Marek Szyprowski @ 2026-03-06 8:50 UTC (permalink / raw)
To: Qinxin Xia, 21cnbao, jonathan.cameron, wangzhou1
Cc: iommu, prime.zeng, fanghao11, linux-kernel, linuxarm
On 25.02.2026 10:37, Qinxin Xia wrote:
> Modify the framework to adapt to more map modes, add benchmark
> support for dma_map_sg, and add support sg map mode in ioctl.
>
> The result:
> [root@localhost]# ./dma_map_benchmark -m 1 -g 8 -t 8 -s 30 -d 2
> dma mapping benchmark(SG_MODE): threads:8 seconds:30 node:-1 dir:FROM_DEVICE granule/sg_nents: 8
> average map latency(us):1.4 standard deviation:0.3
> average unmap latency(us):1.3 standard deviation:0.3
> [root@localhost]# ./dma_map_benchmark -m 0 -g 8 -t 8 -s 30 -d 2
> dma mapping benchmark(SINGLE_MODE): threads:8 seconds:30 node:-1 dir:FROM_DEVICE granule/sg_nents: 8
> average map latency(us):1.0 standard deviation:0.3
> average unmap latency(us):1.3 standard deviation:0.5
Applied to dma-mapping-next. Thanks! Next time please CC: me directly.
> ---
> Changes since V6:
> - Address the comments from Barry, update the comment of the granule.
> - Link: https://lore.kernel.org/lkml/20260112093436.3456315-1-xiaqinxin@huawei.com/
>
> Changes since V5:
> - Address the comments from Barry, the incorrect and unnecessary 'prepare_data' judgment
> is deleted and renamed 'prepare_data()' to 'initialize_data()'. The mode and
> other parameters in the output print are placed in the same print. In addition, the incorrect
> map_mode judgment in ioctrl and the wrong path in sg:prepare() does not release 'params' is fixed.
> - Link: https://lore.kernel.org/all/20251222153246.2220659-1-xiaqinxin@huawei.com/
>
> Changes since V4:
> - Address the comments from Barry and Jonathan, irrelevant patches are deleted and
> the original cache processing logic is restored with add prepare_data().
> - Link: https://lore.kernel.org/all/20250614143454.2927363-3-xiaqinxin@huawei.com/
>
> Changes since V3:
> - Address the comments from Barry, change mode to a more specific namespace.
> - Link: https://lore.kernel.org/all/20250509020238.3378396-1-xiaqinxin@huawei.com/
>
> Changes since V2:
> - Address the comments from Barry and ALOK, some commit information and function
> input parameter names are modified to make them more accurate.
> - Link: https://lore.kernel.org/all/20250506030100.394376-1-xiaqinxin@huawei.com/
>
> Changes since V1:
> - Address the comments from Barry, added some comments and changed the unmap type to void.
> - Link: https://lore.kernel.org/lkml/20250212022718.1995504-1-xiaqinxin@huawei.com/
>
> Qinxin Xia (3):
> dma-mapping: benchmark: modify the framework to adapt to more map
> modes
> dma-mapping: benchmark: add support for dma_map_sg
> tools/dma: Add dma_map_sg support
>
> include/uapi/linux/map_benchmark.h | 13 +-
> kernel/dma/map_benchmark.c | 246 ++++++++++++++++++++++++++---
> tools/dma/dma_map_benchmark.c | 23 ++-
> 3 files changed, 254 insertions(+), 28 deletions(-)
>
Best regards
--
Marek Szyprowski, PhD
Samsung R&D Institute Poland
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH v7 0/3] dma-mapping: benchmark: Add support for dma_map_sg
2026-03-06 8:50 ` Marek Szyprowski
@ 2026-03-07 7:46 ` Qinxin Xia
0 siblings, 0 replies; 8+ messages in thread
From: Qinxin Xia @ 2026-03-07 7:46 UTC (permalink / raw)
To: Marek Szyprowski, 21cnbao, jonathan.cameron, wangzhou1
Cc: iommu, prime.zeng, fanghao11, linux-kernel, linuxarm
On 2026/3/6 16:50:56, Marek Szyprowski <m.szyprowski@samsung.com> wrote:
>> dma mapping benchmark(SG_MODE): threads:8 seconds:30 node:-1 dir:FROM_DEVICE granule/sg_nents: 8
>> average map latency(us):1.4 standard deviation:0.3
>> average unmap latency(us):1.3 standard deviation:0.3
>> [root@localhost]# ./dma_map_benchmark -m 0 -g 8 -t 8 -s 30 -d 2
>> dma mapping benchmark(SINGLE_MODE): threads:8 seconds:30 node:-1 dir:FROM_DEVICE granule/sg_nents: 8
>> average map latency(us):1.0 standard deviation:0.3
>> average unmap latency(us):1.3 standard deviation:0.5
> Applied to dma-mapping-next. Thanks! Next time please CC: me directly.
Thank you for applying the patch. I apologize for not CC-ing you
directly earlier; I'll make sure to do so next time.:-)
--
Thanks,
Qinxin
^ permalink raw reply [flat|nested] 8+ messages in thread
end of thread, other threads:[~2026-03-07 7:46 UTC | newest]
Thread overview: 8+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-02-25 9:37 [PATCH v7 0/3] dma-mapping: benchmark: Add support for dma_map_sg Qinxin Xia
2026-02-25 9:37 ` [PATCH v7 1/3] dma-mapping: benchmark: modify the framework to adapt to more map modes Qinxin Xia
2026-02-25 9:37 ` [PATCH v7 2/3] dma-mapping: benchmark: add support for dma_map_sg Qinxin Xia
2026-02-25 9:38 ` [PATCH v7 3/3] tools/dma: Add dma_map_sg support Qinxin Xia
2026-02-25 10:42 ` [PATCH v7 0/3] dma-mapping: benchmark: Add support for dma_map_sg Barry Song
2026-03-06 7:54 ` Marek Szyprowski
[not found] ` <CGME20260306085057eucas1p176a39f28514d3c7481f5813d597086ca@eucas1p1.samsung.com>
2026-03-06 8:50 ` Marek Szyprowski
2026-03-07 7:46 ` Qinxin Xia
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®