* [PATCH] perf bench: Add atomic CAS benchmark
@ 2026-09-30 9:16 Changbin Du
2026-10-05 22:30 ` Namhyung Kim
2026-10-08 7:44 ` David Laight
0 siblings, 2 replies; 5+ messages in thread
From: Changbin Du @ 2026-09-30 9:16 UTC (permalink / raw)
To: Peter Zijlstra, Ingo Molnar, Arnaldo Carvalho de Melo, Namhyung Kim
Cc: Mark Rutland, Alexander Shishkin, Jiri Olsa, Ian Rogers,
Adrian Hunter, James Clark, linux-kernel, linux-perf-users,
Changbin Du
Add a new 'atomic' collection to perf bench for benchmarking
compare-and-swap (CAS) atomic operations with multi-threaded
contention testing.
The benchmark tests __atomic_compare_exchange_n operations
with configurable thread count and iteration count to measure
atomic contention effects.
Why this benchmark is needed:
- CAS operations are fundamental to lock-free algorithms and data
structures. Understanding their performance characteristics under
contention is critical for designing high-performance concurrent
applications.
- The benchmark helps identify atomic operation latency and
scalability issues across different thread counts, revealing
contention patterns that are not visible in single-threaded tests.
- Useful for evaluating atomic implementation quality on different
architectures and for regression testing after changes to atomic
primitives or memory ordering.
Measurement methodology:
- Each thread starts a private timer (clock_gettime CLOCK_MONOTONIC)
after synchronizing on a pthread_barrier, ensuring all threads
begin simultaneously.
- Each thread performs a hot loop of atomic compare-and-swap on a
shared u64 counter, incrementing from 0 to iterations.
- The shared counter is cache-line aligned (64 bytes) to isolate
contention to the target cache line and avoid false sharing.
- The wall-clock time is measured as the max of all per-thread
runtimes (the time for the slowest thread to finish).
- The first repeat is excluded from statistics as a warmup phase
to avoid cache-cold effects.
Example usage:
$ perf bench atomic cas --threads 2
# Running 'atomic/cas' benchmark:
Threads: 2, iterations/thread: 100000000, repeats: 10 (warmup: 1)
Avg wall-clock time: 7365.480 msec (stddev 66.014 msec)
Total ops: 200,000,000
Throughput total: 27,153,697 ops/sec
Per-thread times and throughput (last repeat):
fastest: 7510.031 msec (13315525 ops/sec)
slowest: 7581.678 msec (13189692 ops/sec)
avg: 7545.854 msec (13252310 ops/sec)
Output fields explained:
- Threads: number of contending threads
- iterations/thread: CAS operations each thread performs
- repeats: number of test runs (first is warmup)
- Avg wall-clock time: mean time for all threads to complete
- stddev: standard deviation across repeats
- Total ops: threads x iterations/thread
- Throughput total: aggregate ops/sec across all threads
- Per-thread times: fastest/slowest/avg thread completion time
- Per-thread throughput: per-thread ops/sec (shows scheduling imbalance)
Assisted-by: opencode:DeepSeek-V4-Pro
Signed-off-by: Changbin Du <changbin.du@gmail.com>
---
tools/perf/Documentation/perf-bench.txt | 22 +++
tools/perf/bench/Build | 1 +
tools/perf/bench/atomic.c | 229 ++++++++++++++++++++++++
tools/perf/bench/bench.h | 1 +
tools/perf/builtin-bench.c | 8 +
5 files changed, 261 insertions(+)
create mode 100644 tools/perf/bench/atomic.c
diff --git a/tools/perf/Documentation/perf-bench.txt b/tools/perf/Documentation/perf-bench.txt
index c5913cf59c98..9cdc19f02bcf 100644
--- a/tools/perf/Documentation/perf-bench.txt
+++ b/tools/perf/Documentation/perf-bench.txt
@@ -58,6 +58,9 @@ SUBSYSTEM
'numa'::
NUMA scheduling and MM benchmarks.
+'atomic'::
+ Atomic operation benchmarks.
+
'futex'::
Futex stressing benchmarks.
@@ -283,6 +286,25 @@ SUITES FOR 'numa'
*mem*::
Suite for evaluating NUMA workloads.
+SUITES FOR 'atomic'
+~~~~~~~~~~~~~~~~~~~
+*cas*::
+Suite for evaluating compare-and-swap (CAS) atomic operations under
+multi-threaded contention.
+
+Options of *cas*
+^^^^^^^^^^^^^^^^
+-t::
+--threads=::
+Number of threads contending for the shared counter (default: 2).
+
+-i::
+--iterations=::
+Number of iterations per thread (default: 100000000).
+
+The per-thread fastest/slowest/avg summary is only printed when
+running with more than one thread.
+
SUITES FOR 'futex'
~~~~~~~~~~~~~~~~~~
*hash*::
diff --git a/tools/perf/bench/Build b/tools/perf/bench/Build
index 67b76fe20ba6..c64c52468d3d 100644
--- a/tools/perf/bench/Build
+++ b/tools/perf/bench/Build
@@ -3,6 +3,7 @@ perf-bench-y += sched-pipe.o
perf-bench-y += sched-seccomp-notify.o
perf-bench-y += syscall.o
perf-bench-y += mem-functions.o
+perf-bench-y += atomic.o
perf-bench-y += futex.o
perf-bench-y += futex-hash.o
perf-bench-y += futex-wake.o
diff --git a/tools/perf/bench/atomic.c b/tools/perf/bench/atomic.c
new file mode 100644
index 000000000000..ce97025018ab
--- /dev/null
+++ b/tools/perf/bench/atomic.c
@@ -0,0 +1,229 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (C) 2026 Changbin Du
+ *
+ * Benchmark for atomic compare-and-swap (CAS) operations.
+ *
+ * Measures throughput (ops/sec) of contended CAS on a single shared
+ * counter across multiple threads. Each thread performs a hot loop
+ * of compare-and-swap, competing with all other threads for the
+ * same cache line.
+ */
+#include "bench.h"
+#include <linux/compiler.h>
+#include "../util/debug.h"
+
+#ifndef HAVE_PTHREAD_BARRIER
+int bench_atomic(int argc __maybe_unused, const char **argv __maybe_unused)
+{
+ pr_err("%s: pthread_barrier_t unavailable, disabling this test...\n", __func__);
+ return 0;
+}
+#else /* HAVE_PTHREAD_BARRIER */
+#include <stdlib.h>
+#include <stdio.h>
+#include <unistd.h>
+#include <pthread.h>
+#include <time.h>
+#include <inttypes.h>
+#include <err.h>
+#include <linux/types.h>
+#include <linux/kernel.h>
+#include <linux/time64.h>
+#include <subcmd/parse-options.h>
+#include "../util/stat.h"
+
+static unsigned int threads = 2;
+static unsigned int iterations = 100000000;
+
+static const struct option options[] = {
+ OPT_UINTEGER('t', "threads", &threads,
+ "Number of threads contending for the counter"),
+ OPT_UINTEGER('i', "iterations", &iterations,
+ "Number of iterations per thread"),
+ OPT_END()
+};
+
+static const char * const bench_usage[] = {
+ "perf bench atomic cas <options>",
+ NULL
+};
+
+static u64 shared_counter __aligned(64);
+static pthread_barrier_t start_barrier __aligned(64);
+
+struct worker_stats {
+ u64 runtime_ns;
+};
+
+static void *worker(void *arg)
+{
+ struct worker_stats *ws = (struct worker_stats *)arg;
+ struct timespec tstart, tend;
+ u64 i;
+
+ pthread_barrier_wait(&start_barrier);
+
+ clock_gettime(CLOCK_MONOTONIC, &tstart);
+ for (i = 0; i < iterations; i++) {
+ u64 old_val, new_val;
+
+ do {
+ old_val = __atomic_load_n(&shared_counter,
+ __ATOMIC_RELAXED);
+ new_val = old_val + 1;
+ } while (!__atomic_compare_exchange_n(&shared_counter,
+ &old_val, new_val,
+ false,
+ __ATOMIC_RELAXED,
+ __ATOMIC_RELAXED));
+ }
+ clock_gettime(CLOCK_MONOTONIC, &tend);
+
+ ws->runtime_ns = (tend.tv_sec - tstart.tv_sec) * NSEC_PER_SEC +
+ (tend.tv_nsec - tstart.tv_nsec);
+
+ return NULL;
+}
+
+static void print_default_format(struct worker_stats *wstats,
+ struct stats *time_stats)
+{
+ double time_avg, time_stddev;
+
+ printf(" Threads: %u, iterations/thread: %u, repeats: %u (warmup: 1)\n",
+ threads, iterations, bench_repeat);
+
+ time_avg = avg_stats(time_stats);
+ time_stddev = stddev_stats(time_stats);
+
+ printf(" Avg wall-clock time: %.3f msec (stddev %.3f msec)\n",
+ time_avg / (double)NSEC_PER_MSEC,
+ time_stddev / (double)NSEC_PER_MSEC);
+ printf(" Total ops: %'" PRIu64 "\n",
+ iterations * (u64)threads);
+ printf(" Throughput total: %'.0f ops/sec\n",
+ (double)(iterations * (u64)threads) /
+ (time_avg / (double)NSEC_PER_SEC));
+
+ /*
+ * Per-thread statistics from the last measured repeat show
+ * fairness / imbalance in CPU scheduling.
+ */
+ if (threads > 1) {
+ u64 min_ns = UINT64_MAX, max_ns = 0, total_ns = 0;
+ double min_thru, max_thru, avg_thru;
+
+ for (unsigned int t = 0; t < threads; t++) {
+ if (wstats[t].runtime_ns < min_ns)
+ min_ns = wstats[t].runtime_ns;
+ if (wstats[t].runtime_ns > max_ns)
+ max_ns = wstats[t].runtime_ns;
+ total_ns += wstats[t].runtime_ns;
+ }
+
+ min_thru = (double)iterations * NSEC_PER_SEC / min_ns;
+ max_thru = (double)iterations * NSEC_PER_SEC / max_ns;
+ avg_thru = (double)threads * iterations * NSEC_PER_SEC / total_ns;
+
+ printf(" Per-thread times and throughput (last repeat):\n");
+ printf(" fastest: %.3f msec (%.0f ops/sec)\n",
+ min_ns / (double)NSEC_PER_MSEC, min_thru);
+ printf(" slowest: %.3f msec (%.0f ops/sec)\n",
+ max_ns / (double)NSEC_PER_MSEC, max_thru);
+ printf(" avg: %.3f msec (%.0f ops/sec)\n",
+ (total_ns / threads) / (double)NSEC_PER_MSEC, avg_thru);
+ }
+}
+
+static int run_cas_benchmark(void)
+{
+ pthread_t *thread_ids;
+ struct worker_stats *wstats;
+ struct stats time_stats;
+ unsigned int r, t;
+
+ thread_ids = calloc(threads, sizeof(pthread_t));
+ if (!thread_ids)
+ return -1;
+
+ wstats = calloc(threads, sizeof(struct worker_stats));
+ if (!wstats) {
+ free(thread_ids);
+ return -1;
+ }
+
+ init_stats(&time_stats);
+
+ for (r = 0; r < bench_repeat + 1; r++) {
+ u64 max_runtime_ns = 0;
+
+ shared_counter = 0;
+ pthread_barrier_init(&start_barrier, NULL, threads);
+
+ for (t = 0; t < threads; t++) {
+ if (pthread_create(&thread_ids[t], NULL,
+ worker, &wstats[t]))
+ err(EXIT_FAILURE, "pthread_create");
+ }
+
+ for (t = 0; t < threads; t++)
+ pthread_join(thread_ids[t], NULL);
+
+ pthread_barrier_destroy(&start_barrier);
+
+ for (t = 0; t < threads; t++) {
+ if (wstats[t].runtime_ns > max_runtime_ns)
+ max_runtime_ns = wstats[t].runtime_ns;
+ }
+
+ /*
+ * Exclude the first repeat (warmup) from statistics so
+ * that cache-cold and lazy-init effects are not counted.
+ */
+ if (r > 0)
+ update_stats(&time_stats, max_runtime_ns);
+ }
+
+ switch (bench_format) {
+ case BENCH_FORMAT_DEFAULT:
+ print_default_format(wstats, &time_stats);
+ break;
+
+ case BENCH_FORMAT_SIMPLE:
+ printf("%.0f\n",
+ (double)(iterations * (u64)threads) /
+ (avg_stats(&time_stats) / (double)NSEC_PER_SEC));
+ break;
+
+ default:
+ fprintf(stderr, "Unknown format: %d\n", bench_format);
+ exit(EXIT_FAILURE);
+ }
+
+ free(thread_ids);
+ free(wstats);
+
+ return 0;
+}
+
+int bench_atomic(int argc, const char **argv)
+{
+ if (parse_options(argc, argv, options, bench_usage, 0)) {
+ usage_with_options(bench_usage, options);
+ exit(EXIT_FAILURE);
+ }
+
+ if (threads < 1) {
+ fprintf(stderr, "Invalid thread count: %u\n", threads);
+ return 1;
+ }
+
+ if (iterations < 1) {
+ fprintf(stderr, "Invalid iteration count: %u\n", iterations);
+ return 1;
+ }
+
+ return run_cas_benchmark();
+}
+#endif /* HAVE_PTHREAD_BARRIER */
diff --git a/tools/perf/bench/bench.h b/tools/perf/bench/bench.h
index 8519eb5a42fa..01eb6ceffb8d 100644
--- a/tools/perf/bench/bench.h
+++ b/tools/perf/bench/bench.h
@@ -30,6 +30,7 @@ int bench_mem_memcpy(int argc, const char **argv);
int bench_mem_memset(int argc, const char **argv);
int bench_mem_mmap(int argc, const char **argv);
int bench_mem_find_bit(int argc, const char **argv);
+int bench_atomic(int argc, const char **argv);
int bench_futex_hash(int argc, const char **argv);
int bench_futex_wake(int argc, const char **argv);
int bench_futex_wake_parallel(int argc, const char **argv);
diff --git a/tools/perf/builtin-bench.c b/tools/perf/builtin-bench.c
index 02d47913cc6a..aeffa0826aea 100644
--- a/tools/perf/builtin-bench.c
+++ b/tools/perf/builtin-bench.c
@@ -14,6 +14,7 @@
* syscall ... System call performance
* mem ... memory access performance
* numa ... NUMA scheduling and MM performance
+ * atomic ... Atomic operation performance
* futex ... Futex performance
* epoll ... Event poll performance
*/
@@ -115,6 +116,12 @@ static const struct bench uprobe_benchmarks[] = {
{ NULL, NULL, NULL },
};
+static const struct bench atomic_benchmarks[] = {
+ { "cas", "Benchmark CAS (compare-and-swap) operations", bench_atomic },
+ { "all", "Run all atomic benchmarks", NULL },
+ { NULL, NULL, NULL }
+};
+
struct collection {
const char *name;
const char *summary;
@@ -128,6 +135,7 @@ static const struct collection collections[] = {
#ifdef HAVE_LIBNUMA_SUPPORT
{ "numa", "NUMA scheduling and MM benchmarks", numa_benchmarks },
#endif
+ { "atomic", "Atomic operation benchmarks", atomic_benchmarks },
{"futex", "Futex stressing benchmarks", futex_benchmarks },
#ifdef HAVE_EVENTFD_SUPPORT
{"epoll", "Epoll stressing benchmarks", epoll_benchmarks },
--
2.53.0
^ permalink raw reply [flat|nested] 5+ messages in thread* Re: [PATCH] perf bench: Add atomic CAS benchmark
2026-09-30 9:16 [PATCH] perf bench: Add atomic CAS benchmark Changbin Du
@ 2026-10-05 22:30 ` Namhyung Kim
2026-10-07 5:27 ` Changbin Du
2026-10-08 7:44 ` David Laight
1 sibling, 1 reply; 5+ messages in thread
From: Namhyung Kim @ 2026-10-05 22:30 UTC (permalink / raw)
To: Changbin Du
Cc: Peter Zijlstra, Ingo Molnar, Arnaldo Carvalho de Melo,
Mark Rutland, Alexander Shishkin, Jiri Olsa, Ian Rogers,
Adrian Hunter, James Clark, linux-kernel, linux-perf-users
Hello,
On Wed, Sep 30, 2026 at 05:16:17PM +0800, Changbin Du wrote:
> Add a new 'atomic' collection to perf bench for benchmarking
> compare-and-swap (CAS) atomic operations with multi-threaded
> contention testing.
>
> The benchmark tests __atomic_compare_exchange_n operations
> with configurable thread count and iteration count to measure
> atomic contention effects.
>
> Why this benchmark is needed:
> - CAS operations are fundamental to lock-free algorithms and data
> structures. Understanding their performance characteristics under
> contention is critical for designing high-performance concurrent
> applications.
> - The benchmark helps identify atomic operation latency and
> scalability issues across different thread counts, revealing
> contention patterns that are not visible in single-threaded tests.
> - Useful for evaluating atomic implementation quality on different
> architectures and for regression testing after changes to atomic
> primitives or memory ordering.
>
> Measurement methodology:
> - Each thread starts a private timer (clock_gettime CLOCK_MONOTONIC)
> after synchronizing on a pthread_barrier, ensuring all threads
> begin simultaneously.
> - Each thread performs a hot loop of atomic compare-and-swap on a
> shared u64 counter, incrementing from 0 to iterations.
> - The shared counter is cache-line aligned (64 bytes) to isolate
> contention to the target cache line and avoid false sharing.
> - The wall-clock time is measured as the max of all per-thread
> runtimes (the time for the slowest thread to finish).
> - The first repeat is excluded from statistics as a warmup phase
> to avoid cache-cold effects.
> Example usage:
> $ perf bench atomic cas --threads 2
> # Running 'atomic/cas' benchmark:
>
> Threads: 2, iterations/thread: 100000000, repeats: 10 (warmup: 1)
> Avg wall-clock time: 7365.480 msec (stddev 66.014 msec)
> Total ops: 200,000,000
> Throughput total: 27,153,697 ops/sec
> Per-thread times and throughput (last repeat):
> fastest: 7510.031 msec (13315525 ops/sec)
> slowest: 7581.678 msec (13189692 ops/sec)
> avg: 7545.854 msec (13252310 ops/sec)
>
> Output fields explained:
> - Threads: number of contending threads
> - iterations/thread: CAS operations each thread performs
> - repeats: number of test runs (first is warmup)
> - Avg wall-clock time: mean time for all threads to complete
> - stddev: standard deviation across repeats
> - Total ops: threads x iterations/thread
> - Throughput total: aggregate ops/sec across all threads
> - Per-thread times: fastest/slowest/avg thread completion time
> - Per-thread throughput: per-thread ops/sec (shows scheduling imbalance)
Thanks for the contribution! I think it's very useful.
Just a few suggestions.
1. it'd be nice to add simple atomic_inc benchmark too.
2. it'd be nice to have an option to try other ordering requirements
than "relaxed".
Thanks,
Namhyung
>
> Assisted-by: opencode:DeepSeek-V4-Pro
> Signed-off-by: Changbin Du <changbin.du@gmail.com>
> ---
> tools/perf/Documentation/perf-bench.txt | 22 +++
> tools/perf/bench/Build | 1 +
> tools/perf/bench/atomic.c | 229 ++++++++++++++++++++++++
> tools/perf/bench/bench.h | 1 +
> tools/perf/builtin-bench.c | 8 +
> 5 files changed, 261 insertions(+)
> create mode 100644 tools/perf/bench/atomic.c
>
> diff --git a/tools/perf/Documentation/perf-bench.txt b/tools/perf/Documentation/perf-bench.txt
> index c5913cf59c98..9cdc19f02bcf 100644
> --- a/tools/perf/Documentation/perf-bench.txt
> +++ b/tools/perf/Documentation/perf-bench.txt
> @@ -58,6 +58,9 @@ SUBSYSTEM
> 'numa'::
> NUMA scheduling and MM benchmarks.
>
> +'atomic'::
> + Atomic operation benchmarks.
> +
> 'futex'::
> Futex stressing benchmarks.
>
> @@ -283,6 +286,25 @@ SUITES FOR 'numa'
> *mem*::
> Suite for evaluating NUMA workloads.
>
> +SUITES FOR 'atomic'
> +~~~~~~~~~~~~~~~~~~~
> +*cas*::
> +Suite for evaluating compare-and-swap (CAS) atomic operations under
> +multi-threaded contention.
> +
> +Options of *cas*
> +^^^^^^^^^^^^^^^^
> +-t::
> +--threads=::
> +Number of threads contending for the shared counter (default: 2).
> +
> +-i::
> +--iterations=::
> +Number of iterations per thread (default: 100000000).
> +
> +The per-thread fastest/slowest/avg summary is only printed when
> +running with more than one thread.
> +
> SUITES FOR 'futex'
> ~~~~~~~~~~~~~~~~~~
> *hash*::
> diff --git a/tools/perf/bench/Build b/tools/perf/bench/Build
> index 67b76fe20ba6..c64c52468d3d 100644
> --- a/tools/perf/bench/Build
> +++ b/tools/perf/bench/Build
> @@ -3,6 +3,7 @@ perf-bench-y += sched-pipe.o
> perf-bench-y += sched-seccomp-notify.o
> perf-bench-y += syscall.o
> perf-bench-y += mem-functions.o
> +perf-bench-y += atomic.o
> perf-bench-y += futex.o
> perf-bench-y += futex-hash.o
> perf-bench-y += futex-wake.o
> diff --git a/tools/perf/bench/atomic.c b/tools/perf/bench/atomic.c
> new file mode 100644
> index 000000000000..ce97025018ab
> --- /dev/null
> +++ b/tools/perf/bench/atomic.c
> @@ -0,0 +1,229 @@
> +// SPDX-License-Identifier: GPL-2.0
> +/*
> + * Copyright (C) 2026 Changbin Du
> + *
> + * Benchmark for atomic compare-and-swap (CAS) operations.
> + *
> + * Measures throughput (ops/sec) of contended CAS on a single shared
> + * counter across multiple threads. Each thread performs a hot loop
> + * of compare-and-swap, competing with all other threads for the
> + * same cache line.
> + */
> +#include "bench.h"
> +#include <linux/compiler.h>
> +#include "../util/debug.h"
> +
> +#ifndef HAVE_PTHREAD_BARRIER
> +int bench_atomic(int argc __maybe_unused, const char **argv __maybe_unused)
> +{
> + pr_err("%s: pthread_barrier_t unavailable, disabling this test...\n", __func__);
> + return 0;
> +}
> +#else /* HAVE_PTHREAD_BARRIER */
> +#include <stdlib.h>
> +#include <stdio.h>
> +#include <unistd.h>
> +#include <pthread.h>
> +#include <time.h>
> +#include <inttypes.h>
> +#include <err.h>
> +#include <linux/types.h>
> +#include <linux/kernel.h>
> +#include <linux/time64.h>
> +#include <subcmd/parse-options.h>
> +#include "../util/stat.h"
> +
> +static unsigned int threads = 2;
> +static unsigned int iterations = 100000000;
> +
> +static const struct option options[] = {
> + OPT_UINTEGER('t', "threads", &threads,
> + "Number of threads contending for the counter"),
> + OPT_UINTEGER('i', "iterations", &iterations,
> + "Number of iterations per thread"),
> + OPT_END()
> +};
> +
> +static const char * const bench_usage[] = {
> + "perf bench atomic cas <options>",
> + NULL
> +};
> +
> +static u64 shared_counter __aligned(64);
> +static pthread_barrier_t start_barrier __aligned(64);
> +
> +struct worker_stats {
> + u64 runtime_ns;
> +};
> +
> +static void *worker(void *arg)
> +{
> + struct worker_stats *ws = (struct worker_stats *)arg;
> + struct timespec tstart, tend;
> + u64 i;
> +
> + pthread_barrier_wait(&start_barrier);
> +
> + clock_gettime(CLOCK_MONOTONIC, &tstart);
> + for (i = 0; i < iterations; i++) {
> + u64 old_val, new_val;
> +
> + do {
> + old_val = __atomic_load_n(&shared_counter,
> + __ATOMIC_RELAXED);
> + new_val = old_val + 1;
> + } while (!__atomic_compare_exchange_n(&shared_counter,
> + &old_val, new_val,
> + false,
> + __ATOMIC_RELAXED,
> + __ATOMIC_RELAXED));
> + }
> + clock_gettime(CLOCK_MONOTONIC, &tend);
> +
> + ws->runtime_ns = (tend.tv_sec - tstart.tv_sec) * NSEC_PER_SEC +
> + (tend.tv_nsec - tstart.tv_nsec);
> +
> + return NULL;
> +}
> +
> +static void print_default_format(struct worker_stats *wstats,
> + struct stats *time_stats)
> +{
> + double time_avg, time_stddev;
> +
> + printf(" Threads: %u, iterations/thread: %u, repeats: %u (warmup: 1)\n",
> + threads, iterations, bench_repeat);
> +
> + time_avg = avg_stats(time_stats);
> + time_stddev = stddev_stats(time_stats);
> +
> + printf(" Avg wall-clock time: %.3f msec (stddev %.3f msec)\n",
> + time_avg / (double)NSEC_PER_MSEC,
> + time_stddev / (double)NSEC_PER_MSEC);
> + printf(" Total ops: %'" PRIu64 "\n",
> + iterations * (u64)threads);
> + printf(" Throughput total: %'.0f ops/sec\n",
> + (double)(iterations * (u64)threads) /
> + (time_avg / (double)NSEC_PER_SEC));
> +
> + /*
> + * Per-thread statistics from the last measured repeat show
> + * fairness / imbalance in CPU scheduling.
> + */
> + if (threads > 1) {
> + u64 min_ns = UINT64_MAX, max_ns = 0, total_ns = 0;
> + double min_thru, max_thru, avg_thru;
> +
> + for (unsigned int t = 0; t < threads; t++) {
> + if (wstats[t].runtime_ns < min_ns)
> + min_ns = wstats[t].runtime_ns;
> + if (wstats[t].runtime_ns > max_ns)
> + max_ns = wstats[t].runtime_ns;
> + total_ns += wstats[t].runtime_ns;
> + }
> +
> + min_thru = (double)iterations * NSEC_PER_SEC / min_ns;
> + max_thru = (double)iterations * NSEC_PER_SEC / max_ns;
> + avg_thru = (double)threads * iterations * NSEC_PER_SEC / total_ns;
> +
> + printf(" Per-thread times and throughput (last repeat):\n");
> + printf(" fastest: %.3f msec (%.0f ops/sec)\n",
> + min_ns / (double)NSEC_PER_MSEC, min_thru);
> + printf(" slowest: %.3f msec (%.0f ops/sec)\n",
> + max_ns / (double)NSEC_PER_MSEC, max_thru);
> + printf(" avg: %.3f msec (%.0f ops/sec)\n",
> + (total_ns / threads) / (double)NSEC_PER_MSEC, avg_thru);
> + }
> +}
> +
> +static int run_cas_benchmark(void)
> +{
> + pthread_t *thread_ids;
> + struct worker_stats *wstats;
> + struct stats time_stats;
> + unsigned int r, t;
> +
> + thread_ids = calloc(threads, sizeof(pthread_t));
> + if (!thread_ids)
> + return -1;
> +
> + wstats = calloc(threads, sizeof(struct worker_stats));
> + if (!wstats) {
> + free(thread_ids);
> + return -1;
> + }
> +
> + init_stats(&time_stats);
> +
> + for (r = 0; r < bench_repeat + 1; r++) {
> + u64 max_runtime_ns = 0;
> +
> + shared_counter = 0;
> + pthread_barrier_init(&start_barrier, NULL, threads);
> +
> + for (t = 0; t < threads; t++) {
> + if (pthread_create(&thread_ids[t], NULL,
> + worker, &wstats[t]))
> + err(EXIT_FAILURE, "pthread_create");
> + }
> +
> + for (t = 0; t < threads; t++)
> + pthread_join(thread_ids[t], NULL);
> +
> + pthread_barrier_destroy(&start_barrier);
> +
> + for (t = 0; t < threads; t++) {
> + if (wstats[t].runtime_ns > max_runtime_ns)
> + max_runtime_ns = wstats[t].runtime_ns;
> + }
> +
> + /*
> + * Exclude the first repeat (warmup) from statistics so
> + * that cache-cold and lazy-init effects are not counted.
> + */
> + if (r > 0)
> + update_stats(&time_stats, max_runtime_ns);
> + }
> +
> + switch (bench_format) {
> + case BENCH_FORMAT_DEFAULT:
> + print_default_format(wstats, &time_stats);
> + break;
> +
> + case BENCH_FORMAT_SIMPLE:
> + printf("%.0f\n",
> + (double)(iterations * (u64)threads) /
> + (avg_stats(&time_stats) / (double)NSEC_PER_SEC));
> + break;
> +
> + default:
> + fprintf(stderr, "Unknown format: %d\n", bench_format);
> + exit(EXIT_FAILURE);
> + }
> +
> + free(thread_ids);
> + free(wstats);
> +
> + return 0;
> +}
> +
> +int bench_atomic(int argc, const char **argv)
> +{
> + if (parse_options(argc, argv, options, bench_usage, 0)) {
> + usage_with_options(bench_usage, options);
> + exit(EXIT_FAILURE);
> + }
> +
> + if (threads < 1) {
> + fprintf(stderr, "Invalid thread count: %u\n", threads);
> + return 1;
> + }
> +
> + if (iterations < 1) {
> + fprintf(stderr, "Invalid iteration count: %u\n", iterations);
> + return 1;
> + }
> +
> + return run_cas_benchmark();
> +}
> +#endif /* HAVE_PTHREAD_BARRIER */
> diff --git a/tools/perf/bench/bench.h b/tools/perf/bench/bench.h
> index 8519eb5a42fa..01eb6ceffb8d 100644
> --- a/tools/perf/bench/bench.h
> +++ b/tools/perf/bench/bench.h
> @@ -30,6 +30,7 @@ int bench_mem_memcpy(int argc, const char **argv);
> int bench_mem_memset(int argc, const char **argv);
> int bench_mem_mmap(int argc, const char **argv);
> int bench_mem_find_bit(int argc, const char **argv);
> +int bench_atomic(int argc, const char **argv);
> int bench_futex_hash(int argc, const char **argv);
> int bench_futex_wake(int argc, const char **argv);
> int bench_futex_wake_parallel(int argc, const char **argv);
> diff --git a/tools/perf/builtin-bench.c b/tools/perf/builtin-bench.c
> index 02d47913cc6a..aeffa0826aea 100644
> --- a/tools/perf/builtin-bench.c
> +++ b/tools/perf/builtin-bench.c
> @@ -14,6 +14,7 @@
> * syscall ... System call performance
> * mem ... memory access performance
> * numa ... NUMA scheduling and MM performance
> + * atomic ... Atomic operation performance
> * futex ... Futex performance
> * epoll ... Event poll performance
> */
> @@ -115,6 +116,12 @@ static const struct bench uprobe_benchmarks[] = {
> { NULL, NULL, NULL },
> };
>
> +static const struct bench atomic_benchmarks[] = {
> + { "cas", "Benchmark CAS (compare-and-swap) operations", bench_atomic },
> + { "all", "Run all atomic benchmarks", NULL },
> + { NULL, NULL, NULL }
> +};
> +
> struct collection {
> const char *name;
> const char *summary;
> @@ -128,6 +135,7 @@ static const struct collection collections[] = {
> #ifdef HAVE_LIBNUMA_SUPPORT
> { "numa", "NUMA scheduling and MM benchmarks", numa_benchmarks },
> #endif
> + { "atomic", "Atomic operation benchmarks", atomic_benchmarks },
> {"futex", "Futex stressing benchmarks", futex_benchmarks },
> #ifdef HAVE_EVENTFD_SUPPORT
> {"epoll", "Epoll stressing benchmarks", epoll_benchmarks },
> --
> 2.53.0
>
^ permalink raw reply [flat|nested] 5+ messages in thread* Re: [PATCH] perf bench: Add atomic CAS benchmark
2026-10-05 22:30 ` Namhyung Kim
@ 2026-10-07 5:27 ` Changbin Du
2026-10-07 21:35 ` Namhyung Kim
0 siblings, 1 reply; 5+ messages in thread
From: Changbin Du @ 2026-10-07 5:27 UTC (permalink / raw)
To: Namhyung Kim
Cc: Changbin Du, Peter Zijlstra, Ingo Molnar,
Arnaldo Carvalho de Melo, Mark Rutland, Alexander Shishkin,
Jiri Olsa, Ian Rogers, Adrian Hunter, James Clark, linux-kernel,
linux-perf-users
Hello,
On Mon, Oct 05, 2026 at 03:30:46PM -0700, Namhyung Kim wrote:
> Hello,
>
> On Wed, Sep 30, 2026 at 05:16:17PM +0800, Changbin Du wrote:
> > Add a new 'atomic' collection to perf bench for benchmarking
> > compare-and-swap (CAS) atomic operations with multi-threaded
> > contention testing.
> >
> > The benchmark tests __atomic_compare_exchange_n operations
> > with configurable thread count and iteration count to measure
> > atomic contention effects.
> >
> > Why this benchmark is needed:
> > - CAS operations are fundamental to lock-free algorithms and data
> > structures. Understanding their performance characteristics under
> > contention is critical for designing high-performance concurrent
> > applications.
> > - The benchmark helps identify atomic operation latency and
> > scalability issues across different thread counts, revealing
> > contention patterns that are not visible in single-threaded tests.
> > - Useful for evaluating atomic implementation quality on different
> > architectures and for regression testing after changes to atomic
> > primitives or memory ordering.
> >
> > Measurement methodology:
> > - Each thread starts a private timer (clock_gettime CLOCK_MONOTONIC)
> > after synchronizing on a pthread_barrier, ensuring all threads
> > begin simultaneously.
> > - Each thread performs a hot loop of atomic compare-and-swap on a
> > shared u64 counter, incrementing from 0 to iterations.
> > - The shared counter is cache-line aligned (64 bytes) to isolate
> > contention to the target cache line and avoid false sharing.
> > - The wall-clock time is measured as the max of all per-thread
> > runtimes (the time for the slowest thread to finish).
> > - The first repeat is excluded from statistics as a warmup phase
> > to avoid cache-cold effects.
> > Example usage:
> > $ perf bench atomic cas --threads 2
> > # Running 'atomic/cas' benchmark:
> >
> > Threads: 2, iterations/thread: 100000000, repeats: 10 (warmup: 1)
> > Avg wall-clock time: 7365.480 msec (stddev 66.014 msec)
> > Total ops: 200,000,000
> > Throughput total: 27,153,697 ops/sec
> > Per-thread times and throughput (last repeat):
> > fastest: 7510.031 msec (13315525 ops/sec)
> > slowest: 7581.678 msec (13189692 ops/sec)
> > avg: 7545.854 msec (13252310 ops/sec)
> >
> > Output fields explained:
> > - Threads: number of contending threads
> > - iterations/thread: CAS operations each thread performs
> > - repeats: number of test runs (first is warmup)
> > - Avg wall-clock time: mean time for all threads to complete
> > - stddev: standard deviation across repeats
> > - Total ops: threads x iterations/thread
> > - Throughput total: aggregate ops/sec across all threads
> > - Per-thread times: fastest/slowest/avg thread completion time
> > - Per-thread throughput: per-thread ops/sec (shows scheduling imbalance)
>
> Thanks for the contribution! I think it's very useful.
> Just a few suggestions.
>
> 1. it'd be nice to add simple atomic_inc benchmark too.
> 2. it'd be nice to have an option to try other ordering requirements
> than "relaxed".
>
> Thanks,
> Namhyung
Thanks for the review!
1. Done in v2. The collection now provides atomic inc alongside
cas, measuring contended __atomic_fetch_add() throughput with the
same skeleton (barrier-synchronized start, wall-clock taken from the
slowest thread, warmup repeat). Note that its ops/sec is not
instruction-level comparable with cas — cas counts successful
compare-and-swaps only, not the retries and loads in between — which
the documentation now points out.
2. Adding memory-order variants is not necessary because ordering
has no semantic role in this benchmark. Memory ordering exists to
constrain the visibility of other memory operations relative to an
atomic access; its cost and effect only become meaningful in an
algorithm with additional accesses to order — for example, ordering
the initialization stores of a new node before publishing a pointer
with a release CAS. This benchmark, however, measures a single shared
counter with no other memory operations in the loop, so a stronger
ordering would not order anything of consequence: it would only
measure the marginal cost of the stronger instruction itself. Those
numbers would say nothing about how orderings behave in real
workloads, and would therefore add a configuration knob that invites
misleading comparisons rather than useful measurement.
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH] perf bench: Add atomic CAS benchmark
2026-10-07 5:27 ` Changbin Du
@ 2026-10-07 21:35 ` Namhyung Kim
0 siblings, 0 replies; 5+ messages in thread
From: Namhyung Kim @ 2026-10-07 21:35 UTC (permalink / raw)
To: Changbin Du
Cc: Peter Zijlstra, Ingo Molnar, Arnaldo Carvalho de Melo,
Mark Rutland, Alexander Shishkin, Jiri Olsa, Ian Rogers,
Adrian Hunter, James Clark, linux-kernel, linux-perf-users
On Wed, Oct 07, 2026 at 01:27:36PM +0800, Changbin Du wrote:
> Hello,
> On Mon, Oct 05, 2026 at 03:30:46PM -0700, Namhyung Kim wrote:
> > Hello,
> >
> > On Wed, Sep 30, 2026 at 05:16:17PM +0800, Changbin Du wrote:
> > > Add a new 'atomic' collection to perf bench for benchmarking
> > > compare-and-swap (CAS) atomic operations with multi-threaded
> > > contention testing.
> > >
> > > The benchmark tests __atomic_compare_exchange_n operations
> > > with configurable thread count and iteration count to measure
> > > atomic contention effects.
> > >
> > > Why this benchmark is needed:
> > > - CAS operations are fundamental to lock-free algorithms and data
> > > structures. Understanding their performance characteristics under
> > > contention is critical for designing high-performance concurrent
> > > applications.
> > > - The benchmark helps identify atomic operation latency and
> > > scalability issues across different thread counts, revealing
> > > contention patterns that are not visible in single-threaded tests.
> > > - Useful for evaluating atomic implementation quality on different
> > > architectures and for regression testing after changes to atomic
> > > primitives or memory ordering.
> > >
> > > Measurement methodology:
> > > - Each thread starts a private timer (clock_gettime CLOCK_MONOTONIC)
> > > after synchronizing on a pthread_barrier, ensuring all threads
> > > begin simultaneously.
> > > - Each thread performs a hot loop of atomic compare-and-swap on a
> > > shared u64 counter, incrementing from 0 to iterations.
> > > - The shared counter is cache-line aligned (64 bytes) to isolate
> > > contention to the target cache line and avoid false sharing.
> > > - The wall-clock time is measured as the max of all per-thread
> > > runtimes (the time for the slowest thread to finish).
> > > - The first repeat is excluded from statistics as a warmup phase
> > > to avoid cache-cold effects.
> > > Example usage:
> > > $ perf bench atomic cas --threads 2
> > > # Running 'atomic/cas' benchmark:
> > >
> > > Threads: 2, iterations/thread: 100000000, repeats: 10 (warmup: 1)
> > > Avg wall-clock time: 7365.480 msec (stddev 66.014 msec)
> > > Total ops: 200,000,000
> > > Throughput total: 27,153,697 ops/sec
> > > Per-thread times and throughput (last repeat):
> > > fastest: 7510.031 msec (13315525 ops/sec)
> > > slowest: 7581.678 msec (13189692 ops/sec)
> > > avg: 7545.854 msec (13252310 ops/sec)
> > >
> > > Output fields explained:
> > > - Threads: number of contending threads
> > > - iterations/thread: CAS operations each thread performs
> > > - repeats: number of test runs (first is warmup)
> > > - Avg wall-clock time: mean time for all threads to complete
> > > - stddev: standard deviation across repeats
> > > - Total ops: threads x iterations/thread
> > > - Throughput total: aggregate ops/sec across all threads
> > > - Per-thread times: fastest/slowest/avg thread completion time
> > > - Per-thread throughput: per-thread ops/sec (shows scheduling imbalance)
> >
> > Thanks for the contribution! I think it's very useful.
> > Just a few suggestions.
> >
> > 1. it'd be nice to add simple atomic_inc benchmark too.
> > 2. it'd be nice to have an option to try other ordering requirements
> > than "relaxed".
> >
> > Thanks,
> > Namhyung
>
> Thanks for the review!
> 1. Done in v2. The collection now provides atomic inc alongside
> cas, measuring contended __atomic_fetch_add() throughput with the
> same skeleton (barrier-synchronized start, wall-clock taken from the
> slowest thread, warmup repeat). Note that its ops/sec is not
> instruction-level comparable with cas — cas counts successful
> compare-and-swaps only, not the retries and loads in between — which
> the documentation now points out.
>
> 2. Adding memory-order variants is not necessary because ordering
> has no semantic role in this benchmark. Memory ordering exists to
> constrain the visibility of other memory operations relative to an
> atomic access; its cost and effect only become meaningful in an
> algorithm with additional accesses to order — for example, ordering
> the initialization stores of a new node before publishing a pointer
> with a release CAS. This benchmark, however, measures a single shared
> counter with no other memory operations in the loop, so a stronger
> ordering would not order anything of consequence: it would only
> measure the marginal cost of the stronger instruction itself. Those
> numbers would say nothing about how orderings behave in real
> workloads, and would therefore add a configuration knob that invites
> misleading comparisons rather than useful measurement.
Agreed that it has no semantics here. Maybe we can add the option with
a meaningful scenario later.
Thanks,
Namhyung
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH] perf bench: Add atomic CAS benchmark
2026-09-30 9:16 [PATCH] perf bench: Add atomic CAS benchmark Changbin Du
2026-10-05 22:30 ` Namhyung Kim
@ 2026-10-08 7:44 ` David Laight
1 sibling, 0 replies; 5+ messages in thread
From: David Laight @ 2026-10-08 7:44 UTC (permalink / raw)
To: Changbin Du
Cc: Peter Zijlstra, Ingo Molnar, Arnaldo Carvalho de Melo,
Namhyung Kim, Mark Rutland, Alexander Shishkin, Jiri Olsa,
Ian Rogers, Adrian Hunter, James Clark, linux-kernel,
linux-perf-users
On Wed, 30 Sep 2026 17:16:17 +0800
Changbin Du <changbin.du@gmail.com> wrote:
> Add a new 'atomic' collection to perf bench for benchmarking
> compare-and-swap (CAS) atomic operations with multi-threaded
> contention testing.
>
> The benchmark tests __atomic_compare_exchange_n operations
> with configurable thread count and iteration count to measure
> atomic contention effects.
>
> Why this benchmark is needed:
> - CAS operations are fundamental to lock-free algorithms and data
> structures. Understanding their performance characteristics under
> contention is critical for designing high-performance concurrent
> applications.
> - The benchmark helps identify atomic operation latency and
> scalability issues across different thread counts, revealing
> contention patterns that are not visible in single-threaded tests.
> - Useful for evaluating atomic implementation quality on different
> architectures and for regression testing after changes to atomic
> primitives or memory ordering.
>
> Measurement methodology:
> - Each thread starts a private timer (clock_gettime CLOCK_MONOTONIC)
> after synchronizing on a pthread_barrier, ensuring all threads
> begin simultaneously.
> - Each thread performs a hot loop of atomic compare-and-swap on a
> shared u64 counter, incrementing from 0 to iterations.
> - The shared counter is cache-line aligned (64 bytes) to isolate
> contention to the target cache line and avoid false sharing.
> - The wall-clock time is measured as the max of all per-thread
> runtimes (the time for the slowest thread to finish).
> - The first repeat is excluded from statistics as a warmup phase
> to avoid cache-cold effects.
> Example usage:
> $ perf bench atomic cas --threads 2
> # Running 'atomic/cas' benchmark:
>
> Threads: 2, iterations/thread: 100000000, repeats: 10 (warmup: 1)
> Avg wall-clock time: 7365.480 msec (stddev 66.014 msec)
> Total ops: 200,000,000
> Throughput total: 27,153,697 ops/sec
> Per-thread times and throughput (last repeat):
> fastest: 7510.031 msec (13315525 ops/sec)
> slowest: 7581.678 msec (13189692 ops/sec)
> avg: 7545.854 msec (13252310 ops/sec)
>
> Output fields explained:
> - Threads: number of contending threads
> - iterations/thread: CAS operations each thread performs
> - repeats: number of test runs (first is warmup)
> - Avg wall-clock time: mean time for all threads to complete
> - stddev: standard deviation across repeats
> - Total ops: threads x iterations/thread
> - Throughput total: aggregate ops/sec across all threads
> - Per-thread times: fastest/slowest/avg thread completion time
> - Per-thread throughput: per-thread ops/sec (shows scheduling imbalance)
>
I think you need to default to one thread per cpu.
Also try to run the test for a fixed time period rather than a very
large count.
You should be able to see that some systems completely fail to make
progress under very heavy contention.
(This isn't one thread getting starved, none of them make progress.)
David
^ permalink raw reply [flat|nested] 5+ messages in thread
end of thread, other threads:[~2026-10-08 7:44 UTC | newest]
Thread overview: 5+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-30 9:16 [PATCH] perf bench: Add atomic CAS benchmark Changbin Du
2026-10-05 22:30 ` Namhyung Kim
2026-10-07 5:27 ` Changbin Du
2026-10-07 21:35 ` Namhyung Kim
2026-10-08 7:44 ` David Laight
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®