mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Namhyung Kim <namhyung@kernel.org>
To: Arnaldo Carvalho de Melo <acme@kernel.org>
Cc: Ingo Molnar <mingo@kernel.org>,
	Thomas Gleixner <tglx@linutronix.de>,
	James Clark <james.clark@linaro.org>,
	Jiri Olsa <jolsa@kernel.org>, Ian Rogers <irogers@google.com>,
	Adrian Hunter <adrian.hunter@intel.com>,
	Clark Williams <williams@redhat.com>,
	linux-kernel@vger.kernel.org, linux-perf-users@vger.kernel.org,
	Arnaldo Carvalho de Melo <acme@redhat.com>
Subject: Re: [PATCH v6 5/5] perf test: Add false_sharing workload exhibiting cross-CPU false sharing
Date: Thu, 1 Oct 2026 00:28:13 -0700	[thread overview]
Message-ID: <ar4LjQ8ryDjWDQ9y@z2> (raw)
In-Reply-To: <20260930213716.2633750-6-acme@kernel.org>

On Wed, Sep 30, 2026 at 11:37:16PM +0200, Arnaldo Carvalho de Melo wrote:
> From: Arnaldo Carvalho de Melo <acme@redhat.com>
> 
> Add a 'perf test -w false_sharing' workload that hammers one shared
> struct from several CPUs, shaped as a TCP connection: a read-mostly
> identity (five-tuple) shares a cacheline with per-packet rx counters
> (the false-sharing line), a second line has packet-path private tx and
> congestion control counters, and a third the connection config.
> 
> The packet path runs in the main thread and up to four lookup threads
> sum the five-tuple and pull the config, reading one volatile shared
> instance directly so the accesses are PC-relative and resolvable by the
> data type profiler.

It'd be great if you can share an output of data type profiling with
cacheline info.  Probably like below?

  $ perf mem record -- perf test -w false_sharing
 
  $ perf report -s type,typecln -H --group --stdio

Thanks,
Namhyung

> 
> Assisted-by: LLM
> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
> ---
>  tools/perf/tests/builtin-test.c               |   1 +
>  tools/perf/tests/shell/data_type_profiling.sh |   9 +-
>  tools/perf/tests/tests.h                      |   1 +
>  tools/perf/tests/workloads/Build              |   2 +
>  tools/perf/tests/workloads/false_sharing.c    | 251 ++++++++++++++++++
>  5 files changed, 262 insertions(+), 2 deletions(-)
>  create mode 100644 tools/perf/tests/workloads/false_sharing.c
> 
> diff --git a/tools/perf/tests/builtin-test.c b/tools/perf/tests/builtin-test.c
> index d2f594921e25bda9..98134b1c74cf80f4 100644
> --- a/tools/perf/tests/builtin-test.c
> +++ b/tools/perf/tests/builtin-test.c
> @@ -175,6 +175,7 @@ static struct test_workload *workloads[] = {
>  	&workload__context_switch_loop,
>  	&workload__deterministic,
>  	&workload__callchain,
> +	&workload__false_sharing,
>  
>  #ifdef HAVE_RUST_SUPPORT
>  	&workload__code_with_type,
> diff --git a/tools/perf/tests/shell/data_type_profiling.sh b/tools/perf/tests/shell/data_type_profiling.sh
> index a916c410274aa888..57203e0e5f857540 100755
> --- a/tools/perf/tests/shell/data_type_profiling.sh
> +++ b/tools/perf/tests/shell/data_type_profiling.sh
> @@ -8,8 +8,8 @@ set -e
>  # data type profiling manifestation
>  
>  # Values in testtypes and testprogs should match
> -testtypes=("# data-type: struct Buf" "# data-type: struct buf")
> -testprogs=("perf test -w code_with_type" "perf test -w datasym")
> +testtypes=("# data-type: struct Buf" "# data-type: struct buf" "# data-type: struct net_conn")
> +testprogs=("perf test -w code_with_type" "perf test -w datasym" "perf test -w false_sharing")
>  
>  err=0
>  perfdata=$(mktemp /tmp/__perf_test.perf.data.XXXXX)
> @@ -59,6 +59,9 @@ test_basic_annotate() {
>  
>      "xC")
>      index=1 ;;
> +
> +    "xFS")
> +    index=2 ;;
>    esac
>  
>    # Under 'set -e' a bare failing command aborts the script through the EXIT
> @@ -114,6 +117,8 @@ test_basic_annotate Basic Rust
>  test_basic_annotate Pipe Rust
>  test_basic_annotate Basic C
>  test_basic_annotate Pipe C
> +test_basic_annotate Basic FS
> +test_basic_annotate Pipe FS
>  
>  cleanup
>  exit $err
> diff --git a/tools/perf/tests/tests.h b/tools/perf/tests/tests.h
> index 9c96f33483d14356..f379620069d34f3a 100644
> --- a/tools/perf/tests/tests.h
> +++ b/tools/perf/tests/tests.h
> @@ -250,6 +250,7 @@ DECLARE_WORKLOAD(jitdump);
>  DECLARE_WORKLOAD(context_switch_loop);
>  DECLARE_WORKLOAD(deterministic);
>  DECLARE_WORKLOAD(callchain);
> +DECLARE_WORKLOAD(false_sharing);
>  
>  #ifdef HAVE_RUST_SUPPORT
>  DECLARE_WORKLOAD(code_with_type);
> diff --git a/tools/perf/tests/workloads/Build b/tools/perf/tests/workloads/Build
> index 048e371eb63e3164..18fceddccb56c013 100644
> --- a/tools/perf/tests/workloads/Build
> +++ b/tools/perf/tests/workloads/Build
> @@ -14,6 +14,7 @@ perf-test-y += jitdump.o
>  perf-test-y += context_switch_loop.o
>  perf-test-y += deterministic.o
>  perf-test-y += callchain.o
> +perf-test-y += false_sharing.o
>  
>  ifeq ($(CONFIG_RUST_SUPPORT),y)
>      perf-test-y += code_with_type.o
> @@ -29,3 +30,4 @@ CFLAGS_inlineloop.o       = -g -O2
>  CFLAGS_deterministic.o    = -g -O0 -fno-inline -U_FORTIFY_SOURCE
>  CFLAGS_named_threads.o    = -g -O0 -fno-inline -U_FORTIFY_SOURCE
>  CFLAGS_callchain.o        = -g -O0 -fno-inline -U_FORTIFY_SOURCE
> +CFLAGS_false_sharing.o    = -g -O0 -fno-inline -U_FORTIFY_SOURCE
> diff --git a/tools/perf/tests/workloads/false_sharing.c b/tools/perf/tests/workloads/false_sharing.c
> new file mode 100644
> index 0000000000000000..e949bb36646a97a5
> --- /dev/null
> +++ b/tools/perf/tests/workloads/false_sharing.c
> @@ -0,0 +1,251 @@
> +// SPDX-License-Identifier: GPL-2.0
> +/*
> + * False-sharing demo for data type profiling, shaped as a TCP
> + * connection: a read-mostly identity shares a cacheline with per-packet
> + * rx counters (the false-sharing line), a second line has packet-path
> + * private tx and congestion control counters, and a third the connection
> + * config.
> + *
> + * 'perf mem record' of this workload followed by 'perf report -s type'
> + * (see tests/shell/data_type_profiling.sh) resolves the accesses to
> + * struct net_conn members, showing the rx counters and the read-mostly
> + * identity sharing cacheline 0.
> + */
> +#include <pthread.h>
> +#include <sched.h>
> +#include <stdint.h>
> +#include <stdlib.h>
> +#include <stdio.h>
> +#include <signal.h>
> +#include <unistd.h>
> +#include <linux/compiler.h>
> +#include "../tests.h"
> +
> +struct net_conn {
> +	/* cacheline 0: identity (read-mostly) + rx counters (per packet) */
> +	uint32_t	saddr;		/*   0 */
> +	uint32_t	daddr;		/*   4 */
> +	uint16_t	sport;		/*   8 */
> +	uint16_t	dport;		/*  10 */
> +	uint8_t		state;		/*  12: 1 == ESTABLISHED */
> +	uint8_t		protocol;	/*  13: 6 == TCP */
> +	uint16_t	__pad0;		/*  14 */
> +	uint64_t	bytes_rx;	/*  16: every packet */
> +	uint64_t	packets_rx;	/*  24: every packet */
> +	uint32_t	rx_queue;	/*  32: backlog depth, fluctuates */
> +	uint8_t		__pad1[24];	/*  36..59 */
> +	uint32_t	last_ack;	/*  60: written per ACK */
> +	/* cacheline 1: tx + congestion control (packet-path private) */
> +	uint64_t	bytes_tx;	/*  64: every packet */
> +	uint64_t	packets_tx;	/*  72: every packet */
> +	uint32_t	cwnd;		/*  80: on every ACK */
> +	uint32_t	ssthresh;	/*  84: on loss */
> +	uint32_t	rtt_us;		/*  88: on every ACK */
> +	uint32_t	retrans;	/*  92: on timeout */
> +	uint32_t	__pad2[8];	/*  96..127 */
> +	/* cacheline 2: config, set at setup, read by everybody */
> +	uint16_t	mss;		/* 128 */
> +	uint8_t		snd_wscale;	/* 130 */
> +	uint8_t		rcv_wscale;	/* 131 */
> +	uint32_t	keepalive_int;	/* 132 */
> +	uint32_t	mark;		/* 136: firewall mark */
> +	uint32_t	priority;	/* 140: traffic class */
> +	uint32_t	__pad3[12];	/* 144..191 */
> +} __attribute__((aligned(64)));
> +
> +/* Volatile so every iteration really loads and stores. */
> +static volatile struct net_conn conn;
> +/* Keeps the reader checksums alive after the threads join. */
> +static volatile unsigned long fs_sink;
> +
> +static volatile sig_atomic_t done;
> +
> +/*
> + * One cacheline each: sum before cpu, or the implicit padding after cpu
> + * pushes the struct past 64 bytes, and aligning it to a cacheline then
> + * rounds it up to 128.
> + */
> +struct fs_reader {
> +	pthread_t	thread;
> +	unsigned long	sum;
> +	int		cpu;
> +	char		__pad[64 - sizeof(pthread_t) - sizeof(unsigned long) - sizeof(int)];
> +} __attribute__((aligned(64)));
> +
> +static void sighandler(int sig __maybe_unused)
> +{
> +	done = 1;
> +}
> +
> +static void pin_to_cpu(int cpu)
> +{
> +	cpu_set_t set;
> +
> +	/* There may be no second CPU in a restricted cpuset. */
> +	if (cpu < 0)
> +		return;
> +
> +	CPU_ZERO(&set);
> +	CPU_SET(cpu, &set);
> +	/* Best effort: in a restricted cpuset this fails and the thread runs unpinned. */
> +	pthread_setaffinity_np(pthread_self(), sizeof(set), &set);
> +}
> +
> +/*
> + * Connection lookup, as a load balancer or 'ss' scrape would do it: the
> + * reads go straight to the global, this file is built -O0 and a local
> + * pointer would be reloaded from a stack slot, a form the data type
> + * resolver does not track.
> + */
> +static void *reader_fn(void *arg)
> +{
> +	struct fs_reader *r = arg;
> +	unsigned long sum = 0;
> +
> +	pthread_setname_np(pthread_self(), "fs-reader");
> +	pin_to_cpu(r->cpu);
> +
> +	while (!done) {
> +		sum += conn.saddr + conn.daddr + conn.sport + conn.dport +
> +		       conn.state + conn.protocol;
> +		sum += conn.mss + conn.snd_wscale + conn.rcv_wscale +
> +		       conn.keepalive_int + conn.mark + conn.priority;
> +	}
> +	r->sum = sum;
> +	return NULL;
> +}
> +
> +static int false_sharing(int argc, const char **argv)
> +{
> +	double sec = 2.0;
> +	int nreaders = 0, nr_allowed = 0, err = 1;
> +	int *allowed = NULL, nallowed = 0;
> +	cpu_set_t set;
> +	int nr_mask_bits = sizeof(set) * 8 < CPU_SETSIZE ? sizeof(set) * 8 : CPU_SETSIZE;
> +	struct fs_reader *readers = NULL;
> +	int i, writer_cpu;
> +	unsigned long n = 0;
> +
> +	pthread_setname_np(pthread_self(), "fs-writer");
> +	if (argc > 0)
> +		sec = atof(argv[0]);
> +	if (!(sec > 0.0)) {
> +		fprintf(stderr, "Error: seconds (%f) must be > 0\n", sec);
> +		return 1;
> +	}
> +	if (argc > 1)
> +		nreaders = atoi(argv[1]);
> +
> +	/*
> +	 * A connection that just got established: identity and config fixed
> +	 * from here on, counters at zero.
> +	 */
> +	conn.saddr = 0x0a000001;	/* 10.0.0.1 */
> +	conn.daddr = 0x0a000002;	/* 10.0.0.2 */
> +	conn.sport = 54321;
> +	conn.dport = 443;
> +	conn.state = 1;			/* ESTABLISHED */
> +	conn.protocol = 6;		/* TCP */
> +	conn.mss = 1448;
> +	conn.snd_wscale = 7;
> +	conn.rcv_wscale = 7;
> +	conn.keepalive_int = 7200;
> +	conn.cwnd = 10;
> +	conn.ssthresh = 65535;
> +	conn.rtt_us = 50;
> +
> +	/*
> +	 * Pin against the allowed set, restricted cpusets still spread the threads.
> +	 * The whole mask is looked at, not the CPU count: the count can be lower
> +	 * than the highest ID in it, as when a cpuset allows only high numbered
> +	 * CPUs, and then no allowed CPU would be found at all.
> +	 */
> +	if (sched_getaffinity(0, sizeof(set), &set) == 0) {
> +		for (i = 0; i < nr_mask_bits; i++) {
> +			if (!CPU_ISSET(i, &set))
> +				continue;
> +			nr_allowed++;
> +		}
> +		allowed = malloc(nr_allowed * sizeof(int));
> +		if (allowed == NULL) {
> +			fprintf(stderr, "Error: malloc failed for CPU list\n");
> +			return 1;
> +		}
> +		for (i = 0; i < nr_mask_bits; i++) {
> +			if (CPU_ISSET(i, &set))
> +				allowed[nallowed++] = i;
> +		}
> +	}
> +	if (nreaders <= 0) {
> +		/* By default leave one CPU for the packet path, up to 4 readers. */
> +		nreaders = nallowed > 1 ? nallowed - 1 : 1;
> +		if (nreaders > 4)
> +			nreaders = 4;
> +	}
> +
> +	signal(SIGINT, sighandler);
> +	signal(SIGALRM, sighandler);
> +
> +	readers = calloc(nreaders, sizeof(*readers));
> +	if (readers == NULL) {
> +		fprintf(stderr, "Error: calloc failed for %d readers\n", nreaders);
> +		goto out;
> +	}
> +	for (i = 0; i < nreaders; i++) {
> +		int cpu = nallowed > 1 ? allowed[(i + 1) % nallowed] : -1;
> +
> +		readers[i].cpu = cpu;
> +		if (pthread_create(&readers[i].thread, NULL, reader_fn, &readers[i])) {
> +			fprintf(stderr, "Error: failed to create reader %d\n", i);
> +			done = 1; // Ensure started threads terminate.
> +			nreaders = i;
> +			goto out_join;
> +		}
> +	}
> +	writer_cpu = nallowed > 0 ? allowed[0] : -1;
> +	if (nallowed == 1)
> +		fprintf(stderr, "Warning: single CPU allowed, no cross-CPU traffic expected\n");
> +	if (writer_cpu >= 0)
> +		pin_to_cpu(writer_cpu);
> +
> +	/*
> +	 * The packet path: receive, acknowledge, transmit, repeat; every 64th
> +	 * packet simulates a loss so ssthresh/retrans get sampled too.
> +	 */
> +	if (sec < 1.0) {
> +		useconds_t usecs = (useconds_t)(sec * 1000000.0);
> +
> +		ualarm(usecs > 0 ? usecs : 1, 0);
> +	} else
> +		alarm((unsigned int)sec);
> +	while (!done) {
> +		conn.bytes_rx += conn.mss;
> +		conn.packets_rx++;
> +		conn.rx_queue = (uint32_t)(n & 0x3f);
> +		conn.last_ack = (uint32_t)n;
> +		conn.bytes_tx += conn.mss;
> +		conn.packets_tx++;
> +		conn.cwnd = 10 + (n & 15);
> +		conn.rtt_us = 50 + (n & 7);
> +		if ((n & 63) == 0) {
> +			conn.ssthresh = conn.cwnd / 2;
> +			conn.retrans++;
> +		}
> +		n++;
> +	}
> +	err = 0;
> +out_join:
> +	for (i = 0; i < nreaders; i++) {
> +		if (readers[i].thread) {
> +			pthread_join(readers[i].thread, /*retval=*/NULL);
> +			fs_sink += readers[i].sum;
> +		}
> +	}
> +	fs_sink += (unsigned long)(conn.bytes_rx + conn.bytes_tx + n);
> +	free(readers);
> +out:
> +	free(allowed);
> +	return err;
> +}
> +
> +DEFINE_WORKLOAD(false_sharing);
> -- 
> 2.55.0
> 

  reply	other threads:[~2026-10-01  7:28 UTC|newest]

Thread overview: 15+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-30 21:37 [PATCH v6 0/5] perf tools: Add progress diagnostics and a false-sharing workload Arnaldo Carvalho de Melo
2026-09-30 21:37 ` [PATCH v6 1/5] perf config: Move perf_config__set_variable() to util/config.c Arnaldo Carvalho de Melo
2026-10-01  7:18   ` Namhyung Kim
2026-10-01  9:19     ` Arnaldo Carvalho de Melo
2026-09-30 21:37 ` [PATCH v6 2/5] perf report: Add --progress option Arnaldo Carvalho de Melo
2026-09-30 21:37 ` [PATCH v6 3/5] perf report: Add --no-progress option Arnaldo Carvalho de Melo
2026-10-01  7:01   ` Namhyung Kim
2026-10-01  9:20     ` Arnaldo Carvalho de Melo
2026-09-30 21:37 ` [PATCH v6 4/5] perf scripts: Add perf-stuck, to tell where a running perf is stuck Arnaldo Carvalho de Melo
2026-10-01  7:24   ` Namhyung Kim
2026-10-01  9:19     ` Arnaldo Carvalho de Melo
2026-09-30 21:37 ` [PATCH v6 5/5] perf test: Add false_sharing workload exhibiting cross-CPU false sharing Arnaldo Carvalho de Melo
2026-10-01  7:28   ` Namhyung Kim [this message]
2026-10-01  9:18     ` Arnaldo Carvalho de Melo
  -- strict thread matches above, loose matches on Subject: below --
2026-09-30 11:24 [PATCH v6 0/5] perf tools: Add progress diagnostics and a false-sharing workload Arnaldo Carvalho de Melo
2026-09-30 11:24 ` [PATCH v6 5/5] perf test: Add false_sharing workload exhibiting cross-CPU false sharing Arnaldo Carvalho de Melo

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=ar4LjQ8ryDjWDQ9y@z2 \
    --to=namhyung@kernel.org \
    --cc=acme@kernel.org \
    --cc=acme@redhat.com \
    --cc=adrian.hunter@intel.com \
    --cc=irogers@google.com \
    --cc=james.clark@linaro.org \
    --cc=jolsa@kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-perf-users@vger.kernel.org \
    --cc=mingo@kernel.org \
    --cc=tglx@linutronix.de \
    --cc=williams@redhat.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®