From: Arnaldo Carvalho de Melo <acme@kernel.org>
To: Namhyung Kim <namhyung@kernel.org>
Cc: Ingo Molnar <mingo@kernel.org>,
Thomas Gleixner <tglx@linutronix.de>,
James Clark <james.clark@linaro.org>,
Jiri Olsa <jolsa@kernel.org>, Ian Rogers <irogers@google.com>,
Adrian Hunter <adrian.hunter@intel.com>,
Clark Williams <williams@redhat.com>,
linux-kernel@vger.kernel.org, linux-perf-users@vger.kernel.org,
Arnaldo Carvalho de Melo <acme@redhat.com>
Subject: Re: [PATCH v6 5/5] perf test: Add false_sharing workload exhibiting cross-CPU false sharing
Date: Thu, 1 Oct 2026 11:18:54 +0200 [thread overview]
Message-ID: <ar4lfjssCX9H2CCL@x2> (raw)
In-Reply-To: <ar4LjQ8ryDjWDQ9y@z2>
On Thu, Oct 01, 2026 at 12:28:13AM -0700, Namhyung Kim wrote:
> On Wed, Sep 30, 2026 at 11:37:16PM +0200, Arnaldo Carvalho de Melo wrote:
> > Add a 'perf test -w false_sharing' workload that hammers one shared
> > struct from several CPUs, shaped as a TCP connection: a read-mostly
> > identity (five-tuple) shares a cacheline with per-packet rx counters
> > (the false-sharing line), a second line has packet-path private tx and
> > congestion control counters, and a third the connection config.
> > The packet path runs in the main thread and up to four lookup threads
> > sum the five-tuple and pull the config, reading one volatile shared
> > instance directly so the accesses are PC-relative and resolvable by the
> > data type profiler.
>
> It'd be great if you can share an output of data type profiling with
> cacheline info. Probably like below?
>
> $ perf mem record -- perf test -w false_sharing
>
> $ perf report -s type,typecln -H --group --stdio
Sure, I should've added it there as I usually do :-\
Here it is, will add to the cset message as well:
root@x2:~# perf mem record -- perf test -r5 -w false_sharing
[ perf record: Woken up 19 times to write data ]
[ perf record: Captured and wrote 5.612 MB perf.data (72454 samples) ]
root@x2:~#
root@x2:~# perf report -s type,typecln -H --group --stdio
# Total Lost Samples: 0
#
# Samples: 72K of events 'cpu/mem-loads,ldlat=30/P, cpu/mem-stores/P'
# Event count (approx.): 1261814616
#
# Overhead Data Type / Data Type Cacheline
# ................... ...............................
#
88.99% 35.88% struct net_conn
56.43% 17.50% struct net_conn: cache-line 0
32.55% 0.00% struct net_conn: cache-line 2
0.01% 18.38% struct net_conn: cache-line 1
10.08% 63.03% (unknown)
10.08% 63.03% (unknown): cache-line 0
0.82% 0.13% int
0.82% 0.13% int: cache-line 0
0.01% 0.02% struct folio
0.01% 0.02% struct folio: cache-line 0
0.01% 0.03% Elf64_Addr
0.01% 0.03% Elf64_Addr: cache-line 0
0.01% 0.03% struct sched_entity
0.01% 0.03% struct sched_entity: cache-line 1
0.00% 0.00% struct sched_entity: cache-line 4
0.00% 0.00% struct sched_entity: cache-line 2
0.00% 0.00% struct sched_entity: cache-line 3
0.00% 0.00% struct sched_entity: cache-line 0
0.01% 0.00% struct task_group
0.01% 0.00% struct task_group: cache-line 4
0.00% 0.00% struct task_group: cache-line 5
0.00% 0.00% struct css_rstat_cpu
0.00% 0.00% struct css_rstat_cpu: cache-line 0
root@x2:~#
root@x2:~# perf report -s type,typecln,typeoff -H --group --stdio
# Total Lost Samples: 0
#
# Samples: 72K of events 'cpu/mem-loads,ldlat=30/P, cpu/mem-stores/P'
# Event count (approx.): 1261814616
#
# Overhead Data Type / Data Type Cacheline / Data Type Offset
# ...................... ..................................................
#
88.99% 35.88% struct net_conn
56.43% 17.50% struct net_conn: cache-line 0
9.81% 0.00% struct net_conn +0xa (dport)
9.54% 0.00% struct net_conn +0x8 (sport)
9.53% 0.00% struct net_conn +0xc (state)
9.48% 0.00% struct net_conn +0xd (protocol)
9.27% 0.00% struct net_conn +0x4 (daddr)
8.79% 0.00% struct net_conn +0 (saddr)
0.00% 12.08% struct net_conn +0x10 (bytes_rx)
0.00% 1.71% struct net_conn +0x18 (packets_rx)
0.00% 2.80% struct net_conn +0x20 (rx_queue)
0.00% 0.91% struct net_conn +0x3c (last_ack)
32.55% 0.00% struct net_conn: cache-line 2
5.70% 0.00% struct net_conn +0x83 (rcv_wscale)
5.46% 0.00% struct net_conn +0x84 (keepalive_int)
5.43% 0.00% struct net_conn +0x82 (snd_wscale)
5.39% 0.00% struct net_conn +0x80 (mss)
5.34% 0.00% struct net_conn +0x88 (mark)
5.23% 0.00% struct net_conn +0x8c (priority)
0.01% 18.38% struct net_conn: cache-line 1
0.01% 0.05% struct net_conn +0x5c (retrans)
0.00% 4.83% struct net_conn +0x40 (bytes_tx)
0.00% 10.63% struct net_conn +0x48 (packets_tx)
0.00% 1.54% struct net_conn +0x58 (rtt_us)
0.00% 1.07% struct net_conn +0x50 (cwnd)
0.00% 0.25% struct net_conn +0x54 (ssthresh)
10.08% 63.03% (unknown)
10.08% 63.03% (unknown): cache-line 0
10.08% 63.03% (unknown)
0.82% 0.13% int
0.82% 0.13% int: cache-line 0
0.82% 0.13% int +0 (no field)
0.01% 0.02% struct folio
0.01% 0.02% struct folio: cache-line 0
0.01% 0.00% struct folio +0 (flags.f)
0.00% 0.01% struct folio +0x34 (_refcount.counter)
0.00% 0.00% struct folio +0x18 (mapping)
:
next prev parent reply other threads:[~2026-10-01 9:18 UTC|newest]
Thread overview: 15+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-30 21:37 [PATCH v6 0/5] perf tools: Add progress diagnostics and a false-sharing workload Arnaldo Carvalho de Melo
2026-09-30 21:37 ` [PATCH v6 1/5] perf config: Move perf_config__set_variable() to util/config.c Arnaldo Carvalho de Melo
2026-10-01 7:18 ` Namhyung Kim
2026-10-01 9:19 ` Arnaldo Carvalho de Melo
2026-09-30 21:37 ` [PATCH v6 2/5] perf report: Add --progress option Arnaldo Carvalho de Melo
2026-09-30 21:37 ` [PATCH v6 3/5] perf report: Add --no-progress option Arnaldo Carvalho de Melo
2026-10-01 7:01 ` Namhyung Kim
2026-10-01 9:20 ` Arnaldo Carvalho de Melo
2026-09-30 21:37 ` [PATCH v6 4/5] perf scripts: Add perf-stuck, to tell where a running perf is stuck Arnaldo Carvalho de Melo
2026-10-01 7:24 ` Namhyung Kim
2026-10-01 9:19 ` Arnaldo Carvalho de Melo
2026-09-30 21:37 ` [PATCH v6 5/5] perf test: Add false_sharing workload exhibiting cross-CPU false sharing Arnaldo Carvalho de Melo
2026-10-01 7:28 ` Namhyung Kim
2026-10-01 9:18 ` Arnaldo Carvalho de Melo [this message]
-- strict thread matches above, loose matches on Subject: below --
2026-09-30 11:24 [PATCH v6 0/5] perf tools: Add progress diagnostics and a false-sharing workload Arnaldo Carvalho de Melo
2026-09-30 11:24 ` [PATCH v6 5/5] perf test: Add false_sharing workload exhibiting cross-CPU false sharing Arnaldo Carvalho de Melo
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=ar4lfjssCX9H2CCL@x2 \
--to=acme@kernel.org \
--cc=acme@redhat.com \
--cc=adrian.hunter@intel.com \
--cc=irogers@google.com \
--cc=james.clark@linaro.org \
--cc=jolsa@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-perf-users@vger.kernel.org \
--cc=mingo@kernel.org \
--cc=namhyung@kernel.org \
--cc=tglx@linutronix.de \
--cc=williams@redhat.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®