mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v6 0/5] perf tools: Add progress diagnostics and a false-sharing workload
@ 2026-09-30 21:37 Arnaldo Carvalho de Melo
  2026-09-30 21:37 ` [PATCH v6 1/5] perf config: Move perf_config__set_variable() to util/config.c Arnaldo Carvalho de Melo
                   ` (4 more replies)
  0 siblings, 5 replies; 15+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-09-30 21:37 UTC (permalink / raw)
  To: Namhyung Kim
  Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
	Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
	Arnaldo Carvalho de Melo

Hi,

This series adds progress and stuck-process diagnostics to perf, and a
workload that makes false sharing visible to data type profiling.

The changes are:

  - move perf_config__set_variable() to util/config.c and serialize config
    parser and read-modify-write state, so non-builtin perf code can persist
    configuration changes safely;

  - add 'perf report --progress' for stdio users, showing the current phase,
    percentage, and counts while a session is processed;

  - add 'perf report --no-progress', the counterpart above for the TUI and
    GTK browsers, whose progress there is no other way to turn off;

  - add the prototype tools/perf/scripts/perf-stuck.sh helper, which samples
    a running perf process and can invoke the accompanying GDB commands when
    progress stops;

  - add 'perf test -w false_sharing', a synthetic TCP-shaped workload with
    identity and packet counters sharing a cacheline, and include it in the
    data type profiling shell test.

Best regards,

- Arnaldo

What changed from v5:

  - tools/perf/util/config.c: check fprintf()/fclose() in
    perf_config_set__write(), a write failure was reported as success;

  - tools/perf/ui/stdio/progress.c: fix the "[42.3%]" changelog example
    and two comments that opened at the wrong indent;

  - tools/perf/scripts/perf-stuck.sh: validate -i/-n, require -g for -x,
    and clear gdb_done on progress so -g can fire more than once;

  - tools/perf/tests/workloads/false_sharing.c: say sum, not hash, and
    describe the actual perf mem record + report -s type verification;

Found in a pre-post read-through of v5, not a list reply.

What changed from v4:

  - tools/perf/builtin-report.c: mark --progress PARSE_OPT_NOAUTONEG,
    parse_long_opt() claimed --no-progress before reaching it.  Sashiko, v4;

  - tools/perf/util/config.c: format the path buffer inside the
    critical section, mkpath() ran outside the lock.  Sashiko, v4;

  - tools/perf/tests/workloads/false_sharing.c: walk every affinity mask
    bit instead of bounding by NPROCESSORS_CONF.  Sashiko, v4;

What changed from v3:

  - tools/perf/util/config.c: keep the config_file_name buffer in
    static storage, a reader outside a parse could race it.  Sashiko, v3;

  - tools/perf/scripts/perf-stuck.gdb: perf-dso walks each candidate,
    resolving the REFCNT_CHECKING proxy instead of aborting.  Sashiko, v3;

  - tools/perf/tests/workloads/false_sharing.c: put sum before cpu in
    fs_reader, alignment was doubling it to 128 bytes.  Sashiko, v3;

  - tools/perf/builtin-report.c: document what --quiet does to
    --progress, asked by Namhyung Kim reviewing v3;

  - tools/perf/builtin-report.c, tools/perf/ui/progress.c: add
    --no-progress, suggested by Namhyung Kim reviewing v3;

  - tools/perf/scripts/perf-stuck.sh: check gdb is present before
    watching instead of failing when -g first fires.  Namhyung, v3;

What changed from v2:

  - tools/perf/util/config.c: make perf_etc_perfconfig() total, it
    returned NULL on allocation failure.  Sashiko, v2;

  - tools/perf/scripts/perf-stuck.gdb: don't deref map_symbol.sym
    without a NULL check, print "(no symbol)" instead.  Sashiko, v2;

  - tools/perf/scripts/perf-stuck.gdb: note the REFCNT_CHECKING proxy
    indirection next to the structure walks;

  - tools/perf/scripts/perf-stuck.sh: count samples with no progress to
    look at, an empty log never fired -g;

  - tools/perf/scripts/perf-stuck.sh: use -- for pgrep and tail, a name
    starting with '-' attached to the wrong process;

  - tools/perf/scripts/perf-stuck.sh: bound the gdb run with
    timeout --signal=INT 30, an inferior call can hang forever;

What changed from v1:

  - avoid calling CPU_SET() with -1 when false_sharing runs with only one
    CPU available in its affinity mask.

      tools/perf/Documentation/perf-report.txt      |  19 ++
  tools/perf/builtin-config.c                   |  70 +------
  tools/perf/builtin-report.c                   |  23 +++
  tools/perf/scripts/perf-stuck.gdb             | 129 +++++++++++++
  tools/perf/scripts/perf-stuck.sh              | 193 ++++++++++++++++++++
  tools/perf/tests/builtin-test.c               |   1 +
  tools/perf/tests/shell/data_type_profiling.sh |   9 +-
  tools/perf/tests/tests.h                      |   1 +
  tools/perf/tests/workloads/Build              |   2 +
  tools/perf/tests/workloads/false_sharing.c    | 251 +++++++++++++++++++++++++
  tools/perf/ui/Build                           |   1 +
  tools/perf/ui/progress.c                      |   6 +
  tools/perf/ui/progress.h                      |   4 +
  tools/perf/ui/stdio/progress.c                | 162 +++++++++++++++
  tools/perf/util/config.c                      | 219 +++++++++++++++++---
  tools/perf/util/config.h                      |   2 +
  tools/perf/util/ordered-events.c              |  16 +-
  tools/perf/util/session.c                     |  12 +-
  18 files changed, 1021 insertions(+), 99 deletions(-)
  create mode 100644 tools/perf/scripts/perf-stuck.gdb
  create mode 100755 tools/perf/scripts/perf-stuck.sh
  create mode 100644 tools/perf/tests/workloads/false_sharing.c
  create mode 100644 tools/perf/ui/stdio/progress.c

base-commit: 0ae6fc78c5ce0dfd
v1-head: 45d7917f7e05e8a29828ed5f0bbdc94fc938f79f
v2-head: d4f84e4de8890194924ccd897a8e6773e7d4240b
v3-head: 485532296710225862daf8ffb19802ae327efe2a
v4-head: 38193635508e0a5f04a6ebf70db27534b68b2d5b
v5-head: 352709b39c0db8789a54d4ecca0251b26b1bf839
--

^ permalink raw reply	[flat|nested] 15+ messages in thread

* [PATCH v6 1/5] perf config: Move perf_config__set_variable() to util/config.c
  2026-09-30 21:37 [PATCH v6 0/5] perf tools: Add progress diagnostics and a false-sharing workload Arnaldo Carvalho de Melo
@ 2026-09-30 21:37 ` Arnaldo Carvalho de Melo
  2026-10-01  7:18   ` Namhyung Kim
  2026-09-30 21:37 ` [PATCH v6 2/5] perf report: Add --progress option Arnaldo Carvalho de Melo
                   ` (3 subsequent siblings)
  4 siblings, 1 reply; 15+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-09-30 21:37 UTC (permalink / raw)
  To: Namhyung Kim
  Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
	Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
	Arnaldo Carvalho de Melo

From: Arnaldo Carvalho de Melo <acme@redhat.com>

Move perf_config__set_variable() out of the 'perf config' builtin so
that opt-in features can persist their choice from outside it, e.g.
util/debuginfo.c writing core.debuginfod=false when the user disables
debuginfod for the rest of the session.  The
set_config() body becomes perf_config_set__write(), with the
system_config choice as an argument, as the builtin's
use_system_config/use_user_config statics are not available outside it.

All config file access shares the static parser state and can now run
on more than one thread, with a feature writing the config while perf
top's display thread reads it, so serialize parsing and
rewriting with a mutex, and the whole read-modify-write of
perf_config__set_variable() against itself.

perf_config_set__write() checked fopen() but none of the fprintf()s or
fclose(), so a write failure after truncating the file was reported as
success.  Harmless for the interactive 'perf config' this came from,
but this makes it an entry point a background feature can call with no
other feedback, so propagate those errors too.

perf_etc_perfconfig() caches what system_path() returns, and that
allocates, so on failure every caller dereferenced NULL, the
pre-existing ones in perf_config_from_file() and in the daemon
included.  Pre-existing, not introduced here, but this rewrites the
accessor anyway, so make it total: the unresolved path is the same
string for the usual absolute ETC_PERFCONFIG.

Assisted-by: LLM
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
---
 tools/perf/builtin-config.c |  70 +-----------
 tools/perf/util/config.c    | 219 ++++++++++++++++++++++++++++++++----
 tools/perf/util/config.h    |   2 +
 3 files changed, 201 insertions(+), 90 deletions(-)

diff --git a/tools/perf/builtin-config.c b/tools/perf/builtin-config.c
index cefd042e4f853466..3b074aca8d344539 100644
--- a/tools/perf/builtin-config.c
+++ b/tools/perf/builtin-config.c
@@ -41,37 +41,7 @@ static struct option config_options[] = {
 
 static int set_config(struct perf_config_set *set, const char *file_name)
 {
-	struct perf_config_section *section = NULL;
-	struct perf_config_item *item = NULL;
-	const char *first_line = "# this file is auto-generated.";
-	FILE *fp;
-
-	if (set == NULL)
-		return -1;
-
-	fp = fopen(file_name, "w");
-	if (!fp)
-		return -1;
-
-	fprintf(fp, "%s\n", first_line);
-
-	/* overwrite configvariables */
-	perf_config_items__for_each_entry(&set->sections, section) {
-		if (!use_system_config && section->from_system_config)
-			continue;
-		fprintf(fp, "[%s]\n", section->name);
-
-		perf_config_items__for_each_entry(&section->items, item) {
-			if (!use_system_config && item->from_system_config)
-				continue;
-			if (item->value)
-				fprintf(fp, "\t%s = %s\n",
-					item->name, item->value);
-		}
-	}
-	fclose(fp);
-
-	return 0;
+	return perf_config_set__write(set, file_name, use_system_config);
 }
 
 static int show_spec_config(struct perf_config_set *set, const char *var)
@@ -158,44 +128,6 @@ static int parse_config_arg(char *arg, char **var, char **value)
 	return 0;
 }
 
-int perf_config__set_variable(const char *var, const char *value)
-{
-	char path[PATH_MAX];
-	char *user_config = mkpath(path, sizeof(path), "%s/.perfconfig", getenv("HOME"));
-	const char *config_filename;
-	struct perf_config_set *set;
-	int ret = -1;
-
-	if (use_system_config)
-		config_exclusive_filename = perf_etc_perfconfig();
-	else if (use_user_config)
-		config_exclusive_filename = user_config;
-
-	if (!config_exclusive_filename)
-		config_filename = user_config;
-	else
-		config_filename = config_exclusive_filename;
-
-	set = perf_config_set__new();
-	if (!set)
-		goto out_err;
-
-	if (perf_config_set__collect(set, config_filename, var, value) < 0) {
-		pr_err("Failed to add '%s=%s'\n", var, value);
-		goto out_err;
-	}
-
-	if (set_config(set, config_filename) < 0) {
-		pr_err("Failed to set the configs on %s\n", config_filename);
-		goto out_err;
-	}
-
-	ret = 0;
-out_err:
-	perf_config_set__delete(set);
-	return ret;
-}
-
 int cmd_config(int argc, const char **argv)
 {
 	int i, ret = -1;
diff --git a/tools/perf/util/config.c b/tools/perf/util/config.c
index 8fe43b032e9af88a..3b6a45569ae3d2eb 100644
--- a/tools/perf/util/config.c
+++ b/tools/perf/util/config.c
@@ -12,6 +12,8 @@
 #include "config.h"
 
 #include <errno.h>
+#include <limits.h>
+#include <pthread.h>
 #include <stdbool.h>
 #include <stdio.h>
 #include <stdlib.h>
@@ -370,10 +372,15 @@ static int perf_parse_long(const char *value, long *ret)
 
 static void bad_config(const char *name)
 {
-	if (config_file_name)
-		pr_warning("bad config value for '%s' in %s, ignoring...\n", name, config_file_name);
-	else
-		pr_warning("bad config value for '%s', ignoring...\n", name);
+	/*
+	 * No file name here: config_file_name is set and cleared by
+	 * whichever thread has a config file in flight, under config_mutex,
+	 * while this runs on the thread dispatching the collected values.
+	 * Reading it would race with that thread: NULL between the check
+	 * and the use, or the file it is parsing, not the one the bad value
+	 * came from.
+	 */
+	pr_warning("bad config value for '%s', ignoring...\n", name);
 }
 
 int perf_config_u64(u64 *dest, const char *name, const char *value)
@@ -549,11 +556,26 @@ int perf_default_config(const char *var, const char *value,
 	return 0;
 }
 
+/*
+ * Serialize config file access: parsing and rewriting share the static
+ * parser state and can run on more than one thread, the debuginfod
+ * fetch writing core.debuginfod=false while perf top reads it.
+ */
+static pthread_mutex_t config_mutex = PTHREAD_MUTEX_INITIALIZER;
+
+/*
+ * perf_config__set_variable() is a read-modify-write of a config file,
+ * and config_mutex serializes just each of its steps: take it for the
+ * whole update so callers can't write over each other's changes.
+ */
+static pthread_mutex_t config_update_mutex = PTHREAD_MUTEX_INITIALIZER;
+
 static int perf_config_from_file(config_fn_t fn, const char *filename, void *data)
 {
 	int ret;
 	FILE *f = fopen(filename, "r");
 
+	pthread_mutex_lock(&config_mutex);
 	ret = -1;
 	if (f) {
 		config_file = f;
@@ -564,15 +586,34 @@ static int perf_config_from_file(config_fn_t fn, const char *filename, void *dat
 		fclose(f);
 		config_file_name = NULL;
 	}
+	pthread_mutex_unlock(&config_mutex);
 	return ret;
 }
 
+/*
+ * Computed once: system_path() allocates, a lazy init racing on two
+ * threads would leak all but one of the strings.
+ */
+static const char *etc_perfconfig;
+
+static void perf_etc_perfconfig__init(void)
+{
+	etc_perfconfig = system_path(ETC_PERFCONFIG);
+	/*
+	 * None of the callers check for NULL, so an allocation failure
+	 * here leaves all of them dereferencing it.  The unresolved path
+	 * is the same string when ETC_PERFCONFIG is absolute anyway.
+	 */
+	if (!etc_perfconfig)
+		etc_perfconfig = ETC_PERFCONFIG;
+}
+
 const char *perf_etc_perfconfig(void)
 {
-	static const char *system_wide;
-	if (!system_wide)
-		system_wide = system_path(ETC_PERFCONFIG);
-	return system_wide;
+	static pthread_once_t once = PTHREAD_ONCE_INIT;
+
+	pthread_once(&once, perf_etc_perfconfig__init);
+	return etc_perfconfig;
 }
 
 static int perf_env_bool(const char *k, int def)
@@ -630,19 +671,25 @@ static char *home_perfconfig(void)
 	return NULL;
 }
 
-const char *perf_home_perfconfig(void)
-{
-	static const char *config;
-	static bool failed;
+/*
+ * Computed once for the same reason as perf_etc_perfconfig() above:
+ * home_perfconfig() allocates, a lazy init racing on two threads would
+ * leak all but one of the strings, and the warnings it may print would
+ * come out more than once.
+ */
+static const char *home_config;
 
-	if (failed || config)
-		return config;
+static void perf_home_perfconfig__init(void)
+{
+	home_config = home_perfconfig();
+}
 
-	config = home_perfconfig();
-	if (!config)
-		failed = true;
+const char *perf_home_perfconfig(void)
+{
+	static pthread_once_t once = PTHREAD_ONCE_INIT;
 
-	return config;
+	pthread_once(&once, perf_home_perfconfig__init);
+	return home_config;
 }
 
 static struct perf_config_section *find_section(struct list_head *sections,
@@ -783,8 +830,18 @@ static int collect_config(const char *var, const char *value,
 int perf_config_set__collect(struct perf_config_set *set, const char *file_name,
 			     const char *var, const char *value)
 {
+	int ret;
+
+	pthread_mutex_lock(&config_mutex);
 	config_file_name = file_name;
-	return collect_config(var, value, set);
+	ret = collect_config(var, value, set);
+	/*
+	 * Don't leave the static parser state pointing at the caller's
+	 * buffer.
+	 */
+	config_file_name = NULL;
+	pthread_mutex_unlock(&config_mutex);
+	return ret;
 }
 
 static int perf_config_set__init(struct perf_config_set *set)
@@ -831,6 +888,16 @@ struct perf_config_set *perf_config_set__load_file(const char *file)
 	return set;
 }
 
+/*
+ * The global config_set is built lazily: two threads in perf_config() at
+ * once would both build one and leak all but the last, and config_set must
+ * not be read while another thread swaps it, so take it one at a time.  Not
+ * with config_mutex: building the set parses the config files, which takes
+ * that one.
+ */
+static pthread_mutex_t config_set_mutex = PTHREAD_MUTEX_INITIALIZER;
+
+/* Called with config_set_mutex held. */
 static int perf_config__init(void)
 {
 	if (config_set == NULL)
@@ -871,16 +938,126 @@ int perf_config_set(struct perf_config_set *set,
 
 int perf_config(config_fn_t fn, void *data)
 {
-	if (config_set == NULL && perf_config__init())
+	struct perf_config_set *set;
+
+	/*
+	 * The mutex is taken just for the lazy init and for the pointer:
+	 * the dispatch below only reads the set, so a callback that called
+	 * perf_config() again would find it unlocked, and only
+	 * perf_config__exit() replaces the set.
+	 */
+	pthread_mutex_lock(&config_set_mutex);
+	if (perf_config__init()) {
+		pthread_mutex_unlock(&config_set_mutex);
 		return -1;
+	}
+	set = config_set;
+	pthread_mutex_unlock(&config_set_mutex);
 
-	return perf_config_set(config_set, fn, data);
+	return perf_config_set(set, fn, data);
 }
 
 void perf_config__exit(void)
 {
+	pthread_mutex_lock(&config_set_mutex);
 	perf_config_set__delete(config_set);
 	config_set = NULL;
+	pthread_mutex_unlock(&config_set_mutex);
+}
+
+int perf_config_set__write(struct perf_config_set *set,
+			   const char *file_name, bool system_config)
+{
+	struct perf_config_section *section = NULL;
+	struct perf_config_item *item = NULL;
+	int ret = 0;
+	FILE *fp;
+
+	pthread_mutex_lock(&config_mutex);
+	fp = fopen(file_name, "w");
+	if (!fp) {
+		pthread_mutex_unlock(&config_mutex);
+		return -1;
+	}
+
+	if (fprintf(fp, "# this file is auto-generated.\n") < 0)
+		ret = -1;
+
+	/* overwrite configvariables */
+	perf_config_sections__for_each_entry(&set->sections, section) {
+		if (!system_config && section->from_system_config)
+			continue;
+		if (fprintf(fp, "[%s]\n", section->name) < 0)
+			ret = -1;
+
+		perf_config_items__for_each_entry(&section->items, item) {
+			if (!system_config && item->from_system_config)
+				continue;
+			if (item->value &&
+			    fprintf(fp, "\t%s = %s\n", item->name, item->value) < 0)
+				ret = -1;
+		}
+	}
+	if (fclose(fp) != 0)
+		ret = -1;
+	pthread_mutex_unlock(&config_mutex);
+
+	return ret;
+}
+
+/*
+ * Set @var=@value in the configuration file perf is using: ~/.perfconfig
+ * or the file named by PERF_CONFIG, which makes perf read only that
+ * file.  The rewrite is the same 'perf config' does, comments are not
+ * preserved.
+ */
+int perf_config__set_variable(const char *var, const char *value)
+{
+	const char *config_filename;
+	bool system_config;
+	struct perf_config_set *set = NULL;
+	int ret = -1;
+
+	pthread_mutex_lock(&config_update_mutex);
+
+	/*
+	 * Not on the stack: the parser publishes this buffer as
+	 * config_file_name, which another thread may still be reading.  It is
+	 * shared by every caller, so it is formatted under the lock above.
+	 */
+	{
+		static char path[PATH_MAX];
+		char *user_config = mkpath(path, sizeof(path), "%s/.perfconfig", getenv("HOME"));
+
+		config_filename = config_exclusive_filename ?: user_config;
+	}
+
+	/*
+	 * When rewriting the system wide file all entries are marked as coming
+	 * from it and must be kept, or it would be truncated down to its
+	 * header.
+	 */
+	system_config = strcmp(config_filename, perf_etc_perfconfig()) == 0;
+
+	set = perf_config_set__new();
+	if (!set)
+		goto out_err;
+
+	if (perf_config_set__collect(set, config_filename, var, value) < 0) {
+		pr_err("Failed to add '%s=%s'\n", var, value);
+		goto out_err;
+	}
+
+	if (perf_config_set__write(set, config_filename, system_config) < 0) {
+		pr_err("Failed to set the configs on %s\n", config_filename);
+		goto out_err;
+	}
+
+	ret = 0;
+out_err:
+	perf_config_set__delete(set);
+	pthread_mutex_unlock(&config_update_mutex);
+	return ret;
 }
 
 static void perf_config_item__delete(struct perf_config_item *item)
diff --git a/tools/perf/util/config.h b/tools/perf/util/config.h
index 987b47cf54c350ba..9098f8a045850c97 100644
--- a/tools/perf/util/config.h
+++ b/tools/perf/util/config.h
@@ -33,6 +33,8 @@ int perf_config_scan(const char *name, const char *fmt, ...) __scanf(2, 3);
 const char *perf_config_get(const char *name);
 int perf_config_set(struct perf_config_set *set,
 		    config_fn_t fn, void *data);
+int perf_config_set__write(struct perf_config_set *set,
+			   const char *file_name, bool system_config);
 int perf_config_int(int *dest, const char *, const char *);
 int perf_config_u8(u8 *dest, const char *name, const char *value);
 int perf_config_u64(u64 *dest, const char *, const char *);
-- 
2.55.0


^ permalink raw reply	[flat|nested] 15+ messages in thread

* [PATCH v6 2/5] perf report: Add --progress option
  2026-09-30 21:37 [PATCH v6 0/5] perf tools: Add progress diagnostics and a false-sharing workload Arnaldo Carvalho de Melo
  2026-09-30 21:37 ` [PATCH v6 1/5] perf config: Move perf_config__set_variable() to util/config.c Arnaldo Carvalho de Melo
@ 2026-09-30 21:37 ` Arnaldo Carvalho de Melo
  2026-09-30 21:37 ` [PATCH v6 3/5] perf report: Add --no-progress option Arnaldo Carvalho de Melo
                   ` (2 subsequent siblings)
  4 siblings, 0 replies; 15+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-09-30 21:37 UTC (permalink / raw)
  To: Namhyung Kim
  Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
	Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
	Arnaldo Carvalho de Melo

From: Arnaldo Carvalho de Melo <acme@redhat.com>

Processing a large session with stdio output gives no feedback about
which phase perf is in or how far along it is: ui_progress updates are
only shown by the TUI.  Add --progress, installing a stdio backend
(ui/stdio/progress.c) that prints the phase title, percentage and
counts:

  Processing events... [ 42.3%] 317M / 746M

Phases can be nested, so the backend tracks the ones started to
complete the right one on ui_progress__finish(); that requires
init()/finish() pairs, fixed in ordered-events.c and the pipe and
directory event processing.

Assisted-by: LLM
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
---
 tools/perf/Documentation/perf-report.txt |  14 ++
 tools/perf/builtin-report.c              |  11 ++
 tools/perf/ui/Build                      |   1 +
 tools/perf/ui/progress.h                 |   2 +
 tools/perf/ui/stdio/progress.c           | 162 +++++++++++++++++++++++
 tools/perf/util/ordered-events.c         |  16 ++-
 tools/perf/util/session.c                |  12 +-
 7 files changed, 211 insertions(+), 7 deletions(-)
 create mode 100644 tools/perf/ui/stdio/progress.c

diff --git a/tools/perf/Documentation/perf-report.txt b/tools/perf/Documentation/perf-report.txt
index 3718ebd297ce0673..a7429a30ec28f903 100644
--- a/tools/perf/Documentation/perf-report.txt
+++ b/tools/perf/Documentation/perf-report.txt
@@ -29,6 +29,20 @@ OPTIONS
 --quiet::
 	Do not show any warnings or messages.  (Suppress -v)
 
+--progress::
+	Show progress for each of the processing phases, printing the
+	percentage done and the current/total counts: the first phase
+	counts the bytes of the perf.data file processed so far, the
+	merge and sort phases count hist entries, e.g.:
+
+	Processing events... [ 42.3%] 4G / 10G
+	Merging related events... [  7.1%] 1024 / 14387
+	Sorting events for output... [ 98.2%] 14132 / 14387
+
+	It is a no-op when using the TUI or GTK browsers, that already
+	present progress information, or when --quiet is used, that asks
+	for no messages at all.
+
 -n::
 --show-nr-samples::
 	Show the number of samples for each symbol
diff --git a/tools/perf/builtin-report.c b/tools/perf/builtin-report.c
index 57225bc87731d074..963808b561e547c3 100644
--- a/tools/perf/builtin-report.c
+++ b/tools/perf/builtin-report.c
@@ -87,6 +87,7 @@ struct report {
 	bool			use_gtk;
 #endif
 	bool			use_stdio;
+	bool			progress;
 	bool			show_full_info;
 	bool			show_threads;
 	bool			inverted_callchain;
@@ -1384,6 +1385,8 @@ int cmd_report(int argc, const char **argv)
 		    "Use the stdio interface"),
 	OPT_BOOLEAN(0, "weights", &symbol_conf.annotate_weight,
 			"Show or hide weight columns in annotation. Default show if non-zero."),
+	OPT_BOOLEAN(0, "progress", &report.progress,
+		    "Show progress while processing the perf.data file"),
 	OPT_BOOLEAN(0, "header", &report.header, "Show data header."),
 	OPT_BOOLEAN(0, "header-only", &report.header_only,
 		    "Show only data header."),
@@ -1790,6 +1793,14 @@ int cmd_report(int argc, const char **argv)
 	else
 		use_browser = 0;
 
+	/*
+	 * For the stdio case: print the percentage of the perf.data file
+	 * processed so far for each processing phase.  --quiet asks for no
+	 * messages at all, so it leaves the phases uncounted.
+	 */
+	if (report.progress && !quiet && use_browser == 0)
+		stdio_progress__init();
+
 	if (report.data_type && use_browser == 1) {
 		symbol_conf.annotate_data_member = true;
 		symbol_conf.annotate_data_sample = true;
diff --git a/tools/perf/ui/Build b/tools/perf/ui/Build
index 6005f813c9e3990c..a7b1740d51c80f23 100644
--- a/tools/perf/ui/Build
+++ b/tools/perf/ui/Build
@@ -4,6 +4,7 @@ perf-ui-y += progress.o
 perf-ui-y += util.o
 perf-ui-y += hist.o
 perf-ui-y += stdio/hist.o
+perf-ui-y += stdio/progress.o
 
 CFLAGS_setup.o += -DLIBDIR="BUILD_STR($(LIBDIR))"
 
diff --git a/tools/perf/ui/progress.h b/tools/perf/ui/progress.h
index 4f52c37b2f099a82..03f1a8bb260ba076 100644
--- a/tools/perf/ui/progress.h
+++ b/tools/perf/ui/progress.h
@@ -23,6 +23,8 @@ void __ui_progress__init(struct ui_progress *p, u64 total,
 
 void ui_progress__update(struct ui_progress *p, u64 adv);
 
+void stdio_progress__init(void);
+
 struct ui_progress_ops {
 	void (*init)(struct ui_progress *p);
 	void (*update)(struct ui_progress *p);
diff --git a/tools/perf/ui/stdio/progress.c b/tools/perf/ui/stdio/progress.c
new file mode 100644
index 0000000000000000..118a6cf6555d59f5
--- /dev/null
+++ b/tools/perf/ui/stdio/progress.c
@@ -0,0 +1,162 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Progress feedback for the stdio (non-TUI/GTK) case, enabled with
+ * 'perf report --progress'.
+ */
+#include <inttypes.h>
+#include <stdio.h>
+#include <unistd.h>
+#include <linux/kernel.h>
+#include "../../util/debug.h"
+#include "../../util/units.h"
+#include "../progress.h"
+
+/*
+ * Phases can be nested, so keep track of the ones started so far to be
+ * able to complete the right one on ui_progress__finish(), which gets
+ * no arguments.
+ */
+#define STDIO_PROGRESS__MAX_DEPTH 8
+
+struct stdio_progress_phase {
+	struct ui_progress	*p;
+	u64			last_printed;
+	size_t			last_len;
+};
+
+static struct stdio_progress_phase stdio_progress__stack[STDIO_PROGRESS__MAX_DEPTH];
+static int stdio_progress__depth;
+static bool stdio_progress__is_tty;
+/* Phases that didn't fit on the stack are not shown. */
+static int stdio_progress__dropped;
+
+static void stdio_progress__print_phase(struct stdio_progress_phase *phase,
+					u64 curr)
+{
+	struct ui_progress *p = phase->p;
+	char buf_cur[20], buf_tot[20], buf[128];
+	double percent = p->total ? 100.0 * (double)curr / (double)p->total : 0.0;
+	size_t len;
+
+	/*
+	 * Only the completion line shows 100.0%: a 99.99% progress would round
+	 * up to it and look like a duplicate at finish time.
+	 */
+	if (curr < p->total && percent > 99.9)
+		percent = 99.9;
+
+	if (p->size) {
+		unit_number__scnprintf(buf_cur, sizeof(buf_cur), curr);
+		unit_number__scnprintf(buf_tot, sizeof(buf_tot), p->total);
+		len = scnprintf(buf, sizeof(buf), "%s [%5.1f%%] %s / %s",
+				p->title, percent, buf_cur, buf_tot);
+	} else {
+		len = scnprintf(buf, sizeof(buf), "%s [%5.1f%%] %" PRIu64 " / %" PRIu64,
+				p->title, percent, curr, p->total);
+	}
+
+	if (!stdio_progress__is_tty) {
+		fprintf(stderr, "%s\n", buf);
+		goto out;
+	}
+
+	/* Pad to the length of the previous line to erase its leftovers. */
+	fprintf(stderr, "\r%s%*s", buf,
+		(int)(len < phase->last_len ? phase->last_len - len : 0), "");
+	phase->last_len = len;
+out:
+	phase->last_printed = curr;
+	fflush(stderr);
+}
+
+static void __stdio_progress__init(struct ui_progress *p)
+{
+	/* The default step is meant for the TUI bar, use 1% steps for stdio. */
+	p->next = p->step = p->total / 100 ?: 1;
+
+	if (stdio_progress__depth == STDIO_PROGRESS__MAX_DEPTH) {
+		/*
+		 * Out of room: don't start this phase, its finish() is
+		 * swallowed and its updates ignored below.
+		 */
+		pr_warning("progress phases nested deeper than %d, not showing progress for %s\n",
+			   STDIO_PROGRESS__MAX_DEPTH, p->title);
+		stdio_progress__dropped++;
+		return;
+	}
+
+	/* Start a nested phase in a line of its own. */
+	if (stdio_progress__depth && stdio_progress__is_tty)
+		fputc('\n', stderr);
+
+	stdio_progress__stack[stdio_progress__depth++] =
+		(struct stdio_progress_phase) {
+			.p = p,
+			.last_printed = 0,
+			.last_len = 0,
+		};
+
+	stdio_progress__print_phase(&stdio_progress__stack[stdio_progress__depth - 1],
+				    p->curr);
+}
+
+static void stdio_progress__update(struct ui_progress *p)
+{
+	/*
+	 * An update that doesn't match the innermost phase means something
+	 * is out of sync: print nothing rather than another phase's
+	 * numbers, or read past the stack.
+	 */
+	if (!stdio_progress__depth ||
+	    stdio_progress__stack[stdio_progress__depth - 1].p != p)
+		return;
+
+	stdio_progress__print_phase(&stdio_progress__stack[stdio_progress__depth - 1],
+				    p->curr);
+}
+
+static void stdio_progress__finish(void)
+{
+	struct stdio_progress_phase *phase;
+
+	/*
+	 * Being the innermost phase, its finish() comes first: swallow it, or
+	 * it would complete the phase that encloses it.
+	 */
+	if (stdio_progress__dropped) {
+		stdio_progress__dropped--;
+		return;
+	}
+
+	if (!stdio_progress__depth)
+		return;
+
+	phase = &stdio_progress__stack[--stdio_progress__depth];
+
+	/*
+	 * The last line may have stopped short of the total, close this phase
+	 * showing it as complete unless that was already printed.
+	 */
+	if (phase->last_printed != phase->p->total)
+		stdio_progress__print_phase(phase, phase->p->total);
+
+	phase->last_printed	= 0;
+	phase->last_len		= 0;
+
+	if (stdio_progress__is_tty)
+		fputc('\n', stderr);
+
+	fflush(stderr);
+}
+
+static struct ui_progress_ops stdio_progress__ops = {
+	.init	= __stdio_progress__init,
+	.update	= stdio_progress__update,
+	.finish	= stdio_progress__finish,
+};
+
+void stdio_progress__init(void)
+{
+	stdio_progress__is_tty = isatty(STDERR_FILENO) == 1;
+	ui_progress__ops = &stdio_progress__ops;
+}
diff --git a/tools/perf/util/ordered-events.c b/tools/perf/util/ordered-events.c
index a5857f9f5af2d3de..54c85663e733be4d 100644
--- a/tools/perf/util/ordered-events.c
+++ b/tools/perf/util/ordered-events.c
@@ -237,14 +237,16 @@ static int do_flush(struct ordered_events *oe, bool show_progress)
 		ui_progress__init(&prog, oe->nr_events, "Processing time ordered events...");
 
 	list_for_each_entry_safe(iter, tmp, head, list) {
-		if (session_done())
-			return 0;
+		if (session_done()) {
+			ret = 0;
+			goto out_progress;
+		}
 
 		if (iter->timestamp > limit)
 			break;
 		ret = oe->deliver(oe, iter);
 		if (ret < 0)
-			return ret;
+			goto out_progress;
 
 		ordered_events__delete(oe, iter);
 		oe->last_flush = iter->timestamp;
@@ -258,10 +260,16 @@ static int do_flush(struct ordered_events *oe, bool show_progress)
 	else if (last_ts <= limit)
 		oe->last = list_entry(head->prev, struct ordered_event, list);
 
+	ret = 0;
+out_progress:
+	/*
+	 * Always pair ui_progress__init() with ui_progress__finish(), the
+	 * stdio backend tracks the phases on a stack.
+	 */
 	if (show_progress)
 		ui_progress__finish();
 
-	return 0;
+	return ret;
 }
 
 static int __ordered_events__flush(struct ordered_events *oe, enum oe_flush how,
diff --git a/tools/perf/util/session.c b/tools/perf/util/session.c
index 7fea9e72726c936c..c4b4c7fb3b864589 100644
--- a/tools/perf/util/session.c
+++ b/tools/perf/util/session.c
@@ -3253,8 +3253,12 @@ static int __perf_session__process_pipe_events(struct perf_session *session)
 	cur_size = sizeof(union perf_event);
 
 	buf = malloc(cur_size);
-	if (!buf)
-		return -errno;
+	if (!buf) {
+		err = -errno;
+		if (update_prog)
+			ui_progress__finish();
+		return err;
+	}
 	ordered_events__set_copy_on_queue(oe, true);
 more:
 	event = buf;
@@ -3753,8 +3757,10 @@ static int __perf_session__process_dir_events(struct perf_session *session)
 	}
 
 	rd = calloc(nr_readers, sizeof(struct reader));
-	if (!rd)
+	if (!rd) {
+		ui_progress__finish();
 		return -ENOMEM;
+	}
 
 	rd[0] = (struct reader) {
 		.fd		 = perf_data__fd(session->data),
-- 
2.55.0


^ permalink raw reply	[flat|nested] 15+ messages in thread

* [PATCH v6 3/5] perf report: Add --no-progress option
  2026-09-30 21:37 [PATCH v6 0/5] perf tools: Add progress diagnostics and a false-sharing workload Arnaldo Carvalho de Melo
  2026-09-30 21:37 ` [PATCH v6 1/5] perf config: Move perf_config__set_variable() to util/config.c Arnaldo Carvalho de Melo
  2026-09-30 21:37 ` [PATCH v6 2/5] perf report: Add --progress option Arnaldo Carvalho de Melo
@ 2026-09-30 21:37 ` Arnaldo Carvalho de Melo
  2026-10-01  7:01   ` Namhyung Kim
  2026-09-30 21:37 ` [PATCH v6 4/5] perf scripts: Add perf-stuck, to tell where a running perf is stuck Arnaldo Carvalho de Melo
  2026-09-30 21:37 ` [PATCH v6 5/5] perf test: Add false_sharing workload exhibiting cross-CPU false sharing Arnaldo Carvalho de Melo
  4 siblings, 1 reply; 15+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-09-30 21:37 UTC (permalink / raw)
  To: Namhyung Kim
  Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
	Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
	Arnaldo Carvalho de Melo

From: Arnaldo Carvalho de Melo <acme@redhat.com>

Now that --progress is being added for the stdio case, wire up its
counterpart for the browsers: the TUI and GTK ones present progress of
their own and there is no way to turn it off.  Install the no-op
ui_progress ops, the ones already used until a backend sets theirs,
after setup_browser() installed the ones of the browser in use.  The
phases are still counted, nothing is shown for them, and no second
option is needed for it: parse-options provides --no-progress as the
negation of --progress, report.progress_set saying that it was asked
for, report.progress being false both when nothing was asked for and
when --no-progress was.

Suggested-by: Namhyung Kim <namhyung@kernel.org>
Assisted-by: LLM
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
---
 tools/perf/Documentation/perf-report.txt |  5 +++++
 tools/perf/builtin-report.c              | 20 ++++++++++++++++----
 tools/perf/ui/progress.c                 |  6 ++++++
 tools/perf/ui/progress.h                 |  2 ++
 4 files changed, 29 insertions(+), 4 deletions(-)

diff --git a/tools/perf/Documentation/perf-report.txt b/tools/perf/Documentation/perf-report.txt
index a7429a30ec28f903..e145124d6c9a1097 100644
--- a/tools/perf/Documentation/perf-report.txt
+++ b/tools/perf/Documentation/perf-report.txt
@@ -43,6 +43,11 @@ OPTIONS
 	present progress information, or when --quiet is used, that asks
 	for no messages at all.
 
+--no-progress::
+	Do not show progress while processing the perf.data file.  It
+	also turns off the progress the TUI and GTK browsers present,
+	which is their own.
+
 -n::
 --show-nr-samples::
 	Show the number of samples for each symbol
diff --git a/tools/perf/builtin-report.c b/tools/perf/builtin-report.c
index 963808b561e547c3..105b2859678673ec 100644
--- a/tools/perf/builtin-report.c
+++ b/tools/perf/builtin-report.c
@@ -88,6 +88,7 @@ struct report {
 #endif
 	bool			use_stdio;
 	bool			progress;
+	bool			progress_set;
 	bool			show_full_info;
 	bool			show_threads;
 	bool			inverted_callchain;
@@ -1385,8 +1386,15 @@ int cmd_report(int argc, const char **argv)
 		    "Use the stdio interface"),
 	OPT_BOOLEAN(0, "weights", &symbol_conf.annotate_weight,
 			"Show or hide weight columns in annotation. Default show if non-zero."),
-	OPT_BOOLEAN(0, "progress", &report.progress,
-		    "Show progress while processing the perf.data file"),
+	/*
+	 * No --no-progress option to add: parse-options provides it as
+	 * the negation of this one, clearing report.progress, which is
+	 * also how it starts out.  progress_set is what tells the hook
+	 * below to stop the TUI and GTK browsers as well, they show
+	 * progress until asked not to.
+	 */
+	OPT_BOOLEAN_SET(0, "progress", &report.progress, &report.progress_set,
+			"Show progress while processing the perf.data file"),
 	OPT_BOOLEAN(0, "header", &report.header, "Show data header."),
 	OPT_BOOLEAN(0, "header-only", &report.header_only,
 		    "Show only data header."),
@@ -1796,9 +1804,13 @@ int cmd_report(int argc, const char **argv)
 	/*
 	 * For the stdio case: print the percentage of the perf.data file
 	 * processed so far for each processing phase.  --quiet asks for no
-	 * messages at all, so it leaves the phases uncounted.
+	 * messages at all, so it leaves the phases uncounted, and
+	 * --no-progress, progress_set with progress cleared, turns off what
+	 * the TUI and GTK browsers show as well.
 	 */
-	if (report.progress && !quiet && use_browser == 0)
+	if (report.progress_set && !report.progress)
+		ui_progress__noop_init();
+	else if (report.progress && !quiet && use_browser == 0)
 		stdio_progress__init();
 
 	if (report.data_type && use_browser == 1) {
diff --git a/tools/perf/ui/progress.c b/tools/perf/ui/progress.c
index 99d60223c74b2957..362680989ace606a 100644
--- a/tools/perf/ui/progress.c
+++ b/tools/perf/ui/progress.c
@@ -13,6 +13,12 @@ static struct ui_progress_ops null_progress__ops =
 
 struct ui_progress_ops *ui_progress__ops = &null_progress__ops;
 
+/* Everything counts but nothing is shown, the way it starts out. */
+void ui_progress__noop_init(void)
+{
+	ui_progress__ops = &null_progress__ops;
+}
+
 void ui_progress__update(struct ui_progress *p, u64 adv)
 {
 	u64 last = p->curr;
diff --git a/tools/perf/ui/progress.h b/tools/perf/ui/progress.h
index 03f1a8bb260ba076..e8c4f9f768aaf12b 100644
--- a/tools/perf/ui/progress.h
+++ b/tools/perf/ui/progress.h
@@ -25,6 +25,8 @@ void ui_progress__update(struct ui_progress *p, u64 adv);
 
 void stdio_progress__init(void);
 
+void ui_progress__noop_init(void);
+
 struct ui_progress_ops {
 	void (*init)(struct ui_progress *p);
 	void (*update)(struct ui_progress *p);
-- 
2.55.0


^ permalink raw reply	[flat|nested] 15+ messages in thread

* [PATCH v6 4/5] perf scripts: Add perf-stuck, to tell where a running perf is stuck
  2026-09-30 21:37 [PATCH v6 0/5] perf tools: Add progress diagnostics and a false-sharing workload Arnaldo Carvalho de Melo
                   ` (2 preceding siblings ...)
  2026-09-30 21:37 ` [PATCH v6 3/5] perf report: Add --no-progress option Arnaldo Carvalho de Melo
@ 2026-09-30 21:37 ` Arnaldo Carvalho de Melo
  2026-10-01  7:24   ` Namhyung Kim
  2026-09-30 21:37 ` [PATCH v6 5/5] perf test: Add false_sharing workload exhibiting cross-CPU false sharing Arnaldo Carvalho de Melo
  4 siblings, 1 reply; 15+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-09-30 21:37 UTC (permalink / raw)
  To: Namhyung Kim
  Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
	Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
	Arnaldo Carvalho de Melo

From: Arnaldo Carvalho de Melo <acme@redhat.com>

A perf that takes forever is hard to tell apart from one stuck in a
loop, and there is no way to see where without attaching gdb.
perf-stuck.sh samples a running process' /proc entries and its progress
line at a fixed interval, and with -g runs gdb (perf-stuck.gdb, adding
the perf-die-chain command) when no progress is made across two
samples, printing the DIE chain a DWARF type chase is stuck in.

It is a prototype: the plan is to turn it into a first class 'perf
stuck' command.  The process name is resolved with pgrep among the
caller's own processes only, as root an unscoped one would attach gdb
to the first process of any user with a matching name.

-i and -n are now validated instead of failing later inside the
sampling loop, -x now requires -g and gets the same readability check
as the default gdb command file, and gdb_done is cleared whenever
progress resumes, so -g can fire again on a later stall instead of at
most once per run.

Assisted-by: LLM
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
---
 tools/perf/scripts/perf-stuck.gdb | 129 ++++++++++++++++++++
 tools/perf/scripts/perf-stuck.sh  | 193 ++++++++++++++++++++++++++++++
 2 files changed, 322 insertions(+)
 create mode 100644 tools/perf/scripts/perf-stuck.gdb
 create mode 100755 tools/perf/scripts/perf-stuck.sh

diff --git a/tools/perf/scripts/perf-stuck.gdb b/tools/perf/scripts/perf-stuck.gdb
new file mode 100644
index 0000000000000000..be7b9e3fcb5d6d98
--- /dev/null
+++ b/tools/perf/scripts/perf-stuck.gdb
@@ -0,0 +1,129 @@
+# SPDX-License-Identifier: GPL-2.0
+#
+# gdb commands for a stuck perf, used by perf-stuck.sh -g and usable directly:
+#
+#   gdb -p $(pgrep -x perf) -batch -x perf-stuck.gdb -ex bt
+#
+# PROTOTYPE: part of the perf-stuck.sh stopgap, wants to become a first
+# class 'perf stuck' command printing these DIE chains without gdb.
+#
+# The commands are for the DWARF type chasers in util/dwarf-aux.c:
+#
+#   perf-die-chain <function> <die variable> [iterations]
+#   perf-die-chain-all [iterations]
+#   perf-dso
+#
+# For each iteration of the chasing loop they print the DIE address, the
+# CU, its file offset, tag and name: a cycle shows as the same (addr, cu)
+# pairs repeating, and a CU changing between iterations means the chase
+# hops between a debug file and its dwz common file.
+
+set pagination off
+set confirm off
+set debuginfod enabled off
+set print pretty on
+set height 0
+set width 0
+
+define perf-die-chain
+  if $argc < 2
+    printf "usage: perf-die-chain <function> <die variable> [iterations]\n"
+  else
+    frame function $arg0
+    if $argc == 3
+      set $perf_die_chain_n = $arg2
+    else
+      set $perf_die_chain_n = 10
+    end
+    set $perf_die_chain_head = $pc
+    set $perf_die_chain_i = 0
+    while $perf_die_chain_i < $perf_die_chain_n
+      # Pointer type DIEs have no DW_AT_name, so dwarf_diename() can
+      # return NULL: printf %s of it would error out and abort this
+      # batch script, handle it.
+      set $perf_die_chain_name = (char *) dwarf_diename($arg1)
+      printf "chain[%d] die=%p addr=%p cu=%p off=0x%lx tag=%d name=", $perf_die_chain_i, $arg1, $arg1->addr, $arg1->cu, ((Dwarf_Off) dwarf_dieoffset($arg1)), ((int) dwarf_tag($arg1))
+      if $perf_die_chain_name == 0
+	printf "(null)\n"
+      else
+	printf "%s\n", $perf_die_chain_name
+      end
+      until *$perf_die_chain_head
+      set $perf_die_chain_i = $perf_die_chain_i + 1
+    end
+  end
+end
+
+document perf-die-chain
+Print the DIE chain being walked by a DWARF type chasing loop.
+usage: perf-die-chain <function> <die variable> [iterations]
+  perf-die-chain die_get_pointer_type type_die
+  perf-die-chain __die_get_real_type vr_die
+  perf-die-chain die_get_real_type vr_die
+end
+
+define perf-die-chain-all
+  if $argc == 0
+    set $perf_die_chain_n = 10
+  else
+    set $perf_die_chain_n = $arg0
+  end
+  if $_any_caller_is("die_get_pointer_type", 20)
+    printf "stuck in die_get_pointer_type():\n"
+    perf-die-chain die_get_pointer_type type_die $perf_die_chain_n
+  else
+    if $_any_caller_is("__die_get_real_type", 20)
+      printf "stuck in __die_get_real_type():\n"
+      perf-die-chain __die_get_real_type vr_die $perf_die_chain_n
+    else
+      if $_any_caller_is("die_get_real_type", 20)
+        printf "stuck in die_get_real_type():\n"
+        perf-die-chain die_get_real_type vr_die $perf_die_chain_n
+      else
+        printf "not in a DWARF type chaser, try: bt\n"
+      end
+    end
+  end
+end
+
+document perf-die-chain-all
+Find which DWARF type chaser the process is in and print the DIE chain.
+usage: perf-die-chain-all [iterations]
+end
+
+define perf-dso
+  if $_any_caller_is("find_data_type", 20)
+    frame function find_data_type
+    # REFCNT_CHECKING, implied by an ASan/LSan build, wraps struct map and
+    # struct dso in a proxy keeping the real object in ->orig, the dso being
+    # map->orig->dso->orig there; struct symbol is not wrapped, so sym->name
+    # needs no ->orig.  There is no way to ask gdb which layout it is looking
+    # at, so walk each candidate until one evaluates.  Requiring gdb's Python
+    # here adds nothing: $_any_caller_is() is one of its functions already.
+    python
+import gdb
+
+dso = None
+for expr in ("dloc->ms->map->dso->name",
+             "dloc->ms->map->orig->dso->orig->name"):
+    try:
+        gdb.parse_and_eval(expr)
+        dso = expr
+        break
+    except gdb.error:
+        pass
+
+if dso is None:
+    print("dso=(no such member in struct map)")
+else:
+    gdb.execute('printf "dso=%s ip=0x%lx sym=%s\\n", ' + dso +
+                ', dloc->ip, dloc->ms->sym ? dloc->ms->sym->name : "(no symbol)"')
+    end
+  else
+    printf "not in find_data_type()\n"
+  end
+end
+
+document perf-dso
+Print the dso, ip and symbol of the data location being resolved.
+end
diff --git a/tools/perf/scripts/perf-stuck.sh b/tools/perf/scripts/perf-stuck.sh
new file mode 100755
index 0000000000000000..f9b545c107905a4e
--- /dev/null
+++ b/tools/perf/scripts/perf-stuck.sh
@@ -0,0 +1,193 @@
+#!/bin/bash
+# SPDX-License-Identifier: GPL-2.0
+#
+# perf-stuck - tell a spinning perf apart from a blocked or recursing one
+#
+# Arnaldo Carvalho de Melo <acme@redhat.com>
+#
+# PROTOTYPE: wants to become a first class 'perf stuck' command, sampling
+# a process from inside perf, with the knowledge of perf's phases and of
+# the DWARF type chasing loops built in, instead of poking /proc and
+# shelling out to gdb.
+#
+# Samples /proc/<pid> at a fixed interval and prints the CPU time used
+# since the previous sample, the [stack] mapping start and size, and the
+# last line of a progress log when one is given, e.g. the stderr of
+# 'perf report --progress': burning a full interval with a constant
+# stack is a loop, a [stack] start moving down is runaway recursion.
+#
+# With -g it runs gdb (perf-stuck.gdb) when no progress is made for two
+# consecutive samples, printing the DIE chain a DWARF type chase is
+# walking.
+#
+# usage: perf-stuck.sh [options] <pid|process-name>
+
+set -u
+
+usage() {
+	cat <<-EOF
+	usage: perf-stuck.sh [options] <pid|process-name>
+
+	  -i <secs>   sampling interval (default: 10)
+	  -n <count>  stop after this many samples (default: watch till it exits)
+	  -l <file>   progress log, its last line is printed with every sample
+	  -g          run gdb with perf-stuck.gdb when no progress is made for
+	              two consecutive samples, writing the output to a temp file
+	  -x <file>   use this gdb command file instead of perf-stuck.gdb
+	  -h          this help
+	EOF
+	exit "${1:-0}"
+}
+
+interval=10
+count=0
+progress_log=
+use_gdb=
+gdb_cmds=
+
+while getopts "i:n:l:gx:h" opt; do
+	case "$opt" in
+	i) interval=$OPTARG ;;
+	n) count=$OPTARG ;;
+	l) progress_log=$OPTARG ;;
+	g) use_gdb=1 ;;
+	x) gdb_cmds=$OPTARG ;;
+	h) usage 0 ;;
+	*) usage 1 ;;
+	esac
+done
+shift $((OPTIND - 1))
+
+[[ "$interval" =~ ^[0-9]+$ ]] && [ "$interval" -gt 0 ] ||
+	{ echo "-i wants a positive integer, got '$interval'"; usage 1; }
+[[ "$count" =~ ^[0-9]+$ ]] ||
+	{ echo "-n wants a non-negative integer, got '$count'"; usage 1; }
+[ -n "$gdb_cmds" ] && [ -z "$use_gdb" ] && { echo "-x needs -g"; usage 1; }
+
+[ $# -eq 1 ] || usage 1
+
+if [[ "$1" =~ ^[0-9]+$ ]]; then
+	pid=$1
+else
+	# Resolve the name against the caller's own processes: as root,
+	# unscoped pgrep picks the first match of any user, e.g. one planted
+	# to get gdb attached to it, use an explicit pid to look at a perf of
+	# another user.
+	pid=$(pgrep -x -u "$(id -u)" -- "$1" | head -1)
+	[ -n "$pid" ] || { echo "no process named '$1' owned by $(id -un)"; exit 1; }
+fi
+
+[ -d /proc/"$pid" ] || { echo "no process $pid"; exit 1; }
+
+if [ -n "$use_gdb" ]; then
+	# Fail before watching, not two samples in when -g would fire.
+	command -v gdb > /dev/null || { echo "gdb not found, -g needs it"; exit 1; }
+	[ -z "$gdb_cmds" ] && gdb_cmds=$(dirname "$0")/perf-stuck.gdb
+	[ -r "$gdb_cmds" ] || { echo "cannot read $gdb_cmds"; exit 1; }
+fi
+
+hz=$(getconf CLK_TCK)
+psz=$(getconf PAGESIZE)
+prev_cpu=
+prev_stack=
+prev_progress=
+stuck=0
+gdb_done=
+nsample=0
+
+# The command line is whatever the process was started with, so drop the
+# control characters from it: a process started with escape sequences in
+# its arguments, e.g. one replaying a log line, would otherwise get them
+# replayed on the terminal of whoever runs this.
+cmdline=$(tr '\0' ' ' < /proc/"$pid"/cmdline | tr -d '[:cntrl:]')
+
+echo "watching $pid ($cmdline) every ${interval}s"
+
+while :; do
+	if [ ! -d /proc/"$pid" ]; then
+		echo "$(date +%T) process gone"
+		break
+	fi
+
+	# Field 2, the command name, is in parentheses and can contain
+	# spaces, so drop it together with the pid before splitting so the
+	# fields line up.  %d keeps the CPU time out of scientific
+	# notation, that bash arithmetic can't parse past six digits.
+	if ! stat_line=$(awk '{ sub(/^[^ ]+ \(.*\) /, "");
+			       printf "%s %d %d\n", $1, $12 + $13, $22 }' \
+			 /proc/"$pid"/stat 2>/dev/null); then
+		echo "$(date +%T) process gone"
+		break
+	fi
+
+	# The process can be gone between the check above and this read, in
+	# which case there is nothing to report: 'set -u' would otherwise
+	# turn the unbound fields into an aborted script.
+	if [ -z "$stat_line" ]; then
+		echo "$(date +%T) process gone"
+		break
+	fi
+
+	stat=($stat_line)
+	state=${stat[0]}
+	cpu=${stat[1]}
+	# field 24 of /proc/<pid>/stat, the resident set size in pages
+	rss=$(( stat[2] * psz / 1024 ))
+
+	stack=$(awk '/\[stack\]/{print $1; exit}' /proc/"$pid"/maps 2>/dev/null)
+	if [ -n "$stack" ]; then
+		stack_start=0x${stack%-*}
+		stack_size=$(( 0x${stack#*-} - stack_start ))
+		stack_txt="$stack size=$((stack_size / 1024))kB"
+	else
+		stack_start=
+		stack_txt="-"
+	fi
+
+	progress=
+	[ -n "$progress_log" ] && [ -s "$progress_log" ] && progress=$(tail -1 -- "$progress_log")
+
+	if [ -n "$prev_cpu" ]; then
+		cpu_delta=$(( cpu - prev_cpu ))
+		# No progress log, or one with nothing written to it yet, leaves
+		# no progress to look at, so count the samples that show none:
+		# -g then looks at where the process is after two intervals.
+		# Otherwise count the repeats of the same last line.
+		if [ -z "$progress" ] || [ "$progress" = "$prev_progress" ]; then
+			stuck=$((stuck + 1))
+		else
+			stuck=0
+			gdb_done=
+		fi
+		stuck_txt="stuck=${stuck}"
+		[ "$stack_start" != "$prev_stack" ] && stuck_txt="$stuck_txt STACK"
+	else
+		cpu_delta=0
+		stuck_txt=""
+	fi
+
+	printf '%s state=%s cpu=+%d (%d.%02ds) rss=%dkB stack=%s %s %s\n' \
+	       "$(date +%T)" "$state" "$cpu_delta" \
+	       $(( cpu_delta / hz )) $(( (cpu_delta % hz) * 100 / hz )) \
+	       "$rss" "$stack_txt" "$stuck_txt" "${progress:-(no progress log)}"
+
+	if [ -n "$use_gdb" ] && [ -z "$gdb_done" ] && [ "$stuck" -ge 2 ]; then
+		gdb_log=$(mktemp /tmp/perf-stuck-gdb.XXXXXX)
+		# The gdb script makes inferior calls, and a call into a perf
+		# wedged in a loop never returns, so bound the run; SIGINT
+		# releases the inferior instead of leaving perf stopped.
+		timeout --signal=INT 30 gdb -p "$pid" -batch -x "$gdb_cmds" -ex bt \
+		    -ex 'perf-die-chain-all' -ex perf-dso -ex detach > "$gdb_log" 2>&1
+		gdb_done=1
+		echo "... gdb output of $pid in $gdb_log"
+	fi
+
+	prev_cpu=$cpu
+	prev_stack=$stack_start
+	prev_progress=$progress
+
+	nsample=$((nsample + 1))
+	[ "$count" -gt 0 ] && [ "$nsample" -ge "$count" ] && break
+
+	sleep "$interval"
+done
-- 
2.55.0


^ permalink raw reply	[flat|nested] 15+ messages in thread

* [PATCH v6 5/5] perf test: Add false_sharing workload exhibiting cross-CPU false sharing
  2026-09-30 21:37 [PATCH v6 0/5] perf tools: Add progress diagnostics and a false-sharing workload Arnaldo Carvalho de Melo
                   ` (3 preceding siblings ...)
  2026-09-30 21:37 ` [PATCH v6 4/5] perf scripts: Add perf-stuck, to tell where a running perf is stuck Arnaldo Carvalho de Melo
@ 2026-09-30 21:37 ` Arnaldo Carvalho de Melo
  2026-10-01  7:28   ` Namhyung Kim
  4 siblings, 1 reply; 15+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-09-30 21:37 UTC (permalink / raw)
  To: Namhyung Kim
  Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
	Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
	Arnaldo Carvalho de Melo

From: Arnaldo Carvalho de Melo <acme@redhat.com>

Add a 'perf test -w false_sharing' workload that hammers one shared
struct from several CPUs, shaped as a TCP connection: a read-mostly
identity (five-tuple) shares a cacheline with per-packet rx counters
(the false-sharing line), a second line has packet-path private tx and
congestion control counters, and a third the connection config.

The packet path runs in the main thread and up to four lookup threads
sum the five-tuple and pull the config, reading one volatile shared
instance directly so the accesses are PC-relative and resolvable by the
data type profiler.

Assisted-by: LLM
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
---
 tools/perf/tests/builtin-test.c               |   1 +
 tools/perf/tests/shell/data_type_profiling.sh |   9 +-
 tools/perf/tests/tests.h                      |   1 +
 tools/perf/tests/workloads/Build              |   2 +
 tools/perf/tests/workloads/false_sharing.c    | 251 ++++++++++++++++++
 5 files changed, 262 insertions(+), 2 deletions(-)
 create mode 100644 tools/perf/tests/workloads/false_sharing.c

diff --git a/tools/perf/tests/builtin-test.c b/tools/perf/tests/builtin-test.c
index d2f594921e25bda9..98134b1c74cf80f4 100644
--- a/tools/perf/tests/builtin-test.c
+++ b/tools/perf/tests/builtin-test.c
@@ -175,6 +175,7 @@ static struct test_workload *workloads[] = {
 	&workload__context_switch_loop,
 	&workload__deterministic,
 	&workload__callchain,
+	&workload__false_sharing,
 
 #ifdef HAVE_RUST_SUPPORT
 	&workload__code_with_type,
diff --git a/tools/perf/tests/shell/data_type_profiling.sh b/tools/perf/tests/shell/data_type_profiling.sh
index a916c410274aa888..57203e0e5f857540 100755
--- a/tools/perf/tests/shell/data_type_profiling.sh
+++ b/tools/perf/tests/shell/data_type_profiling.sh
@@ -8,8 +8,8 @@ set -e
 # data type profiling manifestation
 
 # Values in testtypes and testprogs should match
-testtypes=("# data-type: struct Buf" "# data-type: struct buf")
-testprogs=("perf test -w code_with_type" "perf test -w datasym")
+testtypes=("# data-type: struct Buf" "# data-type: struct buf" "# data-type: struct net_conn")
+testprogs=("perf test -w code_with_type" "perf test -w datasym" "perf test -w false_sharing")
 
 err=0
 perfdata=$(mktemp /tmp/__perf_test.perf.data.XXXXX)
@@ -59,6 +59,9 @@ test_basic_annotate() {
 
     "xC")
     index=1 ;;
+
+    "xFS")
+    index=2 ;;
   esac
 
   # Under 'set -e' a bare failing command aborts the script through the EXIT
@@ -114,6 +117,8 @@ test_basic_annotate Basic Rust
 test_basic_annotate Pipe Rust
 test_basic_annotate Basic C
 test_basic_annotate Pipe C
+test_basic_annotate Basic FS
+test_basic_annotate Pipe FS
 
 cleanup
 exit $err
diff --git a/tools/perf/tests/tests.h b/tools/perf/tests/tests.h
index 9c96f33483d14356..f379620069d34f3a 100644
--- a/tools/perf/tests/tests.h
+++ b/tools/perf/tests/tests.h
@@ -250,6 +250,7 @@ DECLARE_WORKLOAD(jitdump);
 DECLARE_WORKLOAD(context_switch_loop);
 DECLARE_WORKLOAD(deterministic);
 DECLARE_WORKLOAD(callchain);
+DECLARE_WORKLOAD(false_sharing);
 
 #ifdef HAVE_RUST_SUPPORT
 DECLARE_WORKLOAD(code_with_type);
diff --git a/tools/perf/tests/workloads/Build b/tools/perf/tests/workloads/Build
index 048e371eb63e3164..18fceddccb56c013 100644
--- a/tools/perf/tests/workloads/Build
+++ b/tools/perf/tests/workloads/Build
@@ -14,6 +14,7 @@ perf-test-y += jitdump.o
 perf-test-y += context_switch_loop.o
 perf-test-y += deterministic.o
 perf-test-y += callchain.o
+perf-test-y += false_sharing.o
 
 ifeq ($(CONFIG_RUST_SUPPORT),y)
     perf-test-y += code_with_type.o
@@ -29,3 +30,4 @@ CFLAGS_inlineloop.o       = -g -O2
 CFLAGS_deterministic.o    = -g -O0 -fno-inline -U_FORTIFY_SOURCE
 CFLAGS_named_threads.o    = -g -O0 -fno-inline -U_FORTIFY_SOURCE
 CFLAGS_callchain.o        = -g -O0 -fno-inline -U_FORTIFY_SOURCE
+CFLAGS_false_sharing.o    = -g -O0 -fno-inline -U_FORTIFY_SOURCE
diff --git a/tools/perf/tests/workloads/false_sharing.c b/tools/perf/tests/workloads/false_sharing.c
new file mode 100644
index 0000000000000000..e949bb36646a97a5
--- /dev/null
+++ b/tools/perf/tests/workloads/false_sharing.c
@@ -0,0 +1,251 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * False-sharing demo for data type profiling, shaped as a TCP
+ * connection: a read-mostly identity shares a cacheline with per-packet
+ * rx counters (the false-sharing line), a second line has packet-path
+ * private tx and congestion control counters, and a third the connection
+ * config.
+ *
+ * 'perf mem record' of this workload followed by 'perf report -s type'
+ * (see tests/shell/data_type_profiling.sh) resolves the accesses to
+ * struct net_conn members, showing the rx counters and the read-mostly
+ * identity sharing cacheline 0.
+ */
+#include <pthread.h>
+#include <sched.h>
+#include <stdint.h>
+#include <stdlib.h>
+#include <stdio.h>
+#include <signal.h>
+#include <unistd.h>
+#include <linux/compiler.h>
+#include "../tests.h"
+
+struct net_conn {
+	/* cacheline 0: identity (read-mostly) + rx counters (per packet) */
+	uint32_t	saddr;		/*   0 */
+	uint32_t	daddr;		/*   4 */
+	uint16_t	sport;		/*   8 */
+	uint16_t	dport;		/*  10 */
+	uint8_t		state;		/*  12: 1 == ESTABLISHED */
+	uint8_t		protocol;	/*  13: 6 == TCP */
+	uint16_t	__pad0;		/*  14 */
+	uint64_t	bytes_rx;	/*  16: every packet */
+	uint64_t	packets_rx;	/*  24: every packet */
+	uint32_t	rx_queue;	/*  32: backlog depth, fluctuates */
+	uint8_t		__pad1[24];	/*  36..59 */
+	uint32_t	last_ack;	/*  60: written per ACK */
+	/* cacheline 1: tx + congestion control (packet-path private) */
+	uint64_t	bytes_tx;	/*  64: every packet */
+	uint64_t	packets_tx;	/*  72: every packet */
+	uint32_t	cwnd;		/*  80: on every ACK */
+	uint32_t	ssthresh;	/*  84: on loss */
+	uint32_t	rtt_us;		/*  88: on every ACK */
+	uint32_t	retrans;	/*  92: on timeout */
+	uint32_t	__pad2[8];	/*  96..127 */
+	/* cacheline 2: config, set at setup, read by everybody */
+	uint16_t	mss;		/* 128 */
+	uint8_t		snd_wscale;	/* 130 */
+	uint8_t		rcv_wscale;	/* 131 */
+	uint32_t	keepalive_int;	/* 132 */
+	uint32_t	mark;		/* 136: firewall mark */
+	uint32_t	priority;	/* 140: traffic class */
+	uint32_t	__pad3[12];	/* 144..191 */
+} __attribute__((aligned(64)));
+
+/* Volatile so every iteration really loads and stores. */
+static volatile struct net_conn conn;
+/* Keeps the reader checksums alive after the threads join. */
+static volatile unsigned long fs_sink;
+
+static volatile sig_atomic_t done;
+
+/*
+ * One cacheline each: sum before cpu, or the implicit padding after cpu
+ * pushes the struct past 64 bytes, and aligning it to a cacheline then
+ * rounds it up to 128.
+ */
+struct fs_reader {
+	pthread_t	thread;
+	unsigned long	sum;
+	int		cpu;
+	char		__pad[64 - sizeof(pthread_t) - sizeof(unsigned long) - sizeof(int)];
+} __attribute__((aligned(64)));
+
+static void sighandler(int sig __maybe_unused)
+{
+	done = 1;
+}
+
+static void pin_to_cpu(int cpu)
+{
+	cpu_set_t set;
+
+	/* There may be no second CPU in a restricted cpuset. */
+	if (cpu < 0)
+		return;
+
+	CPU_ZERO(&set);
+	CPU_SET(cpu, &set);
+	/* Best effort: in a restricted cpuset this fails and the thread runs unpinned. */
+	pthread_setaffinity_np(pthread_self(), sizeof(set), &set);
+}
+
+/*
+ * Connection lookup, as a load balancer or 'ss' scrape would do it: the
+ * reads go straight to the global, this file is built -O0 and a local
+ * pointer would be reloaded from a stack slot, a form the data type
+ * resolver does not track.
+ */
+static void *reader_fn(void *arg)
+{
+	struct fs_reader *r = arg;
+	unsigned long sum = 0;
+
+	pthread_setname_np(pthread_self(), "fs-reader");
+	pin_to_cpu(r->cpu);
+
+	while (!done) {
+		sum += conn.saddr + conn.daddr + conn.sport + conn.dport +
+		       conn.state + conn.protocol;
+		sum += conn.mss + conn.snd_wscale + conn.rcv_wscale +
+		       conn.keepalive_int + conn.mark + conn.priority;
+	}
+	r->sum = sum;
+	return NULL;
+}
+
+static int false_sharing(int argc, const char **argv)
+{
+	double sec = 2.0;
+	int nreaders = 0, nr_allowed = 0, err = 1;
+	int *allowed = NULL, nallowed = 0;
+	cpu_set_t set;
+	int nr_mask_bits = sizeof(set) * 8 < CPU_SETSIZE ? sizeof(set) * 8 : CPU_SETSIZE;
+	struct fs_reader *readers = NULL;
+	int i, writer_cpu;
+	unsigned long n = 0;
+
+	pthread_setname_np(pthread_self(), "fs-writer");
+	if (argc > 0)
+		sec = atof(argv[0]);
+	if (!(sec > 0.0)) {
+		fprintf(stderr, "Error: seconds (%f) must be > 0\n", sec);
+		return 1;
+	}
+	if (argc > 1)
+		nreaders = atoi(argv[1]);
+
+	/*
+	 * A connection that just got established: identity and config fixed
+	 * from here on, counters at zero.
+	 */
+	conn.saddr = 0x0a000001;	/* 10.0.0.1 */
+	conn.daddr = 0x0a000002;	/* 10.0.0.2 */
+	conn.sport = 54321;
+	conn.dport = 443;
+	conn.state = 1;			/* ESTABLISHED */
+	conn.protocol = 6;		/* TCP */
+	conn.mss = 1448;
+	conn.snd_wscale = 7;
+	conn.rcv_wscale = 7;
+	conn.keepalive_int = 7200;
+	conn.cwnd = 10;
+	conn.ssthresh = 65535;
+	conn.rtt_us = 50;
+
+	/*
+	 * Pin against the allowed set, restricted cpusets still spread the threads.
+	 * The whole mask is looked at, not the CPU count: the count can be lower
+	 * than the highest ID in it, as when a cpuset allows only high numbered
+	 * CPUs, and then no allowed CPU would be found at all.
+	 */
+	if (sched_getaffinity(0, sizeof(set), &set) == 0) {
+		for (i = 0; i < nr_mask_bits; i++) {
+			if (!CPU_ISSET(i, &set))
+				continue;
+			nr_allowed++;
+		}
+		allowed = malloc(nr_allowed * sizeof(int));
+		if (allowed == NULL) {
+			fprintf(stderr, "Error: malloc failed for CPU list\n");
+			return 1;
+		}
+		for (i = 0; i < nr_mask_bits; i++) {
+			if (CPU_ISSET(i, &set))
+				allowed[nallowed++] = i;
+		}
+	}
+	if (nreaders <= 0) {
+		/* By default leave one CPU for the packet path, up to 4 readers. */
+		nreaders = nallowed > 1 ? nallowed - 1 : 1;
+		if (nreaders > 4)
+			nreaders = 4;
+	}
+
+	signal(SIGINT, sighandler);
+	signal(SIGALRM, sighandler);
+
+	readers = calloc(nreaders, sizeof(*readers));
+	if (readers == NULL) {
+		fprintf(stderr, "Error: calloc failed for %d readers\n", nreaders);
+		goto out;
+	}
+	for (i = 0; i < nreaders; i++) {
+		int cpu = nallowed > 1 ? allowed[(i + 1) % nallowed] : -1;
+
+		readers[i].cpu = cpu;
+		if (pthread_create(&readers[i].thread, NULL, reader_fn, &readers[i])) {
+			fprintf(stderr, "Error: failed to create reader %d\n", i);
+			done = 1; // Ensure started threads terminate.
+			nreaders = i;
+			goto out_join;
+		}
+	}
+	writer_cpu = nallowed > 0 ? allowed[0] : -1;
+	if (nallowed == 1)
+		fprintf(stderr, "Warning: single CPU allowed, no cross-CPU traffic expected\n");
+	if (writer_cpu >= 0)
+		pin_to_cpu(writer_cpu);
+
+	/*
+	 * The packet path: receive, acknowledge, transmit, repeat; every 64th
+	 * packet simulates a loss so ssthresh/retrans get sampled too.
+	 */
+	if (sec < 1.0) {
+		useconds_t usecs = (useconds_t)(sec * 1000000.0);
+
+		ualarm(usecs > 0 ? usecs : 1, 0);
+	} else
+		alarm((unsigned int)sec);
+	while (!done) {
+		conn.bytes_rx += conn.mss;
+		conn.packets_rx++;
+		conn.rx_queue = (uint32_t)(n & 0x3f);
+		conn.last_ack = (uint32_t)n;
+		conn.bytes_tx += conn.mss;
+		conn.packets_tx++;
+		conn.cwnd = 10 + (n & 15);
+		conn.rtt_us = 50 + (n & 7);
+		if ((n & 63) == 0) {
+			conn.ssthresh = conn.cwnd / 2;
+			conn.retrans++;
+		}
+		n++;
+	}
+	err = 0;
+out_join:
+	for (i = 0; i < nreaders; i++) {
+		if (readers[i].thread) {
+			pthread_join(readers[i].thread, /*retval=*/NULL);
+			fs_sink += readers[i].sum;
+		}
+	}
+	fs_sink += (unsigned long)(conn.bytes_rx + conn.bytes_tx + n);
+	free(readers);
+out:
+	free(allowed);
+	return err;
+}
+
+DEFINE_WORKLOAD(false_sharing);
-- 
2.55.0


^ permalink raw reply	[flat|nested] 15+ messages in thread

* Re: [PATCH v6 3/5] perf report: Add --no-progress option
  2026-09-30 21:37 ` [PATCH v6 3/5] perf report: Add --no-progress option Arnaldo Carvalho de Melo
@ 2026-10-01  7:01   ` Namhyung Kim
  2026-10-01  9:20     ` Arnaldo Carvalho de Melo
  0 siblings, 1 reply; 15+ messages in thread
From: Namhyung Kim @ 2026-10-01  7:01 UTC (permalink / raw)
  To: Arnaldo Carvalho de Melo
  Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
	Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
	Arnaldo Carvalho de Melo

On Wed, Sep 30, 2026 at 11:37:14PM +0200, Arnaldo Carvalho de Melo wrote:
> From: Arnaldo Carvalho de Melo <acme@redhat.com>
> 
> Now that --progress is being added for the stdio case, wire up its
> counterpart for the browsers: the TUI and GTK ones present progress of
> their own and there is no way to turn it off.  Install the no-op
> ui_progress ops, the ones already used until a backend sets theirs,
> after setup_browser() installed the ones of the browser in use.  The
> phases are still counted, nothing is shown for them, and no second
> option is needed for it: parse-options provides --no-progress as the
> negation of --progress, report.progress_set saying that it was asked
> for, report.progress being false both when nothing was asked for and
> when --no-progress was.
> 
> Suggested-by: Namhyung Kim <namhyung@kernel.org>
> Assisted-by: LLM
> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
> ---
>  tools/perf/Documentation/perf-report.txt |  5 +++++
>  tools/perf/builtin-report.c              | 20 ++++++++++++++++----
>  tools/perf/ui/progress.c                 |  6 ++++++
>  tools/perf/ui/progress.h                 |  2 ++
>  4 files changed, 29 insertions(+), 4 deletions(-)
> 
> diff --git a/tools/perf/Documentation/perf-report.txt b/tools/perf/Documentation/perf-report.txt
> index a7429a30ec28f903..e145124d6c9a1097 100644
> --- a/tools/perf/Documentation/perf-report.txt
> +++ b/tools/perf/Documentation/perf-report.txt
> @@ -43,6 +43,11 @@ OPTIONS
>  	present progress information, or when --quiet is used, that asks
>  	for no messages at all.
>  
> +--no-progress::
> +	Do not show progress while processing the perf.data file.  It
> +	also turns off the progress the TUI and GTK browsers present,
> +	which is their own.
> +
>  -n::
>  --show-nr-samples::
>  	Show the number of samples for each symbol
> diff --git a/tools/perf/builtin-report.c b/tools/perf/builtin-report.c
> index 963808b561e547c3..105b2859678673ec 100644
> --- a/tools/perf/builtin-report.c
> +++ b/tools/perf/builtin-report.c
> @@ -88,6 +88,7 @@ struct report {
>  #endif
>  	bool			use_stdio;
>  	bool			progress;
> +	bool			progress_set;
>  	bool			show_full_info;
>  	bool			show_threads;
>  	bool			inverted_callchain;
> @@ -1385,8 +1386,15 @@ int cmd_report(int argc, const char **argv)
>  		    "Use the stdio interface"),
>  	OPT_BOOLEAN(0, "weights", &symbol_conf.annotate_weight,
>  			"Show or hide weight columns in annotation. Default show if non-zero."),
> -	OPT_BOOLEAN(0, "progress", &report.progress,
> -		    "Show progress while processing the perf.data file"),
> +	/*
> +	 * No --no-progress option to add: parse-options provides it as
> +	 * the negation of this one, clearing report.progress, which is
> +	 * also how it starts out.  progress_set is what tells the hook
> +	 * below to stop the TUI and GTK browsers as well, they show
> +	 * progress until asked not to.
> +	 */

Nit: I think we can drop this comment.


> +	OPT_BOOLEAN_SET(0, "progress", &report.progress, &report.progress_set,
> +			"Show progress while processing the perf.data file"),
>  	OPT_BOOLEAN(0, "header", &report.header, "Show data header."),
>  	OPT_BOOLEAN(0, "header-only", &report.header_only,
>  		    "Show only data header."),
> @@ -1796,9 +1804,13 @@ int cmd_report(int argc, const char **argv)
>  	/*
>  	 * For the stdio case: print the percentage of the perf.data file
>  	 * processed so far for each processing phase.  --quiet asks for no
> -	 * messages at all, so it leaves the phases uncounted.
> +	 * messages at all, so it leaves the phases uncounted, and
> +	 * --no-progress, progress_set with progress cleared, turns off what
> +	 * the TUI and GTK browsers show as well.
>  	 */

Maybe this one too..

Thanks,
Namhyung


> -	if (report.progress && !quiet && use_browser == 0)
> +	if (report.progress_set && !report.progress)
> +		ui_progress__noop_init();
> +	else if (report.progress && !quiet && use_browser == 0)
>  		stdio_progress__init();
>  
>  	if (report.data_type && use_browser == 1) {
> diff --git a/tools/perf/ui/progress.c b/tools/perf/ui/progress.c
> index 99d60223c74b2957..362680989ace606a 100644
> --- a/tools/perf/ui/progress.c
> +++ b/tools/perf/ui/progress.c
> @@ -13,6 +13,12 @@ static struct ui_progress_ops null_progress__ops =
>  
>  struct ui_progress_ops *ui_progress__ops = &null_progress__ops;
>  
> +/* Everything counts but nothing is shown, the way it starts out. */
> +void ui_progress__noop_init(void)
> +{
> +	ui_progress__ops = &null_progress__ops;
> +}
> +
>  void ui_progress__update(struct ui_progress *p, u64 adv)
>  {
>  	u64 last = p->curr;
> diff --git a/tools/perf/ui/progress.h b/tools/perf/ui/progress.h
> index 03f1a8bb260ba076..e8c4f9f768aaf12b 100644
> --- a/tools/perf/ui/progress.h
> +++ b/tools/perf/ui/progress.h
> @@ -25,6 +25,8 @@ void ui_progress__update(struct ui_progress *p, u64 adv);
>  
>  void stdio_progress__init(void);
>  
> +void ui_progress__noop_init(void);
> +
>  struct ui_progress_ops {
>  	void (*init)(struct ui_progress *p);
>  	void (*update)(struct ui_progress *p);
> -- 
> 2.55.0
> 

^ permalink raw reply	[flat|nested] 15+ messages in thread

* Re: [PATCH v6 1/5] perf config: Move perf_config__set_variable() to util/config.c
  2026-09-30 21:37 ` [PATCH v6 1/5] perf config: Move perf_config__set_variable() to util/config.c Arnaldo Carvalho de Melo
@ 2026-10-01  7:18   ` Namhyung Kim
  2026-10-01  9:19     ` Arnaldo Carvalho de Melo
  0 siblings, 1 reply; 15+ messages in thread
From: Namhyung Kim @ 2026-10-01  7:18 UTC (permalink / raw)
  To: Arnaldo Carvalho de Melo
  Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
	Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
	Arnaldo Carvalho de Melo

On Wed, Sep 30, 2026 at 11:37:12PM +0200, Arnaldo Carvalho de Melo wrote:
> From: Arnaldo Carvalho de Melo <acme@redhat.com>
> 
> Move perf_config__set_variable() out of the 'perf config' builtin so
> that opt-in features can persist their choice from outside it, e.g.
> util/debuginfo.c writing core.debuginfod=false when the user disables
> debuginfod for the rest of the session.  The
> set_config() body becomes perf_config_set__write(), with the
> system_config choice as an argument, as the builtin's
> use_system_config/use_user_config statics are not available outside it.
> 
> All config file access shares the static parser state and can now run
> on more than one thread, with a feature writing the config while perf
> top's display thread reads it, so serialize parsing and
> rewriting with a mutex, and the whole read-modify-write of
> perf_config__set_variable() against itself.
> 
> perf_config_set__write() checked fopen() but none of the fprintf()s or
> fclose(), so a write failure after truncating the file was reported as
> success.  Harmless for the interactive 'perf config' this came from,
> but this makes it an entry point a background feature can call with no
> other feedback, so propagate those errors too.
> 
> perf_etc_perfconfig() caches what system_path() returns, and that
> allocates, so on failure every caller dereferenced NULL, the
> pre-existing ones in perf_config_from_file() and in the daemon
> included.  Pre-existing, not introduced here, but this rewrites the
> accessor anyway, so make it total: the unresolved path is the same
> string for the usual absolute ETC_PERFCONFIG.

I think this patch does many things.  Can we split them? 

> 
> Assisted-by: LLM
> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
> ---
>  tools/perf/builtin-config.c |  70 +-----------
>  tools/perf/util/config.c    | 219 ++++++++++++++++++++++++++++++++----
>  tools/perf/util/config.h    |   2 +
>  3 files changed, 201 insertions(+), 90 deletions(-)
> 
> diff --git a/tools/perf/builtin-config.c b/tools/perf/builtin-config.c
> index cefd042e4f853466..3b074aca8d344539 100644
> --- a/tools/perf/builtin-config.c
> +++ b/tools/perf/builtin-config.c
> @@ -41,37 +41,7 @@ static struct option config_options[] = {
>  
>  static int set_config(struct perf_config_set *set, const char *file_name)
>  {
> -	struct perf_config_section *section = NULL;
> -	struct perf_config_item *item = NULL;
> -	const char *first_line = "# this file is auto-generated.";
> -	FILE *fp;
> -
> -	if (set == NULL)
> -		return -1;
> -
> -	fp = fopen(file_name, "w");
> -	if (!fp)
> -		return -1;
> -
> -	fprintf(fp, "%s\n", first_line);
> -
> -	/* overwrite configvariables */
> -	perf_config_items__for_each_entry(&set->sections, section) {
> -		if (!use_system_config && section->from_system_config)
> -			continue;
> -		fprintf(fp, "[%s]\n", section->name);
> -
> -		perf_config_items__for_each_entry(&section->items, item) {
> -			if (!use_system_config && item->from_system_config)
> -				continue;
> -			if (item->value)
> -				fprintf(fp, "\t%s = %s\n",
> -					item->name, item->value);
> -		}
> -	}
> -	fclose(fp);
> -
> -	return 0;
> +	return perf_config_set__write(set, file_name, use_system_config);
>  }
>  
>  static int show_spec_config(struct perf_config_set *set, const char *var)
> @@ -158,44 +128,6 @@ static int parse_config_arg(char *arg, char **var, char **value)
>  	return 0;
>  }
>  
> -int perf_config__set_variable(const char *var, const char *value)
> -{
> -	char path[PATH_MAX];
> -	char *user_config = mkpath(path, sizeof(path), "%s/.perfconfig", getenv("HOME"));
> -	const char *config_filename;
> -	struct perf_config_set *set;
> -	int ret = -1;
> -
> -	if (use_system_config)
> -		config_exclusive_filename = perf_etc_perfconfig();
> -	else if (use_user_config)
> -		config_exclusive_filename = user_config;
> -
> -	if (!config_exclusive_filename)
> -		config_filename = user_config;
> -	else
> -		config_filename = config_exclusive_filename;
> -
> -	set = perf_config_set__new();
> -	if (!set)
> -		goto out_err;
> -
> -	if (perf_config_set__collect(set, config_filename, var, value) < 0) {
> -		pr_err("Failed to add '%s=%s'\n", var, value);
> -		goto out_err;
> -	}
> -
> -	if (set_config(set, config_filename) < 0) {
> -		pr_err("Failed to set the configs on %s\n", config_filename);
> -		goto out_err;
> -	}
> -
> -	ret = 0;
> -out_err:
> -	perf_config_set__delete(set);
> -	return ret;
> -}
> -
>  int cmd_config(int argc, const char **argv)
>  {
>  	int i, ret = -1;
> diff --git a/tools/perf/util/config.c b/tools/perf/util/config.c
> index 8fe43b032e9af88a..3b6a45569ae3d2eb 100644
> --- a/tools/perf/util/config.c
> +++ b/tools/perf/util/config.c
> @@ -12,6 +12,8 @@
>  #include "config.h"
>  
>  #include <errno.h>
> +#include <limits.h>
> +#include <pthread.h>
>  #include <stdbool.h>
>  #include <stdio.h>
>  #include <stdlib.h>
> @@ -370,10 +372,15 @@ static int perf_parse_long(const char *value, long *ret)
>  
>  static void bad_config(const char *name)
>  {
> -	if (config_file_name)
> -		pr_warning("bad config value for '%s' in %s, ignoring...\n", name, config_file_name);
> -	else
> -		pr_warning("bad config value for '%s', ignoring...\n", name);
> +	/*
> +	 * No file name here: config_file_name is set and cleared by
> +	 * whichever thread has a config file in flight, under config_mutex,
> +	 * while this runs on the thread dispatching the collected values.
> +	 * Reading it would race with that thread: NULL between the check
> +	 * and the use, or the file it is parsing, not the one the bad value
> +	 * came from.
> +	 */
> +	pr_warning("bad config value for '%s', ignoring...\n", name);
>  }
>  
>  int perf_config_u64(u64 *dest, const char *name, const char *value)
> @@ -549,11 +556,26 @@ int perf_default_config(const char *var, const char *value,
>  	return 0;
>  }
>  
> +/*
> + * Serialize config file access: parsing and rewriting share the static
> + * parser state and can run on more than one thread, the debuginfod
> + * fetch writing core.debuginfod=false while perf top reads it.
> + */
> +static pthread_mutex_t config_mutex = PTHREAD_MUTEX_INITIALIZER;

Nit: perf has its own mutex type and helper functions.  But I suspect it
doesn't have the static initializer yet.

Thanks,
Namhyung


> +
> +/*
> + * perf_config__set_variable() is a read-modify-write of a config file,
> + * and config_mutex serializes just each of its steps: take it for the
> + * whole update so callers can't write over each other's changes.
> + */
> +static pthread_mutex_t config_update_mutex = PTHREAD_MUTEX_INITIALIZER;
> +
>  static int perf_config_from_file(config_fn_t fn, const char *filename, void *data)
>  {
>  	int ret;
>  	FILE *f = fopen(filename, "r");
>  
> +	pthread_mutex_lock(&config_mutex);
>  	ret = -1;
>  	if (f) {
>  		config_file = f;
> @@ -564,15 +586,34 @@ static int perf_config_from_file(config_fn_t fn, const char *filename, void *dat
>  		fclose(f);
>  		config_file_name = NULL;
>  	}
> +	pthread_mutex_unlock(&config_mutex);
>  	return ret;
>  }
>  
> +/*
> + * Computed once: system_path() allocates, a lazy init racing on two
> + * threads would leak all but one of the strings.
> + */
> +static const char *etc_perfconfig;
> +
> +static void perf_etc_perfconfig__init(void)
> +{
> +	etc_perfconfig = system_path(ETC_PERFCONFIG);
> +	/*
> +	 * None of the callers check for NULL, so an allocation failure
> +	 * here leaves all of them dereferencing it.  The unresolved path
> +	 * is the same string when ETC_PERFCONFIG is absolute anyway.
> +	 */
> +	if (!etc_perfconfig)
> +		etc_perfconfig = ETC_PERFCONFIG;
> +}
> +
>  const char *perf_etc_perfconfig(void)
>  {
> -	static const char *system_wide;
> -	if (!system_wide)
> -		system_wide = system_path(ETC_PERFCONFIG);
> -	return system_wide;
> +	static pthread_once_t once = PTHREAD_ONCE_INIT;
> +
> +	pthread_once(&once, perf_etc_perfconfig__init);
> +	return etc_perfconfig;
>  }
>  
>  static int perf_env_bool(const char *k, int def)
> @@ -630,19 +671,25 @@ static char *home_perfconfig(void)
>  	return NULL;
>  }
>  
> -const char *perf_home_perfconfig(void)
> -{
> -	static const char *config;
> -	static bool failed;
> +/*
> + * Computed once for the same reason as perf_etc_perfconfig() above:
> + * home_perfconfig() allocates, a lazy init racing on two threads would
> + * leak all but one of the strings, and the warnings it may print would
> + * come out more than once.
> + */
> +static const char *home_config;
>  
> -	if (failed || config)
> -		return config;
> +static void perf_home_perfconfig__init(void)
> +{
> +	home_config = home_perfconfig();
> +}
>  
> -	config = home_perfconfig();
> -	if (!config)
> -		failed = true;
> +const char *perf_home_perfconfig(void)
> +{
> +	static pthread_once_t once = PTHREAD_ONCE_INIT;
>  
> -	return config;
> +	pthread_once(&once, perf_home_perfconfig__init);
> +	return home_config;
>  }
>  
>  static struct perf_config_section *find_section(struct list_head *sections,
> @@ -783,8 +830,18 @@ static int collect_config(const char *var, const char *value,
>  int perf_config_set__collect(struct perf_config_set *set, const char *file_name,
>  			     const char *var, const char *value)
>  {
> +	int ret;
> +
> +	pthread_mutex_lock(&config_mutex);
>  	config_file_name = file_name;
> -	return collect_config(var, value, set);
> +	ret = collect_config(var, value, set);
> +	/*
> +	 * Don't leave the static parser state pointing at the caller's
> +	 * buffer.
> +	 */
> +	config_file_name = NULL;
> +	pthread_mutex_unlock(&config_mutex);
> +	return ret;
>  }
>  
>  static int perf_config_set__init(struct perf_config_set *set)
> @@ -831,6 +888,16 @@ struct perf_config_set *perf_config_set__load_file(const char *file)
>  	return set;
>  }
>  
> +/*
> + * The global config_set is built lazily: two threads in perf_config() at
> + * once would both build one and leak all but the last, and config_set must
> + * not be read while another thread swaps it, so take it one at a time.  Not
> + * with config_mutex: building the set parses the config files, which takes
> + * that one.
> + */
> +static pthread_mutex_t config_set_mutex = PTHREAD_MUTEX_INITIALIZER;
> +
> +/* Called with config_set_mutex held. */
>  static int perf_config__init(void)
>  {
>  	if (config_set == NULL)
> @@ -871,16 +938,126 @@ int perf_config_set(struct perf_config_set *set,
>  
>  int perf_config(config_fn_t fn, void *data)
>  {
> -	if (config_set == NULL && perf_config__init())
> +	struct perf_config_set *set;
> +
> +	/*
> +	 * The mutex is taken just for the lazy init and for the pointer:
> +	 * the dispatch below only reads the set, so a callback that called
> +	 * perf_config() again would find it unlocked, and only
> +	 * perf_config__exit() replaces the set.
> +	 */
> +	pthread_mutex_lock(&config_set_mutex);
> +	if (perf_config__init()) {
> +		pthread_mutex_unlock(&config_set_mutex);
>  		return -1;
> +	}
> +	set = config_set;
> +	pthread_mutex_unlock(&config_set_mutex);
>  
> -	return perf_config_set(config_set, fn, data);
> +	return perf_config_set(set, fn, data);
>  }
>  
>  void perf_config__exit(void)
>  {
> +	pthread_mutex_lock(&config_set_mutex);
>  	perf_config_set__delete(config_set);
>  	config_set = NULL;
> +	pthread_mutex_unlock(&config_set_mutex);
> +}
> +
> +int perf_config_set__write(struct perf_config_set *set,
> +			   const char *file_name, bool system_config)
> +{
> +	struct perf_config_section *section = NULL;
> +	struct perf_config_item *item = NULL;
> +	int ret = 0;
> +	FILE *fp;
> +
> +	pthread_mutex_lock(&config_mutex);
> +	fp = fopen(file_name, "w");
> +	if (!fp) {
> +		pthread_mutex_unlock(&config_mutex);
> +		return -1;
> +	}
> +
> +	if (fprintf(fp, "# this file is auto-generated.\n") < 0)
> +		ret = -1;
> +
> +	/* overwrite configvariables */
> +	perf_config_sections__for_each_entry(&set->sections, section) {
> +		if (!system_config && section->from_system_config)
> +			continue;
> +		if (fprintf(fp, "[%s]\n", section->name) < 0)
> +			ret = -1;
> +
> +		perf_config_items__for_each_entry(&section->items, item) {
> +			if (!system_config && item->from_system_config)
> +				continue;
> +			if (item->value &&
> +			    fprintf(fp, "\t%s = %s\n", item->name, item->value) < 0)
> +				ret = -1;
> +		}
> +	}
> +	if (fclose(fp) != 0)
> +		ret = -1;
> +	pthread_mutex_unlock(&config_mutex);
> +
> +	return ret;
> +}
> +
> +/*
> + * Set @var=@value in the configuration file perf is using: ~/.perfconfig
> + * or the file named by PERF_CONFIG, which makes perf read only that
> + * file.  The rewrite is the same 'perf config' does, comments are not
> + * preserved.
> + */
> +int perf_config__set_variable(const char *var, const char *value)
> +{
> +	const char *config_filename;
> +	bool system_config;
> +	struct perf_config_set *set = NULL;
> +	int ret = -1;
> +
> +	pthread_mutex_lock(&config_update_mutex);
> +
> +	/*
> +	 * Not on the stack: the parser publishes this buffer as
> +	 * config_file_name, which another thread may still be reading.  It is
> +	 * shared by every caller, so it is formatted under the lock above.
> +	 */
> +	{
> +		static char path[PATH_MAX];
> +		char *user_config = mkpath(path, sizeof(path), "%s/.perfconfig", getenv("HOME"));
> +
> +		config_filename = config_exclusive_filename ?: user_config;
> +	}
> +
> +	/*
> +	 * When rewriting the system wide file all entries are marked as coming
> +	 * from it and must be kept, or it would be truncated down to its
> +	 * header.
> +	 */
> +	system_config = strcmp(config_filename, perf_etc_perfconfig()) == 0;
> +
> +	set = perf_config_set__new();
> +	if (!set)
> +		goto out_err;
> +
> +	if (perf_config_set__collect(set, config_filename, var, value) < 0) {
> +		pr_err("Failed to add '%s=%s'\n", var, value);
> +		goto out_err;
> +	}
> +
> +	if (perf_config_set__write(set, config_filename, system_config) < 0) {
> +		pr_err("Failed to set the configs on %s\n", config_filename);
> +		goto out_err;
> +	}
> +
> +	ret = 0;
> +out_err:
> +	perf_config_set__delete(set);
> +	pthread_mutex_unlock(&config_update_mutex);
> +	return ret;
>  }
>  
>  static void perf_config_item__delete(struct perf_config_item *item)
> diff --git a/tools/perf/util/config.h b/tools/perf/util/config.h
> index 987b47cf54c350ba..9098f8a045850c97 100644
> --- a/tools/perf/util/config.h
> +++ b/tools/perf/util/config.h
> @@ -33,6 +33,8 @@ int perf_config_scan(const char *name, const char *fmt, ...) __scanf(2, 3);
>  const char *perf_config_get(const char *name);
>  int perf_config_set(struct perf_config_set *set,
>  		    config_fn_t fn, void *data);
> +int perf_config_set__write(struct perf_config_set *set,
> +			   const char *file_name, bool system_config);
>  int perf_config_int(int *dest, const char *, const char *);
>  int perf_config_u8(u8 *dest, const char *name, const char *value);
>  int perf_config_u64(u64 *dest, const char *, const char *);
> -- 
> 2.55.0
> 

^ permalink raw reply	[flat|nested] 15+ messages in thread

* Re: [PATCH v6 4/5] perf scripts: Add perf-stuck, to tell where a running perf is stuck
  2026-09-30 21:37 ` [PATCH v6 4/5] perf scripts: Add perf-stuck, to tell where a running perf is stuck Arnaldo Carvalho de Melo
@ 2026-10-01  7:24   ` Namhyung Kim
  2026-10-01  9:19     ` Arnaldo Carvalho de Melo
  0 siblings, 1 reply; 15+ messages in thread
From: Namhyung Kim @ 2026-10-01  7:24 UTC (permalink / raw)
  To: Arnaldo Carvalho de Melo
  Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
	Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
	Arnaldo Carvalho de Melo

On Wed, Sep 30, 2026 at 11:37:15PM +0200, Arnaldo Carvalho de Melo wrote:
> From: Arnaldo Carvalho de Melo <acme@redhat.com>
> 
> A perf that takes forever is hard to tell apart from one stuck in a
> loop, and there is no way to see where without attaching gdb.
> perf-stuck.sh samples a running process' /proc entries and its progress
> line at a fixed interval, and with -g runs gdb (perf-stuck.gdb, adding
> the perf-die-chain command) when no progress is made across two
> samples, printing the DIE chain a DWARF type chase is stuck in.
> 
> It is a prototype: the plan is to turn it into a first class 'perf
> stuck' command.  The process name is resolved with pgrep among the
> caller's own processes only, as root an unscoped one would attach gdb
> to the first process of any user with a matching name.
> 
> -i and -n are now validated instead of failing later inside the
> sampling loop, -x now requires -g and gets the same readability check
> as the default gdb command file, and gdb_done is cleared whenever
> progress resumes, so -g can fire again on a later stall instead of at
> most once per run.
> 
> Assisted-by: LLM
> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
> ---
>  tools/perf/scripts/perf-stuck.gdb | 129 ++++++++++++++++++++
>  tools/perf/scripts/perf-stuck.sh  | 193 ++++++++++++++++++++++++++++++
>  2 files changed, 322 insertions(+)
>  create mode 100644 tools/perf/scripts/perf-stuck.gdb
>  create mode 100755 tools/perf/scripts/perf-stuck.sh
> 
> diff --git a/tools/perf/scripts/perf-stuck.gdb b/tools/perf/scripts/perf-stuck.gdb
> new file mode 100644
> index 0000000000000000..be7b9e3fcb5d6d98
> --- /dev/null
> +++ b/tools/perf/scripts/perf-stuck.gdb
> @@ -0,0 +1,129 @@
> +# SPDX-License-Identifier: GPL-2.0
> +#
> +# gdb commands for a stuck perf, used by perf-stuck.sh -g and usable directly:
> +#
> +#   gdb -p $(pgrep -x perf) -batch -x perf-stuck.gdb -ex bt
> +#
> +# PROTOTYPE: part of the perf-stuck.sh stopgap, wants to become a first
> +# class 'perf stuck' command printing these DIE chains without gdb.
> +#
> +# The commands are for the DWARF type chasers in util/dwarf-aux.c:
> +#
> +#   perf-die-chain <function> <die variable> [iterations]
> +#   perf-die-chain-all [iterations]
> +#   perf-dso
> +#
> +# For each iteration of the chasing loop they print the DIE address, the
> +# CU, its file offset, tag and name: a cycle shows as the same (addr, cu)
> +# pairs repeating, and a CU changing between iterations means the chase
> +# hops between a debug file and its dwz common file.
> +
> +set pagination off
> +set confirm off
> +set debuginfod enabled off
> +set print pretty on
> +set height 0
> +set width 0
> +
> +define perf-die-chain
> +  if $argc < 2
> +    printf "usage: perf-die-chain <function> <die variable> [iterations]\n"
> +  else
> +    frame function $arg0
> +    if $argc == 3
> +      set $perf_die_chain_n = $arg2
> +    else
> +      set $perf_die_chain_n = 10
> +    end
> +    set $perf_die_chain_head = $pc
> +    set $perf_die_chain_i = 0
> +    while $perf_die_chain_i < $perf_die_chain_n
> +      # Pointer type DIEs have no DW_AT_name, so dwarf_diename() can
> +      # return NULL: printf %s of it would error out and abort this
> +      # batch script, handle it.
> +      set $perf_die_chain_name = (char *) dwarf_diename($arg1)
> +      printf "chain[%d] die=%p addr=%p cu=%p off=0x%lx tag=%d name=", $perf_die_chain_i, $arg1, $arg1->addr, $arg1->cu, ((Dwarf_Off) dwarf_dieoffset($arg1)), ((int) dwarf_tag($arg1))

Nit: can you please break this line?

Thanks,
Namhyung


> +      if $perf_die_chain_name == 0
> +	printf "(null)\n"
> +      else
> +	printf "%s\n", $perf_die_chain_name
> +      end
> +      until *$perf_die_chain_head
> +      set $perf_die_chain_i = $perf_die_chain_i + 1
> +    end
> +  end
> +end
> +
> +document perf-die-chain
> +Print the DIE chain being walked by a DWARF type chasing loop.
> +usage: perf-die-chain <function> <die variable> [iterations]
> +  perf-die-chain die_get_pointer_type type_die
> +  perf-die-chain __die_get_real_type vr_die
> +  perf-die-chain die_get_real_type vr_die
> +end
> +
> +define perf-die-chain-all
> +  if $argc == 0
> +    set $perf_die_chain_n = 10
> +  else
> +    set $perf_die_chain_n = $arg0
> +  end
> +  if $_any_caller_is("die_get_pointer_type", 20)
> +    printf "stuck in die_get_pointer_type():\n"
> +    perf-die-chain die_get_pointer_type type_die $perf_die_chain_n
> +  else
> +    if $_any_caller_is("__die_get_real_type", 20)
> +      printf "stuck in __die_get_real_type():\n"
> +      perf-die-chain __die_get_real_type vr_die $perf_die_chain_n
> +    else
> +      if $_any_caller_is("die_get_real_type", 20)
> +        printf "stuck in die_get_real_type():\n"
> +        perf-die-chain die_get_real_type vr_die $perf_die_chain_n
> +      else
> +        printf "not in a DWARF type chaser, try: bt\n"
> +      end
> +    end
> +  end
> +end
> +
> +document perf-die-chain-all
> +Find which DWARF type chaser the process is in and print the DIE chain.
> +usage: perf-die-chain-all [iterations]
> +end
> +
> +define perf-dso
> +  if $_any_caller_is("find_data_type", 20)
> +    frame function find_data_type
> +    # REFCNT_CHECKING, implied by an ASan/LSan build, wraps struct map and
> +    # struct dso in a proxy keeping the real object in ->orig, the dso being
> +    # map->orig->dso->orig there; struct symbol is not wrapped, so sym->name
> +    # needs no ->orig.  There is no way to ask gdb which layout it is looking
> +    # at, so walk each candidate until one evaluates.  Requiring gdb's Python
> +    # here adds nothing: $_any_caller_is() is one of its functions already.
> +    python
> +import gdb
> +
> +dso = None
> +for expr in ("dloc->ms->map->dso->name",
> +             "dloc->ms->map->orig->dso->orig->name"):
> +    try:
> +        gdb.parse_and_eval(expr)
> +        dso = expr
> +        break
> +    except gdb.error:
> +        pass
> +
> +if dso is None:
> +    print("dso=(no such member in struct map)")
> +else:
> +    gdb.execute('printf "dso=%s ip=0x%lx sym=%s\\n", ' + dso +
> +                ', dloc->ip, dloc->ms->sym ? dloc->ms->sym->name : "(no symbol)"')
> +    end
> +  else
> +    printf "not in find_data_type()\n"
> +  end
> +end
> +
> +document perf-dso
> +Print the dso, ip and symbol of the data location being resolved.
> +end
> diff --git a/tools/perf/scripts/perf-stuck.sh b/tools/perf/scripts/perf-stuck.sh
> new file mode 100755
> index 0000000000000000..f9b545c107905a4e
> --- /dev/null
> +++ b/tools/perf/scripts/perf-stuck.sh
> @@ -0,0 +1,193 @@
> +#!/bin/bash
> +# SPDX-License-Identifier: GPL-2.0
> +#
> +# perf-stuck - tell a spinning perf apart from a blocked or recursing one
> +#
> +# Arnaldo Carvalho de Melo <acme@redhat.com>
> +#
> +# PROTOTYPE: wants to become a first class 'perf stuck' command, sampling
> +# a process from inside perf, with the knowledge of perf's phases and of
> +# the DWARF type chasing loops built in, instead of poking /proc and
> +# shelling out to gdb.
> +#
> +# Samples /proc/<pid> at a fixed interval and prints the CPU time used
> +# since the previous sample, the [stack] mapping start and size, and the
> +# last line of a progress log when one is given, e.g. the stderr of
> +# 'perf report --progress': burning a full interval with a constant
> +# stack is a loop, a [stack] start moving down is runaway recursion.
> +#
> +# With -g it runs gdb (perf-stuck.gdb) when no progress is made for two
> +# consecutive samples, printing the DIE chain a DWARF type chase is
> +# walking.
> +#
> +# usage: perf-stuck.sh [options] <pid|process-name>
> +
> +set -u
> +
> +usage() {
> +	cat <<-EOF
> +	usage: perf-stuck.sh [options] <pid|process-name>
> +
> +	  -i <secs>   sampling interval (default: 10)
> +	  -n <count>  stop after this many samples (default: watch till it exits)
> +	  -l <file>   progress log, its last line is printed with every sample
> +	  -g          run gdb with perf-stuck.gdb when no progress is made for
> +	              two consecutive samples, writing the output to a temp file
> +	  -x <file>   use this gdb command file instead of perf-stuck.gdb
> +	  -h          this help
> +	EOF
> +	exit "${1:-0}"
> +}
> +
> +interval=10
> +count=0
> +progress_log=
> +use_gdb=
> +gdb_cmds=
> +
> +while getopts "i:n:l:gx:h" opt; do
> +	case "$opt" in
> +	i) interval=$OPTARG ;;
> +	n) count=$OPTARG ;;
> +	l) progress_log=$OPTARG ;;
> +	g) use_gdb=1 ;;
> +	x) gdb_cmds=$OPTARG ;;
> +	h) usage 0 ;;
> +	*) usage 1 ;;
> +	esac
> +done
> +shift $((OPTIND - 1))
> +
> +[[ "$interval" =~ ^[0-9]+$ ]] && [ "$interval" -gt 0 ] ||
> +	{ echo "-i wants a positive integer, got '$interval'"; usage 1; }
> +[[ "$count" =~ ^[0-9]+$ ]] ||
> +	{ echo "-n wants a non-negative integer, got '$count'"; usage 1; }
> +[ -n "$gdb_cmds" ] && [ -z "$use_gdb" ] && { echo "-x needs -g"; usage 1; }
> +
> +[ $# -eq 1 ] || usage 1
> +
> +if [[ "$1" =~ ^[0-9]+$ ]]; then
> +	pid=$1
> +else
> +	# Resolve the name against the caller's own processes: as root,
> +	# unscoped pgrep picks the first match of any user, e.g. one planted
> +	# to get gdb attached to it, use an explicit pid to look at a perf of
> +	# another user.
> +	pid=$(pgrep -x -u "$(id -u)" -- "$1" | head -1)
> +	[ -n "$pid" ] || { echo "no process named '$1' owned by $(id -un)"; exit 1; }
> +fi
> +
> +[ -d /proc/"$pid" ] || { echo "no process $pid"; exit 1; }
> +
> +if [ -n "$use_gdb" ]; then
> +	# Fail before watching, not two samples in when -g would fire.
> +	command -v gdb > /dev/null || { echo "gdb not found, -g needs it"; exit 1; }
> +	[ -z "$gdb_cmds" ] && gdb_cmds=$(dirname "$0")/perf-stuck.gdb
> +	[ -r "$gdb_cmds" ] || { echo "cannot read $gdb_cmds"; exit 1; }
> +fi
> +
> +hz=$(getconf CLK_TCK)
> +psz=$(getconf PAGESIZE)
> +prev_cpu=
> +prev_stack=
> +prev_progress=
> +stuck=0
> +gdb_done=
> +nsample=0
> +
> +# The command line is whatever the process was started with, so drop the
> +# control characters from it: a process started with escape sequences in
> +# its arguments, e.g. one replaying a log line, would otherwise get them
> +# replayed on the terminal of whoever runs this.
> +cmdline=$(tr '\0' ' ' < /proc/"$pid"/cmdline | tr -d '[:cntrl:]')
> +
> +echo "watching $pid ($cmdline) every ${interval}s"
> +
> +while :; do
> +	if [ ! -d /proc/"$pid" ]; then
> +		echo "$(date +%T) process gone"
> +		break
> +	fi
> +
> +	# Field 2, the command name, is in parentheses and can contain
> +	# spaces, so drop it together with the pid before splitting so the
> +	# fields line up.  %d keeps the CPU time out of scientific
> +	# notation, that bash arithmetic can't parse past six digits.
> +	if ! stat_line=$(awk '{ sub(/^[^ ]+ \(.*\) /, "");
> +			       printf "%s %d %d\n", $1, $12 + $13, $22 }' \
> +			 /proc/"$pid"/stat 2>/dev/null); then
> +		echo "$(date +%T) process gone"
> +		break
> +	fi
> +
> +	# The process can be gone between the check above and this read, in
> +	# which case there is nothing to report: 'set -u' would otherwise
> +	# turn the unbound fields into an aborted script.
> +	if [ -z "$stat_line" ]; then
> +		echo "$(date +%T) process gone"
> +		break
> +	fi
> +
> +	stat=($stat_line)
> +	state=${stat[0]}
> +	cpu=${stat[1]}
> +	# field 24 of /proc/<pid>/stat, the resident set size in pages
> +	rss=$(( stat[2] * psz / 1024 ))
> +
> +	stack=$(awk '/\[stack\]/{print $1; exit}' /proc/"$pid"/maps 2>/dev/null)
> +	if [ -n "$stack" ]; then
> +		stack_start=0x${stack%-*}
> +		stack_size=$(( 0x${stack#*-} - stack_start ))
> +		stack_txt="$stack size=$((stack_size / 1024))kB"
> +	else
> +		stack_start=
> +		stack_txt="-"
> +	fi
> +
> +	progress=
> +	[ -n "$progress_log" ] && [ -s "$progress_log" ] && progress=$(tail -1 -- "$progress_log")
> +
> +	if [ -n "$prev_cpu" ]; then
> +		cpu_delta=$(( cpu - prev_cpu ))
> +		# No progress log, or one with nothing written to it yet, leaves
> +		# no progress to look at, so count the samples that show none:
> +		# -g then looks at where the process is after two intervals.
> +		# Otherwise count the repeats of the same last line.
> +		if [ -z "$progress" ] || [ "$progress" = "$prev_progress" ]; then
> +			stuck=$((stuck + 1))
> +		else
> +			stuck=0
> +			gdb_done=
> +		fi
> +		stuck_txt="stuck=${stuck}"
> +		[ "$stack_start" != "$prev_stack" ] && stuck_txt="$stuck_txt STACK"
> +	else
> +		cpu_delta=0
> +		stuck_txt=""
> +	fi
> +
> +	printf '%s state=%s cpu=+%d (%d.%02ds) rss=%dkB stack=%s %s %s\n' \
> +	       "$(date +%T)" "$state" "$cpu_delta" \
> +	       $(( cpu_delta / hz )) $(( (cpu_delta % hz) * 100 / hz )) \
> +	       "$rss" "$stack_txt" "$stuck_txt" "${progress:-(no progress log)}"
> +
> +	if [ -n "$use_gdb" ] && [ -z "$gdb_done" ] && [ "$stuck" -ge 2 ]; then
> +		gdb_log=$(mktemp /tmp/perf-stuck-gdb.XXXXXX)
> +		# The gdb script makes inferior calls, and a call into a perf
> +		# wedged in a loop never returns, so bound the run; SIGINT
> +		# releases the inferior instead of leaving perf stopped.
> +		timeout --signal=INT 30 gdb -p "$pid" -batch -x "$gdb_cmds" -ex bt \
> +		    -ex 'perf-die-chain-all' -ex perf-dso -ex detach > "$gdb_log" 2>&1
> +		gdb_done=1
> +		echo "... gdb output of $pid in $gdb_log"
> +	fi
> +
> +	prev_cpu=$cpu
> +	prev_stack=$stack_start
> +	prev_progress=$progress
> +
> +	nsample=$((nsample + 1))
> +	[ "$count" -gt 0 ] && [ "$nsample" -ge "$count" ] && break
> +
> +	sleep "$interval"
> +done
> -- 
> 2.55.0
> 

^ permalink raw reply	[flat|nested] 15+ messages in thread

* Re: [PATCH v6 5/5] perf test: Add false_sharing workload exhibiting cross-CPU false sharing
  2026-09-30 21:37 ` [PATCH v6 5/5] perf test: Add false_sharing workload exhibiting cross-CPU false sharing Arnaldo Carvalho de Melo
@ 2026-10-01  7:28   ` Namhyung Kim
  2026-10-01  9:18     ` Arnaldo Carvalho de Melo
  0 siblings, 1 reply; 15+ messages in thread
From: Namhyung Kim @ 2026-10-01  7:28 UTC (permalink / raw)
  To: Arnaldo Carvalho de Melo
  Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
	Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
	Arnaldo Carvalho de Melo

On Wed, Sep 30, 2026 at 11:37:16PM +0200, Arnaldo Carvalho de Melo wrote:
> From: Arnaldo Carvalho de Melo <acme@redhat.com>
> 
> Add a 'perf test -w false_sharing' workload that hammers one shared
> struct from several CPUs, shaped as a TCP connection: a read-mostly
> identity (five-tuple) shares a cacheline with per-packet rx counters
> (the false-sharing line), a second line has packet-path private tx and
> congestion control counters, and a third the connection config.
> 
> The packet path runs in the main thread and up to four lookup threads
> sum the five-tuple and pull the config, reading one volatile shared
> instance directly so the accesses are PC-relative and resolvable by the
> data type profiler.

It'd be great if you can share an output of data type profiling with
cacheline info.  Probably like below?

  $ perf mem record -- perf test -w false_sharing
 
  $ perf report -s type,typecln -H --group --stdio

Thanks,
Namhyung

> 
> Assisted-by: LLM
> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
> ---
>  tools/perf/tests/builtin-test.c               |   1 +
>  tools/perf/tests/shell/data_type_profiling.sh |   9 +-
>  tools/perf/tests/tests.h                      |   1 +
>  tools/perf/tests/workloads/Build              |   2 +
>  tools/perf/tests/workloads/false_sharing.c    | 251 ++++++++++++++++++
>  5 files changed, 262 insertions(+), 2 deletions(-)
>  create mode 100644 tools/perf/tests/workloads/false_sharing.c
> 
> diff --git a/tools/perf/tests/builtin-test.c b/tools/perf/tests/builtin-test.c
> index d2f594921e25bda9..98134b1c74cf80f4 100644
> --- a/tools/perf/tests/builtin-test.c
> +++ b/tools/perf/tests/builtin-test.c
> @@ -175,6 +175,7 @@ static struct test_workload *workloads[] = {
>  	&workload__context_switch_loop,
>  	&workload__deterministic,
>  	&workload__callchain,
> +	&workload__false_sharing,
>  
>  #ifdef HAVE_RUST_SUPPORT
>  	&workload__code_with_type,
> diff --git a/tools/perf/tests/shell/data_type_profiling.sh b/tools/perf/tests/shell/data_type_profiling.sh
> index a916c410274aa888..57203e0e5f857540 100755
> --- a/tools/perf/tests/shell/data_type_profiling.sh
> +++ b/tools/perf/tests/shell/data_type_profiling.sh
> @@ -8,8 +8,8 @@ set -e
>  # data type profiling manifestation
>  
>  # Values in testtypes and testprogs should match
> -testtypes=("# data-type: struct Buf" "# data-type: struct buf")
> -testprogs=("perf test -w code_with_type" "perf test -w datasym")
> +testtypes=("# data-type: struct Buf" "# data-type: struct buf" "# data-type: struct net_conn")
> +testprogs=("perf test -w code_with_type" "perf test -w datasym" "perf test -w false_sharing")
>  
>  err=0
>  perfdata=$(mktemp /tmp/__perf_test.perf.data.XXXXX)
> @@ -59,6 +59,9 @@ test_basic_annotate() {
>  
>      "xC")
>      index=1 ;;
> +
> +    "xFS")
> +    index=2 ;;
>    esac
>  
>    # Under 'set -e' a bare failing command aborts the script through the EXIT
> @@ -114,6 +117,8 @@ test_basic_annotate Basic Rust
>  test_basic_annotate Pipe Rust
>  test_basic_annotate Basic C
>  test_basic_annotate Pipe C
> +test_basic_annotate Basic FS
> +test_basic_annotate Pipe FS
>  
>  cleanup
>  exit $err
> diff --git a/tools/perf/tests/tests.h b/tools/perf/tests/tests.h
> index 9c96f33483d14356..f379620069d34f3a 100644
> --- a/tools/perf/tests/tests.h
> +++ b/tools/perf/tests/tests.h
> @@ -250,6 +250,7 @@ DECLARE_WORKLOAD(jitdump);
>  DECLARE_WORKLOAD(context_switch_loop);
>  DECLARE_WORKLOAD(deterministic);
>  DECLARE_WORKLOAD(callchain);
> +DECLARE_WORKLOAD(false_sharing);
>  
>  #ifdef HAVE_RUST_SUPPORT
>  DECLARE_WORKLOAD(code_with_type);
> diff --git a/tools/perf/tests/workloads/Build b/tools/perf/tests/workloads/Build
> index 048e371eb63e3164..18fceddccb56c013 100644
> --- a/tools/perf/tests/workloads/Build
> +++ b/tools/perf/tests/workloads/Build
> @@ -14,6 +14,7 @@ perf-test-y += jitdump.o
>  perf-test-y += context_switch_loop.o
>  perf-test-y += deterministic.o
>  perf-test-y += callchain.o
> +perf-test-y += false_sharing.o
>  
>  ifeq ($(CONFIG_RUST_SUPPORT),y)
>      perf-test-y += code_with_type.o
> @@ -29,3 +30,4 @@ CFLAGS_inlineloop.o       = -g -O2
>  CFLAGS_deterministic.o    = -g -O0 -fno-inline -U_FORTIFY_SOURCE
>  CFLAGS_named_threads.o    = -g -O0 -fno-inline -U_FORTIFY_SOURCE
>  CFLAGS_callchain.o        = -g -O0 -fno-inline -U_FORTIFY_SOURCE
> +CFLAGS_false_sharing.o    = -g -O0 -fno-inline -U_FORTIFY_SOURCE
> diff --git a/tools/perf/tests/workloads/false_sharing.c b/tools/perf/tests/workloads/false_sharing.c
> new file mode 100644
> index 0000000000000000..e949bb36646a97a5
> --- /dev/null
> +++ b/tools/perf/tests/workloads/false_sharing.c
> @@ -0,0 +1,251 @@
> +// SPDX-License-Identifier: GPL-2.0
> +/*
> + * False-sharing demo for data type profiling, shaped as a TCP
> + * connection: a read-mostly identity shares a cacheline with per-packet
> + * rx counters (the false-sharing line), a second line has packet-path
> + * private tx and congestion control counters, and a third the connection
> + * config.
> + *
> + * 'perf mem record' of this workload followed by 'perf report -s type'
> + * (see tests/shell/data_type_profiling.sh) resolves the accesses to
> + * struct net_conn members, showing the rx counters and the read-mostly
> + * identity sharing cacheline 0.
> + */
> +#include <pthread.h>
> +#include <sched.h>
> +#include <stdint.h>
> +#include <stdlib.h>
> +#include <stdio.h>
> +#include <signal.h>
> +#include <unistd.h>
> +#include <linux/compiler.h>
> +#include "../tests.h"
> +
> +struct net_conn {
> +	/* cacheline 0: identity (read-mostly) + rx counters (per packet) */
> +	uint32_t	saddr;		/*   0 */
> +	uint32_t	daddr;		/*   4 */
> +	uint16_t	sport;		/*   8 */
> +	uint16_t	dport;		/*  10 */
> +	uint8_t		state;		/*  12: 1 == ESTABLISHED */
> +	uint8_t		protocol;	/*  13: 6 == TCP */
> +	uint16_t	__pad0;		/*  14 */
> +	uint64_t	bytes_rx;	/*  16: every packet */
> +	uint64_t	packets_rx;	/*  24: every packet */
> +	uint32_t	rx_queue;	/*  32: backlog depth, fluctuates */
> +	uint8_t		__pad1[24];	/*  36..59 */
> +	uint32_t	last_ack;	/*  60: written per ACK */
> +	/* cacheline 1: tx + congestion control (packet-path private) */
> +	uint64_t	bytes_tx;	/*  64: every packet */
> +	uint64_t	packets_tx;	/*  72: every packet */
> +	uint32_t	cwnd;		/*  80: on every ACK */
> +	uint32_t	ssthresh;	/*  84: on loss */
> +	uint32_t	rtt_us;		/*  88: on every ACK */
> +	uint32_t	retrans;	/*  92: on timeout */
> +	uint32_t	__pad2[8];	/*  96..127 */
> +	/* cacheline 2: config, set at setup, read by everybody */
> +	uint16_t	mss;		/* 128 */
> +	uint8_t		snd_wscale;	/* 130 */
> +	uint8_t		rcv_wscale;	/* 131 */
> +	uint32_t	keepalive_int;	/* 132 */
> +	uint32_t	mark;		/* 136: firewall mark */
> +	uint32_t	priority;	/* 140: traffic class */
> +	uint32_t	__pad3[12];	/* 144..191 */
> +} __attribute__((aligned(64)));
> +
> +/* Volatile so every iteration really loads and stores. */
> +static volatile struct net_conn conn;
> +/* Keeps the reader checksums alive after the threads join. */
> +static volatile unsigned long fs_sink;
> +
> +static volatile sig_atomic_t done;
> +
> +/*
> + * One cacheline each: sum before cpu, or the implicit padding after cpu
> + * pushes the struct past 64 bytes, and aligning it to a cacheline then
> + * rounds it up to 128.
> + */
> +struct fs_reader {
> +	pthread_t	thread;
> +	unsigned long	sum;
> +	int		cpu;
> +	char		__pad[64 - sizeof(pthread_t) - sizeof(unsigned long) - sizeof(int)];
> +} __attribute__((aligned(64)));
> +
> +static void sighandler(int sig __maybe_unused)
> +{
> +	done = 1;
> +}
> +
> +static void pin_to_cpu(int cpu)
> +{
> +	cpu_set_t set;
> +
> +	/* There may be no second CPU in a restricted cpuset. */
> +	if (cpu < 0)
> +		return;
> +
> +	CPU_ZERO(&set);
> +	CPU_SET(cpu, &set);
> +	/* Best effort: in a restricted cpuset this fails and the thread runs unpinned. */
> +	pthread_setaffinity_np(pthread_self(), sizeof(set), &set);
> +}
> +
> +/*
> + * Connection lookup, as a load balancer or 'ss' scrape would do it: the
> + * reads go straight to the global, this file is built -O0 and a local
> + * pointer would be reloaded from a stack slot, a form the data type
> + * resolver does not track.
> + */
> +static void *reader_fn(void *arg)
> +{
> +	struct fs_reader *r = arg;
> +	unsigned long sum = 0;
> +
> +	pthread_setname_np(pthread_self(), "fs-reader");
> +	pin_to_cpu(r->cpu);
> +
> +	while (!done) {
> +		sum += conn.saddr + conn.daddr + conn.sport + conn.dport +
> +		       conn.state + conn.protocol;
> +		sum += conn.mss + conn.snd_wscale + conn.rcv_wscale +
> +		       conn.keepalive_int + conn.mark + conn.priority;
> +	}
> +	r->sum = sum;
> +	return NULL;
> +}
> +
> +static int false_sharing(int argc, const char **argv)
> +{
> +	double sec = 2.0;
> +	int nreaders = 0, nr_allowed = 0, err = 1;
> +	int *allowed = NULL, nallowed = 0;
> +	cpu_set_t set;
> +	int nr_mask_bits = sizeof(set) * 8 < CPU_SETSIZE ? sizeof(set) * 8 : CPU_SETSIZE;
> +	struct fs_reader *readers = NULL;
> +	int i, writer_cpu;
> +	unsigned long n = 0;
> +
> +	pthread_setname_np(pthread_self(), "fs-writer");
> +	if (argc > 0)
> +		sec = atof(argv[0]);
> +	if (!(sec > 0.0)) {
> +		fprintf(stderr, "Error: seconds (%f) must be > 0\n", sec);
> +		return 1;
> +	}
> +	if (argc > 1)
> +		nreaders = atoi(argv[1]);
> +
> +	/*
> +	 * A connection that just got established: identity and config fixed
> +	 * from here on, counters at zero.
> +	 */
> +	conn.saddr = 0x0a000001;	/* 10.0.0.1 */
> +	conn.daddr = 0x0a000002;	/* 10.0.0.2 */
> +	conn.sport = 54321;
> +	conn.dport = 443;
> +	conn.state = 1;			/* ESTABLISHED */
> +	conn.protocol = 6;		/* TCP */
> +	conn.mss = 1448;
> +	conn.snd_wscale = 7;
> +	conn.rcv_wscale = 7;
> +	conn.keepalive_int = 7200;
> +	conn.cwnd = 10;
> +	conn.ssthresh = 65535;
> +	conn.rtt_us = 50;
> +
> +	/*
> +	 * Pin against the allowed set, restricted cpusets still spread the threads.
> +	 * The whole mask is looked at, not the CPU count: the count can be lower
> +	 * than the highest ID in it, as when a cpuset allows only high numbered
> +	 * CPUs, and then no allowed CPU would be found at all.
> +	 */
> +	if (sched_getaffinity(0, sizeof(set), &set) == 0) {
> +		for (i = 0; i < nr_mask_bits; i++) {
> +			if (!CPU_ISSET(i, &set))
> +				continue;
> +			nr_allowed++;
> +		}
> +		allowed = malloc(nr_allowed * sizeof(int));
> +		if (allowed == NULL) {
> +			fprintf(stderr, "Error: malloc failed for CPU list\n");
> +			return 1;
> +		}
> +		for (i = 0; i < nr_mask_bits; i++) {
> +			if (CPU_ISSET(i, &set))
> +				allowed[nallowed++] = i;
> +		}
> +	}
> +	if (nreaders <= 0) {
> +		/* By default leave one CPU for the packet path, up to 4 readers. */
> +		nreaders = nallowed > 1 ? nallowed - 1 : 1;
> +		if (nreaders > 4)
> +			nreaders = 4;
> +	}
> +
> +	signal(SIGINT, sighandler);
> +	signal(SIGALRM, sighandler);
> +
> +	readers = calloc(nreaders, sizeof(*readers));
> +	if (readers == NULL) {
> +		fprintf(stderr, "Error: calloc failed for %d readers\n", nreaders);
> +		goto out;
> +	}
> +	for (i = 0; i < nreaders; i++) {
> +		int cpu = nallowed > 1 ? allowed[(i + 1) % nallowed] : -1;
> +
> +		readers[i].cpu = cpu;
> +		if (pthread_create(&readers[i].thread, NULL, reader_fn, &readers[i])) {
> +			fprintf(stderr, "Error: failed to create reader %d\n", i);
> +			done = 1; // Ensure started threads terminate.
> +			nreaders = i;
> +			goto out_join;
> +		}
> +	}
> +	writer_cpu = nallowed > 0 ? allowed[0] : -1;
> +	if (nallowed == 1)
> +		fprintf(stderr, "Warning: single CPU allowed, no cross-CPU traffic expected\n");
> +	if (writer_cpu >= 0)
> +		pin_to_cpu(writer_cpu);
> +
> +	/*
> +	 * The packet path: receive, acknowledge, transmit, repeat; every 64th
> +	 * packet simulates a loss so ssthresh/retrans get sampled too.
> +	 */
> +	if (sec < 1.0) {
> +		useconds_t usecs = (useconds_t)(sec * 1000000.0);
> +
> +		ualarm(usecs > 0 ? usecs : 1, 0);
> +	} else
> +		alarm((unsigned int)sec);
> +	while (!done) {
> +		conn.bytes_rx += conn.mss;
> +		conn.packets_rx++;
> +		conn.rx_queue = (uint32_t)(n & 0x3f);
> +		conn.last_ack = (uint32_t)n;
> +		conn.bytes_tx += conn.mss;
> +		conn.packets_tx++;
> +		conn.cwnd = 10 + (n & 15);
> +		conn.rtt_us = 50 + (n & 7);
> +		if ((n & 63) == 0) {
> +			conn.ssthresh = conn.cwnd / 2;
> +			conn.retrans++;
> +		}
> +		n++;
> +	}
> +	err = 0;
> +out_join:
> +	for (i = 0; i < nreaders; i++) {
> +		if (readers[i].thread) {
> +			pthread_join(readers[i].thread, /*retval=*/NULL);
> +			fs_sink += readers[i].sum;
> +		}
> +	}
> +	fs_sink += (unsigned long)(conn.bytes_rx + conn.bytes_tx + n);
> +	free(readers);
> +out:
> +	free(allowed);
> +	return err;
> +}
> +
> +DEFINE_WORKLOAD(false_sharing);
> -- 
> 2.55.0
> 

^ permalink raw reply	[flat|nested] 15+ messages in thread

* Re: [PATCH v6 5/5] perf test: Add false_sharing workload exhibiting cross-CPU false sharing
  2026-10-01  7:28   ` Namhyung Kim
@ 2026-10-01  9:18     ` Arnaldo Carvalho de Melo
  0 siblings, 0 replies; 15+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-10-01  9:18 UTC (permalink / raw)
  To: Namhyung Kim
  Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
	Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
	Arnaldo Carvalho de Melo

On Thu, Oct 01, 2026 at 12:28:13AM -0700, Namhyung Kim wrote:
> On Wed, Sep 30, 2026 at 11:37:16PM +0200, Arnaldo Carvalho de Melo wrote:
> > Add a 'perf test -w false_sharing' workload that hammers one shared
> > struct from several CPUs, shaped as a TCP connection: a read-mostly
> > identity (five-tuple) shares a cacheline with per-packet rx counters
> > (the false-sharing line), a second line has packet-path private tx and
> > congestion control counters, and a third the connection config.

> > The packet path runs in the main thread and up to four lookup threads
> > sum the five-tuple and pull the config, reading one volatile shared
> > instance directly so the accesses are PC-relative and resolvable by the
> > data type profiler.
> 
> It'd be great if you can share an output of data type profiling with
> cacheline info.  Probably like below?
> 
>   $ perf mem record -- perf test -w false_sharing
>  
>   $ perf report -s type,typecln -H --group --stdio

Sure, I should've added it there as I usually do :-\

Here it is, will add to the cset message as well:

root@x2:~# perf mem record -- perf test -r5 -w false_sharing
[ perf record: Woken up 19 times to write data ]
[ perf record: Captured and wrote 5.612 MB perf.data (72454 samples) ]
root@x2:~#
root@x2:~# perf report -s type,typecln -H --group --stdio
# Total Lost Samples: 0
#
# Samples: 72K of events 'cpu/mem-loads,ldlat=30/P, cpu/mem-stores/P'
# Event count (approx.): 1261814616
#
#            Overhead  Data Type / Data Type Cacheline
# ...................  ...............................
#
    88.99%  35.88%     struct net_conn
       56.43%  17.50%     struct net_conn: cache-line 0
       32.55%   0.00%     struct net_conn: cache-line 2
        0.01%  18.38%     struct net_conn: cache-line 1
    10.08%  63.03%     (unknown)
       10.08%  63.03%     (unknown): cache-line 0
     0.82%   0.13%     int
        0.82%   0.13%     int: cache-line 0
     0.01%   0.02%     struct folio
        0.01%   0.02%     struct folio: cache-line 0
     0.01%   0.03%     Elf64_Addr
        0.01%   0.03%     Elf64_Addr: cache-line 0
     0.01%   0.03%     struct sched_entity
        0.01%   0.03%     struct sched_entity: cache-line 1
        0.00%   0.00%     struct sched_entity: cache-line 4
        0.00%   0.00%     struct sched_entity: cache-line 2
        0.00%   0.00%     struct sched_entity: cache-line 3
        0.00%   0.00%     struct sched_entity: cache-line 0
     0.01%   0.00%     struct task_group
        0.01%   0.00%     struct task_group: cache-line 4
        0.00%   0.00%     struct task_group: cache-line 5
     0.00%   0.00%     struct css_rstat_cpu
        0.00%   0.00%     struct css_rstat_cpu: cache-line 0
root@x2:~#
root@x2:~# perf report -s type,typecln,typeoff -H --group --stdio
# Total Lost Samples: 0
#
# Samples: 72K of events 'cpu/mem-loads,ldlat=30/P, cpu/mem-stores/P'
# Event count (approx.): 1261814616
#
#               Overhead  Data Type / Data Type Cacheline / Data Type Offset
# ......................  ..................................................
#
    88.99%  35.88%        struct net_conn
       56.43%  17.50%        struct net_conn: cache-line 0
           9.81%   0.00%        struct net_conn +0xa (dport)
           9.54%   0.00%        struct net_conn +0x8 (sport)
           9.53%   0.00%        struct net_conn +0xc (state)
           9.48%   0.00%        struct net_conn +0xd (protocol)
           9.27%   0.00%        struct net_conn +0x4 (daddr)
           8.79%   0.00%        struct net_conn +0 (saddr)
           0.00%  12.08%        struct net_conn +0x10 (bytes_rx)
           0.00%   1.71%        struct net_conn +0x18 (packets_rx)
           0.00%   2.80%        struct net_conn +0x20 (rx_queue)
           0.00%   0.91%        struct net_conn +0x3c (last_ack)
       32.55%   0.00%        struct net_conn: cache-line 2
           5.70%   0.00%        struct net_conn +0x83 (rcv_wscale)
           5.46%   0.00%        struct net_conn +0x84 (keepalive_int)
           5.43%   0.00%        struct net_conn +0x82 (snd_wscale)
           5.39%   0.00%        struct net_conn +0x80 (mss)
           5.34%   0.00%        struct net_conn +0x88 (mark)
           5.23%   0.00%        struct net_conn +0x8c (priority)
        0.01%  18.38%        struct net_conn: cache-line 1
           0.01%   0.05%        struct net_conn +0x5c (retrans)
           0.00%   4.83%        struct net_conn +0x40 (bytes_tx)
           0.00%  10.63%        struct net_conn +0x48 (packets_tx)
           0.00%   1.54%        struct net_conn +0x58 (rtt_us)
           0.00%   1.07%        struct net_conn +0x50 (cwnd)
           0.00%   0.25%        struct net_conn +0x54 (ssthresh)
    10.08%  63.03%        (unknown)
       10.08%  63.03%        (unknown): cache-line 0
          10.08%  63.03%        (unknown)
     0.82%   0.13%        int      
        0.82%   0.13%        int: cache-line 0
           0.82%   0.13%        int +0 (no field)
     0.01%   0.02%        struct folio
        0.01%   0.02%        struct folio: cache-line 0
           0.01%   0.00%        struct folio +0 (flags.f)
           0.00%   0.01%        struct folio +0x34 (_refcount.counter)
           0.00%   0.00%        struct folio +0x18 (mapping)
:

^ permalink raw reply	[flat|nested] 15+ messages in thread

* Re: [PATCH v6 4/5] perf scripts: Add perf-stuck, to tell where a running perf is stuck
  2026-10-01  7:24   ` Namhyung Kim
@ 2026-10-01  9:19     ` Arnaldo Carvalho de Melo
  0 siblings, 0 replies; 15+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-10-01  9:19 UTC (permalink / raw)
  To: Namhyung Kim
  Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
	Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
	Arnaldo Carvalho de Melo

On Thu, Oct 01, 2026 at 12:24:43AM -0700, Namhyung Kim wrote:
> On Wed, Sep 30, 2026 at 11:37:15PM +0200, Arnaldo Carvalho de Melo wrote:
> > +    set $perf_die_chain_i = 0
> > +    while $perf_die_chain_i < $perf_die_chain_n
> > +      # Pointer type DIEs have no DW_AT_name, so dwarf_diename() can
> > +      # return NULL: printf %s of it would error out and abort this
> > +      # batch script, handle it.
> > +      set $perf_die_chain_name = (char *) dwarf_diename($arg1)
> > +      printf "chain[%d] die=%p addr=%p cu=%p off=0x%lx tag=%d name=", $perf_die_chain_i, $arg1, $arg1->addr, $arg1->cu, ((Dwarf_Off) dwarf_dieoffset($arg1)), ((int) dwarf_tag($arg1))
> 
> Nit: can you please break this line?

Will do,

- Arnaldo

^ permalink raw reply	[flat|nested] 15+ messages in thread

* Re: [PATCH v6 1/5] perf config: Move perf_config__set_variable() to util/config.c
  2026-10-01  7:18   ` Namhyung Kim
@ 2026-10-01  9:19     ` Arnaldo Carvalho de Melo
  0 siblings, 0 replies; 15+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-10-01  9:19 UTC (permalink / raw)
  To: Namhyung Kim
  Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
	Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
	Arnaldo Carvalho de Melo

On Thu, Oct 01, 2026 at 12:18:55AM -0700, Namhyung Kim wrote:
> On Wed, Sep 30, 2026 at 11:37:12PM +0200, Arnaldo Carvalho de Melo wrote:
> > perf_etc_perfconfig() caches what system_path() returns, and that
> > allocates, so on failure every caller dereferenced NULL, the
> > pre-existing ones in perf_config_from_file() and in the daemon
> > included.  Pre-existing, not introduced here, but this rewrites the
> > accessor anyway, so make it total: the unresolved path is the same
> > string for the usual absolute ETC_PERFCONFIG.
> 
> I think this patch does many things.  Can we split them? 

Sure, will split it into three patches.

- Arnaldo

^ permalink raw reply	[flat|nested] 15+ messages in thread

* Re: [PATCH v6 3/5] perf report: Add --no-progress option
  2026-10-01  7:01   ` Namhyung Kim
@ 2026-10-01  9:20     ` Arnaldo Carvalho de Melo
  0 siblings, 0 replies; 15+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-10-01  9:20 UTC (permalink / raw)
  To: Namhyung Kim
  Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
	Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
	Arnaldo Carvalho de Melo

On Thu, Oct 01, 2026 at 12:01:25AM -0700, Namhyung Kim wrote:
> On Wed, Sep 30, 2026 at 11:37:14PM +0200, Arnaldo Carvalho de Melo wrote:
> > +++ b/tools/perf/builtin-report.c
> > @@ -88,6 +88,7 @@ struct report {
> >  #endif
> >  	bool			use_stdio;
> >  	bool			progress;
> > +	bool			progress_set;
> >  	bool			show_full_info;
> >  	bool			show_threads;
> >  	bool			inverted_callchain;
> > @@ -1385,8 +1386,15 @@ int cmd_report(int argc, const char **argv)
> >  		    "Use the stdio interface"),
> >  	OPT_BOOLEAN(0, "weights", &symbol_conf.annotate_weight,
> >  			"Show or hide weight columns in annotation. Default show if non-zero."),
> > -	OPT_BOOLEAN(0, "progress", &report.progress,
> > -		    "Show progress while processing the perf.data file"),
> > +	/*
> > +	 * No --no-progress option to add: parse-options provides it as
> > +	 * the negation of this one, clearing report.progress, which is
> > +	 * also how it starts out.  progress_set is what tells the hook
> > +	 * below to stop the TUI and GTK browsers as well, they show
> > +	 * progress until asked not to.
> > +	 */
 
> Nit: I think we can drop this comment.
 
> > @@ -1796,9 +1804,13 @@ int cmd_report(int argc, const char **argv)
> >  	/*
> >  	 * For the stdio case: print the percentage of the perf.data file
> >  	 * processed so far for each processing phase.  --quiet asks for no
> > -	 * messages at all, so it leaves the phases uncounted.
> > +	 * messages at all, so it leaves the phases uncounted, and
> > +	 * --no-progress, progress_set with progress cleared, turns off what
> > +	 * the TUI and GTK browsers show as well.
> >  	 */
 
> Maybe this one too..

Will drop both.

- Arnaldo

^ permalink raw reply	[flat|nested] 15+ messages in thread

* [PATCH v6 3/5] perf report: Add --no-progress option
  2026-09-30 11:24 [PATCH v6 0/5] perf tools: Add progress diagnostics and a false-sharing workload Arnaldo Carvalho de Melo
@ 2026-09-30 11:24 ` Arnaldo Carvalho de Melo
  0 siblings, 0 replies; 15+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-09-30 11:24 UTC (permalink / raw)
  To: Namhyung Kim
  Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
	Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
	Arnaldo Carvalho de Melo

From: Arnaldo Carvalho de Melo <acme@redhat.com>

Now that --progress is being added for the stdio case, wire up its
counterpart for the browsers: the TUI and GTK ones present progress of
their own and there is no way to turn it off.

Suggested-by: Namhyung Kim <namhyung@kernel.org>
Assisted-by: LLM
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
---
 tools/perf/Documentation/perf-report.txt |  5 +++++
 tools/perf/builtin-report.c              | 20 ++++++++++++++++----
 tools/perf/ui/progress.c                 |  6 ++++++
 tools/perf/ui/progress.h                 |  2 ++
 4 files changed, 29 insertions(+), 4 deletions(-)

diff --git a/tools/perf/Documentation/perf-report.txt b/tools/perf/Documentation/perf-report.txt
index a7429a30ec28f903..e145124d6c9a1097 100644
--- a/tools/perf/Documentation/perf-report.txt
+++ b/tools/perf/Documentation/perf-report.txt
@@ -43,6 +43,11 @@ OPTIONS
 	present progress information, or when --quiet is used, that asks
 	for no messages at all.
 
+--no-progress::
+	Do not show progress while processing the perf.data file.  It
+	also turns off the progress the TUI and GTK browsers present,
+	which is their own.
+
 -n::
 --show-nr-samples::
 	Show the number of samples for each symbol
diff --git a/tools/perf/builtin-report.c b/tools/perf/builtin-report.c
index 963808b561e547c3..105b2859678673ec 100644
--- a/tools/perf/builtin-report.c
+++ b/tools/perf/builtin-report.c
@@ -88,6 +88,7 @@ struct report {
 #endif
 	bool			use_stdio;
 	bool			progress;
+	bool			progress_set;
 	bool			show_full_info;
 	bool			show_threads;
 	bool			inverted_callchain;
@@ -1385,8 +1386,15 @@ int cmd_report(int argc, const char **argv)
 		    "Use the stdio interface"),
 	OPT_BOOLEAN(0, "weights", &symbol_conf.annotate_weight,
 			"Show or hide weight columns in annotation. Default show if non-zero."),
-	OPT_BOOLEAN(0, "progress", &report.progress,
-		    "Show progress while processing the perf.data file"),
+	/*
+	 * No --no-progress option to add: parse-options provides it as
+	 * the negation of this one, clearing report.progress, which is
+	 * also how it starts out.  progress_set is what tells the hook
+	 * below to stop the TUI and GTK browsers as well, they show
+	 * progress until asked not to.
+	 */
+	OPT_BOOLEAN_SET(0, "progress", &report.progress, &report.progress_set,
+			"Show progress while processing the perf.data file"),
 	OPT_BOOLEAN(0, "header", &report.header, "Show data header."),
 	OPT_BOOLEAN(0, "header-only", &report.header_only,
 		    "Show only data header."),
@@ -1796,9 +1804,13 @@ int cmd_report(int argc, const char **argv)
 	/*
 	 * For the stdio case: print the percentage of the perf.data file
 	 * processed so far for each processing phase.  --quiet asks for no
-	 * messages at all, so it leaves the phases uncounted.
+	 * messages at all, so it leaves the phases uncounted, and
+	 * --no-progress, progress_set with progress cleared, turns off what
+	 * the TUI and GTK browsers show as well.
 	 */
-	if (report.progress && !quiet && use_browser == 0)
+	if (report.progress_set && !report.progress)
+		ui_progress__noop_init();
+	else if (report.progress && !quiet && use_browser == 0)
 		stdio_progress__init();
 
 	if (report.data_type && use_browser == 1) {
diff --git a/tools/perf/ui/progress.c b/tools/perf/ui/progress.c
index 99d60223c74b2957..362680989ace606a 100644
--- a/tools/perf/ui/progress.c
+++ b/tools/perf/ui/progress.c
@@ -13,6 +13,12 @@ static struct ui_progress_ops null_progress__ops =
 
 struct ui_progress_ops *ui_progress__ops = &null_progress__ops;
 
+/* Everything counts but nothing is shown, the way it starts out. */
+void ui_progress__noop_init(void)
+{
+	ui_progress__ops = &null_progress__ops;
+}
+
 void ui_progress__update(struct ui_progress *p, u64 adv)
 {
 	u64 last = p->curr;
diff --git a/tools/perf/ui/progress.h b/tools/perf/ui/progress.h
index 03f1a8bb260ba076..e8c4f9f768aaf12b 100644
--- a/tools/perf/ui/progress.h
+++ b/tools/perf/ui/progress.h
@@ -25,6 +25,8 @@ void ui_progress__update(struct ui_progress *p, u64 adv);
 
 void stdio_progress__init(void);
 
+void ui_progress__noop_init(void);
+
 struct ui_progress_ops {
 	void (*init)(struct ui_progress *p);
 	void (*update)(struct ui_progress *p);
-- 
2.55.0


^ permalink raw reply	[flat|nested] 15+ messages in thread

end of thread, other threads:[~2026-10-01  9:21 UTC | newest]

Thread overview: 15+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-30 21:37 [PATCH v6 0/5] perf tools: Add progress diagnostics and a false-sharing workload Arnaldo Carvalho de Melo
2026-09-30 21:37 ` [PATCH v6 1/5] perf config: Move perf_config__set_variable() to util/config.c Arnaldo Carvalho de Melo
2026-10-01  7:18   ` Namhyung Kim
2026-10-01  9:19     ` Arnaldo Carvalho de Melo
2026-09-30 21:37 ` [PATCH v6 2/5] perf report: Add --progress option Arnaldo Carvalho de Melo
2026-09-30 21:37 ` [PATCH v6 3/5] perf report: Add --no-progress option Arnaldo Carvalho de Melo
2026-10-01  7:01   ` Namhyung Kim
2026-10-01  9:20     ` Arnaldo Carvalho de Melo
2026-09-30 21:37 ` [PATCH v6 4/5] perf scripts: Add perf-stuck, to tell where a running perf is stuck Arnaldo Carvalho de Melo
2026-10-01  7:24   ` Namhyung Kim
2026-10-01  9:19     ` Arnaldo Carvalho de Melo
2026-09-30 21:37 ` [PATCH v6 5/5] perf test: Add false_sharing workload exhibiting cross-CPU false sharing Arnaldo Carvalho de Melo
2026-10-01  7:28   ` Namhyung Kim
2026-10-01  9:18     ` Arnaldo Carvalho de Melo
  -- strict thread matches above, loose matches on Subject: below --
2026-09-30 11:24 [PATCH v6 0/5] perf tools: Add progress diagnostics and a false-sharing workload Arnaldo Carvalho de Melo
2026-09-30 11:24 ` [PATCH v6 3/5] perf report: Add --no-progress option Arnaldo Carvalho de Melo

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®