From: Chun-Tse Shao <ctshao@google.com>
To: acme@kernel.org, namhyung@kernel.org, irogers@google.com
Cc: peterz@infradead.org, mingo@redhat.com, mark.rutland@arm.com,
alexander.shishkin@linux.intel.com, jolsa@kernel.org,
adrian.hunter@intel.com, james.clark@linaro.org,
bwicaksono@nvidia.com, linux-perf-users@vger.kernel.org,
linux-kernel@vger.kernel.org, Chun-Tse Shao <ctshao@google.com>
Subject: [PATCH 1/2] perf jevents: Add NVIDIA Tegra410 uncore DDR metrics
Date: Mon, 28 Sep 2026 10:51:26 -0700 [thread overview]
Message-ID: <20260928175127.1032535-2-ctshao@google.com> (raw)
In-Reply-To: <20260928175127.1032535-1-ctshao@google.com>
Add DDR bandwidth and latency metrics for the NVIDIA Tegra410 SoC. The
formulas follow Documentation/admin-guide/perf/nvidia-tegra410-pmu.rst:
AVG_MEM_READ_BANDWIDTH_IN_GBPS = MEM_BYTES_RD / ELAPSED_TIME_IN_NS
AVG_MEM_WRITE_BANDWIDTH_IN_GBPS = MEM_BYTES_WR / ELAPSED_TIME_IN_NS
FREQ_IN_GHZ = CYCLES / ELAPSED_TIME_IN_NS
AVG_LATENCY_IN_CYCLES = RD_CUM_OUTS / RD_REQ
AVERAGE_LATENCY_IN_NS = AVG_LATENCY_IN_CYCLES / FREQ_IN_GHZ
The bandwidth metrics use the mem_bytes_rd/mem_bytes_wr events of the
Unified Coherence Fabric (UCF) PMU with source/destination filters:
lpm_ddr_bw, lpm_ddr_rd_bw, lpm_ddr_wr_bw: all sources to local CMEM.
lpm_ddr_loc_rd_bw, lpm_ddr_loc_wr_bw: local CPU and non-CPU sources to
local CMEM.
lpm_ddr_rem_rd_bw, lpm_ddr_rem_wr_bw: local CPU and non-CPU sources to
remote memory.
The latency metrics use the CPU Memory (CMEM) latency PMU:
lpm_ddr_lat, lpm_ddr_rd_lat: average DDR read latency in ns.
lpm_ddr_wr_lat: the PMU only measures reads, so this is always 0.
There is one CMEM latency PMU per socket and perf sums the counts of the
PMUs in an aggregation, e.g. of both sockets in the default aggregation
mode. Use aggr_nr() to divide the summed cycles by the number of PMUs,
otherwise the frequency is too high and the latency too low by that
factor.
The events are sysfs events of the nvidia_ucf_pmu_<socket-id> and
nvidia_cmem_latency_pmu_<socket-id> PMUs rather than json events, so add
the "nvidia" PMU prefix to the prefixes that skip the named json event
check in metric.py.
pmu-events/Build only runs arm64_metrics.py for the "arm" vendor models,
so also run it for the "nvidia" vendor models to generate the Tegra410
metrics.
Signed-off-by: Chun-Tse Shao <ctshao@google.com>
Assisted-by: Gemini:gemini-3.1-pro-preview
---
tools/perf/pmu-events/Build | 2 +-
tools/perf/pmu-events/arm64_metrics.py | 66 +++++++++++++++++++++++++-
tools/perf/pmu-events/metric.py | 2 +-
3 files changed, 66 insertions(+), 4 deletions(-)
diff --git a/tools/perf/pmu-events/Build b/tools/perf/pmu-events/Build
index 1ac067be9d34..8dc8cfcb8b04 100644
--- a/tools/perf/pmu-events/Build
+++ b/tools/perf/pmu-events/Build
@@ -78,7 +78,7 @@ endif
ifeq ($(JEVENTS_ARCH),$(filter $(JEVENTS_ARCH),arm64 all))
# Generate ARM Json
-ARMS := $(shell ls -d pmu-events/arch/arm64/arm/*|grep -v cmn)
+ARMS := $(shell ls -d pmu-events/arch/arm64/arm/* pmu-events/arch/arm64/nvidia/*|grep -v cmn)
ARM_METRICS = $(foreach x,$(ARMS),$(OUTPUT)$(x)/extra-metrics.json)
ARM_METRICGROUPS = $(foreach x,$(ARMS),$(OUTPUT)$(x)/extra-metricgroups.json)
GEN_JSON += $(ARM_METRICS) $(ARM_METRICGROUPS)
diff --git a/tools/perf/pmu-events/arm64_metrics.py b/tools/perf/pmu-events/arm64_metrics.py
index 4ecda96d11fa..b0d0651fec12 100755
--- a/tools/perf/pmu-events/arm64_metrics.py
+++ b/tools/perf/pmu-events/arm64_metrics.py
@@ -2,14 +2,75 @@
# SPDX-License-Identifier: (LGPL-2.1 OR BSD-2-Clause)
import argparse
import os
-from metric import (JsonEncodeMetric, JsonEncodeMetricGroupDescriptions, LoadEvents,
- MetricGroup)
+from typing import Optional
+from metric import (aggr_nr, d_ratio, Event, JsonEncodeMetric,
+ JsonEncodeMetricGroupDescriptions, LoadEvents, Metric, MetricGroup)
from common_metrics import Cycles
# Global command line arguments.
_args = None
+def NvidiaT410Uncore() -> Optional[MetricGroup]:
+ """Uncore metrics for the NVIDIA Tegra410 SoC.
+
+ The formulas follow Documentation/admin-guide/perf/nvidia-tegra410-pmu.rst.
+ """
+ assert _args is not None
+ if _args.vendor != "nvidia" or _args.model != "t410":
+ return None
+
+ interval_sec = Event("duration_time")
+ mb_per_sec = f"{1/1e6}MB/s"
+
+ def DdrBw() -> MetricGroup:
+ # Unified Coherence Fabric (UCF) PMU:
+ # AVG_MEM_{READ,WRITE}_BANDWIDTH = MEM_BYTES_{RD,WR} / ELAPSED_TIME
+ loc_src = "src_loc_cpu=0x1,src_loc_noncpu=0x1"
+ rd = Event("nvidia_ucf_pmu/mem_bytes_rd,dst_loc_cmem=0x1/")
+ wr = Event("nvidia_ucf_pmu/mem_bytes_wr,dst_loc_cmem=0x1/")
+ loc_rd = Event(f"nvidia_ucf_pmu/mem_bytes_rd,{loc_src},dst_loc_cmem=0x1/")
+ loc_wr = Event(f"nvidia_ucf_pmu/mem_bytes_wr,{loc_src},dst_loc_cmem=0x1/")
+ rem_rd = Event(f"nvidia_ucf_pmu/mem_bytes_rd,{loc_src},dst_rem=0x1/")
+ rem_wr = Event(f"nvidia_ucf_pmu/mem_bytes_wr,{loc_src},dst_rem=0x1/")
+ return MetricGroup("lpm_ddr_bw", [
+ Metric("lpm_ddr_bw", "Total DDR bandwidth",
+ d_ratio(rd + wr, interval_sec), mb_per_sec),
+ Metric("lpm_ddr_rd_bw", "DDR read bandwidth",
+ d_ratio(rd, interval_sec), mb_per_sec),
+ Metric("lpm_ddr_wr_bw", "DDR write bandwidth",
+ d_ratio(wr, interval_sec), mb_per_sec),
+ Metric("lpm_ddr_loc_rd_bw", "Local DDR read bandwidth",
+ d_ratio(loc_rd, interval_sec), mb_per_sec),
+ Metric("lpm_ddr_loc_wr_bw", "Local DDR write bandwidth",
+ d_ratio(loc_wr, interval_sec), mb_per_sec),
+ Metric("lpm_ddr_rem_rd_bw", "Remote DDR read bandwidth",
+ d_ratio(rem_rd, interval_sec), mb_per_sec),
+ Metric("lpm_ddr_rem_wr_bw", "Remote DDR write bandwidth",
+ d_ratio(rem_wr, interval_sec), mb_per_sec),
+ ], description="NVIDIA Tegra410 DDR bandwidth")
+
+ def DdrLat() -> MetricGroup:
+ # CPU Memory (CMEM) latency PMU, one per socket:
+ # AVERAGE_LATENCY_IN_NS = (RD_CUM_OUTS / RD_REQ) / (CYCLES / ELAPSED_TIME_IN_NS)
+ rd_cum_outs = Event("nvidia_cmem_latency_pmu/rd_cum_outs/")
+ rd_req = Event("nvidia_cmem_latency_pmu/rd_req/")
+ cycles = Event("nvidia_cmem_latency_pmu/cycles/")
+ # perf sums the counts of the PMUs in an aggregation, e.g. both sockets
+ # by default, so divide cycles by their number to get the frequency.
+ rd_lat = d_ratio(1e9 * interval_sec * rd_cum_outs * aggr_nr(cycles), rd_req * cycles)
+ return MetricGroup("lpm_ddr_lat", [
+ Metric("lpm_ddr_lat", "Average DDR read latency in ns", rd_lat, "1ns"),
+ Metric("lpm_ddr_rd_lat", "DDR read latency in ns", rd_lat, "1ns"),
+ # The CMEM latency PMU only measures reads, write latency is always 0.
+ Metric("lpm_ddr_wr_lat", "DDR write latency in ns",
+ rd_cum_outs - rd_cum_outs, "1ns"),
+ ], description="NVIDIA Tegra410 DDR latency")
+
+ return MetricGroup("lpm_uncore", [DdrBw(), DdrLat()],
+ description="NVIDIA Tegra410 uncore metrics")
+
+
def main() -> None:
global _args
@@ -37,6 +98,7 @@ def main() -> None:
all_metrics = MetricGroup("", [
Cycles(),
+ NvidiaT410Uncore(),
])
if _args.metricgroups:
diff --git a/tools/perf/pmu-events/metric.py b/tools/perf/pmu-events/metric.py
index f5d81eccd5c3..6fb69024d1fd 100644
--- a/tools/perf/pmu-events/metric.py
+++ b/tools/perf/pmu-events/metric.py
@@ -91,7 +91,7 @@ def CheckEveryEvent(*names: str) -> None:
name = name[:name.find(':')]
elif '/' in name:
name = name[:name.find('/')]
- if any([name.startswith(x) for x in ['amd', 'arm', 'cpu', 'msr', 'power', 'cha', 'uncore']]):
+ if any([name.startswith(x) for x in ['amd', 'arm', 'cpu', 'msr', 'power', 'cha', 'uncore', 'nvidia']]):
continue
if name not in all_events_all_models:
raise ValueError(f"Is {name} a named json event?")
--
2.56.0.rc1.315.gc6ed9934b7-goog
next prev parent reply other threads:[~2026-09-28 17:51 UTC|newest]
Thread overview: 3+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-28 17:51 [PATCH 0/2] perf jevents: Add NVIDIA Tegra410 uncore DDR and PCIe metrics Chun-Tse Shao
2026-09-28 17:51 ` Chun-Tse Shao [this message]
2026-09-28 17:51 ` [PATCH 2/2] perf jevents: Add NVIDIA Tegra410 uncore " Chun-Tse Shao
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260928175127.1032535-2-ctshao@google.com \
--to=ctshao@google.com \
--cc=acme@kernel.org \
--cc=adrian.hunter@intel.com \
--cc=alexander.shishkin@linux.intel.com \
--cc=bwicaksono@nvidia.com \
--cc=irogers@google.com \
--cc=james.clark@linaro.org \
--cc=jolsa@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-perf-users@vger.kernel.org \
--cc=mark.rutland@arm.com \
--cc=mingo@redhat.com \
--cc=namhyung@kernel.org \
--cc=peterz@infradead.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®