From: Chun-Tse Shao <ctshao@google.com>
To: acme@kernel.org, namhyung@kernel.org, irogers@google.com
Cc: peterz@infradead.org, mingo@redhat.com, mark.rutland@arm.com,
alexander.shishkin@linux.intel.com, jolsa@kernel.org,
adrian.hunter@intel.com, james.clark@linaro.org,
bwicaksono@nvidia.com, linux-perf-users@vger.kernel.org,
linux-kernel@vger.kernel.org, Chun-Tse Shao <ctshao@google.com>
Subject: [PATCH 2/2] perf jevents: Add NVIDIA Tegra410 uncore PCIe metrics
Date: Mon, 28 Sep 2026 10:51:27 -0700 [thread overview]
Message-ID: <20260928175127.1032535-3-ctshao@google.com> (raw)
In-Reply-To: <20260928175127.1032535-1-ctshao@google.com>
Add PCIe bandwidth metrics for the NVIDIA Tegra410 SoC. The formulas
follow Documentation/admin-guide/perf/nvidia-tegra410-pmu.rst:
AVG_RD_BANDWIDTH_IN_GBPS = RD_BYTES / ELAPSED_TIME_IN_NS
AVG_WR_BANDWIDTH_IN_GBPS = WR_BYTES / ELAPSED_TIME_IN_NS
There is one PCIE PMU per socket and PCIe Root Complex (RC), named
nvidia_pcie_pmu_<socket-id>_rc_<pcie-rc-id>, with RCs 0 to 5 in each
socket. The metrics sum the rd_bytes/wr_bytes events of both sockets:
lpm_pcie_bw, lpm_pcie_rd_bw, lpm_pcie_wr_bw: all RCs.
lpm_pcie_bw_<rc>, lpm_pcie_rd_bw_<rc>, lpm_pcie_wr_bw_<rc>: one RC.
lpm_pcie_loc_rd_bw_<rc>, lpm_pcie_loc_wr_bw_<rc>: one RC to local
CMEM, GMEM, PCIe peer and CXL memory.
lpm_pcie_rem_rd_bw_<rc>, lpm_pcie_rem_wr_bw_<rc>: one RC to remote
memory.
The per RC metrics are in the lpm_pcie_rc<rc>_bw metric groups.
A socket's event counts as 0 if its PMU isn't present (has_event() is
0), e.g. on a single socket system, or if it isn't counted in the
aggregation (source_count() is 0), e.g. the other socket's event with
--per-socket. Otherwise the metric fails to parse or is nan.
Signed-off-by: Chun-Tse Shao <ctshao@google.com>
Assisted-by: Gemini:gemini-3.1-pro-preview
---
tools/perf/pmu-events/arm64_metrics.py | 68 ++++++++++++++++++++++++--
1 file changed, 64 insertions(+), 4 deletions(-)
diff --git a/tools/perf/pmu-events/arm64_metrics.py b/tools/perf/pmu-events/arm64_metrics.py
index b0d0651fec12..1ca3d8ddb4b7 100755
--- a/tools/perf/pmu-events/arm64_metrics.py
+++ b/tools/perf/pmu-events/arm64_metrics.py
@@ -2,9 +2,10 @@
# SPDX-License-Identifier: (LGPL-2.1 OR BSD-2-Clause)
import argparse
import os
-from typing import Optional
-from metric import (aggr_nr, d_ratio, Event, JsonEncodeMetric,
- JsonEncodeMetricGroupDescriptions, LoadEvents, Metric, MetricGroup)
+from typing import List, Optional, Union
+from metric import (aggr_nr, d_ratio, has_event, source_count, Event, Expression,
+ JsonEncodeMetric, JsonEncodeMetricGroupDescriptions, LoadEvents, Metric,
+ MetricGroup, Select)
from common_metrics import Cycles
# Global command line arguments.
@@ -67,7 +68,66 @@ def NvidiaT410Uncore() -> Optional[MetricGroup]:
rd_cum_outs - rd_cum_outs, "1ns"),
], description="NVIDIA Tegra410 DDR latency")
- return MetricGroup("lpm_uncore", [DdrBw(), DdrLat()],
+ def PcieBw() -> MetricGroup:
+ # PCIE PMU, one per socket and PCIe Root Complex (RC):
+ # AVG_{RD,WR}_BANDWIDTH = {RD,WR}_BYTES / ELAPSED_TIME
+ loc_dst = "dst_loc_cmem=0x1,dst_loc_gmem=0x1,dst_loc_pcie_p2p=0x1,dst_loc_pcie_cxl=0x1"
+ rem_dst = "dst_rem=0x1"
+
+ def SocketSum(rc: str, event: str) -> Expression:
+ # Sum the event over both sockets. A PMU counts as 0 if it isn't
+ # present (has_event() is 0) or isn't counted in the aggregation,
+ # e.g. the other socket with --per-socket (source_count() is 0).
+ def Count(e: Event) -> Expression:
+ return Select(Select(e, source_count(e), 0), has_event(e), 0)
+
+ return (Count(Event(f"nvidia_pcie_pmu_0_{rc}/{event}/")) +
+ Count(Event(f"nvidia_pcie_pmu_1_{rc}/{event}/")))
+
+ # nvidia_pcie_pmu_<socket-id>_rc matches the PMUs of all RCs in the socket.
+ rd = SocketSum("rc", "rd_bytes")
+ wr = SocketSum("rc", "wr_bytes")
+ metrics: List[Union[Metric, MetricGroup]] = [
+ Metric("lpm_pcie_bw", "Total PCIe bandwidth across all Root Complexes",
+ d_ratio(rd + wr, interval_sec), mb_per_sec),
+ Metric("lpm_pcie_rd_bw", "Total PCIe read bandwidth across all Root Complexes",
+ d_ratio(rd, interval_sec), mb_per_sec),
+ Metric("lpm_pcie_wr_bw", "Total PCIe write bandwidth across all Root Complexes",
+ d_ratio(wr, interval_sec), mb_per_sec),
+ ]
+ for i in range(6):
+ rc = f"rc_{i}"
+ rc_rd = SocketSum(rc, "rd_bytes")
+ rc_wr = SocketSum(rc, "wr_bytes")
+ loc_rd = SocketSum(rc, f"rd_bytes,{loc_dst}")
+ loc_wr = SocketSum(rc, f"wr_bytes,{loc_dst}")
+ rem_rd = SocketSum(rc, f"rd_bytes,{rem_dst}")
+ rem_wr = SocketSum(rc, f"wr_bytes,{rem_dst}")
+ metrics.append(MetricGroup(f"lpm_pcie_rc{i}_bw", [
+ Metric(f"lpm_pcie_bw_{i}", f"PCIe total bandwidth on Root Complex {i}",
+ d_ratio(rc_rd + rc_wr, interval_sec), mb_per_sec),
+ Metric(f"lpm_pcie_rd_bw_{i}", f"PCIe read bandwidth on Root Complex {i}",
+ d_ratio(rc_rd, interval_sec), mb_per_sec),
+ Metric(f"lpm_pcie_wr_bw_{i}", f"PCIe write bandwidth on Root Complex {i}",
+ d_ratio(rc_wr, interval_sec), mb_per_sec),
+ Metric(f"lpm_pcie_loc_rd_bw_{i}",
+ f"PCIe local read bandwidth on Root Complex {i}",
+ d_ratio(loc_rd, interval_sec), mb_per_sec),
+ Metric(f"lpm_pcie_loc_wr_bw_{i}",
+ f"PCIe local write bandwidth on Root Complex {i}",
+ d_ratio(loc_wr, interval_sec), mb_per_sec),
+ Metric(f"lpm_pcie_rem_rd_bw_{i}",
+ f"PCIe remote read bandwidth on Root Complex {i}",
+ d_ratio(rem_rd, interval_sec), mb_per_sec),
+ Metric(f"lpm_pcie_rem_wr_bw_{i}",
+ f"PCIe remote write bandwidth on Root Complex {i}",
+ d_ratio(rem_wr, interval_sec), mb_per_sec),
+ ], description=f"PCIe bandwidth on Root Complex {i}"))
+
+ return MetricGroup("lpm_pcie_bw", metrics,
+ description="NVIDIA Tegra410 PCIe Root Complex bandwidth")
+
+ return MetricGroup("lpm_uncore", [DdrBw(), DdrLat(), PcieBw()],
description="NVIDIA Tegra410 uncore metrics")
--
2.56.0.rc1.315.gc6ed9934b7-goog
prev parent reply other threads:[~2026-09-28 17:51 UTC|newest]
Thread overview: 3+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-28 17:51 [PATCH 0/2] perf jevents: Add NVIDIA Tegra410 uncore DDR and " Chun-Tse Shao
2026-09-28 17:51 ` [PATCH 1/2] perf jevents: Add NVIDIA Tegra410 uncore DDR metrics Chun-Tse Shao
2026-09-28 17:51 ` Chun-Tse Shao [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260928175127.1032535-3-ctshao@google.com \
--to=ctshao@google.com \
--cc=acme@kernel.org \
--cc=adrian.hunter@intel.com \
--cc=alexander.shishkin@linux.intel.com \
--cc=bwicaksono@nvidia.com \
--cc=irogers@google.com \
--cc=james.clark@linaro.org \
--cc=jolsa@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-perf-users@vger.kernel.org \
--cc=mark.rutland@arm.com \
--cc=mingo@redhat.com \
--cc=namhyung@kernel.org \
--cc=peterz@infradead.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®