mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH 0/2] perf jevents: Add NVIDIA Tegra410 uncore DDR and PCIe metrics
@ 2026-09-28 17:51 Chun-Tse Shao
  2026-09-28 17:51 ` [PATCH 1/2] perf jevents: Add NVIDIA Tegra410 uncore DDR metrics Chun-Tse Shao
  2026-09-28 17:51 ` [PATCH 2/2] perf jevents: Add NVIDIA Tegra410 uncore PCIe metrics Chun-Tse Shao
  0 siblings, 2 replies; 3+ messages in thread
From: Chun-Tse Shao @ 2026-09-28 17:51 UTC (permalink / raw)
  To: acme, namhyung, irogers
  Cc: peterz, mingo, mark.rutland, alexander.shishkin, jolsa,
	adrian.hunter, james.clark, bwicaksono, linux-perf-users,
	linux-kernel, Chun-Tse Shao

Add DDR bandwidth/latency and PCIe bandwidth metrics for the NVIDIA
Tegra410 SoC to arm64_metrics.py. The metrics use the sysfs events of
the Tegra410 UCF, CMEM latency and PCIE PMUs, and follow the formulas in
Documentation/admin-guide/perf/nvidia-tegra410-pmu.rst.

Patch 1 adds the DDR metrics and makes pmu-events/Build also run
arm64_metrics.py for the nvidia vendor models. Patch 2 adds the PCIe
metrics, in total and per PCIe Root Complex (RC).

Tested on a 2-socket Tegra410 system:

 - memcpy with multiload (from multichase) on socket 0 CPUs and memory:
   lpm_ddr_bw is 740-744 GB/s at steady state (perf stat -I 2000),
   multiload reports 708063 MiB/s (742 GB/s).
 - Read only (stream-sum) and write only (memset) loads on socket 1:
   lpm_ddr_rd_bw is 608 GB/s with 2.4 GB/s of writes, lpm_ddr_wr_bw is
   628 GB/s with 1.1 GB/s of reads.
 - memcpy on socket 0 CPUs with memory on socket 1: socket 0's
   lpm_ddr_rem_rd_bw and socket 1's lpm_ddr_rd_bw count nearly the same
   bytes (258.03 vs 258.73 GB).
 - lpm_ddr_lat idle is 150 ns in the default aggregation, and 132 ns and
   194 ns for socket 0 and 1 with --per-socket. Under the memcpy load it
   is 555 ns by default and 563 ns for socket 0 with --per-socket, and
   stays at 551-558 ns per interval with -I 1000.
 - 8 GiB O_DIRECT dd read from an NVMe drive behind RC 0 of socket 0:
   lpm_pcie_wr_bw_0 counts the 8 GiB at 8.7 GB/s (dd: 8.7 GB/s). With
   the buffer on node 0 it shows up in lpm_pcie_loc_wr_bw_0, with the
   buffer on node 1 in lpm_pcie_rem_wr_bw_0. With --per-socket, socket 0
   shows the same and no metric of either socket is nan.
 - With the nvidia_t410 metrics forced on a machine without these PMUs
   (PERF_CPUID=0x000000004e0f0100 on an x86 JEVENTS_ARCH=all build), the
   PCIe metrics read 0 rather than failing to parse, also with
   --per-socket.
 - "perf test 10" (PMU JSON event tests) passes on the Tegra410 system
   and with the x86 JEVENTS_ARCH=all build.

Chun-Tse Shao (2):
  perf jevents: Add NVIDIA Tegra410 uncore DDR metrics
  perf jevents: Add NVIDIA Tegra410 uncore PCIe metrics

 tools/perf/pmu-events/Build            |   2 +-
 tools/perf/pmu-events/arm64_metrics.py | 126 ++++++++++++++++++++++++-
 tools/perf/pmu-events/metric.py        |   2 +-
 3 files changed, 126 insertions(+), 4 deletions(-)


base-commit: 0ae6fc78c5ce0dfd18d8712a50f0fd4602eff103
--
2.56.0.rc1.315.gc6ed9934b7-goog


^ permalink raw reply	[flat|nested] 3+ messages in thread

end of thread, other threads:[~2026-09-28 17:51 UTC | newest]

Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-28 17:51 [PATCH 0/2] perf jevents: Add NVIDIA Tegra410 uncore DDR and PCIe metrics Chun-Tse Shao
2026-09-28 17:51 ` [PATCH 1/2] perf jevents: Add NVIDIA Tegra410 uncore DDR metrics Chun-Tse Shao
2026-09-28 17:51 ` [PATCH 2/2] perf jevents: Add NVIDIA Tegra410 uncore PCIe metrics Chun-Tse Shao

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®