From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-dl1-f70.google.com (mail-dl1-f70.google.com [74.125.82.70]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1A09450B429 for ; Wed, 23 Sep 2026 18:13:38 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.82.70 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790187226; cv=none; b=TUue/Sj5uP30YUo3ya8Uag8LOJfPpvFx9s5JTvfOwUegLBNHyUBfKr0vY3bSUKgYBK45qg0yFjSJQYC1y/EEKKRODh2hJ2IQ2jnZcSmUCrh+qgUpLFD3CROnBUeL3aI0zU3L8GdfdGSNOl8BLAyh+qcP3CMeDU0WCHv9AR/GvkI= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790187226; c=relaxed/simple; bh=4b0imgbFdYz+4JCKfONwFUgxUK4QQzPGLp4FygSrSe8=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=U/Clp+Ur/Ldofi+CiBNHYPtIfk/9R3fhp816WW5k54T5oTVI6U8pHZ/5lWPt9v/8PM61tZGS4nQupz+tEOEUBXTEfR9PZjzDeETsMLZcGTWFgFv36vt6b+DkJhfANKi06gRkRgUkIJVPqTepOhoecrJF/H5kTIDC7V4jwh2la6c= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--irogers.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=KiGffeV2; arc=none smtp.client-ip=74.125.82.70 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--irogers.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="KiGffeV2" Received: by mail-dl1-f70.google.com with SMTP id a92af1059eb24-1438492fb40so1869895c88.1 for ; Wed, 23 Sep 2026 11:13:37 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1790187214; x=1790792014; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=GdUrPLeidZ9sUt9zVmAl34CruFIaflRaMNzsOxLYKi0=; b=KiGffeV2Z1taDBcZok6Q938vffq48oeY3s4b8ZLt3dGg2tosKbznJQZQGjSOMSrrxc ZHQdO2dt2VmCumx7+LGIPP3UkX9pLYxYwWM9N/95ALmSmaXu5Q6udXU+ZbUr87rkNgFZ pK+7nd6At2r9o/dB9o8ITLWAbcWe/NE8bAF44cCAITI6FZkKg49voHJWvWpO6yQC2lIP M/T4/pvLHGYNGj/fKFiQkYS/7xOLXQx3FaQmMxrPtGhy9NMDACsV5KQbP/VvXRhxPfNd TuuWmXXg7Ind1591A4MBICQ96CtVFVD9JlcsGtIdP1nq3FN1TW/sMxYdOhf75nblehDV Zuvw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790187214; x=1790792014; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=GdUrPLeidZ9sUt9zVmAl34CruFIaflRaMNzsOxLYKi0=; b=HA4soOR57H3MJ4RxRBPl/LBUb4BDQp/baaFZ9tEzj8hI/UY/mJfppFyzwnzowASPyA Xh3VQOkOgqQ9m+W1E3Sr57b/Pqytc0eLBmmwzJcqHYMDGNBKqNQxqNG+A/eOgbIs6HSx E0FuQE7sipKJzNf/de0Ms9l+mYZa3MQCdmmdgIqIy/lFEk/GGH6F6lGgklqEi3V+PuUS q87l50WmmQoZURlgFduF2gHr7D2wXait5k4CCadGkColOa8RysbLZ2JPUGrnzmicr7TN MHp6CO2Og/gM/gY1roj0TYoEDveJy/I+v1YEkftoN5CTMFXL3kvp0eBdEBkSpktkA6pB cTnw== X-Forwarded-Encrypted: i=1; AKwUvBzfef632c6Ys9X485/MNcl6z/LqrQGYYmpcckRSD/nayVy+LY6aEBiNm2lQ38N5b0FC6rCqfw45SwKZG5I=@vger.kernel.org X-Gm-Message-State: AFuF++nyZfar9YTmR7aZn2B/PQOzageAaRUvrZ635aornTBhFTSKPUAU yG6cJ+HMYacxTOnNddNBTRfciiCOXWCFytOclVFTGj3Q0y+jViA06nN5eQ8pTe1xR55zw3J8dRr 4naiD8BHZcA== X-Received: from dldnz10.prod.google.com ([2002:a05:701a:ca0a:b0:144:dc81:9760]) (user=irogers job=prod-delivery.src-stubby-dispatcher) by 2002:a05:7022:ea30:b0:143:271a:308 with SMTP id a92af1059eb24-144f91b5db4mr5035124c88.46.1790187213775; Wed, 23 Sep 2026 11:13:33 -0700 (PDT) Date: Wed, 23 Sep 2026 11:11:39 -0700 In-Reply-To: <20260923181213.3032038-1-irogers@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260923181213.3032038-1-irogers@google.com> X-Mailer: git-send-email 2.56.0.rc1.310.g51773c2048-goog Message-ID: <20260923181213.3032038-17-irogers@google.com> Subject: [PATCH v3 16/49] perf python: Port stat-cpi to perf module From: Ian Rogers To: irogers@google.com, acme@kernel.org, alice.mei.rogers@gmail.com, namhyung@kernel.org Cc: adrian.hunter@intel.com, dapeng1.mi@linux.intel.com, james.clark@linaro.org, leo.yan@linux.dev, linux-kernel@vger.kernel.org, linux-perf-users@vger.kernel.org, mingo@redhat.com, peterz@infradead.org, tmricht@linux.ibm.com Content-Type: text/plain; charset="UTF-8" Port stat-cpi.py from the legacy embedded scripting framework to a standalone Python script in tools/perf/python/ to calculate Cycles Per Instruction (CPI) per interval per CPU or thread. Improvements compared to the legacy script: - Support both perf.data file mode (via perf.session stat callbacks) and live counter collection mode (using perf.parse_events, evlist.open, and evsel.read across intervals), with automatic fallback to user-space (:u) and self-process monitoring when perf_event_paranoid restricts system-wide events (EACCES). - Compute per-interval counter deltas (val, ena, run) keyed by raw event name so cumulative PERF_RECORD_STAT snapshots and hybrid PMU events (e.g. cpu_core/cycles/, cpu_atom/cycles/) are accumulated accurately, and scale counts by time_enabled / time_running when multiplexed. - Replace hard-coded CPU ([0, 1]) and thread ([0]) arrays with dynamic CPU and thread discovery so arbitrary system topologies work automatically. - Add CLI option handling (-i, -I, -p) via argparse and type annotations passing mypy and pylint. Add a shell test (test_stat_cpi_python.sh) to verify the standalone script. Assisted-by: Antigravity:gemini-3.1-pro Signed-off-by: Ian Rogers --- tools/perf/python/stat-cpi.py | 210 ++++++++++++++++++ .../perf/tests/shell/test_stat_cpi_python.sh | 115 ++++++++++ 2 files changed, 325 insertions(+) create mode 100755 tools/perf/python/stat-cpi.py create mode 100755 tools/perf/tests/shell/test_stat_cpi_python.sh diff --git a/tools/perf/python/stat-cpi.py b/tools/perf/python/stat-cpi.py new file mode 100755 index 000000000000..58ffa0b4e274 --- /dev/null +++ b/tools/perf/python/stat-cpi.py @@ -0,0 +1,210 @@ +#!/usr/bin/env python3 +# SPDX-License-Identifier: GPL-2.0 +"""Calculate CPI from perf stat data or live.""" +from __future__ import annotations + +import argparse +import os +import signal +import sys +import time +from typing import Any, Optional +import perf + +class StatCpiAnalyzer: + """Accumulates cycles and instructions and calculates CPI.""" + + def __init__(self, args: argparse.Namespace) -> None: + self.args = args + self.data: dict[str, float] = {} + self.prev_data: dict[str, tuple[int, int, int]] = {} + self.recorded_pairs: set[tuple[int, int]] = set() + + def get_key(self, event: str, cpu: int, thread: int) -> str: + """Get key for data dictionary.""" + return f"{event}-{cpu}-{thread}" + + def store_key(self, cpu: int, thread: int) -> None: + """Store CPU and thread IDs.""" + self.recorded_pairs.add((cpu, thread)) + + def store(self, event: str, cpu: int, thread: int, + counts: tuple[int, int, int], is_delta: bool = False, + raw_name: Optional[str] = None) -> None: + """Store counter values, computing difference from previous + absolute values if not already deltas.""" + self.store_key(cpu, thread) + key = self.get_key(event, cpu, thread) + prev_key = self.get_key(raw_name or event, cpu, thread) + + val, ena, run = counts + if is_delta: + # counts are already deltas + cur_val = val + cur_ena = ena + cur_run = run + else: + if prev_key in self.prev_data: + prev_val, prev_ena, prev_run = self.prev_data[prev_key] + cur_val = val - prev_val + cur_ena = ena - prev_ena + cur_run = run - prev_run + else: + cur_val = val + cur_ena = ena + cur_run = run + self.prev_data[prev_key] = counts # Store absolute value for next time + + # Scale each raw event's delta by its own multiplexing ratio before + # summing across PMUs (e.g. cpu_core and cpu_atom on hybrid systems) so + # enabled time from an idle PMU does not inflate the multiplier of an + # active PMU. + scaled_val = cur_val * (cur_ena / float(cur_run)) if cur_run > 0 else float(cur_val) + self.data[key] = self.data.get(key, 0.0) + scaled_val + + def get(self, event: str, cpu: int, thread: int) -> float: + """Get scaled counter value.""" + key = self.get_key(event, cpu, thread) + return self.data.get(key, 0.0) + + def process_stat_event(self, event: Any, name: Optional[str] = None) -> None: + """Process PERF_RECORD_STAT and PERF_RECORD_STAT_ROUND events.""" + if event.type == perf.RECORD_STAT: + if name: + if "cycles" in name: + event_name = "cycles" + elif "instructions" in name: + event_name = "instructions" + else: + return + self.store(event_name, event.cpu, event.thread, + (event.val, event.ena, event.run), raw_name=name) + elif event.type == perf.RECORD_STAT_ROUND: + timestamp = getattr(event, "time", 0) + self.print_interval(timestamp) + self.data.clear() + self.recorded_pairs.clear() + + def print_interval(self, timestamp: int) -> None: + """Print CPI for the current interval.""" + for cpu, thread in sorted(self.recorded_pairs): + cyc = self.get("cycles", cpu, thread) + ins = self.get("instructions", cpu, thread) + cpi = 0.0 + if ins != 0: + cpi = cyc / float(ins) + t_sec = timestamp / 1000000000.0 + print(f"{t_sec:15f}: cpu {cpu}, thread {thread} -> cpi {cpi:f} ({cyc:.0f}/{ins:.0f})") + + def read_counters(self, evlist: Any) -> None: + """Read counters live.""" + for evsel in evlist: + name = str(evsel) + if "cycles" in name: + event_name = "cycles" + elif "instructions" in name: + event_name = "instructions" + else: + continue + + for cpu in evsel.cpus(): + for thread in evsel.threads(): + try: + counts = evsel.read(cpu, thread) + self.store(event_name, cpu, thread, + (counts.val, counts.ena, counts.run), + is_delta=True, raw_name=name) + except OSError: + pass + + def run_file(self) -> None: + """Process events from file.""" + session: Optional[perf.session] = perf.session( + perf.data(self.args.input), stat=self.process_stat_event + ) + try: + assert session is not None + session.process_events() + finally: + session = None + + def _open_live_evlist(self) -> Any: + """Open evlist for live mode, falling back to user-space or process scope on EACCES.""" + threads = perf.thread_map(self.args.pid) if self.args.pid else None + candidates = [ + ("cycles,instructions", threads), + ("cycles:u,instructions:u", threads), + ] + if threads is None: + self_threads = perf.thread_map(os.getpid()) + candidates.append(("cycles,instructions", self_threads)) + candidates.append(("cycles:u,instructions:u", self_threads)) + + last_err: Optional[OSError] = None + for events, tmap in candidates: + try: + evlist = perf.parse_events(events, None, tmap) + for evsel in evlist: + evsel.read_format |= ( + perf.FORMAT_TOTAL_TIME_ENABLED | perf.FORMAT_TOTAL_TIME_RUNNING + ) + evlist.open() + evlist.enable() + return evlist + except PermissionError as e: + last_err = e + except OSError as e: + if e.errno == 13: + last_err = e + else: + raise + if last_err is not None: + raise last_err + raise RuntimeError("Failed to open events") + + def run_live(self) -> None: + """Read counters live.""" + try: + evlist = self._open_live_evlist() + except OSError as e: + print(f"Failed to open events: {e}", file=sys.stderr) + sys.exit(1) + + def handle_signal(_signum: int, _frame: Any) -> None: + raise KeyboardInterrupt + + signal.signal(signal.SIGINT, signal.default_int_handler) + signal.signal(signal.SIGTERM, handle_signal) + + print("Live mode started. Press Ctrl+C to stop.") + try: + while True: + time.sleep(self.args.interval) + timestamp = time.time_ns() + self.read_counters(evlist) + self.print_interval(timestamp) + self.data.clear() + self.recorded_pairs.clear() + except KeyboardInterrupt: + print("\nStopped.") + finally: + evlist.close() + +def main() -> None: + """Main function.""" + ap = argparse.ArgumentParser(description="Calculate CPI from perf stat data or live") + ap.add_argument("-i", "--input", help="Input file name (enables file mode)") + ap.add_argument("-I", "--interval", type=float, default=1.0, + help="Interval in seconds for live mode") + ap.add_argument("-p", "--pid", type=int, + help="Monitor specific process ID in live mode") + args = ap.parse_args() + + analyzer = StatCpiAnalyzer(args) + if args.input: + analyzer.run_file() + else: + analyzer.run_live() + +if __name__ == "__main__": + main() diff --git a/tools/perf/tests/shell/test_stat_cpi_python.sh b/tools/perf/tests/shell/test_stat_cpi_python.sh new file mode 100755 index 000000000000..0960afa4c797 --- /dev/null +++ b/tools/perf/tests/shell/test_stat_cpi_python.sh @@ -0,0 +1,115 @@ +#!/bin/bash +# SPDX-License-Identifier: GPL-2.0 +# stat-cpi python test + +set -e + +shelldir=$(dirname "$0") +# shellcheck source=lib/setup_python.sh +. "${shelldir}"/lib/setup_python.sh + +# If we don't have the perf python module, we can't test +if ! "$PYTHON" -c 'import perf' > /dev/null 2>&1; then + echo "Skipping test, perf python module not found" + return 2 2>/dev/null || exit 2 +fi + +script_dir="$(dirname "$0")/../../python" +script_path="${script_dir}/stat-cpi.py" + +if [ ! -f "$script_path" ]; then + echo "Skipping test, stat-cpi.py not found at $script_path" + return 2 2>/dev/null || exit 2 +fi + +err=0 +ran=0 +temp_data="" +temp_out="" + +cleanup() { + [ -n "${pid}" ] && kill "$pid" 2>/dev/null || true + [ -n "${workload_pid}" ] && kill "$workload_pid" 2>/dev/null || true + rm -f "${temp_data}" "${temp_out}" + trap - exit term int +} + +trap_cleanup() { + cleanup + exit 1 +} +trap trap_cleanup exit term int + +temp_data=$(mktemp /tmp/perf.data.XXXXXX) +temp_out=$(mktemp /tmp/perf.out.XXXXXX) + +test_live_mode() { + echo "Testing stat-cpi.py live mode..." + if ! perf stat -e cycles,instructions -- sleep 0.1 2>/dev/null && \ + ! perf stat -e cycles:u,instructions:u -- sleep 0.1 2>/dev/null; then + echo "perf stat failed (permissions?), skipping live mode test." + return 0 + fi + perf test -w noploop & + workload_pid=$! + if ! perf stat -e cycles,instructions -p "$workload_pid" -- sleep 0.05 2>/dev/null && \ + ! perf stat -e cycles:u,instructions:u -p "$workload_pid" -- sleep 0.05 2>/dev/null; then + kill "$workload_pid" 2>/dev/null || true + workload_pid="" + echo "perf stat -p failed (ptrace_scope?), skipping live mode test." + return 0 + fi + ran=1 + + # Run live mode for 1 interval in the background, give it a tiny sleep, then interrupt + "$PYTHON" "$script_path" -I 0.1 -p "$workload_pid" > "${temp_out}" & + pid=$! + sleep 0.5 + kill -INT "$pid" 2>/dev/null || true + set +e + wait "$pid" + res=$? + set -e + pid="" + kill "$workload_pid" 2>/dev/null || true + workload_pid="" + if [ $res -ne 0 ] && [ $res -ne 130 ] && [ $res -ne 143 ]; then + echo "Live mode failed or crashed" + err=1 + elif ! grep -q "cpi" "${temp_out}"; then + echo "Live mode produced no cpi output" + err=1 + else + echo "Live mode test passed." + fi +} + +test_file_mode() { + echo "Testing stat-cpi.py file mode..." + # Generate some stat events - perf stat -I represents interval reporting + if ! perf stat -e cycles,instructions -I 100 record -o "${temp_data}" \ + -- sleep 0.5 2>/dev/null && \ + ! perf stat -e cycles:u,instructions:u -I 100 record -o "${temp_data}" \ + -- sleep 0.5 2>/dev/null; then + echo "perf stat failed (permissions?), skipping file mode test." + return + fi + ran=1 + + out=$("$PYTHON" "$script_path" -i "${temp_data}") + if ! echo "$out" | grep -q "cpi"; then + echo "File mode test failed." + err=1 + else + echo "File mode test passed." + fi +} + +test_live_mode +test_file_mode + +cleanup +if [ $ran -eq 0 ]; then + exit 2 +fi +exit $err -- 2.56.0.rc1.310.g51773c2048-goog