From: Wei Huang <wei.huang2@amd.com>
To: <bhelgaas@google.com>
Cc: <linux-pci@vger.kernel.org>, <linux-doc@vger.kernel.org>,
<linux-kernel@vger.kernel.org>, <skhan@linuxfoundation.org>,
<corbet@lwn.net>, <rdunlap@infradead.org>,
<fengchengwen@huawei.com>, <wei.huang2@amd.com>
Subject: [PATCH RFC 3/3] Expose the CPU/Steering Tag mapping via debugfs
Date: Thu, 1 Oct 2026 10:55:51 -0500 [thread overview]
Message-ID: <20261001155551.4182899-4-wei.huang2@amd.com> (raw)
In-Reply-To: <20261001155551.4182899-1-wei.huang2@amd.com>
Steering Tags (ST) are opaque, platform specific values obtained from the
ACPI ST _DSM, and are only visible to the driver that asks for them. That
makes it hard to validate whether firmware reports a sane mapping at all.
Expose the mapping read-only in the debugfs directory of each Root Port
that has ST _DSM support:
# cat /sys/kernel/debug/pci/0000:00:01.1/tph_cpu_st
cpu vm_st vm_xst vm_ph_ignore pm_st pm_xst pm_ph_ignore
0 0x03 0x0103 0 - - 0
1 0x03 0x0103 0 - - 0
Each line holds the raw ST info the _DSM returns for one online CPU;
invalid tags and failed queries are reported as "-".
The interface is per Root Port because the _DSM is evaluated on the ACPI
device of the host bridge above it. A system with multiple host bridges may
return different tags for the same CPU, and the value that matters to a
driver is the one for the Root Port its device sits below. Root Ports below
the same host bridge report identical contents.
The file is created during enumeration and removed with the device's
debugfs directory in pci_destroy_dev(), which returns only once no read is
running, so the pci_dev the file refers to stays valid.
Signed-off-by: Wei Huang <wei.huang2@amd.com>
---
Documentation/PCI/tph.rst | 32 +++++++++
drivers/pci/tph.c | 136 ++++++++++++++++++++++++++++++++++++++
2 files changed, 168 insertions(+)
diff --git a/Documentation/PCI/tph.rst b/Documentation/PCI/tph.rst
index b6cf22b9bd90..19239546af3b 100644
--- a/Documentation/PCI/tph.rst
+++ b/Documentation/PCI/tph.rst
@@ -125,6 +125,38 @@ been changed. Here is a sample code for IRQ affinity notifier:
return;
}
+Inspect the firmware ST mapping
+-------------------------------
+
+When firmware provides the ST _DSM, the CPU to ST mapping it returns is
+exposed read-only under debugfs, in the directory of each Root Port::
+
+ /sys/kernel/debug/pci/<domain:bus:device.function>/tph_cpu_st
+
+Each file holds one line per online CPU::
+
+ # cat /sys/kernel/debug/pci/0000:00:01.1/tph_cpu_st
+ cpu vm_st vm_xst vm_ph_ignore pm_st pm_xst pm_ph_ignore
+ 0 0x03 0x0103 0 - - 0
+ 1 0x03 0x0103 0 - - 0
+
+``vm_*`` and ``pm_*`` refer to volatile and persistent memory, ``*_st`` and
+``*_xst`` are the 8-bit and 16-bit (extended) Steering Tags, and
+``*_ph_ignore`` reports whether the Processing Hint is ignored by the
+platform. Steering Tags that firmware marks as invalid, and CPUs whose
+query fails, are shown as ``-``.
+
+The _DSM is evaluated on the ACPI device of the host bridge above the Root
+Port, so all Root Ports below the same host bridge report the same mapping.
+The file is per Root Port because that is the granularity drivers see: the
+mapping that applies to a device is the one in the directory of the Root
+Port the device sits below.
+
+Tags are queried on read rather than cached, so the output always matches
+what pcie_tph_get_cpu_st() would return at that moment. The list of CPUs
+is a snapshot; CPUs that come or go while the file is being read may be
+missing from the output.
+
Disable TPH system-wide
-----------------------
diff --git a/drivers/pci/tph.c b/drivers/pci/tph.c
index d2c6d3dda202..bd682185ffbb 100644
--- a/drivers/pci/tph.c
+++ b/drivers/pci/tph.c
@@ -6,11 +6,14 @@
* Eric Van Tassell <Eric.VanTassell@amd.com>
* Wei Huang <wei.huang2@amd.com>
*/
+#include <linux/cpu.h>
+#include <linux/debugfs.h>
#include <linux/pci.h>
#include <linux/pci-acpi.h>
#include <linux/msi.h>
#include <linux/bitfield.h>
#include <linux/pci-tph.h>
+#include <linux/seq_file.h>
#include "pci.h"
@@ -171,6 +174,133 @@ static int tph_get_cpu_st_info(struct pci_dev *pdev, unsigned int cpu,
}
#endif
+#if defined(CONFIG_ACPI) && defined(CONFIG_DEBUG_FS)
+/* Longest Steering Tag string is "0xffff" */
+#define TPH_ST_STR_SIZE 7
+
+static bool tph_dsm_supported(struct pci_dev *pdev)
+{
+ acpi_handle handle = tph_dsm_handle(pdev);
+
+ if (!handle)
+ return false;
+
+ return acpi_check_dsm(handle, &pci_acpi_dsm_guid, TPH_ST_DSM_REV,
+ BIT(TPH_ST_DSM_FUNC_INDEX));
+}
+
+static const char *tph_st_str(char *buf, bool valid, u16 st, int digits)
+{
+ if (!valid)
+ return "-";
+
+ snprintf(buf, TPH_ST_STR_SIZE, "0x%0*x", digits, st);
+
+ return buf;
+}
+
+/*
+ * Position 0 is the header, position N > 0 is the first online CPU with an
+ * id >= N - 1. Encoding the CPU id in the position, instead of an index
+ * into the online mask, keeps the remaining output stable if a CPU goes
+ * offline between two read() calls.
+ *
+ * The iterator itself is the CPU id biased by 2, because 0 is the end of
+ * the sequence and 1 is SEQ_START_TOKEN.
+ */
+static void *tph_cpu_st_pos(loff_t *pos)
+{
+ unsigned int cpu;
+
+ if (!*pos)
+ return SEQ_START_TOKEN;
+
+ if (*pos > nr_cpu_ids)
+ return NULL;
+
+ cpu = cpumask_next((int)*pos - 2, cpu_online_mask);
+ if (cpu >= nr_cpu_ids)
+ return NULL;
+
+ *pos = cpu + 1;
+
+ return (void *)(unsigned long)(cpu + 2);
+}
+
+static void *tph_cpu_st_start(struct seq_file *s, loff_t *pos)
+{
+ return tph_cpu_st_pos(pos);
+}
+
+static void *tph_cpu_st_next(struct seq_file *s, void *v, loff_t *pos)
+{
+ ++*pos;
+
+ return tph_cpu_st_pos(pos);
+}
+
+static void tph_cpu_st_stop(struct seq_file *s, void *v)
+{
+}
+
+static int tph_cpu_st_show(struct seq_file *s, void *v)
+{
+ char vm_st[TPH_ST_STR_SIZE], vm_xst[TPH_ST_STR_SIZE];
+ char pm_st[TPH_ST_STR_SIZE], pm_xst[TPH_ST_STR_SIZE];
+ struct pci_dev *pdev = s->private;
+ union st_info info;
+ unsigned int cpu;
+
+ if (v == SEQ_START_TOKEN) {
+ seq_printf(s, "%-4s %-6s %-6s %-12s %-6s %-6s %s\n",
+ "cpu", "vm_st", "vm_xst", "vm_ph_ignore",
+ "pm_st", "pm_xst", "pm_ph_ignore");
+ return 0;
+ }
+
+ cpu = (unsigned long)v - 2;
+
+ if (tph_get_cpu_st_info(pdev, cpu, &info)) {
+ seq_printf(s, "%-4u %-6s %-6s %-12s %-6s %-6s %s\n", cpu,
+ "-", "-", "-", "-", "-", "-");
+ return 0;
+ }
+
+ seq_printf(s, "%-4u %-6s %-6s %-12u %-6s %-6s %u\n", cpu,
+ tph_st_str(vm_st, info.vm_st_valid, info.vm_st, 2),
+ tph_st_str(vm_xst, info.vm_xst_valid, info.vm_xst, 4),
+ (unsigned int)info.vm_ph_ignore,
+ tph_st_str(pm_st, info.pm_st_valid, info.pm_st, 2),
+ tph_st_str(pm_xst, info.pm_xst_valid, info.pm_xst, 4),
+ (unsigned int)info.pm_ph_ignore);
+
+ return 0;
+}
+
+static const struct seq_operations tph_cpu_st_sops = {
+ .start = tph_cpu_st_start,
+ .next = tph_cpu_st_next,
+ .stop = tph_cpu_st_stop,
+ .show = tph_cpu_st_show,
+};
+DEFINE_SEQ_ATTRIBUTE(tph_cpu_st);
+
+static void tph_debugfs_add_cpu_st(struct pci_dev *pdev)
+{
+ if (!pci_is_pcie(pdev) ||
+ pci_pcie_type(pdev) != PCI_EXP_TYPE_ROOT_PORT)
+ return;
+
+ if (!tph_dsm_supported(pdev))
+ return;
+
+ debugfs_create_file("tph_cpu_st", 0444, pci_dev_debugfs_dir(pdev),
+ pdev, &tph_cpu_st_fops);
+}
+#else
+static void tph_debugfs_add_cpu_st(struct pci_dev *pdev) { }
+#endif /* CONFIG_ACPI && CONFIG_DEBUG_FS */
+
/* Update the TPH Requester Enable field of TPH Control Register */
static void set_ctrl_reg_req_en(struct pci_dev *pdev, u8 req_type)
{
@@ -536,6 +666,12 @@ void pci_tph_init(struct pci_dev *pdev)
int num_entries;
u32 save_size;
+ /*
+ * A Root Port has no ST table of its own, but it is where the
+ * firmware ST mapping of its host bridge is exposed.
+ */
+ tph_debugfs_add_cpu_st(pdev);
+
pdev->tph_cap = pci_find_ext_capability(pdev, PCI_EXT_CAP_ID_TPH);
if (!pdev->tph_cap)
return;
--
2.55.0
prev parent reply other threads:[~2026-10-01 15:56 UTC|newest]
Thread overview: 4+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-01 15:55 [PATCH RFC 0/3] PCI/TPH: Expose CPU-to-Steering Tag mappings " Wei Huang
2026-10-01 15:55 ` [PATCH RFC 1/3] PCI/TPH: Factor out the ST _DSM evaluation Wei Huang
2026-10-01 15:55 ` [PATCH RFC 2/3] PCI: Add per-device debugfs directories Wei Huang
2026-10-01 15:55 ` Wei Huang [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261001155551.4182899-4-wei.huang2@amd.com \
--to=wei.huang2@amd.com \
--cc=bhelgaas@google.com \
--cc=corbet@lwn.net \
--cc=fengchengwen@huawei.com \
--cc=linux-doc@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-pci@vger.kernel.org \
--cc=rdunlap@infradead.org \
--cc=skhan@linuxfoundation.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®