* [PATCH v9 00/18] PCI/P2PDMA: Route peer-to-peer DMA by TLP class
@ 2026-10-01 11:55 Leon Romanovsky
2026-10-01 11:55 ` [PATCH v9 01/18] PCI/P2PDMA: Document the TLP attribute assumptions Leon Romanovsky
` (18 more replies)
0 siblings, 19 replies; 38+ messages in thread
From: Leon Romanovsky @ 2026-10-01 11:55 UTC (permalink / raw)
To: Bjorn Helgaas, Logan Gunthorpe, Leon Romanovsky, Jason Gunthorpe,
Joerg Roedel (AMD),
Will Deacon, Robin Murphy, Christian König,
Thomas Hellström
Cc: linux-pci, linux-kernel, linux-doc, iommu, Tushar Dave,
linux-media, dri-devel, linaro-mm-sig, linux-rdma, kvm,
Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Sumit Semwal
PCI P2PDMA applies Request and Completion Redirect throughout both paths.
This misclassifies asymmetric and nested switches, and reports one answer
for every kind of TLP.
Three ACS controls act on TLP attributes the client chooses rather than
on the topology: Translation Blocking and Direct Translated P2P act on
a Request's Address Type, and Completion Redirect skips Completions carrying
Relaxed Ordering.
Evaluate each direction at the path divergence, decide every class from the
one walk, and treat a client with ATS enabled as translating unless its
driver declares per-mapping ATS.
This completes the P2PDMA side; dma-buf and mlx5 follow separately.
Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
---
Changes in v9:
- Removed dmabuf patches for now.
- Removed mlx5 per-mapping ATS declaration for now; it comes back with
the dmabuf patches.
- Link to v8: https://patch.msgid.link/20260928-fix-p2p-acs-v4-0-v8-0-404453b9c435@nvidia.com
Changes in v8:
- Added extra Reviewed-by from Logan.
- Added patches to take care of ATS per-device vs. per-mapping option
- Link to v7: https://patch.msgid.link/20260920-fix-p2p-acs-v4-0-v7-0-ca0828ab697c@nvidia.com
Changes in v7:
- Removed "Document pdev->p2pdma lifetime rules" patch, it gives nothing
after p2pmem fix.
- Split dmabuf patch.
- Tushar retested the series, so added his Tested-by.
- Added Logan's ROB tags and fixed minor documentation issues pointed
by him.
- Link to v6: https://lore.kernel.org/all/20260914-fix-p2p-acs-v4-0-v6-0-5ef07ec9ef06@nvidia.com/
Changes in v6:
- Changed DMABUF to use callbacks and not directly stored pointer.
- Link to v5: https://patch.msgid.link/20260910-fix-p2p-acs-v4-0-v5-0-856087f63c0d@nvidia.com
Changes in v5:
- Rebase on the posted fixes.
- Dropped tags from changed patches.
- Remove Egress Control Vector interpretation and coverage.
- Keep enabled Egress Control conservative as a Request redirect.
- Use pci_dbg()/dev_dbg() for diagnostics and drop the "debug" prefix.
- Removed code comments from "Document the pdev->p2pdma lifetime and RCU
rules" patch and reduced description to actual lifetime explanation.
- Added note that Linux assumes that TLPs are in strict-ordering and
untranslated.
- Added code to calculate p2p paths per-TLP type.
- Converted mlx5 to use that new proposed API.
- Link to https://patch.msgid.link/20260821-fix-p2p-acs-v4-0-v4-0-94426b96de73@nvidia.com
Changes in v4:
- Reject ACS Violations and unreadable routing state instead of treating
them as host-bridge redirects
- Added Tested-by tags from Tushar Dave
- Added support to asymmetric ACS routing
- Limited redirect checks to the two ports at the path divergence
- Added standalone ACS routing diagnostics for hardware retesting
- Dropped " PCI: Account for Direct Translated P2P in ACS isolation checks" patch
- Link to v3: https://patch.msgid.link/20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com
Changes in v3:
- Fixed pci_p2pdma_add_resource() error unwinding
- Made pdev->p2pdma teardown wait unconditionally for RCU readers
- Restricted pci_p2pmem_find_many() to pool-backed providers
- Documented the pdev->p2pdma lifetime and RCU rules
- Fixed calc_map_type_and_dist() handling of the verbose argument
- Required the ACS port and target to share a bus before indexing the
Egress Control Vector
- Gave pci_acs_enabled() and pci_acs_path_enabled() a scope, so the ACS
Direct Translated P2P rule no longer stops pci_enable_pasid() from
enabling PASID
- Dropped "Report ACS ports when the paths share no upstream bridge":
the mapping type cannot change without a shared upstream bridge, so
the pci=disable_acs_redir= hint was not actionable there and the ACS
walk only cost config space reads
- Folded the Request Redirect rule into pci_acs_rr_ineffective(), so
pci_acs_flags_enabled() and the Intel SPT PCH quirk share one copy
- Renamed pci_acs_egress_ctrl_set() to pci_acs_egress_ctrl_is_set(), it
reads the bit rather than setting it
- Reworded the blocked-path warning: ACS may also leave the direct route
indeterminate rather than blocked
- Added KUnit coverage for the shared-bus guard, a device with no ACS
capability and an unreadable ACS Control register
- Added the missing Fixes: tags, a second one on the
pci_p2pdma_add_resource() unwinding fix (the dangling devres action
dates to f58ef9d1d135) and one on the Egress Control isolation change
- Link to v2: https://patch.msgid.link/20260806-fix-p2p-acs-v2-0-0cec14812965@nvidia.com
Changes in v2:
- Added Logan's ROB tags
- Added commas in Documentation patch
- Link to v1: https://patch.msgid.link/20260802-fix-p2p-acs-v1-0-a7c5eb64fff6@nvidia.com
---
Leon Romanovsky (18):
PCI/P2PDMA: Document the TLP attribute assumptions
PCI/P2PDMA: Derive routing from directional ACS controls
PCI: Reject unreadable ACS controls in isolation checks
PCI/P2PDMA: Evaluate ACS controls at the path divergence
PCI/P2PDMA: Document directional ACS routing
PCI/P2PDMA: Collect the path's ACS controls before deciding
PCI/P2PDMA: Answer routing per TLP class
PCI/P2PDMA: Route Relaxed Ordering Completions directly
PCI/P2PDMA: Reject Translated Requests blocked by Translation Blocking
PCI/P2PDMA: Route Translated Requests under Direct Translated P2P
PCI/P2PDMA: Log detailed ACS routing diagnostics
PCI/P2PDMA: Add KUnit tests for the ACS routing decisions
PCI/P2PDMA: Test the ACS P2P routing walk
PCI: Add KUnit coverage for ACS isolation checks
PCI/P2PDMA: Document TLP-class routing
PCI/P2PDMA: Let a client declare that it selects ATS per mapping
PCI/P2PDMA: Evaluate the ATS path for clients with ATS enabled
PCI/P2PDMA: Test the routing of clients with ATS enabled
Documentation/admin-guide/kernel-parameters.txt | 15 +-
Documentation/driver-api/pci/p2pdma.rst | 80 +++
drivers/pci/Kconfig | 15 +
drivers/pci/Makefile | 1 +
drivers/pci/p2pdma.c | 719 ++++++++++++++++++--
drivers/pci/pci.c | 7 +-
drivers/pci/pci.h | 58 ++
drivers/pci/pci_acs_test.c | 860 ++++++++++++++++++++++++
drivers/pci/quirks.c | 6 +-
include/linux/pci-p2pdma.h | 12 +-
10 files changed, 1687 insertions(+), 86 deletions(-)
---
base-commit: 08dbfad3f5040f5bdb6c529da20d6d4e81fefd72
change-id: 20260821-fix-p2p-acs-v4-0-e72455e3a261
prerequisite-message-id: <20260830-batch-p2p-fixes-v1-0-5044e8dfbe2e@nvidia.com>
prerequisite-patch-id: 6b25c7fcf164cdfc14e9fac5b908d97fcf6509d7
prerequisite-patch-id: 0d083c281001365aae4b35544cf28891a6ab9a96
prerequisite-patch-id: bfd9dabf271f3cc9a3a61f46387d20c20311363d
prerequisite-patch-id: fad0275efc722830fc591509506c0a5e4f581073
prerequisite-patch-id: 0c83bee688fec1f6d1564654df7c630fa6a4a978
Best regards,
--
Leon Romanovsky <leonro@nvidia.com>
^ permalink raw reply [flat|nested] 38+ messages in thread
* [PATCH v9 01/18] PCI/P2PDMA: Document the TLP attribute assumptions
2026-10-01 11:55 [PATCH v9 00/18] PCI/P2PDMA: Route peer-to-peer DMA by TLP class Leon Romanovsky
@ 2026-10-01 11:55 ` Leon Romanovsky
2026-10-07 20:12 ` Bjorn Helgaas
2026-10-01 11:55 ` [PATCH v9 02/18] PCI/P2PDMA: Derive routing from directional ACS controls Leon Romanovsky
` (17 subsequent siblings)
18 siblings, 1 reply; 38+ messages in thread
From: Leon Romanovsky @ 2026-10-01 11:55 UTC (permalink / raw)
To: Bjorn Helgaas, Logan Gunthorpe, Leon Romanovsky, Jason Gunthorpe,
Joerg Roedel (AMD),
Will Deacon, Robin Murphy, Christian König,
Thomas Hellström
Cc: linux-pci, linux-kernel, linux-doc, iommu, Tushar Dave,
linux-media, dri-devel, linaro-mm-sig, linux-rdma, kvm,
Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Sumit Semwal
From: Leon Romanovsky <leonro@nvidia.com>
P2PDMA selects a mapping without receiving the Request's ordering or
Address Type attributes. Its ACS handles only strictly ordered Requests
carrying an Untranslated address.
Document that the result is not defined for Relaxed Ordering or
ATS-translated Requests because those TLP attributes can select different
routes through the fabric.
Reviewed-by: Logan Gunthorpe <logang@deltatee.com>
Tested-by: Tushar Dave <tdave@nvidia.com>
Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
---
Documentation/driver-api/pci/p2pdma.rst | 7 +++++++
1 file changed, 7 insertions(+)
diff --git a/Documentation/driver-api/pci/p2pdma.rst b/Documentation/driver-api/pci/p2pdma.rst
index 63cff9e4d2c9..80f8fec9b0e9 100644
--- a/Documentation/driver-api/pci/p2pdma.rst
+++ b/Documentation/driver-api/pci/p2pdma.rst
@@ -15,6 +15,13 @@ then based on the ACS settings the transaction can route entirely within
the PCIe hierarchy and never reach the root port. The kernel will evaluate
the PCIe topology and always permit P2P in these well-defined cases.
+This evaluation assumes clients issue strictly ordered Requests carrying an
+Untranslated address. Its result is not defined when clients use Relaxed
+Ordering or issue ATS-translated Requests because those TLP attributes can
+select different routes through the fabric. Unless ACS Translation Blocking
+is enabled, a Port with ACS Direct Translated P2P enabled routes a
+Translated Request directly to the peer regardless of the redirect controls.
+
However, if the P2P transaction reaches the host bridge then it might have to
hairpin back out the same root port, be routed inside the CPU SOC to another
PCIe root port, or routed internally to the SOC.
--
2.55.0
^ permalink raw reply [flat|nested] 38+ messages in thread
* [PATCH v9 02/18] PCI/P2PDMA: Derive routing from directional ACS controls
2026-10-01 11:55 [PATCH v9 00/18] PCI/P2PDMA: Route peer-to-peer DMA by TLP class Leon Romanovsky
2026-10-01 11:55 ` [PATCH v9 01/18] PCI/P2PDMA: Document the TLP attribute assumptions Leon Romanovsky
@ 2026-10-01 11:55 ` Leon Romanovsky
2026-10-06 20:49 ` Bjorn Helgaas
2026-10-01 11:55 ` [PATCH v9 03/18] PCI: Reject unreadable ACS controls in isolation checks Leon Romanovsky
` (16 subsequent siblings)
18 siblings, 1 reply; 38+ messages in thread
From: Leon Romanovsky @ 2026-10-01 11:55 UTC (permalink / raw)
To: Bjorn Helgaas, Logan Gunthorpe, Leon Romanovsky, Jason Gunthorpe,
Joerg Roedel (AMD),
Will Deacon, Robin Murphy, Christian König,
Thomas Hellström
Cc: linux-pci, linux-kernel, linux-doc, iommu, Tushar Dave,
linux-media, dri-devel, linaro-mm-sig, linux-rdma, kvm,
Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Sumit Semwal
From: Leon Romanovsky <leonro@nvidia.com>
pci_bridge_has_acs_redir() treats Request and Completion Redirect as
interchangeable. On asymmetric fabrics, a control for only the reverse TLP
direction can unnecessarily force P2PDMA through the host bridge.
Evaluate Request Redirect for client Requests and Completion Redirect for
provider read Completions. Continue treating enabled Egress Control
conservatively as a Request redirect.
Fixes: 52916982af48 ("PCI/P2PDMA: Support peer-to-peer memory")
Reviewed-by: Logan Gunthorpe <logang@deltatee.com>
Tested-by: Tushar Dave <tdave@nvidia.com>
Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
---
drivers/pci/p2pdma.c | 75 ++++++++++++++++++++++++++++++++++++++++------------
1 file changed, 58 insertions(+), 17 deletions(-)
diff --git a/drivers/pci/p2pdma.c b/drivers/pci/p2pdma.c
index 4e4d2df17a45..12612b82d80d 100644
--- a/drivers/pci/p2pdma.c
+++ b/drivers/pci/p2pdma.c
@@ -21,6 +21,8 @@
#include <linux/seq_buf.h>
#include <linux/xarray.h>
+#include "pci.h"
+
struct pci_p2pdma {
struct gen_pool *pool;
bool p2pmem_published;
@@ -490,26 +492,56 @@ static struct pci_dev *find_parent_pci_dev(struct device *dev)
return NULL;
}
+enum pci_acs_p2pdma_state {
+ PCI_ACS_P2PDMA_DIRECT,
+ PCI_ACS_P2PDMA_REDIRECT,
+};
+
/*
- * Check if a PCI bridge has its ACS redirection bits set to redirect P2P
- * TLPs upstream via ACS. Returns 1 if the packets will be redirected
- * upstream, 0 otherwise.
+ * Decide how a peer-to-peer Request at an ACS-capable ingress port routes,
+ * from that port's ACS Control register.
+ *
+ * Linux does not read the Egress Control Vector, so Egress Control is treated
+ * conservatively as a redirect. Per PCIe r7.0 Table 6-11 the outcomes it
+ * selects are a direct route and an ACS Violation, and neither one lets peer
+ * bus addressing be assumed.
*/
-static int pci_bridge_has_acs_redir(struct pci_dev *pdev)
+static enum pci_acs_p2pdma_state
+pci_acs_p2pdma_request(u16 ctrl)
{
- int pos;
- u16 ctrl;
+ return ctrl & (PCI_ACS_RR | PCI_ACS_EC) ?
+ PCI_ACS_P2PDMA_REDIRECT : PCI_ACS_P2PDMA_DIRECT;
+}
- pos = pdev->acs_cap;
- if (!pos)
- return 0;
+/*
+ * Decide how a peer-to-peer Completion at an ACS-capable ingress port routes.
+ * PCIe r7.0 sec 6.12.1.1: no ACS control other than P2P Completion Redirect
+ * affects a Completion.
+ */
+static enum pci_acs_p2pdma_state
+pci_acs_p2pdma_completion(u16 ctrl)
+{
+ return ctrl & PCI_ACS_CR ? PCI_ACS_P2PDMA_REDIRECT :
+ PCI_ACS_P2PDMA_DIRECT;
+}
- pci_read_config_word(pdev, pos + PCI_ACS_CTRL, &ctrl);
+/*
+ * Read @pdev's ACS Control register. A device without an ACS capability has
+ * no peer-to-peer controls at all, which routes the same as having them all
+ * clear. Returns false when the register is present but cannot be read; @ctrl
+ * is then meaningless.
+ */
+static bool pci_acs_p2pdma_ctrl(struct pci_dev *pdev, u16 *ctrl)
+{
+ int pos;
- if (ctrl & (PCI_ACS_RR | PCI_ACS_CR | PCI_ACS_EC))
- return 1;
+ pos = pdev->acs_cap;
+ if (!pos) {
+ *ctrl = 0;
+ return true;
+ }
- return 0;
+ return !pci_read_config_word(pdev, pos + PCI_ACS_CTRL, ctrl);
}
static void seq_buf_print_bus_devfn(struct seq_buf *buf, struct pci_dev *pdev)
@@ -698,6 +730,10 @@ static unsigned long map_types_idx(struct pci_dev *client)
* then to Device B. The mapping type returned depends on the ACS
* redirection setting of the ports along the path.
*
+ * The client initiates Requests to provider memory. Check Request Redirect
+ * on the client path and Completion Redirect for read Completions on the
+ * provider path.
+ *
* If ACS redirect is set on any port in the path, traffic between the
* devices will go through the host bridge, so return
* PCI_P2PDMA_MAP_THRU_HOST_BRIDGE; otherwise return
@@ -721,6 +757,7 @@ calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client,
int dist_a = 0;
int dist_b = 0;
char buf[128];
+ u16 ctrl;
seq_buf_init(&acs_list, buf, sizeof(buf));
@@ -732,7 +769,9 @@ calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client,
while (a) {
dist_b = 0;
- if (pci_bridge_has_acs_redir(a)) {
+ if (!pci_acs_p2pdma_ctrl(a, &ctrl) ||
+ pci_acs_p2pdma_completion(ctrl) ==
+ PCI_ACS_P2PDMA_REDIRECT) {
seq_buf_print_bus_devfn(&acs_list, a);
acs_cnt++;
}
@@ -761,7 +800,9 @@ calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client,
if (a == bb)
break;
- if (pci_bridge_has_acs_redir(bb)) {
+ if (!pci_acs_p2pdma_ctrl(bb, &ctrl) ||
+ pci_acs_p2pdma_request(ctrl) ==
+ PCI_ACS_P2PDMA_REDIRECT) {
seq_buf_print_bus_devfn(&acs_list, bb);
acs_cnt++;
}
@@ -1109,10 +1150,10 @@ EXPORT_SYMBOL_GPL(pci_p2pmem_publish);
/**
* pci_p2pdma_map_type - Determine the mapping type for P2PDMA transfers
* @provider: P2PDMA provider structure
- * @dev: Target device for the transfer
+ * @dev: Client device that initiates the transfer
*
* Determines how peer-to-peer DMA transfers should be mapped between
- * the provider and the target device. The mapping type indicates whether
+ * the provider and the client device. The mapping type indicates whether
* the transfer can be done directly through PCI switches or must go
* through the host bridge.
*/
--
2.55.0
^ permalink raw reply [flat|nested] 38+ messages in thread
* [PATCH v9 03/18] PCI: Reject unreadable ACS controls in isolation checks
2026-10-01 11:55 [PATCH v9 00/18] PCI/P2PDMA: Route peer-to-peer DMA by TLP class Leon Romanovsky
2026-10-01 11:55 ` [PATCH v9 01/18] PCI/P2PDMA: Document the TLP attribute assumptions Leon Romanovsky
2026-10-01 11:55 ` [PATCH v9 02/18] PCI/P2PDMA: Derive routing from directional ACS controls Leon Romanovsky
@ 2026-10-01 11:55 ` Leon Romanovsky
2026-10-01 11:55 ` [PATCH v9 04/18] PCI/P2PDMA: Evaluate ACS controls at the path divergence Leon Romanovsky
` (15 subsequent siblings)
18 siblings, 0 replies; 38+ messages in thread
From: Leon Romanovsky @ 2026-10-01 11:55 UTC (permalink / raw)
To: Bjorn Helgaas, Logan Gunthorpe, Leon Romanovsky, Jason Gunthorpe,
Joerg Roedel (AMD),
Will Deacon, Robin Murphy, Christian König,
Thomas Hellström
Cc: linux-pci, linux-kernel, linux-doc, iommu, Tushar Dave,
linux-media, dri-devel, linaro-mm-sig, linux-rdma, kvm,
Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Sumit Semwal
From: Leon Romanovsky <leonro@nvidia.com>
pci_acs_flags_enabled() and the Intel SPT PCH quirk use ACS registers
without checking config-space read errors. A failed read may leave control
state indeterminate yet allow the device to satisfy requested isolation
controls.
Return false when either ACS capability or control state cannot be read.
An unknown state cannot prove isolation.
Reviewed-by: Logan Gunthorpe <logang@deltatee.com>
Tested-by: Tushar Dave <tdave@nvidia.com>
Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
---
drivers/pci/pci.c | 3 ++-
drivers/pci/quirks.c | 6 ++++--
2 files changed, 6 insertions(+), 3 deletions(-)
diff --git a/drivers/pci/pci.c b/drivers/pci/pci.c
index b2879a6be5f8..f7d94ecf9157 100644
--- a/drivers/pci/pci.c
+++ b/drivers/pci/pci.c
@@ -3594,7 +3594,8 @@ static bool pci_acs_flags_enabled(struct pci_dev *pdev, u16 acs_flags)
*/
acs_flags &= (pdev->acs_capabilities | PCI_ACS_EC);
- pci_read_config_word(pdev, pos + PCI_ACS_CTRL, &ctrl);
+ if (pci_read_config_word(pdev, pos + PCI_ACS_CTRL, &ctrl))
+ return false;
return (ctrl & acs_flags) == acs_flags;
}
diff --git a/drivers/pci/quirks.c b/drivers/pci/quirks.c
index de9bbccda21f..d5c3e6802840 100644
--- a/drivers/pci/quirks.c
+++ b/drivers/pci/quirks.c
@@ -4992,10 +4992,12 @@ static int pci_quirk_intel_spt_pch_acs(struct pci_dev *dev, u16 acs_flags)
return -ENOTTY;
/* see pci_acs_flags_enabled() */
- pci_read_config_dword(dev, pos + PCI_ACS_CAP, &cap);
+ if (pci_read_config_dword(dev, pos + PCI_ACS_CAP, &cap))
+ return 0;
acs_flags &= (cap | PCI_ACS_EC);
- pci_read_config_dword(dev, pos + INTEL_SPT_ACS_CTRL, &ctrl);
+ if (pci_read_config_dword(dev, pos + INTEL_SPT_ACS_CTRL, &ctrl))
+ return 0;
return pci_acs_ctrl_enabled(acs_flags, ctrl);
}
--
2.55.0
^ permalink raw reply [flat|nested] 38+ messages in thread
* [PATCH v9 04/18] PCI/P2PDMA: Evaluate ACS controls at the path divergence
2026-10-01 11:55 [PATCH v9 00/18] PCI/P2PDMA: Route peer-to-peer DMA by TLP class Leon Romanovsky
` (2 preceding siblings ...)
2026-10-01 11:55 ` [PATCH v9 03/18] PCI: Reject unreadable ACS controls in isolation checks Leon Romanovsky
@ 2026-10-01 11:55 ` Leon Romanovsky
2026-10-06 21:08 ` Bjorn Helgaas
2026-10-01 11:55 ` [PATCH v9 05/18] PCI/P2PDMA: Document directional ACS routing Leon Romanovsky
` (14 subsequent siblings)
18 siblings, 1 reply; 38+ messages in thread
From: Leon Romanovsky @ 2026-10-01 11:55 UTC (permalink / raw)
To: Bjorn Helgaas, Logan Gunthorpe, Leon Romanovsky, Jason Gunthorpe,
Joerg Roedel (AMD),
Will Deacon, Robin Murphy, Christian König,
Thomas Hellström
Cc: linux-pci, linux-kernel, linux-doc, iommu, Tushar Dave,
linux-media, dri-devel, linaro-mm-sig, linux-rdma, kvm,
Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Sumit Semwal
From: Leon Romanovsky <leonro@nvidia.com>
ACS redirect controls choose between peer and upstream routes only at the
path divergence. Applying them below that point rejects valid nested
topologies because traffic already has only an upstream route.
Evaluate Request controls on the client-side divergence port and Completion
Redirect on the provider-side port and reject an unreadable ACS Control
register.
Fixes: 52916982af48 ("PCI/P2PDMA: Support peer-to-peer memory")
Reviewed-by: Logan Gunthorpe <logang@deltatee.com>
Tested-by: Tushar Dave <tdave@nvidia.com>
Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
---
drivers/pci/p2pdma.c | 101 +++++++++++++++++++++++++++++----------------
include/linux/pci-p2pdma.h | 8 ++--
2 files changed, 70 insertions(+), 39 deletions(-)
diff --git a/drivers/pci/p2pdma.c b/drivers/pci/p2pdma.c
index 12612b82d80d..550e6c7346ef 100644
--- a/drivers/pci/p2pdma.c
+++ b/drivers/pci/p2pdma.c
@@ -493,6 +493,7 @@ static struct pci_dev *find_parent_pci_dev(struct device *dev)
}
enum pci_acs_p2pdma_state {
+ PCI_ACS_P2PDMA_NOT_SUPPORTED,
PCI_ACS_P2PDMA_DIRECT,
PCI_ACS_P2PDMA_REDIRECT,
};
@@ -730,13 +731,13 @@ static unsigned long map_types_idx(struct pci_dev *client)
* then to Device B. The mapping type returned depends on the ACS
* redirection setting of the ports along the path.
*
- * The client initiates Requests to provider memory. Check Request Redirect
- * on the client path and Completion Redirect for read Completions on the
- * provider path.
+ * The client initiates Requests to provider memory. At the path divergence,
+ * check Request Redirect and Egress Control on the client-side port, and
+ * Completion Redirect for read Completions on the provider-side port.
*
- * If ACS redirect is set on any port in the path, traffic between the
- * devices will go through the host bridge, so return
- * PCI_P2PDMA_MAP_THRU_HOST_BRIDGE; otherwise return
+ * If ACS redirects traffic at either divergence port, return
+ * PCI_P2PDMA_MAP_THRU_HOST_BRIDGE. If the ACS Control register cannot be
+ * read, return PCI_P2PDMA_MAP_NOT_SUPPORTED. Otherwise, return
* PCI_P2PDMA_MAP_BUS_ADDR.
*
* Any two devices that have a data path that goes through the host bridge
@@ -750,10 +751,13 @@ calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client,
int *dist, bool verbose)
{
enum pci_p2pdma_map_type map_type = PCI_P2PDMA_MAP_THRU_HOST_BRIDGE;
+ enum pci_acs_p2pdma_state state = PCI_ACS_P2PDMA_NOT_SUPPORTED;
struct pci_dev *a = provider, *b = client, *bb;
+ struct pci_dev *a_child = NULL, *b_child = NULL;
+ struct pci_dev *acs_unreadable = NULL;
struct pci_p2pdma *p2pdma;
struct seq_buf acs_list;
- int acs_cnt = 0;
+ int acs_redirect_cnt = 0;
int dist_a = 0;
int dist_b = 0;
char buf[128];
@@ -768,51 +772,67 @@ calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client,
*/
while (a) {
dist_b = 0;
-
- if (!pci_acs_p2pdma_ctrl(a, &ctrl) ||
- pci_acs_p2pdma_completion(ctrl) ==
- PCI_ACS_P2PDMA_REDIRECT) {
- seq_buf_print_bus_devfn(&acs_list, a);
- acs_cnt++;
- }
-
+ b_child = NULL;
bb = b;
while (bb) {
if (a == bb)
- goto check_b_path_acs;
+ goto check_paths_acs;
+ b_child = bb;
bb = pci_upstream_bridge(bb);
dist_b++;
}
+ a_child = a;
a = pci_upstream_bridge(a);
dist_a++;
}
+ /*
+ * The paths share no upstream bridge, so there is no direct path for
+ * ACS to gate: PCI_P2PDMA_MAP_BUS_ADDR is not reachable here and the
+ * request can only get to the peer through the host bridge.
+ */
*dist = dist_a + dist_b;
goto map_through_host_bridge;
-check_b_path_acs:
- bb = b;
-
- while (bb) {
- if (a == bb)
- break;
+check_paths_acs:
+ *dist = dist_a + dist_b;
- if (!pci_acs_p2pdma_ctrl(bb, &ctrl) ||
- pci_acs_p2pdma_request(ctrl) ==
- PCI_ACS_P2PDMA_REDIRECT) {
- seq_buf_print_bus_devfn(&acs_list, bb);
- acs_cnt++;
+ /*
+ * ACS P2P routing controls apply where a TLP can route toward the peer
+ * or upstream. Below that divergence, its only route toward the other
+ * branch is upstream, so redirect controls do not affect the path.
+ */
+ if (a_child && b_child) {
+ if (pci_acs_p2pdma_ctrl(a_child, &ctrl))
+ state = pci_acs_p2pdma_completion(ctrl);
+ if (state != PCI_ACS_P2PDMA_DIRECT) {
+ seq_buf_print_bus_devfn(&acs_list, a_child);
+ if (state == PCI_ACS_P2PDMA_REDIRECT)
+ acs_redirect_cnt++;
+ else if (!acs_unreadable)
+ acs_unreadable = a_child;
}
- bb = pci_upstream_bridge(bb);
+ state = PCI_ACS_P2PDMA_NOT_SUPPORTED;
+ if (pci_acs_p2pdma_ctrl(b_child, &ctrl))
+ state = pci_acs_p2pdma_request(ctrl);
+ if (state != PCI_ACS_P2PDMA_DIRECT) {
+ seq_buf_print_bus_devfn(&acs_list, b_child);
+ if (state == PCI_ACS_P2PDMA_REDIRECT)
+ acs_redirect_cnt++;
+ else if (!acs_unreadable)
+ acs_unreadable = b_child;
+ }
}
- *dist = dist_a + dist_b;
-
- if (!acs_cnt) {
+ /*
+ * Below a shared upstream bridge, a path whose divergence ports do not
+ * redirect routes the request directly.
+ */
+ if (!acs_unreadable && !acs_redirect_cnt) {
map_type = PCI_P2PDMA_MAP_BUS_ADDR;
goto done;
}
@@ -821,10 +841,21 @@ calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client,
/* Drop the final semicolon; the list is not empty here. */
if (!seq_buf_has_overflowed(&acs_list))
acs_list.buffer[acs_list.len - 1] = '\0';
- pci_warn(client, "ACS redirect is set between the client and provider (%s)\n",
- pci_name(provider));
- pci_warn(client, "to disable ACS redirect for this path, add the kernel parameter: pci=disable_acs_redir=%s\n",
- seq_buf_str(&acs_list));
+ if (acs_unreadable)
+ pci_warn(client, "ACS Control is unreadable for provider %s at %s\n",
+ pci_name(provider), pci_name(acs_unreadable));
+ else {
+ pci_warn(client, "ACS redirect is set between the client and provider (%s)\n",
+ pci_name(provider));
+ pci_warn(client, "to disable ACS controls for this path, add the kernel parameter: pci=disable_acs_redir=%s\n",
+ seq_buf_str(&acs_list));
+ }
+ }
+
+ /* An unreadable control does not establish an upstream redirect. */
+ if (acs_unreadable) {
+ map_type = PCI_P2PDMA_MAP_NOT_SUPPORTED;
+ goto done;
}
map_through_host_bridge:
diff --git a/include/linux/pci-p2pdma.h b/include/linux/pci-p2pdma.h
index 873de20a2247..dd17501ba1b6 100644
--- a/include/linux/pci-p2pdma.h
+++ b/include/linux/pci-p2pdma.h
@@ -42,10 +42,10 @@ enum pci_p2pdma_map_type {
PCI_P2PDMA_MAP_NONE,
/*
- * PCI_P2PDMA_MAP_NOT_SUPPORTED: Indicates the transaction will
- * traverse the host bridge and the host bridge is not in the
- * allowlist. DMA Mapping routines should return an error when
- * this is returned.
+ * PCI_P2PDMA_MAP_NOT_SUPPORTED: Indicates no safe mapping is available,
+ * for example because ACS blocks the direct path or the required host
+ * bridge is not in the allowlist. DMA Mapping routines should return an
+ * error when this is returned.
*/
PCI_P2PDMA_MAP_NOT_SUPPORTED,
--
2.55.0
^ permalink raw reply [flat|nested] 38+ messages in thread
* [PATCH v9 05/18] PCI/P2PDMA: Document directional ACS routing
2026-10-01 11:55 [PATCH v9 00/18] PCI/P2PDMA: Route peer-to-peer DMA by TLP class Leon Romanovsky
` (3 preceding siblings ...)
2026-10-01 11:55 ` [PATCH v9 04/18] PCI/P2PDMA: Evaluate ACS controls at the path divergence Leon Romanovsky
@ 2026-10-01 11:55 ` Leon Romanovsky
2026-10-06 21:32 ` Bjorn Helgaas
2026-10-01 11:55 ` [PATCH v9 06/18] PCI/P2PDMA: Collect the path's ACS controls before deciding Leon Romanovsky
` (13 subsequent siblings)
18 siblings, 1 reply; 38+ messages in thread
From: Leon Romanovsky @ 2026-10-01 11:55 UTC (permalink / raw)
To: Bjorn Helgaas, Logan Gunthorpe, Leon Romanovsky, Jason Gunthorpe,
Joerg Roedel (AMD),
Will Deacon, Robin Murphy, Christian König,
Thomas Hellström
Cc: linux-pci, linux-kernel, linux-doc, iommu, Tushar Dave,
linux-media, dri-devel, linaro-mm-sig, linux-rdma, kvm,
Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Sumit Semwal
From: Leon Romanovsky <leonro@nvidia.com>
P2PDMA documentation describes ACS controls as path-wide, although Request
and Completion controls apply to different transaction directions and only
affect peer-versus-upstream decisions at the path divergence.
Document the fixed client and provider roles, the divergence port checked
for each TLP direction, and the conservative handling of unreadable ACS
state. Clarify which controls disable_acs_redir changes.
Reviewed-by: Logan Gunthorpe <logang@deltatee.com>
Tested-by: Tushar Dave <tdave@nvidia.com>
Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
---
Documentation/admin-guide/kernel-parameters.txt | 15 +++++++++------
Documentation/driver-api/pci/p2pdma.rst | 13 +++++++++++++
2 files changed, 22 insertions(+), 6 deletions(-)
diff --git a/Documentation/admin-guide/kernel-parameters.txt b/Documentation/admin-guide/kernel-parameters.txt
index 68647ff4bdd2..bc83e07dd5fc 100644
--- a/Documentation/admin-guide/kernel-parameters.txt
+++ b/Documentation/admin-guide/kernel-parameters.txt
@@ -5291,12 +5291,15 @@ Kernel parameters
disable_acs_redir=<pci_dev>[; ...]
Specify one or more PCI devices (in the format
specified above) separated by semicolons.
- Each device specified will have the PCI ACS
- redirect capabilities forced off which will
- allow P2P traffic between devices through
- bridges without forcing it upstream. Note:
- this removes isolation between devices and
- may put more devices in an IOMMU group.
+ Each device specified will have the PCI ACS P2P
+ Request Redirect, Completion Redirect, and Egress
+ Control features forced off. This may allow P2P
+ traffic through bridges that would otherwise be
+ redirected upstream. This may allow P2P traffic
+ through bridges that would otherwise be redirected
+ upstream and thus this removes isolation between
+ devices and may cause affected devices to share
+ an IOMMU group.
config_acs=
Format:
<ACS flags>@<pci_dev>[; ...]
diff --git a/Documentation/driver-api/pci/p2pdma.rst b/Documentation/driver-api/pci/p2pdma.rst
index 80f8fec9b0e9..42b18610bf7d 100644
--- a/Documentation/driver-api/pci/p2pdma.rst
+++ b/Documentation/driver-api/pci/p2pdma.rst
@@ -15,6 +15,19 @@ then based on the ACS settings the transaction can route entirely within
the PCIe hierarchy and never reach the root port. The kernel will evaluate
the PCIe topology and always permit P2P in these well-defined cases.
+The client remains the PCIe requester when it reads or writes provider memory.
+Where the paths diverge, the kernel therefore evaluates P2P Request Redirect
+and Egress Control on the client-side port, and P2P Completion Redirect on the
+provider-side port for completions from a read. An enabled Egress Control is
+conservatively treated as a Request redirect.
+
+Below the divergence, the route toward the other branch is already upstream,
+so those P2P redirect controls do not affect it. Redirect controls for the
+reverse transaction directions do not affect the mapping. P2P DMA is routed
+through the host bridge when either applicable port redirects. If an ACS
+Control register cannot be read, P2P DMA is rejected because the kernel cannot
+establish a usable route.
+
This evaluation assumes clients issue strictly ordered Requests carrying an
Untranslated address. Its result is not defined when clients use Relaxed
Ordering or issue ATS-translated Requests because those TLP attributes can
--
2.55.0
^ permalink raw reply [flat|nested] 38+ messages in thread
* [PATCH v9 06/18] PCI/P2PDMA: Collect the path's ACS controls before deciding
2026-10-01 11:55 [PATCH v9 00/18] PCI/P2PDMA: Route peer-to-peer DMA by TLP class Leon Romanovsky
` (4 preceding siblings ...)
2026-10-01 11:55 ` [PATCH v9 05/18] PCI/P2PDMA: Document directional ACS routing Leon Romanovsky
@ 2026-10-01 11:55 ` Leon Romanovsky
2026-10-06 21:48 ` Bjorn Helgaas
2026-10-01 11:55 ` [PATCH v9 07/18] PCI/P2PDMA: Answer routing per TLP class Leon Romanovsky
` (12 subsequent siblings)
18 siblings, 1 reply; 38+ messages in thread
From: Leon Romanovsky @ 2026-10-01 11:55 UTC (permalink / raw)
To: Bjorn Helgaas, Logan Gunthorpe, Leon Romanovsky, Jason Gunthorpe,
Joerg Roedel (AMD),
Will Deacon, Robin Murphy, Christian König,
Thomas Hellström
Cc: linux-pci, linux-kernel, linux-doc, iommu, Tushar Dave,
linux-media, dri-devel, linaro-mm-sig, linux-rdma, kvm,
Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Sumit Semwal
From: Leon Romanovsky <leonro@nvidia.com>
calc_map_type_and_dist() reads each divergence port's ACS Control register
and folds the result into running counters as it goes. Any routing property
that depends on the kind of TLP being routed would have to be threaded
through that code, so there is nowhere to put one without reading the
registers again for each kind.
Collect the two ports' ACS Control values into struct pci_p2pdma_acs_path
first, then decide from it. pci_p2pdma_route() applies the same rule as
before: a path routes directly only when both directions do.
Reviewed-by: Logan Gunthorpe <logang@deltatee.com>
Tested-by: Tushar Dave <tdave@nvidia.com>
Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
---
drivers/pci/p2pdma.c | 148 +++++++++++++++++++++++++++++++++------------------
1 file changed, 96 insertions(+), 52 deletions(-)
diff --git a/drivers/pci/p2pdma.c b/drivers/pci/p2pdma.c
index 550e6c7346ef..841c86be31bb 100644
--- a/drivers/pci/p2pdma.c
+++ b/drivers/pci/p2pdma.c
@@ -553,6 +553,80 @@ static void seq_buf_print_bus_devfn(struct seq_buf *buf, struct pci_dev *pdev)
seq_buf_printf(buf, "%s;", pci_name(pdev));
}
+/*
+ * What the topology walk found out about one provider/client path. Producing
+ * this costs a walk and one config read per divergence port, none of which
+ * depends on the TLP being routed.
+ *
+ * @req_ctrl: ACS Control of the client-side divergence port. That is the
+ * first port at which a Request can route toward the peer rather
+ * than upstream, so it is where the Request controls apply.
+ * @cpl_ctrl: ACS Control of the provider-side divergence port, likewise for
+ * the Completions travelling back.
+ * @unreadable: First port whose ACS Control could not be read, if any.
+ */
+struct pci_p2pdma_acs_path {
+ u16 req_ctrl;
+ u16 cpl_ctrl;
+ struct pci_dev *unreadable;
+};
+
+/*
+ * Combine both directions into a mapping type. Only a path that routes the
+ * Request and the Completions it generates directly can be programmed with
+ * the peer's bus addresses.
+ */
+static enum pci_p2pdma_map_type
+pci_p2pdma_route(const struct pci_p2pdma_acs_path *path)
+{
+ if (path->unreadable)
+ return PCI_P2PDMA_MAP_NOT_SUPPORTED;
+
+ if (pci_acs_p2pdma_request(path->req_ctrl) == PCI_ACS_P2PDMA_DIRECT &&
+ pci_acs_p2pdma_completion(path->cpl_ctrl) == PCI_ACS_P2PDMA_DIRECT)
+ return PCI_P2PDMA_MAP_BUS_ADDR;
+
+ return PCI_P2PDMA_MAP_THRU_HOST_BRIDGE;
+}
+
+/*
+ * Name the ports that keep this path off a direct route, so that the admin
+ * can hand them to pci=disable_acs_redir=.
+ */
+static void pci_p2pdma_warn_path(struct pci_dev *client,
+ struct pci_dev *provider,
+ const struct pci_p2pdma_acs_path *path,
+ struct pci_dev *a_child,
+ struct pci_dev *b_child)
+{
+ struct seq_buf acs_list;
+ char buf[128];
+
+ if (path->unreadable) {
+ pci_warn(client,
+ "ACS Control is unreadable for provider %s at %s\n",
+ pci_name(provider), pci_name(path->unreadable));
+ return;
+ }
+
+ seq_buf_init(&acs_list, buf, sizeof(buf));
+ if (pci_acs_p2pdma_completion(path->cpl_ctrl) != PCI_ACS_P2PDMA_DIRECT)
+ seq_buf_print_bus_devfn(&acs_list, a_child);
+ if (pci_acs_p2pdma_request(path->req_ctrl) != PCI_ACS_P2PDMA_DIRECT)
+ seq_buf_print_bus_devfn(&acs_list, b_child);
+
+ /* Drop the final semicolon; the list is not empty here. */
+ if (!seq_buf_has_overflowed(&acs_list))
+ acs_list.buffer[acs_list.len - 1] = '\0';
+
+ pci_warn(client,
+ "ACS redirect is set between the client and provider (%s)\n",
+ pci_name(provider));
+ pci_warn(client,
+ "to disable ACS controls for this path, add the kernel parameter: pci=disable_acs_redir=%s\n",
+ seq_buf_str(&acs_list));
+}
+
static bool cpu_supports_p2pdma(void)
{
#ifdef CONFIG_X86
@@ -751,19 +825,13 @@ calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client,
int *dist, bool verbose)
{
enum pci_p2pdma_map_type map_type = PCI_P2PDMA_MAP_THRU_HOST_BRIDGE;
- enum pci_acs_p2pdma_state state = PCI_ACS_P2PDMA_NOT_SUPPORTED;
struct pci_dev *a = provider, *b = client, *bb;
struct pci_dev *a_child = NULL, *b_child = NULL;
- struct pci_dev *acs_unreadable = NULL;
+ struct pci_p2pdma_acs_path path = {};
struct pci_p2pdma *p2pdma;
- struct seq_buf acs_list;
- int acs_redirect_cnt = 0;
+ bool cpu_p2pdma, host_whitelisted = false;
int dist_a = 0;
int dist_b = 0;
- char buf[128];
- u16 ctrl;
-
- seq_buf_init(&acs_list, buf, sizeof(buf));
/*
* Note, we don't need to take references to devices returned by
@@ -806,61 +874,35 @@ calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client,
* branch is upstream, so redirect controls do not affect the path.
*/
if (a_child && b_child) {
- if (pci_acs_p2pdma_ctrl(a_child, &ctrl))
- state = pci_acs_p2pdma_completion(ctrl);
- if (state != PCI_ACS_P2PDMA_DIRECT) {
- seq_buf_print_bus_devfn(&acs_list, a_child);
- if (state == PCI_ACS_P2PDMA_REDIRECT)
- acs_redirect_cnt++;
- else if (!acs_unreadable)
- acs_unreadable = a_child;
- }
-
- state = PCI_ACS_P2PDMA_NOT_SUPPORTED;
- if (pci_acs_p2pdma_ctrl(b_child, &ctrl))
- state = pci_acs_p2pdma_request(ctrl);
- if (state != PCI_ACS_P2PDMA_DIRECT) {
- seq_buf_print_bus_devfn(&acs_list, b_child);
- if (state == PCI_ACS_P2PDMA_REDIRECT)
- acs_redirect_cnt++;
- else if (!acs_unreadable)
- acs_unreadable = b_child;
- }
+ if (!pci_acs_p2pdma_ctrl(a_child, &path.cpl_ctrl))
+ path.unreadable = a_child;
+ if (!pci_acs_p2pdma_ctrl(b_child, &path.req_ctrl) &&
+ !path.unreadable)
+ path.unreadable = b_child;
}
/*
* Below a shared upstream bridge, a path whose divergence ports do not
* redirect routes the request directly.
*/
- if (!acs_unreadable && !acs_redirect_cnt) {
- map_type = PCI_P2PDMA_MAP_BUS_ADDR;
+ map_type = pci_p2pdma_route(&path);
+ if (map_type == PCI_P2PDMA_MAP_BUS_ADDR)
goto done;
- }
- if (verbose) {
- /* Drop the final semicolon; the list is not empty here. */
- if (!seq_buf_has_overflowed(&acs_list))
- acs_list.buffer[acs_list.len - 1] = '\0';
- if (acs_unreadable)
- pci_warn(client, "ACS Control is unreadable for provider %s at %s\n",
- pci_name(provider), pci_name(acs_unreadable));
- else {
- pci_warn(client, "ACS redirect is set between the client and provider (%s)\n",
- pci_name(provider));
- pci_warn(client, "to disable ACS controls for this path, add the kernel parameter: pci=disable_acs_redir=%s\n",
- seq_buf_str(&acs_list));
- }
- }
+ if (verbose)
+ pci_p2pdma_warn_path(client, provider, &path, a_child, b_child);
/* An unreadable control does not establish an upstream redirect. */
- if (acs_unreadable) {
- map_type = PCI_P2PDMA_MAP_NOT_SUPPORTED;
+ if (path.unreadable)
goto done;
- }
map_through_host_bridge:
- if (!cpu_supports_p2pdma() &&
- !host_bridge_whitelist(provider, client, verbose)) {
+ cpu_p2pdma = cpu_supports_p2pdma();
+ if (!cpu_p2pdma)
+ host_whitelisted = host_bridge_whitelist(provider, client,
+ verbose);
+
+ if (!cpu_p2pdma && !host_whitelisted) {
if (verbose)
pci_warn(client, "cannot be used for peer-to-peer DMA as the client and provider (%s) do not share an upstream bridge or whitelisted host bridge\n",
pci_name(provider));
@@ -1193,8 +1235,9 @@ enum pci_p2pdma_map_type pci_p2pdma_map_type(struct p2pdma_provider *provider,
{
enum pci_p2pdma_map_type type = PCI_P2PDMA_MAP_NOT_SUPPORTED;
struct pci_dev *pdev = to_pci_dev(provider->owner);
- struct pci_dev *client;
struct pci_p2pdma *p2pdma;
+ unsigned long cache_index;
+ struct pci_dev *client;
int dist;
if (!pdev->p2pdma)
@@ -1204,13 +1247,14 @@ enum pci_p2pdma_map_type pci_p2pdma_map_type(struct p2pdma_provider *provider,
return PCI_P2PDMA_MAP_NOT_SUPPORTED;
client = to_pci_dev(dev);
+ cache_index = map_types_idx(client);
rcu_read_lock();
p2pdma = rcu_dereference(pdev->p2pdma);
if (p2pdma)
type = xa_to_value(xa_load(&p2pdma->map_types,
- map_types_idx(client)));
+ cache_index));
rcu_read_unlock();
if (type == PCI_P2PDMA_MAP_UNKNOWN)
--
2.55.0
^ permalink raw reply [flat|nested] 38+ messages in thread
* [PATCH v9 07/18] PCI/P2PDMA: Answer routing per TLP class
2026-10-01 11:55 [PATCH v9 00/18] PCI/P2PDMA: Route peer-to-peer DMA by TLP class Leon Romanovsky
` (5 preceding siblings ...)
2026-10-01 11:55 ` [PATCH v9 06/18] PCI/P2PDMA: Collect the path's ACS controls before deciding Leon Romanovsky
@ 2026-10-01 11:55 ` Leon Romanovsky
2026-10-01 11:55 ` [PATCH v9 08/18] PCI/P2PDMA: Route Relaxed Ordering Completions directly Leon Romanovsky
` (11 subsequent siblings)
18 siblings, 0 replies; 38+ messages in thread
From: Leon Romanovsky @ 2026-10-01 11:55 UTC (permalink / raw)
To: Bjorn Helgaas, Logan Gunthorpe, Leon Romanovsky, Jason Gunthorpe,
Joerg Roedel (AMD),
Will Deacon, Robin Murphy, Christian König,
Thomas Hellström
Cc: linux-pci, linux-kernel, linux-doc, iommu, Tushar Dave,
linux-media, dri-devel, linaro-mm-sig, linux-rdma, kvm,
Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Sumit Semwal
From: Leon Romanovsky <leonro@nvidia.com>
calc_map_type_and_dist() returns one mapping type per provider and client,
valid only for strictly ordered Requests carrying an Untranslated address.
Clients that use Relaxed Ordering or ATS cannot ask what the fabric would
do with their traffic.
Add enum pci_p2pdma_tlp_flags to name a class and pci_p2pdma_map_type_tlp()
to ask about one. The topology walk and the ACS Control reads do not depend
on the class, so decide all of them from the one walk and cache them
together, four bits each. Every class still answers alike; the controls
that tell them apart come next.
Tested-by: Tushar Dave <tdave@nvidia.com>
Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
---
drivers/pci/p2pdma.c | 159 +++++++++++++++++++++++++++++++++++++++++----------
1 file changed, 128 insertions(+), 31 deletions(-)
diff --git a/drivers/pci/p2pdma.c b/drivers/pci/p2pdma.c
index 841c86be31bb..43e225cc5735 100644
--- a/drivers/pci/p2pdma.c
+++ b/drivers/pci/p2pdma.c
@@ -492,6 +492,33 @@ static struct pci_dev *find_parent_pci_dev(struct device *dev)
return NULL;
}
+/**
+ * enum pci_p2pdma_tlp_flags - Properties of the TLPs a client will issue
+ *
+ * These describe the traffic rather than the topology, and select which ACS
+ * controls apply along the peer-to-peer path. A value of 0 means strictly
+ * ordered Requests carrying an Untranslated address.
+ *
+ * @PCI_P2PDMA_TLP_TRANSLATED: Requests carry an ATS Translated address. PCIe
+ * r7.0 sec 6.12.3 routes those to the peer regardless of ACS P2P Request
+ * Redirect and ACS P2P Egress Control wherever ACS Direct Translated P2P
+ * is enabled.
+ * @PCI_P2PDMA_TLP_RELAXED_CPL: The provider returns Completions with the
+ * Relaxed Ordering attribute set. PCIe r7.0 sec 6.12.1.1 never redirects
+ * those, so ACS P2P Completion Redirect does not gate the path. The
+ * Completer chooses this attribute and the specification does not require
+ * it to copy Relaxed Ordering from the Request into the Completion, so a
+ * caller passing this flag asserts that its provider does.
+ */
+enum pci_p2pdma_tlp_flags {
+ PCI_P2PDMA_TLP_TRANSLATED = 1 << 0,
+ PCI_P2PDMA_TLP_RELAXED_CPL = 1 << 1,
+};
+
+/* Every combination of the flags above selects one routing class. */
+#define PCI_P2PDMA_TLP_CLASSES \
+ ((PCI_P2PDMA_TLP_TRANSLATED | PCI_P2PDMA_TLP_RELAXED_CPL) + 1)
+
enum pci_acs_p2pdma_state {
PCI_ACS_P2PDMA_NOT_SUPPORTED,
PCI_ACS_P2PDMA_DIRECT,
@@ -508,7 +535,7 @@ enum pci_acs_p2pdma_state {
* bus addressing be assumed.
*/
static enum pci_acs_p2pdma_state
-pci_acs_p2pdma_request(u16 ctrl)
+pci_acs_p2pdma_request(u16 ctrl, unsigned int tlp_flags)
{
return ctrl & (PCI_ACS_RR | PCI_ACS_EC) ?
PCI_ACS_P2PDMA_REDIRECT : PCI_ACS_P2PDMA_DIRECT;
@@ -520,7 +547,7 @@ pci_acs_p2pdma_request(u16 ctrl)
* affects a Completion.
*/
static enum pci_acs_p2pdma_state
-pci_acs_p2pdma_completion(u16 ctrl)
+pci_acs_p2pdma_completion(u16 ctrl, unsigned int tlp_flags)
{
return ctrl & PCI_ACS_CR ? PCI_ACS_P2PDMA_REDIRECT :
PCI_ACS_P2PDMA_DIRECT;
@@ -577,13 +604,16 @@ struct pci_p2pdma_acs_path {
* the peer's bus addresses.
*/
static enum pci_p2pdma_map_type
-pci_p2pdma_route(const struct pci_p2pdma_acs_path *path)
+pci_p2pdma_route(const struct pci_p2pdma_acs_path *path,
+ unsigned int tlp_flags)
{
if (path->unreadable)
return PCI_P2PDMA_MAP_NOT_SUPPORTED;
- if (pci_acs_p2pdma_request(path->req_ctrl) == PCI_ACS_P2PDMA_DIRECT &&
- pci_acs_p2pdma_completion(path->cpl_ctrl) == PCI_ACS_P2PDMA_DIRECT)
+ if (pci_acs_p2pdma_request(path->req_ctrl, tlp_flags) ==
+ PCI_ACS_P2PDMA_DIRECT &&
+ pci_acs_p2pdma_completion(path->cpl_ctrl, tlp_flags) ==
+ PCI_ACS_P2PDMA_DIRECT)
return PCI_P2PDMA_MAP_BUS_ADDR;
return PCI_P2PDMA_MAP_THRU_HOST_BRIDGE;
@@ -591,7 +621,9 @@ pci_p2pdma_route(const struct pci_p2pdma_acs_path *path)
/*
* Name the ports that keep this path off a direct route, so that the admin
- * can hand them to pci=disable_acs_redir=.
+ * can hand them to pci=disable_acs_redir=. That parameter clears the
+ * redirect controls, which decide the default class, so that is the class
+ * this reports on.
*/
static void pci_p2pdma_warn_path(struct pci_dev *client,
struct pci_dev *provider,
@@ -610,9 +642,10 @@ static void pci_p2pdma_warn_path(struct pci_dev *client,
}
seq_buf_init(&acs_list, buf, sizeof(buf));
- if (pci_acs_p2pdma_completion(path->cpl_ctrl) != PCI_ACS_P2PDMA_DIRECT)
+ if (pci_acs_p2pdma_completion(path->cpl_ctrl, 0) !=
+ PCI_ACS_P2PDMA_DIRECT)
seq_buf_print_bus_devfn(&acs_list, a_child);
- if (pci_acs_p2pdma_request(path->req_ctrl) != PCI_ACS_P2PDMA_DIRECT)
+ if (pci_acs_p2pdma_request(path->req_ctrl, 0) != PCI_ACS_P2PDMA_DIRECT)
seq_buf_print_bus_devfn(&acs_list, b_child);
/* Drop the final semicolon; the list is not empty here. */
@@ -780,6 +813,31 @@ static unsigned long map_types_idx(struct pci_dev *client)
return (pci_domain_nr(client->bus) << 16) | pci_dev_id(client);
}
+/*
+ * One cache entry holds the routing of every TLP class, four bits each,
+ * indexed by the &enum pci_p2pdma_tlp_flags combination that selects it. An
+ * absent entry reads back as PCI_P2PDMA_MAP_UNKNOWN in every class.
+ */
+static_assert(PCI_P2PDMA_MAP_THRU_HOST_BRIDGE < 16);
+
+static unsigned long
+pci_p2pdma_map_types_pack(const enum pci_p2pdma_map_type *type)
+{
+ unsigned long val = 0;
+ unsigned int flags;
+
+ for (flags = 0; flags < PCI_P2PDMA_TLP_CLASSES; flags++)
+ val |= (unsigned long)type[flags] << (flags * 4);
+
+ return val;
+}
+
+static enum pci_p2pdma_map_type
+pci_p2pdma_map_types_unpack(unsigned long val, unsigned int tlp_flags)
+{
+ return (val >> (tlp_flags * 4)) & 0xf;
+}
+
/*
* Calculate the P2PDMA mapping type and distance between two PCI devices.
*
@@ -809,6 +867,10 @@ static unsigned long map_types_idx(struct pci_dev *client)
* check Request Redirect and Egress Control on the client-side port, and
* Completion Redirect for read Completions on the provider-side port.
*
+ * Those controls apply to different TLPs, so every class named by &enum
+ * pci_p2pdma_tlp_flags is decided from the one walk and cached together;
+ * @tlp_flags selects which one is returned.
+ *
* If ACS redirects traffic at either divergence port, return
* PCI_P2PDMA_MAP_THRU_HOST_BRIDGE. If the ACS Control register cannot be
* read, return PCI_P2PDMA_MAP_NOT_SUPPORTED. Otherwise, return
@@ -822,14 +884,16 @@ static unsigned long map_types_idx(struct pci_dev *client)
*/
static enum pci_p2pdma_map_type
calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client,
- int *dist, bool verbose)
+ int *dist, unsigned int tlp_flags, bool verbose)
{
- enum pci_p2pdma_map_type map_type = PCI_P2PDMA_MAP_THRU_HOST_BRIDGE;
+ enum pci_p2pdma_map_type map_type[PCI_P2PDMA_TLP_CLASSES];
struct pci_dev *a = provider, *b = client, *bb;
struct pci_dev *a_child = NULL, *b_child = NULL;
struct pci_p2pdma_acs_path path = {};
struct pci_p2pdma *p2pdma;
bool cpu_p2pdma, host_whitelisted = false;
+ bool host_fallback = false;
+ unsigned int flags;
int dist_a = 0;
int dist_b = 0;
@@ -863,6 +927,8 @@ calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client,
* request can only get to the peer through the host bridge.
*/
*dist = dist_a + dist_b;
+ for (flags = 0; flags < PCI_P2PDMA_TLP_CLASSES; flags++)
+ map_type[flags] = PCI_P2PDMA_MAP_THRU_HOST_BRIDGE;
goto map_through_host_bridge;
check_paths_acs:
@@ -882,18 +948,24 @@ calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client,
}
/*
- * Below a shared upstream bridge, a path whose divergence ports do not
- * redirect routes the request directly.
+ * The walk and the config reads above serve every class; only the
+ * decision below depends on the kind of TLP being routed.
*/
- map_type = pci_p2pdma_route(&path);
- if (map_type == PCI_P2PDMA_MAP_BUS_ADDR)
- goto done;
+ for (flags = 0; flags < PCI_P2PDMA_TLP_CLASSES; flags++) {
+ map_type[flags] = pci_p2pdma_route(&path, flags);
+ if (map_type[flags] == PCI_P2PDMA_MAP_THRU_HOST_BRIDGE)
+ host_fallback = true;
+ }
- if (verbose)
- pci_p2pdma_warn_path(client, provider, &path, a_child, b_child);
+ if (verbose && map_type[0] != PCI_P2PDMA_MAP_BUS_ADDR)
+ pci_p2pdma_warn_path(client, provider, &path, a_child,
+ b_child);
- /* An unreadable control does not establish an upstream redirect. */
- if (path.unreadable)
+ /*
+ * Nothing needs the host bridge: the classes that did not get a direct
+ * route have no fallback that would use it.
+ */
+ if (!host_fallback)
goto done;
map_through_host_bridge:
@@ -906,16 +978,19 @@ calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client,
if (verbose)
pci_warn(client, "cannot be used for peer-to-peer DMA as the client and provider (%s) do not share an upstream bridge or whitelisted host bridge\n",
pci_name(provider));
- map_type = PCI_P2PDMA_MAP_NOT_SUPPORTED;
+ for (flags = 0; flags < PCI_P2PDMA_TLP_CLASSES; flags++)
+ if (map_type[flags] == PCI_P2PDMA_MAP_THRU_HOST_BRIDGE)
+ map_type[flags] = PCI_P2PDMA_MAP_NOT_SUPPORTED;
}
done:
rcu_read_lock();
p2pdma = rcu_dereference(provider->p2pdma);
if (p2pdma)
xa_store(&p2pdma->map_types, map_types_idx(client),
- xa_mk_value(map_type), GFP_ATOMIC);
+ xa_mk_value(pci_p2pdma_map_types_pack(map_type)),
+ GFP_ATOMIC);
rcu_read_unlock();
- return map_type;
+ return map_type[tlp_flags];
}
/**
@@ -956,7 +1031,7 @@ int pci_p2pdma_distance_many(struct pci_dev *provider, struct device **clients,
return -1;
}
- map = calc_map_type_and_dist(provider, pci_client, &distance,
+ map = calc_map_type_and_dist(provider, pci_client, &distance, 0,
verbose);
pci_dev_put(pci_client);
@@ -1221,22 +1296,28 @@ void pci_p2pmem_publish(struct pci_dev *pdev, bool publish)
EXPORT_SYMBOL_GPL(pci_p2pmem_publish);
/**
- * pci_p2pdma_map_type - Determine the mapping type for P2PDMA transfers
+ * pci_p2pdma_map_type_tlp - Determine the mapping type for P2PDMA transfers
* @provider: P2PDMA provider structure
* @dev: Client device that initiates the transfer
+ * @tlp_flags: &enum pci_p2pdma_tlp_flags describing the TLPs @dev will issue
*
* Determines how peer-to-peer DMA transfers should be mapped between
* the provider and the client device. The mapping type indicates whether
* the transfer can be done directly through PCI switches or must go
* through the host bridge.
+ *
+ * ACS routes a peer-to-peer transaction by the attributes its TLPs carry, so
+ * the answer depends on @tlp_flags. A caller that passes flags its traffic
+ * does not match gets a mapping the fabric will not deliver.
*/
-enum pci_p2pdma_map_type pci_p2pdma_map_type(struct p2pdma_provider *provider,
- struct device *dev)
+static enum pci_p2pdma_map_type
+pci_p2pdma_map_type_tlp(struct p2pdma_provider *provider, struct device *dev,
+ unsigned int tlp_flags)
{
- enum pci_p2pdma_map_type type = PCI_P2PDMA_MAP_NOT_SUPPORTED;
struct pci_dev *pdev = to_pci_dev(provider->owner);
+ unsigned long cache_index, cached = 0;
+ enum pci_p2pdma_map_type type;
struct pci_p2pdma *p2pdma;
- unsigned long cache_index;
struct pci_dev *client;
int dist;
@@ -1253,16 +1334,32 @@ enum pci_p2pdma_map_type pci_p2pdma_map_type(struct p2pdma_provider *provider,
p2pdma = rcu_dereference(pdev->p2pdma);
if (p2pdma)
- type = xa_to_value(xa_load(&p2pdma->map_types,
- cache_index));
+ cached = xa_to_value(xa_load(&p2pdma->map_types,
+ cache_index));
rcu_read_unlock();
+ type = pci_p2pdma_map_types_unpack(cached, tlp_flags);
if (type == PCI_P2PDMA_MAP_UNKNOWN)
- return calc_map_type_and_dist(pdev, client, &dist, true);
+ return calc_map_type_and_dist(pdev, client, &dist, tlp_flags,
+ true);
return type;
}
+/**
+ * pci_p2pdma_map_type - Determine the mapping type for P2PDMA transfers
+ * @provider: P2PDMA provider structure
+ * @dev: Client device that initiates the transfer
+ *
+ * Same as pci_p2pdma_map_type_tlp() for a client issuing strictly ordered
+ * Requests that carry an Untranslated address.
+ */
+enum pci_p2pdma_map_type pci_p2pdma_map_type(struct p2pdma_provider *provider,
+ struct device *dev)
+{
+ return pci_p2pdma_map_type_tlp(provider, dev, 0);
+}
+
void __pci_p2pdma_update_state(struct pci_p2pdma_map_state *state,
struct device *dev, struct page *page)
{
--
2.55.0
^ permalink raw reply [flat|nested] 38+ messages in thread
* [PATCH v9 08/18] PCI/P2PDMA: Route Relaxed Ordering Completions directly
2026-10-01 11:55 [PATCH v9 00/18] PCI/P2PDMA: Route peer-to-peer DMA by TLP class Leon Romanovsky
` (6 preceding siblings ...)
2026-10-01 11:55 ` [PATCH v9 07/18] PCI/P2PDMA: Answer routing per TLP class Leon Romanovsky
@ 2026-10-01 11:55 ` Leon Romanovsky
2026-10-06 22:21 ` Bjorn Helgaas
2026-10-01 11:55 ` [PATCH v9 09/18] PCI/P2PDMA: Reject Translated Requests blocked by Translation Blocking Leon Romanovsky
` (10 subsequent siblings)
18 siblings, 1 reply; 38+ messages in thread
From: Leon Romanovsky @ 2026-10-01 11:55 UTC (permalink / raw)
To: Bjorn Helgaas, Logan Gunthorpe, Leon Romanovsky, Jason Gunthorpe,
Joerg Roedel (AMD),
Will Deacon, Robin Murphy, Christian König,
Thomas Hellström
Cc: linux-pci, linux-kernel, linux-doc, iommu, Tushar Dave,
linux-media, dri-devel, linaro-mm-sig, linux-rdma, kvm,
Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Sumit Semwal
From: Leon Romanovsky <leonro@nvidia.com>
ACS P2P Completion Redirect leaves Completions carrying the Relaxed
Ordering attribute alone. PCIe r7.0 sec 6.12.1.1 redirects only those "that
do not have the Relaxed Ordering Attribute bit set", and sec 7.7.12.5
describes the enable bit as "applicable only to Completions whose Relaxed
Ordering Attribute is clear". P2PDMA reports one answer for every kind of
TLP, so a client whose provider returns such Completions is sent through
the host bridge for a redirect that never happens to it.
Add enum pci_p2pdma_tlp_flags and let a caller state that property.
Reviewed-by: Logan Gunthorpe <logang@deltatee.com>
Tested-by: Tushar Dave <tdave@nvidia.com>
Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
---
drivers/pci/p2pdma.c | 6 +++++-
1 file changed, 5 insertions(+), 1 deletion(-)
diff --git a/drivers/pci/p2pdma.c b/drivers/pci/p2pdma.c
index 43e225cc5735..1fadef6d0609 100644
--- a/drivers/pci/p2pdma.c
+++ b/drivers/pci/p2pdma.c
@@ -544,11 +544,15 @@ pci_acs_p2pdma_request(u16 ctrl, unsigned int tlp_flags)
/*
* Decide how a peer-to-peer Completion at an ACS-capable ingress port routes.
* PCIe r7.0 sec 6.12.1.1: no ACS control other than P2P Completion Redirect
- * affects a Completion.
+ * affects a Completion, and that one leaves Completions carrying the Relaxed
+ * Ordering attribute alone.
*/
static enum pci_acs_p2pdma_state
pci_acs_p2pdma_completion(u16 ctrl, unsigned int tlp_flags)
{
+ if (tlp_flags & PCI_P2PDMA_TLP_RELAXED_CPL)
+ return PCI_ACS_P2PDMA_DIRECT;
+
return ctrl & PCI_ACS_CR ? PCI_ACS_P2PDMA_REDIRECT :
PCI_ACS_P2PDMA_DIRECT;
}
--
2.55.0
^ permalink raw reply [flat|nested] 38+ messages in thread
* [PATCH v9 09/18] PCI/P2PDMA: Reject Translated Requests blocked by Translation Blocking
2026-10-01 11:55 [PATCH v9 00/18] PCI/P2PDMA: Route peer-to-peer DMA by TLP class Leon Romanovsky
` (7 preceding siblings ...)
2026-10-01 11:55 ` [PATCH v9 08/18] PCI/P2PDMA: Route Relaxed Ordering Completions directly Leon Romanovsky
@ 2026-10-01 11:55 ` Leon Romanovsky
2026-10-01 11:55 ` [PATCH v9 10/18] PCI/P2PDMA: Route Translated Requests under Direct Translated P2P Leon Romanovsky
` (9 subsequent siblings)
18 siblings, 0 replies; 38+ messages in thread
From: Leon Romanovsky @ 2026-10-01 11:55 UTC (permalink / raw)
To: Bjorn Helgaas, Logan Gunthorpe, Leon Romanovsky, Jason Gunthorpe,
Joerg Roedel (AMD),
Will Deacon, Robin Murphy, Christian König,
Thomas Hellström
Cc: linux-pci, linux-kernel, linux-doc, iommu, Tushar Dave,
linux-media, dri-devel, linaro-mm-sig, linux-rdma, kvm,
Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Sumit Semwal
From: Leon Romanovsky <leonro@nvidia.com>
A Downstream Port with ACS Translation Blocking enabled treats every
Upstream Memory Request whose Address Type is not Untranslated as an ACS
Violation, ahead of "any applicable ACS P2P control mechanisms" per PCIe
r7.0 sec 6.12.1.1. P2PDMA never looks at that bit, so it reports a
bus-addressable path where an ATS client's Requests would be rejected.
Add PCI_ACS_P2PDMA_BLOCKED, and because blocking is not a routing control,
scan the path for it rather than the divergence port alone. A blocked
Request has no host bridge fallback, since the Address Type is rejected
wherever the Request is addressed.
The two routes do not pass the same ports, so scan them separately. A
direct route turns around at the divergence and only passes the ports
below it, while the host bridge route keeps climbing and passes that port
and everything above it as well.
Reviewed-by: Logan Gunthorpe <logang@deltatee.com>
Tested-by: Tushar Dave <tdave@nvidia.com>
Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
---
drivers/pci/p2pdma.c | 120 +++++++++++++++++++++++++++++++++++++++++++++++----
1 file changed, 112 insertions(+), 8 deletions(-)
diff --git a/drivers/pci/p2pdma.c b/drivers/pci/p2pdma.c
index 1fadef6d0609..569a74de3b3a 100644
--- a/drivers/pci/p2pdma.c
+++ b/drivers/pci/p2pdma.c
@@ -523,11 +523,12 @@ enum pci_acs_p2pdma_state {
PCI_ACS_P2PDMA_NOT_SUPPORTED,
PCI_ACS_P2PDMA_DIRECT,
PCI_ACS_P2PDMA_REDIRECT,
+ PCI_ACS_P2PDMA_BLOCKED,
};
/*
* Decide how a peer-to-peer Request at an ACS-capable ingress port routes,
- * from that port's ACS Control register.
+ * from that port's ACS Control register and the Request's Address Type.
*
* Linux does not read the Egress Control Vector, so Egress Control is treated
* conservatively as a redirect. Per PCIe r7.0 Table 6-11 the outcomes it
@@ -537,6 +538,18 @@ enum pci_acs_p2pdma_state {
static enum pci_acs_p2pdma_state
pci_acs_p2pdma_request(u16 ctrl, unsigned int tlp_flags)
{
+ if (tlp_flags & PCI_P2PDMA_TLP_TRANSLATED) {
+ /*
+ * PCIe r7.0 sec 6.12.1.1: Translation Blocking makes every
+ * Upstream Memory Request whose Address Type is not
+ * Untranslated an ACS Violation, taking precedence over the
+ * P2P controls. Sec 7.7.12.5: Direct Translated P2P "is
+ * ignored if ACS Translation Blocking Enable is 1b".
+ */
+ if (ctrl & PCI_ACS_TB)
+ return PCI_ACS_P2PDMA_BLOCKED;
+ }
+
return ctrl & (PCI_ACS_RR | PCI_ACS_EC) ?
PCI_ACS_P2PDMA_REDIRECT : PCI_ACS_P2PDMA_DIRECT;
}
@@ -576,6 +589,39 @@ static bool pci_acs_p2pdma_ctrl(struct pci_dev *pdev, u16 *ctrl)
return !pci_read_config_word(pdev, pos + PCI_ACS_CTRL, ctrl);
}
+/*
+ * Report whether any port between @client and @common rejects Translated
+ * addresses. If @common is NULL, every port up to the host bridge is walked.
+ * A port whose ACS Control cannot be read counts as blocking, which withdraws
+ * only the Translated classes because an Untranslated Request is routed by the
+ * redirect controls instead.
+ */
+static bool pci_p2pdma_path_blocks_translation(struct pci_dev *client,
+ struct pci_dev *common)
+{
+ struct pci_dev *pdev;
+ u16 ctrl;
+
+ /*
+ * @common is @client itself when the provider is the client or sits
+ * below it. The Request never travels upstream then, so no port sees
+ * it and none can reject its Address Type.
+ */
+ if (client == common)
+ return false;
+
+ for (pdev = pci_upstream_bridge(client); pdev && pdev != common;
+ pdev = pci_upstream_bridge(pdev)) {
+ if (!pci_acs_p2pdma_ctrl(pdev, &ctrl))
+ return true;
+
+ if (ctrl & PCI_ACS_TB)
+ return true;
+ }
+
+ return false;
+}
+
static void seq_buf_print_bus_devfn(struct seq_buf *buf, struct pci_dev *pdev)
{
if (!buf)
@@ -594,14 +640,42 @@ static void seq_buf_print_bus_devfn(struct seq_buf *buf, struct pci_dev *pdev)
* than upstream, so it is where the Request controls apply.
* @cpl_ctrl: ACS Control of the provider-side divergence port, likewise for
* the Completions travelling back.
+ * @tb_on_path: A port below the divergence blocks Translated addresses.
+ * Every route out of @client passes those, so none carries them.
+ * @tb_above_divergence: A port at or above the divergence blocks Translated
+ * addresses. Only a Request continuing to the host bridge passes
+ * those, so a direct route is still open to them.
+ * @no_common_bridge: The two paths share no upstream bridge, so no direct
+ * route exists for ACS to gate.
* @unreadable: First port whose ACS Control could not be read, if any.
*/
struct pci_p2pdma_acs_path {
u16 req_ctrl;
u16 cpl_ctrl;
+ bool tb_on_path;
+ bool tb_above_divergence;
+ bool no_common_bridge;
struct pci_dev *unreadable;
};
+/*
+ * ACS Translation Blocking is not a routing control, so unlike the redirect
+ * controls it is not decided at the divergence alone. PCIe r7.0 sec 6.12.1.1
+ * has every Downstream Port check the Address Type of each Upstream Memory
+ * Request it receives, ahead of "any applicable ACS P2P control mechanisms".
+ * A port below the divergence cannot redirect the Request anywhere it was not
+ * already going, but it can still reject a Translated address.
+ */
+static enum pci_acs_p2pdma_state
+pci_p2pdma_request_state(const struct pci_p2pdma_acs_path *path,
+ unsigned int tlp_flags)
+{
+ if (tlp_flags & PCI_P2PDMA_TLP_TRANSLATED && path->tb_on_path)
+ return PCI_ACS_P2PDMA_BLOCKED;
+
+ return pci_acs_p2pdma_request(path->req_ctrl, tlp_flags);
+}
+
/*
* Combine both directions into a mapping type. Only a path that routes the
* Request and the Completions it generates directly can be programmed with
@@ -611,15 +685,35 @@ static enum pci_p2pdma_map_type
pci_p2pdma_route(const struct pci_p2pdma_acs_path *path,
unsigned int tlp_flags)
{
+ enum pci_acs_p2pdma_state req;
+
if (path->unreadable)
return PCI_P2PDMA_MAP_NOT_SUPPORTED;
- if (pci_acs_p2pdma_request(path->req_ctrl, tlp_flags) ==
- PCI_ACS_P2PDMA_DIRECT &&
+ req = pci_p2pdma_request_state(path, tlp_flags);
+
+ /*
+ * Translation Blocking rejects the Address Type rather than the
+ * target, so a blocked Request stays blocked however it is addressed.
+ * No host bridge fallback keeps a Translated address working; the
+ * caller has to issue a different kind of Request instead.
+ */
+ if (req == PCI_ACS_P2PDMA_BLOCKED)
+ return PCI_P2PDMA_MAP_NOT_SUPPORTED;
+
+ if (!path->no_common_bridge && req == PCI_ACS_P2PDMA_DIRECT &&
pci_acs_p2pdma_completion(path->cpl_ctrl, tlp_flags) ==
PCI_ACS_P2PDMA_DIRECT)
return PCI_P2PDMA_MAP_BUS_ADDR;
+ /*
+ * A Request that turns around at the divergence never reaches the
+ * ports above it, but one that keeps climbing to the host bridge
+ * does, so that route has to clear their Translation Blocking too.
+ */
+ if (tlp_flags & PCI_P2PDMA_TLP_TRANSLATED && path->tb_above_divergence)
+ return PCI_P2PDMA_MAP_NOT_SUPPORTED;
+
return PCI_P2PDMA_MAP_THRU_HOST_BRIDGE;
}
@@ -868,8 +962,11 @@ pci_p2pdma_map_types_unpack(unsigned long val, unsigned int tlp_flags)
* redirection setting of the ports along the path.
*
* The client initiates Requests to provider memory. At the path divergence,
- * check Request Redirect and Egress Control on the client-side port, and
- * Completion Redirect for read Completions on the provider-side port.
+ * check Request Redirect, Egress Control, Translation Blocking and Direct
+ * Translated P2P on the client-side port, and Completion Redirect for read
+ * Completions on the provider-side port. Translation Blocking is checked on
+ * every client-side port instead, because it rejects a Request rather than
+ * routing it.
*
* Those controls apply to different TLPs, so every class named by &enum
* pci_p2pdma_tlp_flags is decided from the one walk and cached together;
@@ -877,8 +974,8 @@ pci_p2pdma_map_types_unpack(unsigned long val, unsigned int tlp_flags)
*
* If ACS redirects traffic at either divergence port, return
* PCI_P2PDMA_MAP_THRU_HOST_BRIDGE. If the ACS Control register cannot be
- * read, return PCI_P2PDMA_MAP_NOT_SUPPORTED. Otherwise, return
- * PCI_P2PDMA_MAP_BUS_ADDR.
+ * read, or Translation Blocking rejects the class being asked about, return
+ * PCI_P2PDMA_MAP_NOT_SUPPORTED. Otherwise, return PCI_P2PDMA_MAP_BUS_ADDR.
*
* Any two devices that have a data path that goes through the host bridge
* will consult a whitelist. If the host bridge is in the whitelist, return
@@ -931,8 +1028,10 @@ calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client,
* request can only get to the peer through the host bridge.
*/
*dist = dist_a + dist_b;
+ path.no_common_bridge = true;
+ path.tb_on_path = pci_p2pdma_path_blocks_translation(client, NULL);
for (flags = 0; flags < PCI_P2PDMA_TLP_CLASSES; flags++)
- map_type[flags] = PCI_P2PDMA_MAP_THRU_HOST_BRIDGE;
+ map_type[flags] = pci_p2pdma_route(&path, flags);
goto map_through_host_bridge;
check_paths_acs:
@@ -951,6 +1050,11 @@ calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client,
path.unreadable = b_child;
}
+ path.tb_on_path = pci_p2pdma_path_blocks_translation(client, a);
+ if (b_child)
+ path.tb_above_divergence =
+ pci_p2pdma_path_blocks_translation(b_child, NULL);
+
/*
* The walk and the config reads above serve every class; only the
* decision below depends on the kind of TLP being routed.
--
2.55.0
^ permalink raw reply [flat|nested] 38+ messages in thread
* [PATCH v9 10/18] PCI/P2PDMA: Route Translated Requests under Direct Translated P2P
2026-10-01 11:55 [PATCH v9 00/18] PCI/P2PDMA: Route peer-to-peer DMA by TLP class Leon Romanovsky
` (8 preceding siblings ...)
2026-10-01 11:55 ` [PATCH v9 09/18] PCI/P2PDMA: Reject Translated Requests blocked by Translation Blocking Leon Romanovsky
@ 2026-10-01 11:55 ` Leon Romanovsky
2026-10-06 22:28 ` Bjorn Helgaas
2026-10-01 11:55 ` [PATCH v9 11/18] PCI/P2PDMA: Log detailed ACS routing diagnostics Leon Romanovsky
` (8 subsequent siblings)
18 siblings, 1 reply; 38+ messages in thread
From: Leon Romanovsky @ 2026-10-01 11:55 UTC (permalink / raw)
To: Bjorn Helgaas, Logan Gunthorpe, Leon Romanovsky, Jason Gunthorpe,
Joerg Roedel (AMD),
Will Deacon, Robin Murphy, Christian König,
Thomas Hellström
Cc: linux-pci, linux-kernel, linux-doc, iommu, Tushar Dave,
linux-media, dri-devel, linaro-mm-sig, linux-rdma, kvm,
Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Sumit Semwal
From: Leon Romanovsky <leonro@nvidia.com>
A Downstream Port with ACS Direct Translated P2P enabled routes a Request
whose Address Type is Translated "to the peer Egress Port without
redirection, regardless of ACS P2P Request Redirect and ACS P2P Egress
Control", per PCIe r7.0 sec 6.12.3. P2PDMA assumes every Request carries an
Untranslated address, so it sends an ATS client through the host bridge
even where the fabric would route it straight to the peer.
Add PCI_P2PDMA_TLP_TRANSLATED and consult Direct Translated P2P for the
Requests it describes.
Reviewed-by: Logan Gunthorpe <logang@deltatee.com>
Tested-by: Tushar Dave <tdave@nvidia.com>
Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
---
drivers/pci/p2pdma.c | 9 +++++++++
1 file changed, 9 insertions(+)
diff --git a/drivers/pci/p2pdma.c b/drivers/pci/p2pdma.c
index 569a74de3b3a..3fd2cb8d16f0 100644
--- a/drivers/pci/p2pdma.c
+++ b/drivers/pci/p2pdma.c
@@ -548,6 +548,15 @@ pci_acs_p2pdma_request(u16 ctrl, unsigned int tlp_flags)
*/
if (ctrl & PCI_ACS_TB)
return PCI_ACS_P2PDMA_BLOCKED;
+
+ /*
+ * PCIe r7.0 sec 6.12.3: ACS Direct Translated P2P routes a
+ * Request carrying a Translated address to the peer "without
+ * redirection, regardless of ACS P2P Request Redirect and ACS
+ * P2P Egress Control settings".
+ */
+ if (ctrl & PCI_ACS_DT)
+ return PCI_ACS_P2PDMA_DIRECT;
}
return ctrl & (PCI_ACS_RR | PCI_ACS_EC) ?
--
2.55.0
^ permalink raw reply [flat|nested] 38+ messages in thread
* [PATCH v9 11/18] PCI/P2PDMA: Log detailed ACS routing diagnostics
2026-10-01 11:55 [PATCH v9 00/18] PCI/P2PDMA: Route peer-to-peer DMA by TLP class Leon Romanovsky
` (9 preceding siblings ...)
2026-10-01 11:55 ` [PATCH v9 10/18] PCI/P2PDMA: Route Translated Requests under Direct Translated P2P Leon Romanovsky
@ 2026-10-01 11:55 ` Leon Romanovsky
2026-10-01 11:55 ` [PATCH v9 12/18] PCI/P2PDMA: Add KUnit tests for the ACS routing decisions Leon Romanovsky
` (7 subsequent siblings)
18 siblings, 0 replies; 38+ messages in thread
From: Leon Romanovsky @ 2026-10-01 11:55 UTC (permalink / raw)
To: Bjorn Helgaas, Logan Gunthorpe, Leon Romanovsky, Jason Gunthorpe,
Joerg Roedel (AMD),
Will Deacon, Robin Murphy, Christian König,
Thomas Hellström
Cc: linux-pci, linux-kernel, linux-doc, iommu, Tushar Dave,
linux-media, dri-devel, linaro-mm-sig, linux-rdma, kvm,
Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Sumit Semwal
From: Leon Romanovsky <leonro@nvidia.com>
When P2PDMA rejects a mapping, existing warnings identify only the final
ACS or host-bridge result. They omit topology, live controls, divergence
ports, cache state, and intermediate routing decisions.
Emit debug-level messages for verbose calculations and cache lookups.
Report both paths, decoded ACS controls, directional decisions, host
fallback, and the final mapping. This keeps incidental unsupported probes
quiet while allowing the diagnostics to be enabled when needed.
Reviewed-by: Logan Gunthorpe <logang@deltatee.com>
Tested-by: Tushar Dave <tdave@nvidia.com>
Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
---
drivers/pci/p2pdma.c | 232 +++++++++++++++++++++++++++++++++++++++++++++++----
1 file changed, 215 insertions(+), 17 deletions(-)
diff --git a/drivers/pci/p2pdma.c b/drivers/pci/p2pdma.c
index 3fd2cb8d16f0..cef486e22f97 100644
--- a/drivers/pci/p2pdma.c
+++ b/drivers/pci/p2pdma.c
@@ -579,23 +579,128 @@ pci_acs_p2pdma_completion(u16 ctrl, unsigned int tlp_flags)
PCI_ACS_P2PDMA_DIRECT;
}
+static const char *pci_acs_p2pdma_state_name(enum pci_acs_p2pdma_state state)
+{
+ switch (state) {
+ case PCI_ACS_P2PDMA_DIRECT:
+ return "direct";
+ case PCI_ACS_P2PDMA_REDIRECT:
+ return "redirect";
+ case PCI_ACS_P2PDMA_BLOCKED:
+ return "blocked";
+ case PCI_ACS_P2PDMA_NOT_SUPPORTED:
+ return "not-supported";
+ }
+
+ return "invalid";
+}
+
+static const char *pci_p2pdma_map_type_name(enum pci_p2pdma_map_type type)
+{
+ switch (type) {
+ case PCI_P2PDMA_MAP_UNKNOWN:
+ return "unknown";
+ case PCI_P2PDMA_MAP_NONE:
+ return "none";
+ case PCI_P2PDMA_MAP_NOT_SUPPORTED:
+ return "not-supported";
+ case PCI_P2PDMA_MAP_BUS_ADDR:
+ return "bus-address";
+ case PCI_P2PDMA_MAP_THRU_HOST_BRIDGE:
+ return "through-host-bridge";
+ }
+
+ return "invalid";
+}
+
/*
* Read @pdev's ACS Control register. A device without an ACS capability has
* no peer-to-peer controls at all, which routes the same as having them all
* clear. Returns false when the register is present but cannot be read; @ctrl
* is then meaningless.
*/
-static bool pci_acs_p2pdma_ctrl(struct pci_dev *pdev, u16 *ctrl)
+static bool pci_acs_p2pdma_ctrl(struct pci_dev *pdev, const char *what,
+ u16 *ctrl, bool verbose)
{
- int pos;
+ int pos, ret;
pos = pdev->acs_cap;
if (!pos) {
+ if (verbose)
+ pci_dbg(pdev,
+ "P2PDMA ACS: %s has no ACS capability\n", what);
*ctrl = 0;
return true;
}
- return !pci_read_config_word(pdev, pos + PCI_ACS_CTRL, ctrl);
+ ret = pci_read_config_word(pdev, pos + PCI_ACS_CTRL, ctrl);
+ if (ret) {
+ if (verbose)
+ pci_dbg(pdev,
+ "P2PDMA ACS: %s ACS Control read failed at %#x: %#x\n",
+ what, pos + PCI_ACS_CTRL, ret);
+ return false;
+ }
+
+ if (verbose) {
+ pci_dbg(pdev,
+ "P2PDMA ACS: %s cap=%#x caps=%#06x ctrl=%#06x\n",
+ what, pos, pdev->acs_capabilities, *ctrl);
+ pci_dbg(pdev,
+ "P2PDMA ACS: control bits SV=%u TB=%u RR=%u CR=%u UF=%u EC=%u DT=%u\n",
+ !!(*ctrl & PCI_ACS_SV), !!(*ctrl & PCI_ACS_TB),
+ !!(*ctrl & PCI_ACS_RR), !!(*ctrl & PCI_ACS_CR),
+ !!(*ctrl & PCI_ACS_UF), !!(*ctrl & PCI_ACS_EC),
+ !!(*ctrl & PCI_ACS_DT));
+ }
+
+ return true;
+}
+
+static void pci_p2pdma_log_path(const char *name, struct pci_dev *start,
+ struct pci_dev *common)
+{
+ struct pci_dev *pdev, *upstream;
+ int hop = 0, ret, type;
+ u16 ctrl;
+
+ for (pdev = start; pdev; pdev = upstream, hop++) {
+ upstream = pci_upstream_bridge(pdev);
+ type = pci_is_pcie(pdev) ? pci_pcie_type(pdev) : -1;
+ pci_dbg(pdev,
+ "P2PDMA ACS: %s path hop=%d common=%u pcie=%u type=%d class=%#08x vendor=%04x device=%04x upstream=%s\n",
+ name, hop, pdev == common, pci_is_pcie(pdev), type,
+ pdev->class, pdev->vendor, pdev->device,
+ upstream ? pci_name(upstream) : "<none>");
+
+ if (pdev->subordinate)
+ pci_dbg(pdev,
+ "P2PDMA ACS: bridge bus range=%02llx-%02llx\n",
+ (unsigned long long)pdev->subordinate->busn_res.start,
+ (unsigned long long)pdev->subordinate->busn_res.end);
+
+ if (!pdev->acs_cap) {
+ pci_dbg(pdev, "P2PDMA ACS: ACS capability absent\n");
+ continue;
+ }
+
+ ret = pci_read_config_word(pdev, pdev->acs_cap + PCI_ACS_CTRL,
+ &ctrl);
+ if (ret) {
+ pci_dbg(pdev,
+ "P2PDMA ACS: ACS cap=%#x caps=%#06x Control read failed: %#x\n",
+ pdev->acs_cap, pdev->acs_capabilities, ret);
+ continue;
+ }
+
+ pci_dbg(pdev,
+ "P2PDMA ACS: ACS cap=%#x caps=%#06x ctrl=%#06x SV=%u TB=%u RR=%u CR=%u UF=%u EC=%u DT=%u\n",
+ pdev->acs_cap, pdev->acs_capabilities, ctrl,
+ !!(ctrl & PCI_ACS_SV), !!(ctrl & PCI_ACS_TB),
+ !!(ctrl & PCI_ACS_RR), !!(ctrl & PCI_ACS_CR),
+ !!(ctrl & PCI_ACS_UF), !!(ctrl & PCI_ACS_EC),
+ !!(ctrl & PCI_ACS_DT));
+ }
}
/*
@@ -606,7 +711,8 @@ static bool pci_acs_p2pdma_ctrl(struct pci_dev *pdev, u16 *ctrl)
* redirect controls instead.
*/
static bool pci_p2pdma_path_blocks_translation(struct pci_dev *client,
- struct pci_dev *common)
+ struct pci_dev *common,
+ bool verbose)
{
struct pci_dev *pdev;
u16 ctrl;
@@ -621,11 +727,15 @@ static bool pci_p2pdma_path_blocks_translation(struct pci_dev *client,
for (pdev = pci_upstream_bridge(client); pdev && pdev != common;
pdev = pci_upstream_bridge(pdev)) {
- if (!pci_acs_p2pdma_ctrl(pdev, &ctrl))
+ if (!pci_acs_p2pdma_ctrl(pdev, "path hop", &ctrl, verbose))
return true;
- if (ctrl & PCI_ACS_TB)
+ if (ctrl & PCI_ACS_TB) {
+ if (verbose)
+ pci_dbg(pdev,
+ "P2PDMA ACS: Translation Blocking rejects Translated Requests on this path\n");
return true;
+ }
}
return false;
@@ -999,14 +1109,22 @@ calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client,
enum pci_p2pdma_map_type map_type[PCI_P2PDMA_TLP_CLASSES];
struct pci_dev *a = provider, *b = client, *bb;
struct pci_dev *a_child = NULL, *b_child = NULL;
+ struct pci_host_bridge *provider_host, *client_host;
struct pci_p2pdma_acs_path path = {};
struct pci_p2pdma *p2pdma;
bool cpu_p2pdma, host_whitelisted = false;
+ bool cache_store = false;
bool host_fallback = false;
unsigned int flags;
int dist_a = 0;
int dist_b = 0;
+ if (verbose)
+ pci_dbg(client,
+ "P2PDMA ACS: begin provider=%s client=%s cache-index=%#lx\n",
+ pci_name(provider), pci_name(client),
+ map_types_idx(client));
+
/*
* Note, we don't need to take references to devices returned by
* pci_upstream_bridge() seeing we hold a reference to a child
@@ -1037,14 +1155,32 @@ calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client,
* request can only get to the peer through the host bridge.
*/
*dist = dist_a + dist_b;
+ if (verbose) {
+ pci_dbg(client,
+ "P2PDMA ACS: no common upstream bridge provider-distance=%d client-distance=%d total=%d\n",
+ dist_a, dist_b, *dist);
+ pci_p2pdma_log_path("provider", provider, NULL);
+ pci_p2pdma_log_path("client", client, NULL);
+ }
path.no_common_bridge = true;
- path.tb_on_path = pci_p2pdma_path_blocks_translation(client, NULL);
+ path.tb_on_path = pci_p2pdma_path_blocks_translation(client, NULL,
+ verbose);
for (flags = 0; flags < PCI_P2PDMA_TLP_CLASSES; flags++)
map_type[flags] = pci_p2pdma_route(&path, flags);
goto map_through_host_bridge;
check_paths_acs:
*dist = dist_a + dist_b;
+ if (verbose) {
+ pci_dbg(client,
+ "P2PDMA ACS: common=%s provider-divergence=%s client-divergence=%s provider-distance=%d client-distance=%d total=%d\n",
+ pci_name(a),
+ a_child ? pci_name(a_child) : "<none>",
+ b_child ? pci_name(b_child) : "<none>",
+ dist_a, dist_b, *dist);
+ pci_p2pdma_log_path("provider", provider, a);
+ pci_p2pdma_log_path("client", client, a);
+ }
/*
* ACS P2P routing controls apply where a TLP can route toward the peer
@@ -1052,17 +1188,33 @@ calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client,
* branch is upstream, so redirect controls do not affect the path.
*/
if (a_child && b_child) {
- if (!pci_acs_p2pdma_ctrl(a_child, &path.cpl_ctrl))
+ if (!pci_acs_p2pdma_ctrl(a_child, "completion", &path.cpl_ctrl,
+ verbose))
path.unreadable = a_child;
- if (!pci_acs_p2pdma_ctrl(b_child, &path.req_ctrl) &&
- !path.unreadable)
+ if (!pci_acs_p2pdma_ctrl(b_child, "request", &path.req_ctrl,
+ verbose) && !path.unreadable)
path.unreadable = b_child;
+ if (verbose && !path.unreadable)
+ pci_dbg(client,
+ "P2PDMA ACS: request=%s at %s completion=%s at %s\n",
+ pci_acs_p2pdma_state_name(
+ pci_p2pdma_request_state(&path, 0)),
+ pci_name(b_child),
+ pci_acs_p2pdma_state_name(
+ pci_acs_p2pdma_completion(path.cpl_ctrl,
+ 0)),
+ pci_name(a_child));
+ } else if (verbose) {
+ pci_dbg(client,
+ "P2PDMA ACS: peer divergence is incomplete; no ACS peer-routing controls evaluated\n");
}
- path.tb_on_path = pci_p2pdma_path_blocks_translation(client, a);
+ path.tb_on_path = pci_p2pdma_path_blocks_translation(client, a,
+ verbose);
if (b_child)
path.tb_above_divergence =
- pci_p2pdma_path_blocks_translation(b_child, NULL);
+ pci_p2pdma_path_blocks_translation(b_child, NULL,
+ verbose);
/*
* The walk and the config reads above serve every class; only the
@@ -1091,6 +1243,19 @@ calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client,
host_whitelisted = host_bridge_whitelist(provider, client,
verbose);
+ if (verbose) {
+ provider_host = pci_find_host_bridge(provider->bus);
+ client_host = pci_find_host_bridge(client->bus);
+ pci_dbg(client,
+ "P2PDMA ACS: host fallback cpu-support=%u whitelist=%s provider-host=%s client-host=%s same-host=%u\n",
+ cpu_p2pdma,
+ cpu_p2pdma ? "not-consulted" :
+ (host_whitelisted ? "yes" : "no"),
+ provider_host ? dev_name(&provider_host->dev) : "<none>",
+ client_host ? dev_name(&client_host->dev) : "<none>",
+ provider_host && provider_host == client_host);
+ }
+
if (!cpu_p2pdma && !host_whitelisted) {
if (verbose)
pci_warn(client, "cannot be used for peer-to-peer DMA as the client and provider (%s) do not share an upstream bridge or whitelisted host bridge\n",
@@ -1102,11 +1267,31 @@ calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client,
done:
rcu_read_lock();
p2pdma = rcu_dereference(provider->p2pdma);
- if (p2pdma)
+ if (p2pdma) {
xa_store(&p2pdma->map_types, map_types_idx(client),
- xa_mk_value(pci_p2pdma_map_types_pack(map_type)),
- GFP_ATOMIC);
+ xa_mk_value(pci_p2pdma_map_types_pack(map_type)), GFP_ATOMIC);
+ cache_store = true;
+ }
rcu_read_unlock();
+ if (verbose) {
+ pci_dbg(client,
+ "P2PDMA ACS: final provider=%s result=%s(%d) tlp-flags=%#x distance=%d unreadable=%s cache-store=%u index=%#lx\n",
+ pci_name(provider),
+ pci_p2pdma_map_type_name(map_type[tlp_flags]),
+ map_type[tlp_flags], tlp_flags, *dist,
+ path.unreadable ? pci_name(path.unreadable) : "<none>",
+ cache_store, map_types_idx(client));
+ pci_dbg(client,
+ "P2PDMA ACS: classes strict=%s relaxed=%s translated=%s translated+relaxed=%s\n",
+ pci_p2pdma_map_type_name(map_type[0]),
+ pci_p2pdma_map_type_name(
+ map_type[PCI_P2PDMA_TLP_RELAXED_CPL]),
+ pci_p2pdma_map_type_name(
+ map_type[PCI_P2PDMA_TLP_TRANSLATED]),
+ pci_p2pdma_map_type_name(
+ map_type[PCI_P2PDMA_TLP_TRANSLATED |
+ PCI_P2PDMA_TLP_RELAXED_CPL]));
+ }
return map_type[tlp_flags];
}
@@ -1436,13 +1621,21 @@ pci_p2pdma_map_type_tlp(struct p2pdma_provider *provider, struct device *dev,
enum pci_p2pdma_map_type type;
struct pci_p2pdma *p2pdma;
struct pci_dev *client;
+ bool provider_state;
int dist;
- if (!pdev->p2pdma)
+ if (!pdev->p2pdma) {
+ pci_dbg(pdev,
+ "P2PDMA ACS: map lookup rejected; provider state is absent\n");
return PCI_P2PDMA_MAP_NOT_SUPPORTED;
+ }
- if (!dev_is_pci(dev))
+ if (!dev_is_pci(dev)) {
+ dev_dbg(dev,
+ "P2PDMA ACS: provider=%s map lookup rejected; client is not PCI\n",
+ pci_name(pdev));
return PCI_P2PDMA_MAP_NOT_SUPPORTED;
+ }
client = to_pci_dev(dev);
cache_index = map_types_idx(client);
@@ -1453,8 +1646,13 @@ pci_p2pdma_map_type_tlp(struct p2pdma_provider *provider, struct device *dev,
if (p2pdma)
cached = xa_to_value(xa_load(&p2pdma->map_types,
cache_index));
+ provider_state = !!p2pdma;
rcu_read_unlock();
type = pci_p2pdma_map_types_unpack(cached, tlp_flags);
+ pci_dbg(client,
+ "P2PDMA ACS: map lookup provider=%s index=%#lx tlp-flags=%#x cached=%s(%d) provider-state=%u\n",
+ pci_name(pdev), cache_index, tlp_flags,
+ pci_p2pdma_map_type_name(type), type, provider_state);
if (type == PCI_P2PDMA_MAP_UNKNOWN)
return calc_map_type_and_dist(pdev, client, &dist, tlp_flags,
--
2.55.0
^ permalink raw reply [flat|nested] 38+ messages in thread
* [PATCH v9 12/18] PCI/P2PDMA: Add KUnit tests for the ACS routing decisions
2026-10-01 11:55 [PATCH v9 00/18] PCI/P2PDMA: Route peer-to-peer DMA by TLP class Leon Romanovsky
` (10 preceding siblings ...)
2026-10-01 11:55 ` [PATCH v9 11/18] PCI/P2PDMA: Log detailed ACS routing diagnostics Leon Romanovsky
@ 2026-10-01 11:55 ` Leon Romanovsky
2026-10-01 11:55 ` [PATCH v9 13/18] PCI/P2PDMA: Test the ACS P2P routing walk Leon Romanovsky
` (6 subsequent siblings)
18 siblings, 0 replies; 38+ messages in thread
From: Leon Romanovsky @ 2026-10-01 11:55 UTC (permalink / raw)
To: Bjorn Helgaas, Logan Gunthorpe, Leon Romanovsky, Jason Gunthorpe,
Joerg Roedel (AMD),
Will Deacon, Robin Murphy, Christian König,
Thomas Hellström
Cc: linux-pci, linux-kernel, linux-doc, iommu, Tushar Dave,
linux-media, dri-devel, linaro-mm-sig, linux-rdma, kvm,
Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Sumit Semwal
From: Leon Romanovsky <leonro@nvidia.com>
pci_acs_p2pdma_request() and pci_acs_p2pdma_completion() turn an ACS
Control register and a TLP class into a routing decision. Which bits apply
to which direction and which class is easy to get wrong, and hardware that
exposes a given combination may not be at hand.
Drive both from a table of register values and classes, covering the
redirect controls per direction and Translation Blocking, Direct Translated
P2P and Relaxed Ordering. Direct Translated P2P gets a case with and
without a redirect to override, since it changes nothing without one.
Exposing the two helpers moves their state enum and the TLP flags into
pci.h.
Tested-by: Tushar Dave <tdave@nvidia.com>
Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
---
drivers/pci/Kconfig | 15 ++++++
drivers/pci/Makefile | 1 +
drivers/pci/p2pdma.c | 40 ++-------------
drivers/pci/pci.h | 43 ++++++++++++++++
drivers/pci/pci_acs_test.c | 121 +++++++++++++++++++++++++++++++++++++++++++++
5 files changed, 184 insertions(+), 36 deletions(-)
diff --git a/drivers/pci/Kconfig b/drivers/pci/Kconfig
index 0c7408509ba2..7a3eb5beb328 100644
--- a/drivers/pci/Kconfig
+++ b/drivers/pci/Kconfig
@@ -226,6 +226,21 @@ config PCI_P2PDMA
If unsure, say N.
+config PCI_ACS_KUNIT_TEST
+ tristate "KUnit tests for PCI ACS P2P routing" if !KUNIT_ALL_TESTS
+ depends on PCI_P2PDMA && KUNIT
+ default KUNIT_ALL_TESTS
+ help
+ Enable KUnit tests for the PCI ACS peer-to-peer routing decision
+ logic, including direction-specific Request and Completion
+ controls that cannot all be exercised on typical peer-to-peer
+ hardware.
+
+ For more information on KUnit and unit tests in general, refer to
+ the KUnit documentation in Documentation/dev-tools/kunit/.
+
+ If unsure, say N.
+
config PCI_LABEL
def_bool y if (DMI || ACPI)
select NLS
diff --git a/drivers/pci/Makefile b/drivers/pci/Makefile
index 41ebc3b9a518..6305d128d3df 100644
--- a/drivers/pci/Makefile
+++ b/drivers/pci/Makefile
@@ -31,6 +31,7 @@ obj-$(CONFIG_PCI_STUB) += pci-stub.o
obj-$(CONFIG_PCI_PF_STUB) += pci-pf-stub.o
obj-$(CONFIG_PCI_ECAM) += ecam.o
obj-$(CONFIG_PCI_P2PDMA) += p2pdma.o
+obj-$(CONFIG_PCI_ACS_KUNIT_TEST) += pci_acs_test.o
obj-$(CONFIG_XEN_PCIDEV_FRONTEND) += xen-pcifront.o
obj-$(CONFIG_VGA_ARB) += vgaarb.o
obj-$(CONFIG_PCI_DOE) += doe.o
diff --git a/drivers/pci/p2pdma.c b/drivers/pci/p2pdma.c
index cef486e22f97..913d1a32a836 100644
--- a/drivers/pci/p2pdma.c
+++ b/drivers/pci/p2pdma.c
@@ -492,40 +492,6 @@ static struct pci_dev *find_parent_pci_dev(struct device *dev)
return NULL;
}
-/**
- * enum pci_p2pdma_tlp_flags - Properties of the TLPs a client will issue
- *
- * These describe the traffic rather than the topology, and select which ACS
- * controls apply along the peer-to-peer path. A value of 0 means strictly
- * ordered Requests carrying an Untranslated address.
- *
- * @PCI_P2PDMA_TLP_TRANSLATED: Requests carry an ATS Translated address. PCIe
- * r7.0 sec 6.12.3 routes those to the peer regardless of ACS P2P Request
- * Redirect and ACS P2P Egress Control wherever ACS Direct Translated P2P
- * is enabled.
- * @PCI_P2PDMA_TLP_RELAXED_CPL: The provider returns Completions with the
- * Relaxed Ordering attribute set. PCIe r7.0 sec 6.12.1.1 never redirects
- * those, so ACS P2P Completion Redirect does not gate the path. The
- * Completer chooses this attribute and the specification does not require
- * it to copy Relaxed Ordering from the Request into the Completion, so a
- * caller passing this flag asserts that its provider does.
- */
-enum pci_p2pdma_tlp_flags {
- PCI_P2PDMA_TLP_TRANSLATED = 1 << 0,
- PCI_P2PDMA_TLP_RELAXED_CPL = 1 << 1,
-};
-
-/* Every combination of the flags above selects one routing class. */
-#define PCI_P2PDMA_TLP_CLASSES \
- ((PCI_P2PDMA_TLP_TRANSLATED | PCI_P2PDMA_TLP_RELAXED_CPL) + 1)
-
-enum pci_acs_p2pdma_state {
- PCI_ACS_P2PDMA_NOT_SUPPORTED,
- PCI_ACS_P2PDMA_DIRECT,
- PCI_ACS_P2PDMA_REDIRECT,
- PCI_ACS_P2PDMA_BLOCKED,
-};
-
/*
* Decide how a peer-to-peer Request at an ACS-capable ingress port routes,
* from that port's ACS Control register and the Request's Address Type.
@@ -535,7 +501,7 @@ enum pci_acs_p2pdma_state {
* selects are a direct route and an ACS Violation, and neither one lets peer
* bus addressing be assumed.
*/
-static enum pci_acs_p2pdma_state
+VISIBLE_IF_KUNIT enum pci_acs_p2pdma_state
pci_acs_p2pdma_request(u16 ctrl, unsigned int tlp_flags)
{
if (tlp_flags & PCI_P2PDMA_TLP_TRANSLATED) {
@@ -562,6 +528,7 @@ pci_acs_p2pdma_request(u16 ctrl, unsigned int tlp_flags)
return ctrl & (PCI_ACS_RR | PCI_ACS_EC) ?
PCI_ACS_P2PDMA_REDIRECT : PCI_ACS_P2PDMA_DIRECT;
}
+EXPORT_SYMBOL_IF_KUNIT(pci_acs_p2pdma_request);
/*
* Decide how a peer-to-peer Completion at an ACS-capable ingress port routes.
@@ -569,7 +536,7 @@ pci_acs_p2pdma_request(u16 ctrl, unsigned int tlp_flags)
* affects a Completion, and that one leaves Completions carrying the Relaxed
* Ordering attribute alone.
*/
-static enum pci_acs_p2pdma_state
+VISIBLE_IF_KUNIT enum pci_acs_p2pdma_state
pci_acs_p2pdma_completion(u16 ctrl, unsigned int tlp_flags)
{
if (tlp_flags & PCI_P2PDMA_TLP_RELAXED_CPL)
@@ -578,6 +545,7 @@ pci_acs_p2pdma_completion(u16 ctrl, unsigned int tlp_flags)
return ctrl & PCI_ACS_CR ? PCI_ACS_P2PDMA_REDIRECT :
PCI_ACS_P2PDMA_DIRECT;
}
+EXPORT_SYMBOL_IF_KUNIT(pci_acs_p2pdma_completion);
static const char *pci_acs_p2pdma_state_name(enum pci_acs_p2pdma_state state)
{
diff --git a/drivers/pci/pci.h b/drivers/pci/pci.h
index ba3c3fddddc2..04c4042dd9b8 100644
--- a/drivers/pci/pci.h
+++ b/drivers/pci/pci.h
@@ -2,6 +2,7 @@
#ifndef DRIVERS_PCI_H
#define DRIVERS_PCI_H
+#include <kunit/visibility.h>
#include <linux/bug.h>
#include <linux/align.h>
#include <linux/bitfield.h>
@@ -1093,6 +1094,48 @@ resource_size_t pci_min_window_alignment(struct pci_bus *bus,
void pci_acs_init(struct pci_dev *dev);
void pci_enable_acs(struct pci_dev *dev);
+
+/**
+ * enum pci_p2pdma_tlp_flags - Properties of the TLPs a client will issue
+ *
+ * These describe the traffic rather than the topology, and select which ACS
+ * controls apply along the peer-to-peer path. A value of 0 means strictly
+ * ordered Requests carrying an Untranslated address.
+ *
+ * @PCI_P2PDMA_TLP_TRANSLATED: Requests carry an ATS Translated address. PCIe
+ * r7.0 sec 6.12.3 routes those to the peer regardless of ACS P2P Request
+ * Redirect and ACS P2P Egress Control wherever ACS Direct Translated P2P
+ * is enabled.
+ * @PCI_P2PDMA_TLP_RELAXED_CPL: The provider returns Completions with the
+ * Relaxed Ordering attribute set. PCIe r7.0 sec 6.12.1.1 never redirects
+ * those, so ACS P2P Completion Redirect does not gate the path. The
+ * Completer chooses this attribute and the specification does not require
+ * it to copy Relaxed Ordering from the Request into the Completion, so a
+ * caller passing this flag asserts that its provider does.
+ */
+enum pci_p2pdma_tlp_flags {
+ PCI_P2PDMA_TLP_TRANSLATED = 1 << 0,
+ PCI_P2PDMA_TLP_RELAXED_CPL = 1 << 1,
+};
+
+/* Every combination of the flags above selects one routing class. */
+#define PCI_P2PDMA_TLP_CLASSES \
+ ((PCI_P2PDMA_TLP_TRANSLATED | PCI_P2PDMA_TLP_RELAXED_CPL) + 1)
+
+enum pci_acs_p2pdma_state {
+ PCI_ACS_P2PDMA_NOT_SUPPORTED,
+ PCI_ACS_P2PDMA_DIRECT,
+ PCI_ACS_P2PDMA_REDIRECT,
+ PCI_ACS_P2PDMA_BLOCKED,
+};
+
+#if IS_ENABLED(CONFIG_KUNIT)
+enum pci_acs_p2pdma_state pci_acs_p2pdma_request(u16 ctrl,
+ unsigned int tlp_flags);
+enum pci_acs_p2pdma_state pci_acs_p2pdma_completion(u16 ctrl,
+ unsigned int tlp_flags);
+#endif
+
#ifdef CONFIG_PCI_QUIRKS
int pci_dev_specific_acs_enabled(struct pci_dev *dev, u16 acs_flags);
int pci_dev_specific_enable_acs(struct pci_dev *dev);
diff --git a/drivers/pci/pci_acs_test.c b/drivers/pci/pci_acs_test.c
new file mode 100644
index 000000000000..ce6b9375da36
--- /dev/null
+++ b/drivers/pci/pci_acs_test.c
@@ -0,0 +1,121 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * KUnit tests for PCI ACS peer-to-peer routing decisions.
+ *
+ * These exercise Request and Completion routing independently of the ACS
+ * settings exposed by available PCIe hardware.
+ */
+#include <kunit/test.h>
+
+#include <linux/pci.h>
+#include <linux/pci-p2pdma.h>
+#include <linux/pci_regs.h>
+
+#include "pci.h"
+
+struct acs_decision_case {
+ const char *desc;
+ u16 ctrl;
+ unsigned int tlp_flags;
+ enum pci_acs_p2pdma_state expect;
+};
+
+/* Shorthands to keep the tables below readable. */
+#define ACS_DIRECT PCI_ACS_P2PDMA_DIRECT
+#define ACS_REDIR PCI_ACS_P2PDMA_REDIRECT
+#define ACS_RO PCI_P2PDMA_TLP_RELAXED_CPL
+#define ACS_AT PCI_P2PDMA_TLP_TRANSLATED
+#define ACS_BLOCK PCI_ACS_P2PDMA_BLOCKED
+
+/* Request routing ignores Completion Redirect. */
+static const struct acs_decision_case acs_request_cases[] = {
+ { "req/none", 0, 0, ACS_DIRECT },
+ { "req/rr", PCI_ACS_RR, 0, ACS_REDIR },
+ { "req/cr", PCI_ACS_CR, 0, ACS_DIRECT },
+ { "req/rr_cr", PCI_ACS_RR | PCI_ACS_CR, 0, ACS_REDIR },
+ { "req/ec", PCI_ACS_EC, 0, ACS_REDIR },
+ { "req/ec_cr", PCI_ACS_EC | PCI_ACS_CR, 0, ACS_REDIR },
+
+ /*
+ * Direct Translated P2P overrides the redirect controls, but only for
+ * a Request that actually carries a Translated address.
+ */
+ { "req/dt", PCI_ACS_DT, 0, ACS_DIRECT },
+ { "req/dt_rr", PCI_ACS_DT | PCI_ACS_RR, 0, ACS_REDIR },
+ { "req/at", 0, ACS_AT, ACS_DIRECT },
+ { "req/at_rr", PCI_ACS_RR, ACS_AT, ACS_REDIR },
+ { "req/at_dt_rr", PCI_ACS_DT | PCI_ACS_RR, ACS_AT, ACS_DIRECT },
+ { "req/at_dt_ec", PCI_ACS_DT | PCI_ACS_EC, ACS_AT, ACS_DIRECT },
+
+ /*
+ * Translation Blocking rejects a Translated address outright, and
+ * makes the port ignore Direct Translated P2P.
+ */
+ { "req/tb", PCI_ACS_TB, 0, ACS_DIRECT },
+ { "req/tb_rr", PCI_ACS_TB | PCI_ACS_RR, 0, ACS_REDIR },
+ { "req/at_tb", PCI_ACS_TB, ACS_AT, ACS_BLOCK },
+ { "req/at_tb_dt", PCI_ACS_TB | PCI_ACS_DT, ACS_AT, ACS_BLOCK },
+};
+
+/* Completion routing depends only on Completion Redirect. */
+static const struct acs_decision_case acs_completion_cases[] = {
+ { "cpl/none", 0, 0, ACS_DIRECT },
+ { "cpl/rr", PCI_ACS_RR, 0, ACS_DIRECT },
+ { "cpl/cr", PCI_ACS_CR, 0, ACS_REDIR },
+ { "cpl/rr_cr", PCI_ACS_RR | PCI_ACS_CR, 0, ACS_REDIR },
+ { "cpl/ec", PCI_ACS_EC, 0, ACS_DIRECT },
+ { "cpl/ec_cr", PCI_ACS_EC | PCI_ACS_CR, 0, ACS_REDIR },
+
+ /* Relaxed Ordering Completions are never redirected. */
+ { "cpl/ro", 0, ACS_RO, ACS_DIRECT },
+ { "cpl/ro_cr", PCI_ACS_CR, ACS_RO, ACS_DIRECT },
+ { "cpl/ro_rr_cr", PCI_ACS_RR | PCI_ACS_CR, ACS_RO, ACS_DIRECT },
+};
+
+#undef ACS_DIRECT
+#undef ACS_REDIR
+#undef ACS_RO
+#undef ACS_AT
+#undef ACS_BLOCK
+
+static void acs_decision_desc(const struct acs_decision_case *c, char *desc)
+{
+ strscpy(desc, c->desc, KUNIT_PARAM_DESC_SIZE);
+}
+
+KUNIT_ARRAY_PARAM(acs_request, acs_request_cases, acs_decision_desc);
+KUNIT_ARRAY_PARAM(acs_completion, acs_completion_cases, acs_decision_desc);
+
+static void pci_acs_p2pdma_request_test(struct kunit *test)
+{
+ const struct acs_decision_case *c = test->param_value;
+
+ KUNIT_EXPECT_EQ(test, pci_acs_p2pdma_request(c->ctrl, c->tlp_flags),
+ c->expect);
+}
+
+static void pci_acs_p2pdma_completion_test(struct kunit *test)
+{
+ const struct acs_decision_case *c = test->param_value;
+
+ KUNIT_EXPECT_EQ(test, pci_acs_p2pdma_completion(c->ctrl, c->tlp_flags),
+ c->expect);
+}
+
+static struct kunit_case pci_acs_test_cases[] = {
+ KUNIT_CASE_PARAM(pci_acs_p2pdma_request_test,
+ acs_request_gen_params),
+ KUNIT_CASE_PARAM(pci_acs_p2pdma_completion_test,
+ acs_completion_gen_params),
+ {}
+};
+
+static struct kunit_suite pci_acs_test_suite = {
+ .name = "pci_acs",
+ .test_cases = pci_acs_test_cases,
+};
+kunit_test_suite(pci_acs_test_suite);
+
+MODULE_IMPORT_NS("EXPORTED_FOR_KUNIT_TESTING");
+MODULE_LICENSE("GPL");
+MODULE_DESCRIPTION("KUnit tests for PCI ACS peer-to-peer routing decisions");
--
2.55.0
^ permalink raw reply [flat|nested] 38+ messages in thread
* [PATCH v9 13/18] PCI/P2PDMA: Test the ACS P2P routing walk
2026-10-01 11:55 [PATCH v9 00/18] PCI/P2PDMA: Route peer-to-peer DMA by TLP class Leon Romanovsky
` (11 preceding siblings ...)
2026-10-01 11:55 ` [PATCH v9 12/18] PCI/P2PDMA: Add KUnit tests for the ACS routing decisions Leon Romanovsky
@ 2026-10-01 11:55 ` Leon Romanovsky
2026-10-01 11:55 ` [PATCH v9 14/18] PCI: Add KUnit coverage for ACS isolation checks Leon Romanovsky
` (5 subsequent siblings)
18 siblings, 0 replies; 38+ messages in thread
From: Leon Romanovsky @ 2026-10-01 11:55 UTC (permalink / raw)
To: Bjorn Helgaas, Logan Gunthorpe, Leon Romanovsky, Jason Gunthorpe,
Joerg Roedel (AMD),
Will Deacon, Robin Murphy, Christian König,
Thomas Hellström
Cc: linux-pci, linux-kernel, linux-doc, iommu, Tushar Dave,
linux-media, dri-devel, linaro-mm-sig, linux-rdma, kvm,
Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Sumit Semwal
From: Leon Romanovsky <leonro@nvidia.com>
calc_map_type_and_dist() decides which ports along a path carry the
routing controls, and holds every class's answer in one cache entry.
Neither depends on a single register, so a table of them cannot reach the
walk itself.
Drive the walk over a fabricated fabric of two devices below a switch,
with fake config space supplying the ACS Control registers. Cover the
ports below the divergence, the three cases where two classes of one path
disagree, and the packed cache, which the fabric has no provider state to
exercise indirectly.
Tested-by: Tushar Dave <tdave@nvidia.com>
Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
---
drivers/pci/p2pdma.c | 9 +-
drivers/pci/pci.h | 9 +
drivers/pci/pci_acs_test.c | 478 +++++++++++++++++++++++++++++++++++++++++++++
3 files changed, 493 insertions(+), 3 deletions(-)
diff --git a/drivers/pci/p2pdma.c b/drivers/pci/p2pdma.c
index 913d1a32a836..c1e9d43bded2 100644
--- a/drivers/pci/p2pdma.c
+++ b/drivers/pci/p2pdma.c
@@ -1005,7 +1005,7 @@ static unsigned long map_types_idx(struct pci_dev *client)
*/
static_assert(PCI_P2PDMA_MAP_THRU_HOST_BRIDGE < 16);
-static unsigned long
+VISIBLE_IF_KUNIT unsigned long
pci_p2pdma_map_types_pack(const enum pci_p2pdma_map_type *type)
{
unsigned long val = 0;
@@ -1016,12 +1016,14 @@ pci_p2pdma_map_types_pack(const enum pci_p2pdma_map_type *type)
return val;
}
+EXPORT_SYMBOL_IF_KUNIT(pci_p2pdma_map_types_pack);
-static enum pci_p2pdma_map_type
+VISIBLE_IF_KUNIT enum pci_p2pdma_map_type
pci_p2pdma_map_types_unpack(unsigned long val, unsigned int tlp_flags)
{
return (val >> (tlp_flags * 4)) & 0xf;
}
+EXPORT_SYMBOL_IF_KUNIT(pci_p2pdma_map_types_unpack);
/*
* Calculate the P2PDMA mapping type and distance between two PCI devices.
@@ -1070,7 +1072,7 @@ pci_p2pdma_map_types_unpack(unsigned long val, unsigned int tlp_flags)
* ports per above. If the device is not in the whitelist, return
* PCI_P2PDMA_MAP_NOT_SUPPORTED.
*/
-static enum pci_p2pdma_map_type
+VISIBLE_IF_KUNIT enum pci_p2pdma_map_type
calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client,
int *dist, unsigned int tlp_flags, bool verbose)
{
@@ -1262,6 +1264,7 @@ calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client,
}
return map_type[tlp_flags];
}
+EXPORT_SYMBOL_IF_KUNIT(calc_map_type_and_dist);
/**
* pci_p2pdma_distance_many - Determine the cumulative distance between
diff --git a/drivers/pci/pci.h b/drivers/pci/pci.h
index 04c4042dd9b8..812a48afdfee 100644
--- a/drivers/pci/pci.h
+++ b/drivers/pci/pci.h
@@ -7,6 +7,7 @@
#include <linux/align.h>
#include <linux/bitfield.h>
#include <linux/pci.h>
+#include <linux/pci-p2pdma.h>
#include <trace/events/pci.h>
struct pcie_tlp_log;
@@ -1134,6 +1135,14 @@ enum pci_acs_p2pdma_state pci_acs_p2pdma_request(u16 ctrl,
unsigned int tlp_flags);
enum pci_acs_p2pdma_state pci_acs_p2pdma_completion(u16 ctrl,
unsigned int tlp_flags);
+unsigned long pci_p2pdma_map_types_pack(const enum pci_p2pdma_map_type *type);
+enum pci_p2pdma_map_type pci_p2pdma_map_types_unpack(unsigned long val,
+ unsigned int tlp_flags);
+enum pci_p2pdma_map_type calc_map_type_and_dist(struct pci_dev *provider,
+ struct pci_dev *client,
+ int *dist,
+ unsigned int tlp_flags,
+ bool verbose);
#endif
#ifdef CONFIG_PCI_QUIRKS
diff --git a/drivers/pci/pci_acs_test.c b/drivers/pci/pci_acs_test.c
index ce6b9375da36..7f4f9cc04bfb 100644
--- a/drivers/pci/pci_acs_test.c
+++ b/drivers/pci/pci_acs_test.c
@@ -102,11 +102,489 @@ static void pci_acs_p2pdma_completion_test(struct kunit *test)
c->expect);
}
+/*
+ * Drive calc_map_type_and_dist() over a fabricated PCIe fabric matching the
+ * canonical topology of two devices below one switch:
+ *
+ * host bridge / root bus
+ * Root Port
+ * Switch Upstream Port
+ * Switch Downstream Port 0
+ * Nested Switch -- provider
+ * Switch Downstream Port 1
+ * Nested Switch -- client
+ *
+ * Fake config-space operations supply the ACS Control registers. This lets
+ * the cases vary both divergence ports and controls below the divergence
+ * without depending on real hardware.
+ */
+struct acs_port_cfg {
+ u16 ctrl;
+ bool fail_read;
+};
+
+struct acs_fabric {
+ struct pci_dev *rootport;
+ struct pci_dev *provider;
+ struct pci_dev *client;
+ struct pci_dev *dn0; /* Downstream Port 0 (provider side) */
+ struct pci_dev *dn1; /* Downstream Port 1 (client side) */
+ struct pci_dev *provider_leaf;
+ struct pci_dev *client_leaf;
+ struct acs_port_cfg dn0_cfg;
+ struct acs_port_cfg dn1_cfg;
+ struct acs_port_cfg provider_leaf_cfg;
+ struct acs_port_cfg client_leaf_cfg;
+ struct acs_port_cfg rootport_cfg;
+};
+
+static int acs_port_read(struct pci_dev *port, struct acs_port_cfg *cfg,
+ int where, int size, u32 *val)
+{
+ if (port->acs_cap && size == 2 &&
+ where == port->acs_cap + PCI_ACS_CTRL) {
+ if (cfg->fail_read)
+ return PCIBIOS_DEVICE_NOT_FOUND;
+ *val = cfg->ctrl;
+ }
+
+ return PCIBIOS_SUCCESSFUL;
+}
+
+static int acs_fabric_read(struct pci_bus *bus, unsigned int devfn,
+ int where, int size, u32 *val)
+{
+ struct acs_fabric *f = bus->sysdata;
+
+ *val = 0;
+ if (bus == f->rootport->bus && devfn == f->rootport->devfn)
+ return acs_port_read(f->rootport, &f->rootport_cfg,
+ where, size, val);
+ if (bus == f->dn0->bus && devfn == f->dn0->devfn)
+ return acs_port_read(f->dn0, &f->dn0_cfg, where, size, val);
+ if (bus == f->dn1->bus && devfn == f->dn1->devfn)
+ return acs_port_read(f->dn1, &f->dn1_cfg, where, size, val);
+ if (bus == f->provider_leaf->bus &&
+ devfn == f->provider_leaf->devfn)
+ return acs_port_read(f->provider_leaf, &f->provider_leaf_cfg,
+ where, size, val);
+ if (bus == f->client_leaf->bus && devfn == f->client_leaf->devfn)
+ return acs_port_read(f->client_leaf, &f->client_leaf_cfg,
+ where, size, val);
+
+ return PCIBIOS_SUCCESSFUL;
+}
+
+static int acs_fabric_write(struct pci_bus *bus, unsigned int devfn,
+ int where, int size, u32 val)
+{
+ return PCIBIOS_SUCCESSFUL;
+}
+
+static struct pci_ops acs_fabric_ops = {
+ .read = acs_fabric_read,
+ .write = acs_fabric_write,
+};
+
+static struct pci_bus *acs_add_bus(struct kunit *test, struct pci_bus *parent,
+ struct pci_dev *self, u8 nr, void *sysdata)
+{
+ struct pci_bus *bus = kunit_kzalloc(test, sizeof(*bus), GFP_KERNEL);
+
+ KUNIT_ASSERT_NOT_NULL(test, bus);
+ bus->parent = parent;
+ bus->self = self;
+ bus->number = nr;
+ bus->ops = &acs_fabric_ops;
+ bus->sysdata = sysdata;
+ INIT_LIST_HEAD(&bus->devices);
+ return bus;
+}
+
+static struct pci_dev *acs_add_dev(struct kunit *test, struct pci_bus *bus,
+ unsigned int devfn, int pcie_type)
+{
+ struct pci_dev *dev = kunit_kzalloc(test, sizeof(*dev), GFP_KERNEL);
+
+ KUNIT_ASSERT_NOT_NULL(test, dev);
+ dev->bus = bus;
+ dev->devfn = devfn;
+ dev->pcie_cap = 0x40;
+ dev->pcie_flags_reg = (pcie_type << 4) | 0x2;
+ list_add_tail(&dev->bus_list, &bus->devices);
+ return dev;
+}
+
+static void acs_build_fabric(struct kunit *test, struct acs_fabric *f)
+{
+ struct pci_bus *bus0, *bus1, *bus2, *bus3, *bus4, *bus5, *bus6;
+ struct pci_bus *bus7, *bus8;
+ struct pci_dev *swup, *provider_swup, *client_swup;
+ struct pci_dev *rootport;
+ struct pci_host_bridge *host;
+
+ host = kunit_kzalloc(test, sizeof(*host), GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, host);
+
+ bus0 = acs_add_bus(test, NULL, NULL, 0, f);
+ /* The Root Port doubles as the whitelisted host-bridge device. */
+ rootport = acs_add_dev(test, bus0, PCI_DEVFN(0, 0),
+ PCI_EXP_TYPE_ROOT_PORT);
+ f->rootport = rootport;
+ rootport->vendor = PCI_VENDOR_ID_GOOGLE;
+ rootport->device = 0x1234;
+ host->bus = bus0;
+ bus0->bridge = &host->dev;
+
+ bus1 = acs_add_bus(test, bus0, rootport, 1, f);
+ swup = acs_add_dev(test, bus1, PCI_DEVFN(0, 0),
+ PCI_EXP_TYPE_UPSTREAM);
+
+ bus2 = acs_add_bus(test, bus1, swup, 2, f);
+ f->dn0 = acs_add_dev(test, bus2, PCI_DEVFN(0, 0),
+ PCI_EXP_TYPE_DOWNSTREAM);
+ f->dn1 = acs_add_dev(test, bus2, PCI_DEVFN(1, 0),
+ PCI_EXP_TYPE_DOWNSTREAM);
+
+ bus3 = acs_add_bus(test, bus2, f->dn0, 3, f);
+ provider_swup = acs_add_dev(test, bus3, PCI_DEVFN(0, 0),
+ PCI_EXP_TYPE_UPSTREAM);
+ bus5 = acs_add_bus(test, bus3, provider_swup, 5, f);
+ f->provider_leaf = acs_add_dev(test, bus5, PCI_DEVFN(0, 0),
+ PCI_EXP_TYPE_DOWNSTREAM);
+ bus7 = acs_add_bus(test, bus5, f->provider_leaf, 7, f);
+ f->provider = acs_add_dev(test, bus7, PCI_DEVFN(0, 0),
+ PCI_EXP_TYPE_ENDPOINT);
+
+ bus4 = acs_add_bus(test, bus2, f->dn1, 4, f);
+ client_swup = acs_add_dev(test, bus4, PCI_DEVFN(0, 0),
+ PCI_EXP_TYPE_UPSTREAM);
+ bus6 = acs_add_bus(test, bus4, client_swup, 6, f);
+ f->client_leaf = acs_add_dev(test, bus6, PCI_DEVFN(0, 0),
+ PCI_EXP_TYPE_DOWNSTREAM);
+ bus8 = acs_add_bus(test, bus6, f->client_leaf, 8, f);
+ f->client = acs_add_dev(test, bus8, PCI_DEVFN(0, 0),
+ PCI_EXP_TYPE_ENDPOINT);
+}
+
+static enum pci_p2pdma_map_type acs_walk_map(struct acs_fabric *f,
+ unsigned int tlp_flags)
+{
+ int dist;
+
+ return calc_map_type_and_dist(f->provider, f->client, &dist, tlp_flags,
+ false);
+}
+
+static void acs_walk_bus_addr_test(struct kunit *test)
+{
+ struct acs_fabric f = {};
+
+ acs_build_fabric(test, &f);
+ KUNIT_EXPECT_EQ(test, acs_walk_map(&f, 0), PCI_P2PDMA_MAP_BUS_ADDR);
+}
+
+static void acs_walk_request_redirect_test(struct kunit *test)
+{
+ struct acs_fabric f = {};
+
+ acs_build_fabric(test, &f);
+ f.dn1->acs_cap = 0x100;
+ f.dn1->acs_capabilities = PCI_ACS_RR;
+ f.dn1_cfg.ctrl = PCI_ACS_RR;
+
+ KUNIT_EXPECT_EQ(test, acs_walk_map(&f, 0),
+ PCI_P2PDMA_MAP_THRU_HOST_BRIDGE);
+}
+
+static void acs_walk_completion_redirect_test(struct kunit *test)
+{
+ struct acs_fabric f = {};
+
+ acs_build_fabric(test, &f);
+ f.dn0->acs_cap = 0x100;
+ f.dn0->acs_capabilities = PCI_ACS_CR;
+ f.dn0_cfg.ctrl = PCI_ACS_CR;
+
+ KUNIT_EXPECT_EQ(test, acs_walk_map(&f, 0),
+ PCI_P2PDMA_MAP_THRU_HOST_BRIDGE);
+}
+
+static void acs_walk_egress_control_test(struct kunit *test)
+{
+ struct acs_fabric f = {};
+
+ acs_build_fabric(test, &f);
+ f.dn1->acs_cap = 0x100;
+ f.dn1->acs_capabilities = PCI_ACS_EC;
+ f.dn1_cfg.ctrl = PCI_ACS_EC;
+
+ KUNIT_EXPECT_EQ(test, acs_walk_map(&f, 0),
+ PCI_P2PDMA_MAP_THRU_HOST_BRIDGE);
+}
+
+static void acs_walk_asymmetric_direct_test(struct kunit *test)
+{
+ struct acs_fabric f = {};
+
+ acs_build_fabric(test, &f);
+ /* These controls affect only the reverse transaction directions. */
+ f.dn0->acs_cap = 0x100;
+ f.dn0->acs_capabilities = PCI_ACS_RR | PCI_ACS_EC;
+ f.dn0_cfg.ctrl = PCI_ACS_RR | PCI_ACS_EC;
+ f.dn1->acs_cap = 0x100;
+ f.dn1->acs_capabilities = PCI_ACS_CR;
+ f.dn1_cfg.ctrl = PCI_ACS_CR;
+
+ KUNIT_EXPECT_EQ(test, acs_walk_map(&f, 0), PCI_P2PDMA_MAP_BUS_ADDR);
+}
+
+static void acs_walk_nested_completion_redirect_test(struct kunit *test)
+{
+ struct acs_fabric f = {};
+
+ acs_build_fabric(test, &f);
+ f.provider_leaf->acs_cap = 0x100;
+ f.provider_leaf->acs_capabilities = PCI_ACS_CR;
+ f.provider_leaf_cfg.ctrl = PCI_ACS_CR;
+
+ KUNIT_EXPECT_EQ(test, acs_walk_map(&f, 0), PCI_P2PDMA_MAP_BUS_ADDR);
+}
+
+static void acs_walk_nested_request_redirect_test(struct kunit *test)
+{
+ struct acs_fabric f = {};
+
+ acs_build_fabric(test, &f);
+ f.client_leaf->acs_cap = 0x100;
+ f.client_leaf->acs_capabilities = PCI_ACS_RR;
+ f.client_leaf_cfg.ctrl = PCI_ACS_RR;
+
+ KUNIT_EXPECT_EQ(test, acs_walk_map(&f, 0), PCI_P2PDMA_MAP_BUS_ADDR);
+}
+
+static void acs_walk_translation_blocking_test(struct kunit *test)
+{
+ struct acs_fabric f = {};
+
+ acs_build_fabric(test, &f);
+ f.client_leaf->acs_cap = 0x100;
+ f.client_leaf->acs_capabilities = PCI_ACS_TB;
+ f.client_leaf_cfg.ctrl = PCI_ACS_TB;
+
+ /* Untranslated Requests are unaffected by Translation Blocking. */
+ KUNIT_EXPECT_EQ(test, acs_walk_map(&f, 0), PCI_P2PDMA_MAP_BUS_ADDR);
+ KUNIT_EXPECT_EQ(test, acs_walk_map(&f, PCI_P2PDMA_TLP_TRANSLATED),
+ PCI_P2PDMA_MAP_NOT_SUPPORTED);
+}
+
+static void acs_walk_relaxed_completion_test(struct kunit *test)
+{
+ struct acs_fabric f = {};
+
+ acs_build_fabric(test, &f);
+ f.dn0->acs_cap = 0x100;
+ f.dn0->acs_capabilities = PCI_ACS_CR;
+ f.dn0_cfg.ctrl = PCI_ACS_CR;
+
+ KUNIT_EXPECT_EQ(test, acs_walk_map(&f, 0),
+ PCI_P2PDMA_MAP_THRU_HOST_BRIDGE);
+ KUNIT_EXPECT_EQ(test, acs_walk_map(&f, PCI_P2PDMA_TLP_RELAXED_CPL),
+ PCI_P2PDMA_MAP_BUS_ADDR);
+}
+
+static void acs_walk_direct_translated_test(struct kunit *test)
+{
+ struct acs_fabric f = {};
+
+ acs_build_fabric(test, &f);
+ f.dn1->acs_cap = 0x100;
+ f.dn1->acs_capabilities = PCI_ACS_RR | PCI_ACS_DT;
+ f.dn1_cfg.ctrl = PCI_ACS_RR | PCI_ACS_DT;
+
+ KUNIT_EXPECT_EQ(test, acs_walk_map(&f, 0),
+ PCI_P2PDMA_MAP_THRU_HOST_BRIDGE);
+ KUNIT_EXPECT_EQ(test, acs_walk_map(&f, PCI_P2PDMA_TLP_TRANSLATED),
+ PCI_P2PDMA_MAP_BUS_ADDR);
+}
+
+/*
+ * The cache stores one packed value per client, so every class has to come
+ * back out under the flags that selected it.
+ */
+static void acs_map_types_pack_test(struct kunit *test)
+{
+ static const enum pci_p2pdma_map_type type[PCI_P2PDMA_TLP_CLASSES] = {
+ [0] = PCI_P2PDMA_MAP_THRU_HOST_BRIDGE,
+ [PCI_P2PDMA_TLP_TRANSLATED] = PCI_P2PDMA_MAP_NOT_SUPPORTED,
+ [PCI_P2PDMA_TLP_RELAXED_CPL] = PCI_P2PDMA_MAP_BUS_ADDR,
+ [PCI_P2PDMA_TLP_TRANSLATED | PCI_P2PDMA_TLP_RELAXED_CPL] =
+ PCI_P2PDMA_MAP_UNKNOWN,
+ };
+ unsigned long packed = pci_p2pdma_map_types_pack(type);
+ unsigned int flags;
+
+ for (flags = 0; flags < PCI_P2PDMA_TLP_CLASSES; flags++)
+ KUNIT_EXPECT_EQ(test,
+ pci_p2pdma_map_types_unpack(packed, flags),
+ type[flags]);
+
+ /* An absent cache entry reads back as unknown in every class. */
+ for (flags = 0; flags < PCI_P2PDMA_TLP_CLASSES; flags++)
+ KUNIT_EXPECT_EQ(test, pci_p2pdma_map_types_unpack(0, flags),
+ PCI_P2PDMA_MAP_UNKNOWN);
+}
+
+/*
+ * The provider can be an ancestor of the client, which leaves no divergence
+ * to evaluate. Translation Blocking still applies to every port the Request
+ * passes on its way up.
+ */
+static void acs_walk_ancestor_provider_tb_test(struct kunit *test)
+{
+ struct acs_fabric f = {};
+ int dist;
+
+ acs_build_fabric(test, &f);
+ f.client_leaf->acs_cap = 0x100;
+ f.client_leaf->acs_capabilities = PCI_ACS_TB;
+ f.client_leaf_cfg.ctrl = PCI_ACS_TB;
+
+ KUNIT_EXPECT_EQ(test,
+ calc_map_type_and_dist(f.dn1, f.client, &dist, 0,
+ false),
+ PCI_P2PDMA_MAP_BUS_ADDR);
+ KUNIT_EXPECT_EQ(test,
+ calc_map_type_and_dist(f.dn1, f.client, &dist,
+ PCI_P2PDMA_TLP_TRANSLATED,
+ false),
+ PCI_P2PDMA_MAP_NOT_SUPPORTED);
+}
+
+/*
+ * Without a common upstream bridge the Request still climbs towards the host
+ * bridge, so Translation Blocking on the way withdraws the Translated classes
+ * there too.
+ */
+static void acs_walk_no_common_bridge_tb_test(struct kunit *test)
+{
+ struct acs_fabric f = {}, g = {};
+ int dist;
+
+ acs_build_fabric(test, &f);
+ acs_build_fabric(test, &g);
+ f.client_leaf->acs_cap = 0x100;
+ f.client_leaf->acs_capabilities = PCI_ACS_TB;
+ f.client_leaf_cfg.ctrl = PCI_ACS_TB;
+
+ KUNIT_EXPECT_EQ(test,
+ calc_map_type_and_dist(g.provider, f.client, &dist, 0,
+ false),
+ PCI_P2PDMA_MAP_THRU_HOST_BRIDGE);
+ KUNIT_EXPECT_EQ(test,
+ calc_map_type_and_dist(g.provider, f.client, &dist,
+ PCI_P2PDMA_TLP_TRANSLATED,
+ false),
+ PCI_P2PDMA_MAP_NOT_SUPPORTED);
+}
+
+/*
+ * A Request between the client and itself never leaves the device, so no port
+ * is in a position to inspect its Address Type. Translation Blocking directly
+ * above the client must not withdraw the Translated classes.
+ */
+static void acs_walk_self_dma_tb_test(struct kunit *test)
+{
+ struct acs_fabric f = {};
+ int dist;
+
+ acs_build_fabric(test, &f);
+ f.client_leaf->acs_cap = 0x100;
+ f.client_leaf->acs_capabilities = PCI_ACS_TB;
+ f.client_leaf_cfg.ctrl = PCI_ACS_TB;
+
+ KUNIT_EXPECT_EQ(test,
+ calc_map_type_and_dist(f.client, f.client, &dist, 0,
+ false),
+ PCI_P2PDMA_MAP_BUS_ADDR);
+ KUNIT_EXPECT_EQ(test,
+ calc_map_type_and_dist(f.client, f.client, &dist,
+ PCI_P2PDMA_TLP_TRANSLATED,
+ false),
+ PCI_P2PDMA_MAP_BUS_ADDR);
+}
+
+/*
+ * A direct route turns around at the divergence, so Translation Blocking above
+ * it does not touch one. The host bridge route keeps climbing past that port,
+ * and the Translated classes have to go without it.
+ */
+static void acs_walk_tb_above_divergence_test(struct kunit *test)
+{
+ struct acs_fabric f = {};
+ int dist;
+
+ acs_build_fabric(test, &f);
+ f.rootport->acs_cap = 0x100;
+ f.rootport->acs_capabilities = PCI_ACS_TB;
+ f.rootport_cfg.ctrl = PCI_ACS_TB;
+
+ /* Nothing redirects yet, so the Request never reaches the Root Port. */
+ KUNIT_EXPECT_EQ(test,
+ calc_map_type_and_dist(f.provider, f.client, &dist,
+ PCI_P2PDMA_TLP_TRANSLATED,
+ false),
+ PCI_P2PDMA_MAP_BUS_ADDR);
+
+ /* Request Redirect sends it up past the Root Port instead. */
+ f.dn1->acs_cap = 0x100;
+ f.dn1->acs_capabilities = PCI_ACS_RR;
+ f.dn1_cfg.ctrl = PCI_ACS_RR;
+
+ KUNIT_EXPECT_EQ(test,
+ calc_map_type_and_dist(f.provider, f.client, &dist, 0,
+ false),
+ PCI_P2PDMA_MAP_THRU_HOST_BRIDGE);
+ KUNIT_EXPECT_EQ(test,
+ calc_map_type_and_dist(f.provider, f.client, &dist,
+ PCI_P2PDMA_TLP_TRANSLATED,
+ false),
+ PCI_P2PDMA_MAP_NOT_SUPPORTED);
+}
+
+static void acs_walk_unreadable_control_test(struct kunit *test)
+{
+ struct acs_fabric f = {};
+
+ acs_build_fabric(test, &f);
+ f.dn1->acs_cap = 0x100;
+ f.dn1_cfg.fail_read = true;
+
+ KUNIT_EXPECT_EQ(test, acs_walk_map(&f, 0),
+ PCI_P2PDMA_MAP_NOT_SUPPORTED);
+}
+
static struct kunit_case pci_acs_test_cases[] = {
KUNIT_CASE_PARAM(pci_acs_p2pdma_request_test,
acs_request_gen_params),
KUNIT_CASE_PARAM(pci_acs_p2pdma_completion_test,
acs_completion_gen_params),
+ KUNIT_CASE(acs_walk_bus_addr_test),
+ KUNIT_CASE(acs_walk_request_redirect_test),
+ KUNIT_CASE(acs_walk_completion_redirect_test),
+ KUNIT_CASE(acs_walk_egress_control_test),
+ KUNIT_CASE(acs_walk_asymmetric_direct_test),
+ KUNIT_CASE(acs_walk_nested_completion_redirect_test),
+ KUNIT_CASE(acs_walk_nested_request_redirect_test),
+ KUNIT_CASE(acs_walk_translation_blocking_test),
+ KUNIT_CASE(acs_walk_ancestor_provider_tb_test),
+ KUNIT_CASE(acs_walk_no_common_bridge_tb_test),
+ KUNIT_CASE(acs_walk_self_dma_tb_test),
+ KUNIT_CASE(acs_walk_relaxed_completion_test),
+ KUNIT_CASE(acs_walk_direct_translated_test),
+ KUNIT_CASE(acs_walk_tb_above_divergence_test),
+ KUNIT_CASE(acs_walk_unreadable_control_test),
+ KUNIT_CASE(acs_map_types_pack_test),
{}
};
--
2.55.0
^ permalink raw reply [flat|nested] 38+ messages in thread
* [PATCH v9 14/18] PCI: Add KUnit coverage for ACS isolation checks
2026-10-01 11:55 [PATCH v9 00/18] PCI/P2PDMA: Route peer-to-peer DMA by TLP class Leon Romanovsky
` (12 preceding siblings ...)
2026-10-01 11:55 ` [PATCH v9 13/18] PCI/P2PDMA: Test the ACS P2P routing walk Leon Romanovsky
@ 2026-10-01 11:55 ` Leon Romanovsky
2026-10-01 11:55 ` [PATCH v9 15/18] PCI/P2PDMA: Document TLP-class routing Leon Romanovsky
` (4 subsequent siblings)
18 siblings, 0 replies; 38+ messages in thread
From: Leon Romanovsky @ 2026-10-01 11:55 UTC (permalink / raw)
To: Bjorn Helgaas, Logan Gunthorpe, Leon Romanovsky, Jason Gunthorpe,
Joerg Roedel (AMD),
Will Deacon, Robin Murphy, Christian König,
Thomas Hellström
Cc: linux-pci, linux-kernel, linux-doc, iommu, Tushar Dave,
linux-media, dri-devel, linaro-mm-sig, linux-rdma, kvm,
Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Sumit Semwal
From: Leon Romanovsky <leonro@nvidia.com>
Direct Translated P2P does not weaken IOMMU isolation because a Translated
Request carries an address supplied by the IOMMU. Config-space read
failures, however, leave ACS state unknown and must not report isolation.
Exercise both cases with fake config-space operations. Also cover missing
and unrequested controls and a missing ACS capability.
Tested-by: Tushar Dave <tdave@nvidia.com>
Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
---
drivers/pci/pci.c | 4 +-
drivers/pci/pci.h | 1 +
drivers/pci/pci_acs_test.c | 138 +++++++++++++++++++++++++++++++++++++++++++++
3 files changed, 142 insertions(+), 1 deletion(-)
diff --git a/drivers/pci/pci.c b/drivers/pci/pci.c
index f7d94ecf9157..4a9ab3882aac 100644
--- a/drivers/pci/pci.c
+++ b/drivers/pci/pci.c
@@ -3578,7 +3578,8 @@ void pci_configure_ari(struct pci_dev *dev)
}
}
-static bool pci_acs_flags_enabled(struct pci_dev *pdev, u16 acs_flags)
+VISIBLE_IF_KUNIT
+bool pci_acs_flags_enabled(struct pci_dev *pdev, u16 acs_flags)
{
int pos;
u16 ctrl;
@@ -3598,6 +3599,7 @@ static bool pci_acs_flags_enabled(struct pci_dev *pdev, u16 acs_flags)
return false;
return (ctrl & acs_flags) == acs_flags;
}
+EXPORT_SYMBOL_IF_KUNIT(pci_acs_flags_enabled);
/**
* pci_acs_enabled - test ACS against required flags for a given device
diff --git a/drivers/pci/pci.h b/drivers/pci/pci.h
index 812a48afdfee..a4a07f31f744 100644
--- a/drivers/pci/pci.h
+++ b/drivers/pci/pci.h
@@ -1131,6 +1131,7 @@ enum pci_acs_p2pdma_state {
};
#if IS_ENABLED(CONFIG_KUNIT)
+bool pci_acs_flags_enabled(struct pci_dev *pdev, u16 acs_flags);
enum pci_acs_p2pdma_state pci_acs_p2pdma_request(u16 ctrl,
unsigned int tlp_flags);
enum pci_acs_p2pdma_state pci_acs_p2pdma_completion(u16 ctrl,
diff --git a/drivers/pci/pci_acs_test.c b/drivers/pci/pci_acs_test.c
index 7f4f9cc04bfb..6a2cd38d968f 100644
--- a/drivers/pci/pci_acs_test.c
+++ b/drivers/pci/pci_acs_test.c
@@ -102,6 +102,140 @@ static void pci_acs_p2pdma_completion_test(struct kunit *test)
c->expect);
}
+/* Flags an IOMMU asks for; see REQ_ACS_FLAGS in drivers/iommu/iommu.c. */
+#define ACS_REQ_FLAGS (PCI_ACS_SV | PCI_ACS_RR | PCI_ACS_CR | PCI_ACS_UF)
+#define ACS_ALL_CAPS (PCI_ACS_SV | PCI_ACS_TB | PCI_ACS_RR | PCI_ACS_CR | \
+ PCI_ACS_UF | PCI_ACS_DT)
+#define ACS_TEST_CAP 0x100
+
+struct acs_ctrl_cfg {
+ unsigned int devfn;
+ u16 cap; /* Offset where the ACS capability responds */
+ u16 ctrl;
+ bool fail_read;
+};
+
+static int acs_ctrl_read(struct pci_bus *bus, unsigned int devfn,
+ int where, int size, u32 *val)
+{
+ struct acs_ctrl_cfg *cfg = bus->sysdata;
+
+ *val = 0;
+ if (cfg->fail_read)
+ return PCIBIOS_DEVICE_NOT_FOUND;
+
+ if (devfn == cfg->devfn && size == 2 &&
+ where == cfg->cap + PCI_ACS_CTRL)
+ *val = cfg->ctrl;
+ return PCIBIOS_SUCCESSFUL;
+}
+
+static int acs_ctrl_write(struct pci_bus *bus, unsigned int devfn,
+ int where, int size, u32 val)
+{
+ return PCIBIOS_SUCCESSFUL;
+}
+
+static struct pci_ops acs_ctrl_ops = {
+ .read = acs_ctrl_read,
+ .write = acs_ctrl_write,
+};
+
+struct acs_isolation_case {
+ const char *desc;
+ u16 ctrl;
+ u16 req;
+ bool expect;
+};
+
+static const struct acs_isolation_case acs_isolation_cases[] = {
+ { "all_enabled", ACS_REQ_FLAGS, ACS_REQ_FLAGS, true },
+ /* Translated Requests remain isolated by their IOMMU translation. */
+ { "dt", ACS_REQ_FLAGS | PCI_ACS_DT, ACS_REQ_FLAGS, true },
+ { "rr_not_enabled", PCI_ACS_SV | PCI_ACS_CR | PCI_ACS_UF,
+ ACS_REQ_FLAGS, false },
+ { "rr_not_required", PCI_ACS_SV | PCI_ACS_CR | PCI_ACS_UF,
+ PCI_ACS_SV | PCI_ACS_CR | PCI_ACS_UF, true },
+};
+
+static void acs_isolation_desc(const struct acs_isolation_case *c, char *desc)
+{
+ strscpy(desc, c->desc, KUNIT_PARAM_DESC_SIZE);
+}
+
+KUNIT_ARRAY_PARAM(acs_isolation, acs_isolation_cases, acs_isolation_desc);
+
+static void pci_acs_flags_enabled_test(struct kunit *test)
+{
+ const struct acs_isolation_case *c = test->param_value;
+ struct acs_ctrl_cfg cfg = {
+ .devfn = PCI_DEVFN(0, 0),
+ .cap = ACS_TEST_CAP,
+ .ctrl = c->ctrl,
+ };
+ struct pci_bus *bus = kunit_kzalloc(test, sizeof(*bus), GFP_KERNEL);
+ struct pci_dev *pdev = kunit_kzalloc(test, sizeof(*pdev), GFP_KERNEL);
+
+ KUNIT_ASSERT_NOT_NULL(test, bus);
+ KUNIT_ASSERT_NOT_NULL(test, pdev);
+
+ bus->ops = &acs_ctrl_ops;
+ bus->sysdata = &cfg;
+
+ pdev->bus = bus;
+ pdev->devfn = cfg.devfn;
+ pdev->acs_cap = ACS_TEST_CAP;
+ pdev->acs_capabilities = ACS_ALL_CAPS;
+
+ KUNIT_EXPECT_EQ(test, pci_acs_flags_enabled(pdev, c->req), c->expect);
+}
+
+static bool acs_isolated(struct kunit *test, struct acs_ctrl_cfg *cfg,
+ u16 acs_cap, u16 acs_flags)
+{
+ struct pci_bus *bus = kunit_kzalloc(test, sizeof(*bus), GFP_KERNEL);
+ struct pci_dev *pdev = kunit_kzalloc(test, sizeof(*pdev), GFP_KERNEL);
+
+ KUNIT_ASSERT_NOT_NULL(test, bus);
+ KUNIT_ASSERT_NOT_NULL(test, pdev);
+
+ bus->ops = &acs_ctrl_ops;
+ bus->sysdata = cfg;
+
+ pdev->bus = bus;
+ pdev->devfn = cfg->devfn;
+ pdev->acs_cap = acs_cap;
+ pdev->acs_capabilities = ACS_ALL_CAPS;
+
+ return pci_acs_flags_enabled(pdev, acs_flags);
+}
+
+static void pci_acs_flags_no_cap_test(struct kunit *test)
+{
+ struct acs_ctrl_cfg cfg = {
+ .devfn = PCI_DEVFN(0, 0),
+ .cap = 0,
+ .ctrl = ACS_REQ_FLAGS,
+ };
+
+ KUNIT_EXPECT_FALSE(test, acs_isolated(test, &cfg, 0, ACS_REQ_FLAGS));
+}
+
+static void pci_acs_flags_read_fails_test(struct kunit *test)
+{
+ u16 no_rr = ACS_REQ_FLAGS & ~PCI_ACS_RR;
+ struct acs_ctrl_cfg cfg = {
+ .devfn = PCI_DEVFN(0, 0),
+ .cap = ACS_TEST_CAP,
+ .ctrl = ACS_REQ_FLAGS,
+ };
+
+ KUNIT_EXPECT_TRUE(test, acs_isolated(test, &cfg, ACS_TEST_CAP, no_rr));
+
+ cfg.fail_read = true;
+ KUNIT_EXPECT_FALSE(test, acs_isolated(test, &cfg, ACS_TEST_CAP, no_rr));
+}
+
/*
* Drive calc_map_type_and_dist() over a fabricated PCIe fabric matching the
* canonical topology of two devices below one switch:
@@ -569,6 +703,10 @@ static struct kunit_case pci_acs_test_cases[] = {
acs_request_gen_params),
KUNIT_CASE_PARAM(pci_acs_p2pdma_completion_test,
acs_completion_gen_params),
+ KUNIT_CASE_PARAM(pci_acs_flags_enabled_test,
+ acs_isolation_gen_params),
+ KUNIT_CASE(pci_acs_flags_no_cap_test),
+ KUNIT_CASE(pci_acs_flags_read_fails_test),
KUNIT_CASE(acs_walk_bus_addr_test),
KUNIT_CASE(acs_walk_request_redirect_test),
KUNIT_CASE(acs_walk_completion_redirect_test),
--
2.55.0
^ permalink raw reply [flat|nested] 38+ messages in thread
* [PATCH v9 15/18] PCI/P2PDMA: Document TLP-class routing
2026-10-01 11:55 [PATCH v9 00/18] PCI/P2PDMA: Route peer-to-peer DMA by TLP class Leon Romanovsky
` (13 preceding siblings ...)
2026-10-01 11:55 ` [PATCH v9 14/18] PCI: Add KUnit coverage for ACS isolation checks Leon Romanovsky
@ 2026-10-01 11:55 ` Leon Romanovsky
2026-10-01 11:55 ` [PATCH v9 16/18] PCI/P2PDMA: Let a client declare that it selects ATS per mapping Leon Romanovsky
` (3 subsequent siblings)
18 siblings, 0 replies; 38+ messages in thread
From: Leon Romanovsky @ 2026-10-01 11:55 UTC (permalink / raw)
To: Bjorn Helgaas, Logan Gunthorpe, Leon Romanovsky, Jason Gunthorpe,
Joerg Roedel (AMD),
Will Deacon, Robin Murphy, Christian König,
Thomas Hellström
Cc: linux-pci, linux-kernel, linux-doc, iommu, Tushar Dave,
linux-media, dri-devel, linaro-mm-sig, linux-rdma, kvm,
Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Sumit Semwal
From: Leon Romanovsky <leonro@nvidia.com>
The P2PDMA documentation stated that the mapping result is not defined for
Relaxed Ordering or ATS-translated Requests. It now is.
Replace that paragraph with what the three TLP-sensitive ACS controls do
and the table of outcomes per class. Record that an answer relying on
Relaxed Ordering Completions holds only for a provider that returns them,
which the PCIe specification leaves optional.
Tested-by: Tushar Dave <tdave@nvidia.com>
Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
---
Documentation/driver-api/pci/p2pdma.rst | 64 +++++++++++++++++++++++++++++----
1 file changed, 58 insertions(+), 6 deletions(-)
diff --git a/Documentation/driver-api/pci/p2pdma.rst b/Documentation/driver-api/pci/p2pdma.rst
index 42b18610bf7d..39a47b1e308e 100644
--- a/Documentation/driver-api/pci/p2pdma.rst
+++ b/Documentation/driver-api/pci/p2pdma.rst
@@ -28,12 +28,64 @@ through the host bridge when either applicable port redirects. If an ACS
Control register cannot be read, P2P DMA is rejected because the kernel cannot
establish a usable route.
-This evaluation assumes clients issue strictly ordered Requests carrying an
-Untranslated address. Its result is not defined when clients use Relaxed
-Ordering or issue ATS-translated Requests because those TLP attributes can
-select different routes through the fabric. Unless ACS Translation Blocking
-is enabled, a Port with ACS Direct Translated P2P enabled routes a
-Translated Request directly to the peer regardless of the redirect controls.
+Three of those controls act on TLP attributes that the client chooses rather
+than on the topology, so the same path routes differently for different
+traffic. ACS Translation Blocking rejects any Request whose Address Type is
+not Untranslated, and takes precedence over every other P2P control. ACS
+Direct Translated P2P routes a Translated Request to the peer regardless of
+Request Redirect and Egress Control. ACS Completion Redirect leaves alone
+Completions that carry the Relaxed Ordering attribute.
+
+P2PDMA therefore decides each class of traffic separately, and
+``pci_p2pdma_map_type()`` answers for the one a DMA mapping is built for:
+strictly ordered Requests carrying an Untranslated address.
+
+The two directions are decided independently. Translation Blocking (TB),
+Direct Translated P2P (DT), Request Redirect (RR) and Egress Control (EC) on
+the client-side port decide the Request:
+
+===== ===== ======= ============ ==========
+TB DT RR/EC TLP class Request
+===== ===== ======= ============ ==========
+set x x translated blocked
+clear set x translated direct
+clear clear clear translated direct
+clear clear set translated redirected
+x x clear untranslated direct
+x x set untranslated redirected
+===== ===== ======= ============ ==========
+
+Completion Redirect (CR) on the provider-side port decides the Completions:
+
+===== ========= ==========
+CR TLP class Completion
+===== ========= ==========
+x relaxed direct
+clear strict direct
+set strict redirected
+===== ========= ==========
+
+Combining the two tables gives the mapping type. A class is bus addressable
+when its Request and its Completions both route directly, and anything else
+goes through the host bridge. A blocked Request has neither route, because
+Translation Blocking rejects the Address Type itself rather than the
+destination. With no ACS control set anywhere, every class routes directly.
+
+Note that DT only matters where RR or EC would otherwise redirect: it
+overrides them for a Translated address rather than granting a direct route
+that was not already there.
+
+Translation Blocking is not a routing control, so it is evaluated on every
+port the Request passes rather than at the divergence alone. A direct route
+turns around at the divergence and only passes the ports below it, while the
+host bridge route keeps climbing and passes that port and everything above it
+as well. Neither route falls back to the other, because the Address Type is
+rejected wherever the Request is addressed.
+
+The Completer chooses whether a Completion carries Relaxed Ordering, and the
+PCIe specification does not require it to copy that attribute from the
+Request, so an answer that relies on Relaxed Ordering Completions holds only
+for a provider that does.
However, if the P2P transaction reaches the host bridge then it might have to
hairpin back out the same root port, be routed inside the CPU SOC to another
--
2.55.0
^ permalink raw reply [flat|nested] 38+ messages in thread
* [PATCH v9 16/18] PCI/P2PDMA: Let a client declare that it selects ATS per mapping
2026-10-01 11:55 [PATCH v9 00/18] PCI/P2PDMA: Route peer-to-peer DMA by TLP class Leon Romanovsky
` (14 preceding siblings ...)
2026-10-01 11:55 ` [PATCH v9 15/18] PCI/P2PDMA: Document TLP-class routing Leon Romanovsky
@ 2026-10-01 11:55 ` Leon Romanovsky
2026-10-01 11:55 ` [PATCH v9 17/18] PCI/P2PDMA: Evaluate the ATS path for clients with ATS enabled Leon Romanovsky
` (2 subsequent siblings)
18 siblings, 0 replies; 38+ messages in thread
From: Leon Romanovsky @ 2026-10-01 11:55 UTC (permalink / raw)
To: Bjorn Helgaas, Logan Gunthorpe, Leon Romanovsky, Jason Gunthorpe,
Joerg Roedel (AMD),
Will Deacon, Robin Murphy, Christian König,
Thomas Hellström
Cc: linux-pci, linux-kernel, linux-doc, iommu, Tushar Dave,
linux-media, dri-devel, linaro-mm-sig, linux-rdma, kvm,
Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Sumit Semwal
From: Leon Romanovsky <leonro@nvidia.com>
The PCIe ATS Enable bit covers the whole device, so P2PDMA cannot tell
from it whether a client will translate a given address. Most devices
translate any address once ATS is enabled, but some choose ATS per DMA
mapping and can still use bus addresses for the rest.
Add pcim_p2pdma_set_ats_per_mapping() so the driver of such a device can
declare that before P2PDMA starts treating clients with ATS enabled as
translating everything. Keep it in the device's P2PDMA state, which
already ends with the driver binding.
Reviewed-by: Thomas Hellström <thomas.hellstrom@linux.intel.com>
Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
---
drivers/pci/p2pdma.c | 27 +++++++++++++++++++++++++++
include/linux/pci-p2pdma.h | 4 ++++
2 files changed, 31 insertions(+)
diff --git a/drivers/pci/p2pdma.c b/drivers/pci/p2pdma.c
index c1e9d43bded2..126e28d2a5dd 100644
--- a/drivers/pci/p2pdma.c
+++ b/drivers/pci/p2pdma.c
@@ -26,6 +26,7 @@
struct pci_p2pdma {
struct gen_pool *pool;
bool p2pmem_published;
+ bool ats_per_mapping;
struct xarray map_types;
struct p2pdma_provider mem[PCI_STD_NUM_BARS];
};
@@ -1568,6 +1569,32 @@ void pci_p2pmem_publish(struct pci_dev *pdev, bool publish)
}
EXPORT_SYMBOL_GPL(pci_p2pmem_publish);
+/**
+ * pcim_p2pdma_set_ats_per_mapping - Declare per-mapping ATS for a client
+ * @pdev: PCI device that initiates peer-to-peer DMA
+ *
+ * Declare that @pdev issues Translated Requests only for the DMA mappings its
+ * driver sets up to use ATS, rather than for any address once ATS is enabled.
+ * P2PDMA then routes this client's Requests with the Address Type each query
+ * asks about, rather than the one its ATS Enable bit implies.
+ */
+void pcim_p2pdma_set_ats_per_mapping(struct pci_dev *pdev)
+{
+ struct pci_p2pdma *p2p;
+
+ p2p = rcu_dereference_protected(pdev->p2pdma, 1);
+ if (!p2p)
+ /*
+ * ats_per_mapping is a performance optimization,
+ * if pcim_p2pdma_init() didn't set pdev->p2pdma pointer
+ * for some reason, let's simply use ATS global settings.
+ */
+ return;
+
+ p2p->ats_per_mapping = true;
+}
+EXPORT_SYMBOL_GPL(pcim_p2pdma_set_ats_per_mapping);
+
/**
* pci_p2pdma_map_type_tlp - Determine the mapping type for P2PDMA transfers
* @provider: P2PDMA provider structure
diff --git a/include/linux/pci-p2pdma.h b/include/linux/pci-p2pdma.h
index dd17501ba1b6..ce389e18d023 100644
--- a/include/linux/pci-p2pdma.h
+++ b/include/linux/pci-p2pdma.h
@@ -82,6 +82,7 @@ struct scatterlist *pci_p2pmem_alloc_sgl(struct pci_dev *pdev,
unsigned int *nents, u32 length);
void pci_p2pmem_free_sgl(struct pci_dev *pdev, struct scatterlist *sgl);
void pci_p2pmem_publish(struct pci_dev *pdev, bool publish);
+void pcim_p2pdma_set_ats_per_mapping(struct pci_dev *pdev);
int pci_p2pdma_enable_store(const char *page, struct pci_dev **p2p_dev,
bool *use_p2pdma);
ssize_t pci_p2pdma_enable_show(char *page, struct pci_dev *p2p_dev,
@@ -138,6 +139,9 @@ static inline void pci_p2pmem_free_sgl(struct pci_dev *pdev,
static inline void pci_p2pmem_publish(struct pci_dev *pdev, bool publish)
{
}
+static inline void pcim_p2pdma_set_ats_per_mapping(struct pci_dev *pdev)
+{
+}
static inline int pci_p2pdma_enable_store(const char *page,
struct pci_dev **p2p_dev, bool *use_p2pdma)
{
--
2.55.0
^ permalink raw reply [flat|nested] 38+ messages in thread
* [PATCH v9 17/18] PCI/P2PDMA: Evaluate the ATS path for clients with ATS enabled
2026-10-01 11:55 [PATCH v9 00/18] PCI/P2PDMA: Route peer-to-peer DMA by TLP class Leon Romanovsky
` (15 preceding siblings ...)
2026-10-01 11:55 ` [PATCH v9 16/18] PCI/P2PDMA: Let a client declare that it selects ATS per mapping Leon Romanovsky
@ 2026-10-01 11:55 ` Leon Romanovsky
2026-10-01 11:55 ` [PATCH v9 18/18] PCI/P2PDMA: Test the routing of " Leon Romanovsky
2026-10-06 19:29 ` [PATCH v9 00/18] PCI/P2PDMA: Route peer-to-peer DMA by TLP class Bjorn Helgaas
18 siblings, 0 replies; 38+ messages in thread
From: Leon Romanovsky @ 2026-10-01 11:55 UTC (permalink / raw)
To: Bjorn Helgaas, Logan Gunthorpe, Leon Romanovsky, Jason Gunthorpe,
Joerg Roedel (AMD),
Will Deacon, Robin Murphy, Christian König,
Thomas Hellström
Cc: linux-pci, linux-kernel, linux-doc, iommu, Tushar Dave,
linux-media, dri-devel, linaro-mm-sig, linux-rdma, kvm,
Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Sumit Semwal
From: Leon Romanovsky <leonro@nvidia.com>
P2PDMA assumes every client issues Untranslated Requests. A device with
ATS enabled may translate any address it is handed, including a bus
address, and Translation Blocking can reject its Translated Requests on
a route the Untranslated answer called usable.
Unless the client declared per-mapping ATS, take the Address Type from
its ATS Enable bit. Never hand such a client bus addresses: only an IOVA
survives translation, and whatever it issues untranslated goes through
the host bridge, whose route must therefore be usable. Apply this after
the cache, which keeps holding answers that depend only on the topology.
Reviewed-by: Thomas Hellström <thomas.hellstrom@linux.intel.com>
Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
---
Documentation/driver-api/pci/p2pdma.rst | 8 +++++
drivers/pci/p2pdma.c | 53 +++++++++++++++++++++++++++++++--
2 files changed, 58 insertions(+), 3 deletions(-)
diff --git a/Documentation/driver-api/pci/p2pdma.rst b/Documentation/driver-api/pci/p2pdma.rst
index 39a47b1e308e..a5030a743e69 100644
--- a/Documentation/driver-api/pci/p2pdma.rst
+++ b/Documentation/driver-api/pci/p2pdma.rst
@@ -40,6 +40,14 @@ P2PDMA therefore decides each class of traffic separately, and
``pci_p2pdma_map_type()`` answers for the one a DMA mapping is built for:
strictly ordered Requests carrying an Untranslated address.
+The PCIe ATS Enable bit covers the whole device, and most devices translate
+any address they are handed once it is set. Unless a driver has declared with
+``pcim_p2pdma_set_ats_per_mapping()`` that its device chooses ATS per mapping,
+that bit decides which Address Type P2PDMA assumes, and a client with ATS
+enabled is never handed a bus address, which it would translate as though it
+were an IOVA. Its Translated Requests may still route directly, but anything
+it issues untranslated reaches the host bridge, so that route has to work too.
+
The two directions are decided independently. Translation Blocking (TB),
Direct Translated P2P (DT), Request Redirect (RR) and Egress Control (EC) on
the client-side port decide the Request:
diff --git a/drivers/pci/p2pdma.c b/drivers/pci/p2pdma.c
index 126e28d2a5dd..9d17a3d2bf1f 100644
--- a/drivers/pci/p2pdma.c
+++ b/drivers/pci/p2pdma.c
@@ -1595,6 +1595,41 @@ void pcim_p2pdma_set_ats_per_mapping(struct pci_dev *pdev)
}
EXPORT_SYMBOL_GPL(pcim_p2pdma_set_ats_per_mapping);
+static unsigned int pci_p2pdma_client_tlp_flags(struct pci_dev *client,
+ unsigned int tlp_flags)
+{
+ if (client->ats_enabled)
+ return tlp_flags | PCI_P2PDMA_TLP_TRANSLATED;
+
+ return tlp_flags & ~PCI_P2PDMA_TLP_TRANSLATED;
+}
+
+/*
+ * A client that translates every address it is handed cannot be handed a bus
+ * address, which it would translate as though it were an IOVA. Its Translated
+ * Requests may still route directly, but a Request it issues without a
+ * translation carries the IOVA to the host bridge, so that route has to work
+ * as well.
+ */
+static enum pci_p2pdma_map_type
+pci_p2pdma_client_map_type(struct pci_dev *provider, struct pci_dev *client,
+ bool per_mapping, enum pci_p2pdma_map_type type)
+{
+ if (per_mapping || !client->ats_enabled ||
+ type != PCI_P2PDMA_MAP_BUS_ADDR)
+ return type;
+
+ pci_dbg(client,
+ "P2PDMA ACS: provider=%s bus address withheld; ATS is enabled for the whole client\n",
+ pci_name(provider));
+
+ if (cpu_supports_p2pdma() ||
+ host_bridge_whitelist(provider, client, false))
+ return PCI_P2PDMA_MAP_THRU_HOST_BRIDGE;
+
+ return PCI_P2PDMA_MAP_NOT_SUPPORTED;
+}
+
/**
* pci_p2pdma_map_type_tlp - Determine the mapping type for P2PDMA transfers
* @provider: P2PDMA provider structure
@@ -1609,6 +1644,11 @@ EXPORT_SYMBOL_GPL(pcim_p2pdma_set_ats_per_mapping);
* ACS routes a peer-to-peer transaction by the attributes its TLPs carry, so
* the answer depends on @tlp_flags. A caller that passes flags its traffic
* does not match gets a mapping the fabric will not deliver.
+ *
+ * Only a client whose driver called pcim_p2pdma_set_ats_per_mapping() takes
+ * the Address Type from @tlp_flags. For any other client its ATS Enable bit
+ * decides, and a client that translates every address it is handed never gets
+ * %PCI_P2PDMA_MAP_BUS_ADDR.
*/
static enum pci_p2pdma_map_type
pci_p2pdma_map_type_tlp(struct p2pdma_provider *provider, struct device *dev,
@@ -1620,6 +1660,7 @@ pci_p2pdma_map_type_tlp(struct p2pdma_provider *provider, struct device *dev,
struct pci_p2pdma *p2pdma;
struct pci_dev *client;
bool provider_state;
+ bool per_mapping;
int dist;
if (!pdev->p2pdma) {
@@ -1639,6 +1680,9 @@ pci_p2pdma_map_type_tlp(struct p2pdma_provider *provider, struct device *dev,
cache_index = map_types_idx(client);
rcu_read_lock();
+ /* The declaration belongs to the client, the cache to the provider. */
+ p2pdma = rcu_dereference(client->p2pdma);
+ per_mapping = p2pdma && p2pdma->ats_per_mapping;
p2pdma = rcu_dereference(pdev->p2pdma);
if (p2pdma)
@@ -1646,6 +1690,8 @@ pci_p2pdma_map_type_tlp(struct p2pdma_provider *provider, struct device *dev,
cache_index));
provider_state = !!p2pdma;
rcu_read_unlock();
+ if (!per_mapping)
+ tlp_flags = pci_p2pdma_client_tlp_flags(client, tlp_flags);
type = pci_p2pdma_map_types_unpack(cached, tlp_flags);
pci_dbg(client,
"P2PDMA ACS: map lookup provider=%s index=%#lx tlp-flags=%#x cached=%s(%d) provider-state=%u\n",
@@ -1653,10 +1699,10 @@ pci_p2pdma_map_type_tlp(struct p2pdma_provider *provider, struct device *dev,
pci_p2pdma_map_type_name(type), type, provider_state);
if (type == PCI_P2PDMA_MAP_UNKNOWN)
- return calc_map_type_and_dist(pdev, client, &dist, tlp_flags,
+ type = calc_map_type_and_dist(pdev, client, &dist, tlp_flags,
true);
- return type;
+ return pci_p2pdma_client_map_type(pdev, client, per_mapping, type);
}
/**
@@ -1665,7 +1711,8 @@ pci_p2pdma_map_type_tlp(struct p2pdma_provider *provider, struct device *dev,
* @dev: Client device that initiates the transfer
*
* Same as pci_p2pdma_map_type_tlp() for a client issuing strictly ordered
- * Requests that carry an Untranslated address.
+ * Requests. Their Address Type is Untranslated unless the client enables ATS
+ * for the whole device.
*/
enum pci_p2pdma_map_type pci_p2pdma_map_type(struct p2pdma_provider *provider,
struct device *dev)
--
2.55.0
^ permalink raw reply [flat|nested] 38+ messages in thread
* [PATCH v9 18/18] PCI/P2PDMA: Test the routing of clients with ATS enabled
2026-10-01 11:55 [PATCH v9 00/18] PCI/P2PDMA: Route peer-to-peer DMA by TLP class Leon Romanovsky
` (16 preceding siblings ...)
2026-10-01 11:55 ` [PATCH v9 17/18] PCI/P2PDMA: Evaluate the ATS path for clients with ATS enabled Leon Romanovsky
@ 2026-10-01 11:55 ` Leon Romanovsky
2026-10-06 19:29 ` [PATCH v9 00/18] PCI/P2PDMA: Route peer-to-peer DMA by TLP class Bjorn Helgaas
18 siblings, 0 replies; 38+ messages in thread
From: Leon Romanovsky @ 2026-10-01 11:55 UTC (permalink / raw)
To: Bjorn Helgaas, Logan Gunthorpe, Leon Romanovsky, Jason Gunthorpe,
Joerg Roedel (AMD),
Will Deacon, Robin Murphy, Christian König,
Thomas Hellström
Cc: linux-pci, linux-kernel, linux-doc, iommu, Tushar Dave,
linux-media, dri-devel, linaro-mm-sig, linux-rdma, kvm,
Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Sumit Semwal
From: Leon Romanovsky <leonro@nvidia.com>
A client that enables ATS for the whole device must never be handed a
bus address, and Translation Blocking must reject its peer-to-peer
traffic even when the caller asks for the default class. That has to
hold for answers taken from the cache as well.
Cover those on the fabricated fabric, together with a client that
declared per-mapping ATS keeping the per-class answers and a client
without ATS answering for the Untranslated Requests it will issue.
Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
---
drivers/pci/p2pdma.c | 8 +--
drivers/pci/pci.h | 5 ++
drivers/pci/pci_acs_test.c | 123 +++++++++++++++++++++++++++++++++++++++++++++
3 files changed, 133 insertions(+), 3 deletions(-)
diff --git a/drivers/pci/p2pdma.c b/drivers/pci/p2pdma.c
index 9d17a3d2bf1f..38c712d77f15 100644
--- a/drivers/pci/p2pdma.c
+++ b/drivers/pci/p2pdma.c
@@ -1595,14 +1595,15 @@ void pcim_p2pdma_set_ats_per_mapping(struct pci_dev *pdev)
}
EXPORT_SYMBOL_GPL(pcim_p2pdma_set_ats_per_mapping);
-static unsigned int pci_p2pdma_client_tlp_flags(struct pci_dev *client,
- unsigned int tlp_flags)
+VISIBLE_IF_KUNIT unsigned int
+pci_p2pdma_client_tlp_flags(struct pci_dev *client, unsigned int tlp_flags)
{
if (client->ats_enabled)
return tlp_flags | PCI_P2PDMA_TLP_TRANSLATED;
return tlp_flags & ~PCI_P2PDMA_TLP_TRANSLATED;
}
+EXPORT_SYMBOL_IF_KUNIT(pci_p2pdma_client_tlp_flags);
/*
* A client that translates every address it is handed cannot be handed a bus
@@ -1611,7 +1612,7 @@ static unsigned int pci_p2pdma_client_tlp_flags(struct pci_dev *client,
* translation carries the IOVA to the host bridge, so that route has to work
* as well.
*/
-static enum pci_p2pdma_map_type
+VISIBLE_IF_KUNIT enum pci_p2pdma_map_type
pci_p2pdma_client_map_type(struct pci_dev *provider, struct pci_dev *client,
bool per_mapping, enum pci_p2pdma_map_type type)
{
@@ -1629,6 +1630,7 @@ pci_p2pdma_client_map_type(struct pci_dev *provider, struct pci_dev *client,
return PCI_P2PDMA_MAP_NOT_SUPPORTED;
}
+EXPORT_SYMBOL_IF_KUNIT(pci_p2pdma_client_map_type);
/**
* pci_p2pdma_map_type_tlp - Determine the mapping type for P2PDMA transfers
diff --git a/drivers/pci/pci.h b/drivers/pci/pci.h
index a4a07f31f744..8b10965ab536 100644
--- a/drivers/pci/pci.h
+++ b/drivers/pci/pci.h
@@ -1144,6 +1144,11 @@ enum pci_p2pdma_map_type calc_map_type_and_dist(struct pci_dev *provider,
int *dist,
unsigned int tlp_flags,
bool verbose);
+unsigned int pci_p2pdma_client_tlp_flags(struct pci_dev *client,
+ unsigned int tlp_flags);
+enum pci_p2pdma_map_type
+pci_p2pdma_client_map_type(struct pci_dev *provider, struct pci_dev *client,
+ bool per_mapping, enum pci_p2pdma_map_type type);
#endif
#ifdef CONFIG_PCI_QUIRKS
diff --git a/drivers/pci/pci_acs_test.c b/drivers/pci/pci_acs_test.c
index 6a2cd38d968f..a9a3b80ea0ac 100644
--- a/drivers/pci/pci_acs_test.c
+++ b/drivers/pci/pci_acs_test.c
@@ -698,6 +698,125 @@ static void acs_walk_unreadable_control_test(struct kunit *test)
PCI_P2PDMA_MAP_NOT_SUPPORTED);
}
+/*
+ * Route the way pci_p2pdma_map_type_tlp() does once it has a client: the
+ * client decides which flags apply, and the answer is then checked against
+ * what the client can be handed.
+ */
+static enum pci_p2pdma_map_type acs_client_map(struct acs_fabric *f,
+ bool per_mapping,
+ unsigned int tlp_flags)
+{
+ enum pci_p2pdma_map_type type;
+ int dist;
+
+ if (!per_mapping)
+ tlp_flags = pci_p2pdma_client_tlp_flags(f->client, tlp_flags);
+ type = calc_map_type_and_dist(f->provider, f->client, &dist, tlp_flags,
+ false);
+ return pci_p2pdma_client_map_type(f->provider, f->client, per_mapping,
+ type);
+}
+
+/*
+ * A client with ATS enabled for the whole device translates whatever it is
+ * handed, so it never gets bus addresses even on a direct route, and
+ * Translation Blocking rejects it even when its caller asks for the default
+ * class.
+ */
+static void acs_client_ats_test(struct kunit *test)
+{
+ struct acs_fabric f = {};
+
+ acs_build_fabric(test, &f);
+ f.client->ats_enabled = 1;
+
+ KUNIT_EXPECT_EQ(test, acs_client_map(&f, false, 0),
+ PCI_P2PDMA_MAP_THRU_HOST_BRIDGE);
+
+ f.client_leaf->acs_cap = 0x100;
+ f.client_leaf->acs_capabilities = PCI_ACS_TB;
+ f.client_leaf_cfg.ctrl = PCI_ACS_TB;
+
+ KUNIT_EXPECT_EQ(test, acs_client_map(&f, false, 0),
+ PCI_P2PDMA_MAP_NOT_SUPPORTED);
+}
+
+/*
+ * The cache holds what the topology allows, filled when the client first
+ * asked. A client whose ATS was enabled after that must still be kept off the
+ * bus addresses the cache recorded.
+ */
+static void acs_client_ats_cached_test(struct kunit *test)
+{
+ enum pci_p2pdma_map_type type[PCI_P2PDMA_TLP_CLASSES], hit;
+ struct acs_fabric f = {};
+ unsigned long cached;
+ unsigned int flags;
+ int dist;
+
+ acs_build_fabric(test, &f);
+ for (flags = 0; flags < PCI_P2PDMA_TLP_CLASSES; flags++)
+ type[flags] = calc_map_type_and_dist(f.provider, f.client,
+ &dist, flags, false);
+ cached = pci_p2pdma_map_types_pack(type);
+
+ f.client->ats_enabled = 1;
+ flags = pci_p2pdma_client_tlp_flags(f.client, 0);
+ hit = pci_p2pdma_map_types_unpack(cached, flags);
+
+ KUNIT_EXPECT_EQ(test,
+ pci_p2pdma_client_map_type(f.provider, f.client, false,
+ hit),
+ PCI_P2PDMA_MAP_THRU_HOST_BRIDGE);
+}
+
+/*
+ * A client that declared per-mapping ATS picks the Address Type per mapping,
+ * so it keeps the per-class answers even with ATS enabled.
+ */
+static void acs_client_ats_per_mapping_test(struct kunit *test)
+{
+ struct acs_fabric f = {};
+
+ acs_build_fabric(test, &f);
+ f.client->ats_enabled = 1;
+
+ KUNIT_EXPECT_EQ(test, acs_client_map(&f, true, 0),
+ PCI_P2PDMA_MAP_BUS_ADDR);
+ KUNIT_EXPECT_EQ(test,
+ acs_client_map(&f, true, PCI_P2PDMA_TLP_TRANSLATED),
+ PCI_P2PDMA_MAP_BUS_ADDR);
+
+ f.client_leaf->acs_cap = 0x100;
+ f.client_leaf->acs_capabilities = PCI_ACS_TB;
+ f.client_leaf_cfg.ctrl = PCI_ACS_TB;
+
+ KUNIT_EXPECT_EQ(test, acs_client_map(&f, true, 0),
+ PCI_P2PDMA_MAP_BUS_ADDR);
+ KUNIT_EXPECT_EQ(test,
+ acs_client_map(&f, true, PCI_P2PDMA_TLP_TRANSLATED),
+ PCI_P2PDMA_MAP_NOT_SUPPORTED);
+}
+
+/*
+ * A client with ATS disabled cannot issue Translated Requests, so a caller
+ * asking about them gets the answer for the Untranslated ones it will issue.
+ */
+static void acs_client_no_ats_test(struct kunit *test)
+{
+ struct acs_fabric f = {};
+
+ acs_build_fabric(test, &f);
+ f.dn1->acs_cap = 0x100;
+ f.dn1->acs_capabilities = PCI_ACS_RR | PCI_ACS_DT;
+ f.dn1_cfg.ctrl = PCI_ACS_RR | PCI_ACS_DT;
+
+ KUNIT_EXPECT_EQ(test,
+ acs_client_map(&f, false, PCI_P2PDMA_TLP_TRANSLATED),
+ PCI_P2PDMA_MAP_THRU_HOST_BRIDGE);
+}
+
static struct kunit_case pci_acs_test_cases[] = {
KUNIT_CASE_PARAM(pci_acs_p2pdma_request_test,
acs_request_gen_params),
@@ -723,6 +842,10 @@ static struct kunit_case pci_acs_test_cases[] = {
KUNIT_CASE(acs_walk_tb_above_divergence_test),
KUNIT_CASE(acs_walk_unreadable_control_test),
KUNIT_CASE(acs_map_types_pack_test),
+ KUNIT_CASE(acs_client_ats_test),
+ KUNIT_CASE(acs_client_ats_cached_test),
+ KUNIT_CASE(acs_client_ats_per_mapping_test),
+ KUNIT_CASE(acs_client_no_ats_test),
{}
};
--
2.55.0
^ permalink raw reply [flat|nested] 38+ messages in thread
* Re: [PATCH v9 00/18] PCI/P2PDMA: Route peer-to-peer DMA by TLP class
2026-10-01 11:55 [PATCH v9 00/18] PCI/P2PDMA: Route peer-to-peer DMA by TLP class Leon Romanovsky
` (17 preceding siblings ...)
2026-10-01 11:55 ` [PATCH v9 18/18] PCI/P2PDMA: Test the routing of " Leon Romanovsky
@ 2026-10-06 19:29 ` Bjorn Helgaas
2026-10-06 21:22 ` Leon Romanovsky
18 siblings, 1 reply; 38+ messages in thread
From: Bjorn Helgaas @ 2026-10-06 19:29 UTC (permalink / raw)
To: Leon Romanovsky
Cc: Bjorn Helgaas, Logan Gunthorpe, Jason Gunthorpe,
Joerg Roedel (AMD),
Will Deacon, Robin Murphy, Christian König,
Thomas Hellström, linux-pci, linux-kernel, linux-doc, iommu,
Tushar Dave, linux-media, dri-devel, linaro-mm-sig, linux-rdma,
kvm, Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Sumit Semwal
On Thu, Oct 01, 2026 at 02:55:08PM +0300, Leon Romanovsky wrote:
> PCI P2PDMA applies Request and Completion Redirect throughout both paths.
> This misclassifies asymmetric and nested switches, and reports one answer
> for every kind of TLP.
>
> Three ACS controls act on TLP attributes the client chooses rather than
> on the topology: Translation Blocking and Direct Translated P2P act on
> a Request's Address Type, and Completion Redirect skips Completions carrying
> Relaxed Ordering.
>
> Evaluate each direction at the path divergence, decide every class from the
> one walk, and treat a client with ATS enabled as translating unless its
> driver declares per-mapping ATS.
>
> This completes the P2PDMA side; dma-buf and mlx5 follow separately.
> ...
> Leon Romanovsky (18):
> PCI/P2PDMA: Document the TLP attribute assumptions
> PCI/P2PDMA: Derive routing from directional ACS controls
> PCI: Reject unreadable ACS controls in isolation checks
> PCI/P2PDMA: Evaluate ACS controls at the path divergence
> PCI/P2PDMA: Document directional ACS routing
> PCI/P2PDMA: Collect the path's ACS controls before deciding
> PCI/P2PDMA: Answer routing per TLP class
> PCI/P2PDMA: Route Relaxed Ordering Completions directly
> PCI/P2PDMA: Reject Translated Requests blocked by Translation Blocking
> PCI/P2PDMA: Route Translated Requests under Direct Translated P2P
> PCI/P2PDMA: Log detailed ACS routing diagnostics
> PCI/P2PDMA: Add KUnit tests for the ACS routing decisions
> PCI/P2PDMA: Test the ACS P2P routing walk
> PCI: Add KUnit coverage for ACS isolation checks
> PCI/P2PDMA: Document TLP-class routing
> PCI/P2PDMA: Let a client declare that it selects ATS per mapping
> PCI/P2PDMA: Evaluate the ATS path for clients with ATS enabled
> PCI/P2PDMA: Test the routing of clients with ATS enabled
>
> Documentation/admin-guide/kernel-parameters.txt | 15 +-
> Documentation/driver-api/pci/p2pdma.rst | 80 +++
> drivers/pci/Kconfig | 15 +
> drivers/pci/Makefile | 1 +
> drivers/pci/p2pdma.c | 719 ++++++++++++++++++--
> drivers/pci/pci.c | 7 +-
> drivers/pci/pci.h | 58 ++
> drivers/pci/pci_acs_test.c | 860 ++++++++++++++++++++++++
> drivers/pci/quirks.c | 6 +-
> include/linux/pci-p2pdma.h | 12 +-
> 10 files changed, 1687 insertions(+), 86 deletions(-)
> ---
> base-commit: 08dbfad3f5040f5bdb6c529da20d6d4e81fefd72
> change-id: 20260821-fix-p2p-acs-v4-0-e72455e3a261
> prerequisite-message-id: <20260830-batch-p2p-fixes-v1-0-5044e8dfbe2e@nvidia.com>
> prerequisite-patch-id: 6b25c7fcf164cdfc14e9fac5b908d97fcf6509d7
> prerequisite-patch-id: 0d083c281001365aae4b35544cf28891a6ab9a96
> prerequisite-patch-id: bfd9dabf271f3cc9a3a61f46387d20c20311363d
> prerequisite-patch-id: fad0275efc722830fc591509506c0a5e4f581073
> prerequisite-patch-id: 0c83bee688fec1f6d1564654df7c630fa6a4a978
This seems like material for the PCI tree, but I'm not sure how to
apply it. It doesn't apply cleanly on the current pci/p2pdma
(https://git.kernel.org/pub/scm/linux/kernel/git/pci/pci.git/log/?h=p2pdma),
which does contain your series from
20260830-batch-p2p-fixes-v1-0-5044e8dfbe2e@nvidia.com.
I could probably fix the conflicts but I don't know why there should
be conflicts, since the only commits on pci/p2pdma other than yours
are a few trivial allow-list updates.
Bjorn
^ permalink raw reply [flat|nested] 38+ messages in thread
* Re: [PATCH v9 02/18] PCI/P2PDMA: Derive routing from directional ACS controls
2026-10-01 11:55 ` [PATCH v9 02/18] PCI/P2PDMA: Derive routing from directional ACS controls Leon Romanovsky
@ 2026-10-06 20:49 ` Bjorn Helgaas
2026-10-07 6:32 ` Leon Romanovsky
0 siblings, 1 reply; 38+ messages in thread
From: Bjorn Helgaas @ 2026-10-06 20:49 UTC (permalink / raw)
To: Leon Romanovsky
Cc: Bjorn Helgaas, Logan Gunthorpe, Jason Gunthorpe,
Joerg Roedel (AMD),
Will Deacon, Robin Murphy, Christian König,
Thomas Hellström, linux-pci, linux-kernel, linux-doc, iommu,
Tushar Dave, linux-media, dri-devel, linaro-mm-sig, linux-rdma,
kvm, Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Sumit Semwal
On Thu, Oct 01, 2026 at 02:55:10PM +0300, Leon Romanovsky wrote:
> From: Leon Romanovsky <leonro@nvidia.com>
>
> pci_bridge_has_acs_redir() treats Request and Completion Redirect as
> interchangeable. On asymmetric fabrics, a control for only the reverse TLP
> direction can unnecessarily force P2PDMA through the host bridge.
Does "the reverse TLP direction" refer to Completions?
> Evaluate Request Redirect for client Requests and Completion Redirect for
> provider read Completions. Continue treating enabled Egress Control
> conservatively as a Request redirect.
Completion Redirect is intended to avoid ordering rule violations
between Completions and Requests when Requests are redirected (PCIe
r7.0, sec 6.12.1.1). I assume this patch preserves the ordering rule,
but does the commit log need to say something about that? I don't
know enough about P2P DMA for it to be obvious to me.
Not really a question for this series, but p2pdma.c and p2pdma.rst
refer to "clients" and "providers", neither of which are mentioned in
the PCIe spec. In this case it sounds like a client is a Requester
and a provider is a Completer in spec terms. Is that always the case?
If so, "client" and "provider" in this paragraph are not adding any
information.
If "client" is not the same concept as "Requester" and "provider" not
the same as "Completer", maybe p2pdma.rst could explain the
difference?
> Fixes: 52916982af48 ("PCI/P2PDMA: Support peer-to-peer memory")
> Reviewed-by: Logan Gunthorpe <logang@deltatee.com>
> Tested-by: Tushar Dave <tdave@nvidia.com>
> Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
> ---
> drivers/pci/p2pdma.c | 75 ++++++++++++++++++++++++++++++++++++++++------------
> 1 file changed, 58 insertions(+), 17 deletions(-)
>
> diff --git a/drivers/pci/p2pdma.c b/drivers/pci/p2pdma.c
> index 4e4d2df17a45..12612b82d80d 100644
> --- a/drivers/pci/p2pdma.c
> +++ b/drivers/pci/p2pdma.c
> @@ -21,6 +21,8 @@
> #include <linux/seq_buf.h>
> #include <linux/xarray.h>
>
> +#include "pci.h"
> +
> struct pci_p2pdma {
> struct gen_pool *pool;
> bool p2pmem_published;
> @@ -490,26 +492,56 @@ static struct pci_dev *find_parent_pci_dev(struct device *dev)
> return NULL;
> }
>
> +enum pci_acs_p2pdma_state {
> + PCI_ACS_P2PDMA_DIRECT,
> + PCI_ACS_P2PDMA_REDIRECT,
> +};
> +
> /*
> - * Check if a PCI bridge has its ACS redirection bits set to redirect P2P
> - * TLPs upstream via ACS. Returns 1 if the packets will be redirected
> - * upstream, 0 otherwise.
> + * Decide how a peer-to-peer Request at an ACS-capable ingress port routes,
> + * from that port's ACS Control register.
> + *
> + * Linux does not read the Egress Control Vector, so Egress Control is treated
> + * conservatively as a redirect. Per PCIe r7.0 Table 6-11 the outcomes it
> + * selects are a direct route and an ACS Violation, and neither one lets peer
> + * bus addressing be assumed.
> */
> -static int pci_bridge_has_acs_redir(struct pci_dev *pdev)
> +static enum pci_acs_p2pdma_state
> +pci_acs_p2pdma_request(u16 ctrl)
> {
> - int pos;
> - u16 ctrl;
> + return ctrl & (PCI_ACS_RR | PCI_ACS_EC) ?
> + PCI_ACS_P2PDMA_REDIRECT : PCI_ACS_P2PDMA_DIRECT;
> +}
>
> - pos = pdev->acs_cap;
> - if (!pos)
> - return 0;
> +/*
> + * Decide how a peer-to-peer Completion at an ACS-capable ingress port routes.
> + * PCIe r7.0 sec 6.12.1.1: no ACS control other than P2P Completion Redirect
> + * affects a Completion.
> + */
> +static enum pci_acs_p2pdma_state
> +pci_acs_p2pdma_completion(u16 ctrl)
> +{
> + return ctrl & PCI_ACS_CR ? PCI_ACS_P2PDMA_REDIRECT :
> + PCI_ACS_P2PDMA_DIRECT;
> +}
>
> - pci_read_config_word(pdev, pos + PCI_ACS_CTRL, &ctrl);
> +/*
> + * Read @pdev's ACS Control register. A device without an ACS capability has
> + * no peer-to-peer controls at all, which routes the same as having them all
> + * clear. Returns false when the register is present but cannot be read; @ctrl
> + * is then meaningless.
> + */
> +static bool pci_acs_p2pdma_ctrl(struct pci_dev *pdev, u16 *ctrl)
> +{
> + int pos;
>
> - if (ctrl & (PCI_ACS_RR | PCI_ACS_CR | PCI_ACS_EC))
> - return 1;
> + pos = pdev->acs_cap;
> + if (!pos) {
> + *ctrl = 0;
> + return true;
> + }
>
> - return 0;
> + return !pci_read_config_word(pdev, pos + PCI_ACS_CTRL, ctrl);
> }
>
> static void seq_buf_print_bus_devfn(struct seq_buf *buf, struct pci_dev *pdev)
> @@ -698,6 +730,10 @@ static unsigned long map_types_idx(struct pci_dev *client)
> * then to Device B. The mapping type returned depends on the ACS
> * redirection setting of the ports along the path.
> *
> + * The client initiates Requests to provider memory. Check Request Redirect
> + * on the client path and Completion Redirect for read Completions on the
> + * provider path.
> + *
> * If ACS redirect is set on any port in the path, traffic between the
> * devices will go through the host bridge, so return
> * PCI_P2PDMA_MAP_THRU_HOST_BRIDGE; otherwise return
> @@ -721,6 +757,7 @@ calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client,
> int dist_a = 0;
> int dist_b = 0;
> char buf[128];
> + u16 ctrl;
>
> seq_buf_init(&acs_list, buf, sizeof(buf));
>
> @@ -732,7 +769,9 @@ calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client,
> while (a) {
> dist_b = 0;
>
> - if (pci_bridge_has_acs_redir(a)) {
> + if (!pci_acs_p2pdma_ctrl(a, &ctrl) ||
> + pci_acs_p2pdma_completion(ctrl) ==
> + PCI_ACS_P2PDMA_REDIRECT) {
> seq_buf_print_bus_devfn(&acs_list, a);
> acs_cnt++;
> }
> @@ -761,7 +800,9 @@ calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client,
> if (a == bb)
> break;
>
> - if (pci_bridge_has_acs_redir(bb)) {
> + if (!pci_acs_p2pdma_ctrl(bb, &ctrl) ||
> + pci_acs_p2pdma_request(ctrl) ==
> + PCI_ACS_P2PDMA_REDIRECT) {
> seq_buf_print_bus_devfn(&acs_list, bb);
> acs_cnt++;
> }
> @@ -1109,10 +1150,10 @@ EXPORT_SYMBOL_GPL(pci_p2pmem_publish);
> /**
> * pci_p2pdma_map_type - Determine the mapping type for P2PDMA transfers
> * @provider: P2PDMA provider structure
> - * @dev: Target device for the transfer
> + * @dev: Client device that initiates the transfer
> *
> * Determines how peer-to-peer DMA transfers should be mapped between
> - * the provider and the target device. The mapping type indicates whether
> + * the provider and the client device. The mapping type indicates whether
> * the transfer can be done directly through PCI switches or must go
> * through the host bridge.
> */
>
> --
> 2.55.0
>
^ permalink raw reply [flat|nested] 38+ messages in thread
* Re: [PATCH v9 04/18] PCI/P2PDMA: Evaluate ACS controls at the path divergence
2026-10-01 11:55 ` [PATCH v9 04/18] PCI/P2PDMA: Evaluate ACS controls at the path divergence Leon Romanovsky
@ 2026-10-06 21:08 ` Bjorn Helgaas
2026-10-07 10:51 ` Leon Romanovsky
0 siblings, 1 reply; 38+ messages in thread
From: Bjorn Helgaas @ 2026-10-06 21:08 UTC (permalink / raw)
To: Leon Romanovsky
Cc: Bjorn Helgaas, Logan Gunthorpe, Jason Gunthorpe,
Joerg Roedel (AMD),
Will Deacon, Robin Murphy, Christian König,
Thomas Hellström, linux-pci, linux-kernel, linux-doc, iommu,
Tushar Dave, linux-media, dri-devel, linaro-mm-sig, linux-rdma,
kvm, Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Sumit Semwal
On Thu, Oct 01, 2026 at 02:55:12PM +0300, Leon Romanovsky wrote:
> From: Leon Romanovsky <leonro@nvidia.com>
>
> ACS redirect controls choose between peer and upstream routes only at the
> path divergence. Applying them below that point rejects valid nested
> topologies because traffic already has only an upstream route.
I guess the point here is that prior to this patch,
calc_map_type_and_dist() returned PCI_P2PDMA_MAP_NOT_SUPPORTED in a
case where it didn't need to? Can you include an example to make this
concrete?
It looks like in v7.3, we only return PCI_P2PDMA_MAP_NOT_SUPPORTED if
a TLP has to go through a host bridge. Do we mistakenly assume that
if a bridge has PCI_ACS_RR set, a Request must go all the way to the
host bridge, even if a bridge closer to the root does not have
PCI_ACS_RR set?
> Evaluate Request controls on the client-side divergence port and Completion
> Redirect on the provider-side port and reject an unreadable ACS Control
> register.
>
> Fixes: 52916982af48 ("PCI/P2PDMA: Support peer-to-peer memory")
> Reviewed-by: Logan Gunthorpe <logang@deltatee.com>
> Tested-by: Tushar Dave <tdave@nvidia.com>
> Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
> ---
> drivers/pci/p2pdma.c | 101 +++++++++++++++++++++++++++++----------------
> include/linux/pci-p2pdma.h | 8 ++--
> 2 files changed, 70 insertions(+), 39 deletions(-)
>
> diff --git a/drivers/pci/p2pdma.c b/drivers/pci/p2pdma.c
> index 12612b82d80d..550e6c7346ef 100644
> --- a/drivers/pci/p2pdma.c
> +++ b/drivers/pci/p2pdma.c
> @@ -493,6 +493,7 @@ static struct pci_dev *find_parent_pci_dev(struct device *dev)
> }
>
> enum pci_acs_p2pdma_state {
> + PCI_ACS_P2PDMA_NOT_SUPPORTED,
> PCI_ACS_P2PDMA_DIRECT,
> PCI_ACS_P2PDMA_REDIRECT,
> };
> @@ -730,13 +731,13 @@ static unsigned long map_types_idx(struct pci_dev *client)
> * then to Device B. The mapping type returned depends on the ACS
> * redirection setting of the ports along the path.
> *
> - * The client initiates Requests to provider memory. Check Request Redirect
> - * on the client path and Completion Redirect for read Completions on the
> - * provider path.
> + * The client initiates Requests to provider memory. At the path divergence,
> + * check Request Redirect and Egress Control on the client-side port, and
> + * Completion Redirect for read Completions on the provider-side port.
> *
> - * If ACS redirect is set on any port in the path, traffic between the
> - * devices will go through the host bridge, so return
> - * PCI_P2PDMA_MAP_THRU_HOST_BRIDGE; otherwise return
> + * If ACS redirects traffic at either divergence port, return
> + * PCI_P2PDMA_MAP_THRU_HOST_BRIDGE. If the ACS Control register cannot be
> + * read, return PCI_P2PDMA_MAP_NOT_SUPPORTED. Otherwise, return
> * PCI_P2PDMA_MAP_BUS_ADDR.
> *
> * Any two devices that have a data path that goes through the host bridge
> @@ -750,10 +751,13 @@ calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client,
> int *dist, bool verbose)
> {
> enum pci_p2pdma_map_type map_type = PCI_P2PDMA_MAP_THRU_HOST_BRIDGE;
> + enum pci_acs_p2pdma_state state = PCI_ACS_P2PDMA_NOT_SUPPORTED;
> struct pci_dev *a = provider, *b = client, *bb;
> + struct pci_dev *a_child = NULL, *b_child = NULL;
> + struct pci_dev *acs_unreadable = NULL;
> struct pci_p2pdma *p2pdma;
> struct seq_buf acs_list;
> - int acs_cnt = 0;
> + int acs_redirect_cnt = 0;
> int dist_a = 0;
> int dist_b = 0;
> char buf[128];
> @@ -768,51 +772,67 @@ calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client,
> */
> while (a) {
> dist_b = 0;
> -
> - if (!pci_acs_p2pdma_ctrl(a, &ctrl) ||
> - pci_acs_p2pdma_completion(ctrl) ==
> - PCI_ACS_P2PDMA_REDIRECT) {
> - seq_buf_print_bus_devfn(&acs_list, a);
> - acs_cnt++;
> - }
> -
> + b_child = NULL;
> bb = b;
>
> while (bb) {
> if (a == bb)
> - goto check_b_path_acs;
> + goto check_paths_acs;
>
> + b_child = bb;
> bb = pci_upstream_bridge(bb);
> dist_b++;
> }
>
> + a_child = a;
> a = pci_upstream_bridge(a);
> dist_a++;
> }
>
> + /*
> + * The paths share no upstream bridge, so there is no direct path for
> + * ACS to gate: PCI_P2PDMA_MAP_BUS_ADDR is not reachable here and the
> + * request can only get to the peer through the host bridge.
> + */
> *dist = dist_a + dist_b;
> goto map_through_host_bridge;
>
> -check_b_path_acs:
> - bb = b;
> -
> - while (bb) {
> - if (a == bb)
> - break;
> +check_paths_acs:
> + *dist = dist_a + dist_b;
>
> - if (!pci_acs_p2pdma_ctrl(bb, &ctrl) ||
> - pci_acs_p2pdma_request(ctrl) ==
> - PCI_ACS_P2PDMA_REDIRECT) {
> - seq_buf_print_bus_devfn(&acs_list, bb);
> - acs_cnt++;
> + /*
> + * ACS P2P routing controls apply where a TLP can route toward the peer
> + * or upstream. Below that divergence, its only route toward the other
> + * branch is upstream, so redirect controls do not affect the path.
> + */
> + if (a_child && b_child) {
> + if (pci_acs_p2pdma_ctrl(a_child, &ctrl))
> + state = pci_acs_p2pdma_completion(ctrl);
> + if (state != PCI_ACS_P2PDMA_DIRECT) {
> + seq_buf_print_bus_devfn(&acs_list, a_child);
> + if (state == PCI_ACS_P2PDMA_REDIRECT)
> + acs_redirect_cnt++;
> + else if (!acs_unreadable)
> + acs_unreadable = a_child;
> }
>
> - bb = pci_upstream_bridge(bb);
> + state = PCI_ACS_P2PDMA_NOT_SUPPORTED;
> + if (pci_acs_p2pdma_ctrl(b_child, &ctrl))
> + state = pci_acs_p2pdma_request(ctrl);
> + if (state != PCI_ACS_P2PDMA_DIRECT) {
> + seq_buf_print_bus_devfn(&acs_list, b_child);
> + if (state == PCI_ACS_P2PDMA_REDIRECT)
> + acs_redirect_cnt++;
> + else if (!acs_unreadable)
> + acs_unreadable = b_child;
> + }
> }
>
> - *dist = dist_a + dist_b;
> -
> - if (!acs_cnt) {
> + /*
> + * Below a shared upstream bridge, a path whose divergence ports do not
> + * redirect routes the request directly.
> + */
> + if (!acs_unreadable && !acs_redirect_cnt) {
> map_type = PCI_P2PDMA_MAP_BUS_ADDR;
> goto done;
> }
> @@ -821,10 +841,21 @@ calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client,
> /* Drop the final semicolon; the list is not empty here. */
> if (!seq_buf_has_overflowed(&acs_list))
> acs_list.buffer[acs_list.len - 1] = '\0';
> - pci_warn(client, "ACS redirect is set between the client and provider (%s)\n",
> - pci_name(provider));
> - pci_warn(client, "to disable ACS redirect for this path, add the kernel parameter: pci=disable_acs_redir=%s\n",
> - seq_buf_str(&acs_list));
> + if (acs_unreadable)
> + pci_warn(client, "ACS Control is unreadable for provider %s at %s\n",
> + pci_name(provider), pci_name(acs_unreadable));
> + else {
> + pci_warn(client, "ACS redirect is set between the client and provider (%s)\n",
> + pci_name(provider));
> + pci_warn(client, "to disable ACS controls for this path, add the kernel parameter: pci=disable_acs_redir=%s\n",
> + seq_buf_str(&acs_list));
> + }
> + }
> +
> + /* An unreadable control does not establish an upstream redirect. */
> + if (acs_unreadable) {
> + map_type = PCI_P2PDMA_MAP_NOT_SUPPORTED;
> + goto done;
> }
>
> map_through_host_bridge:
> diff --git a/include/linux/pci-p2pdma.h b/include/linux/pci-p2pdma.h
> index 873de20a2247..dd17501ba1b6 100644
> --- a/include/linux/pci-p2pdma.h
> +++ b/include/linux/pci-p2pdma.h
> @@ -42,10 +42,10 @@ enum pci_p2pdma_map_type {
> PCI_P2PDMA_MAP_NONE,
>
> /*
> - * PCI_P2PDMA_MAP_NOT_SUPPORTED: Indicates the transaction will
> - * traverse the host bridge and the host bridge is not in the
> - * allowlist. DMA Mapping routines should return an error when
> - * this is returned.
> + * PCI_P2PDMA_MAP_NOT_SUPPORTED: Indicates no safe mapping is available,
> + * for example because ACS blocks the direct path or the required host
> + * bridge is not in the allowlist. DMA Mapping routines should return an
> + * error when this is returned.
> */
> PCI_P2PDMA_MAP_NOT_SUPPORTED,
>
>
> --
> 2.55.0
>
^ permalink raw reply [flat|nested] 38+ messages in thread
* Re: [PATCH v9 00/18] PCI/P2PDMA: Route peer-to-peer DMA by TLP class
2026-10-06 19:29 ` [PATCH v9 00/18] PCI/P2PDMA: Route peer-to-peer DMA by TLP class Bjorn Helgaas
@ 2026-10-06 21:22 ` Leon Romanovsky
2026-10-06 21:36 ` Bjorn Helgaas
0 siblings, 1 reply; 38+ messages in thread
From: Leon Romanovsky @ 2026-10-06 21:22 UTC (permalink / raw)
To: Bjorn Helgaas
Cc: Bjorn Helgaas, Logan Gunthorpe, Jason Gunthorpe,
Joerg Roedel (AMD),
Will Deacon, Robin Murphy, Christian König,
Thomas Hellström, linux-pci, linux-kernel, linux-doc, iommu,
Tushar Dave, linux-media, dri-devel, linaro-mm-sig, linux-rdma,
kvm, Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Sumit Semwal
On Tue, Oct 06, 2026 at 02:29:55PM -0500, Bjorn Helgaas wrote:
> On Thu, Oct 01, 2026 at 02:55:08PM +0300, Leon Romanovsky wrote:
> > PCI P2PDMA applies Request and Completion Redirect throughout both paths.
> > This misclassifies asymmetric and nested switches, and reports one answer
> > for every kind of TLP.
> >
> > Three ACS controls act on TLP attributes the client chooses rather than
> > on the topology: Translation Blocking and Direct Translated P2P act on
> > a Request's Address Type, and Completion Redirect skips Completions carrying
> > Relaxed Ordering.
> >
> > Evaluate each direction at the path divergence, decide every class from the
> > one walk, and treat a client with ATS enabled as translating unless its
> > driver declares per-mapping ATS.
> >
> > This completes the P2PDMA side; dma-buf and mlx5 follow separately.
> > ...
>
> > Leon Romanovsky (18):
> > PCI/P2PDMA: Document the TLP attribute assumptions
> > PCI/P2PDMA: Derive routing from directional ACS controls
> > PCI: Reject unreadable ACS controls in isolation checks
> > PCI/P2PDMA: Evaluate ACS controls at the path divergence
> > PCI/P2PDMA: Document directional ACS routing
> > PCI/P2PDMA: Collect the path's ACS controls before deciding
> > PCI/P2PDMA: Answer routing per TLP class
> > PCI/P2PDMA: Route Relaxed Ordering Completions directly
> > PCI/P2PDMA: Reject Translated Requests blocked by Translation Blocking
> > PCI/P2PDMA: Route Translated Requests under Direct Translated P2P
> > PCI/P2PDMA: Log detailed ACS routing diagnostics
> > PCI/P2PDMA: Add KUnit tests for the ACS routing decisions
> > PCI/P2PDMA: Test the ACS P2P routing walk
> > PCI: Add KUnit coverage for ACS isolation checks
> > PCI/P2PDMA: Document TLP-class routing
> > PCI/P2PDMA: Let a client declare that it selects ATS per mapping
> > PCI/P2PDMA: Evaluate the ATS path for clients with ATS enabled
> > PCI/P2PDMA: Test the routing of clients with ATS enabled
> >
> > Documentation/admin-guide/kernel-parameters.txt | 15 +-
> > Documentation/driver-api/pci/p2pdma.rst | 80 +++
> > drivers/pci/Kconfig | 15 +
> > drivers/pci/Makefile | 1 +
> > drivers/pci/p2pdma.c | 719 ++++++++++++++++++--
> > drivers/pci/pci.c | 7 +-
> > drivers/pci/pci.h | 58 ++
> > drivers/pci/pci_acs_test.c | 860 ++++++++++++++++++++++++
> > drivers/pci/quirks.c | 6 +-
> > include/linux/pci-p2pdma.h | 12 +-
> > 10 files changed, 1687 insertions(+), 86 deletions(-)
> > ---
> > base-commit: 08dbfad3f5040f5bdb6c529da20d6d4e81fefd72
> > change-id: 20260821-fix-p2p-acs-v4-0-e72455e3a261
> > prerequisite-message-id: <20260830-batch-p2p-fixes-v1-0-5044e8dfbe2e@nvidia.com>
> > prerequisite-patch-id: 6b25c7fcf164cdfc14e9fac5b908d97fcf6509d7
> > prerequisite-patch-id: 0d083c281001365aae4b35544cf28891a6ab9a96
> > prerequisite-patch-id: bfd9dabf271f3cc9a3a61f46387d20c20311363d
> > prerequisite-patch-id: fad0275efc722830fc591509506c0a5e4f581073
> > prerequisite-patch-id: 0c83bee688fec1f6d1564654df7c630fa6a4a978
>
> This seems like material for the PCI tree, but I'm not sure how to
> apply it. It doesn't apply cleanly on the current pci/p2pdma
> (https://git.kernel.org/pub/scm/linux/kernel/git/pci/pci.git/log/?h=p2pdma),
> which does contain your series from
> 20260830-batch-p2p-fixes-v1-0-5044e8dfbe2e@nvidia.com.
>
> I could probably fix the conflicts but I don't know why there should
> be conflicts, since the only commits on pci/p2pdma other than yours
> are a few trivial allow-list updates.
So maybe this is related to these trivial updates, my series was based
on the clean p2pdma branch:
https://git.kernel.org/pub/scm/linux/kernel/git/leon/linux-rdma.git/log/?h=fix-p2p-acs-v9
Thanks
>
> Bjorn
^ permalink raw reply [flat|nested] 38+ messages in thread
* Re: [PATCH v9 05/18] PCI/P2PDMA: Document directional ACS routing
2026-10-01 11:55 ` [PATCH v9 05/18] PCI/P2PDMA: Document directional ACS routing Leon Romanovsky
@ 2026-10-06 21:32 ` Bjorn Helgaas
2026-10-07 13:46 ` Leon Romanovsky
0 siblings, 1 reply; 38+ messages in thread
From: Bjorn Helgaas @ 2026-10-06 21:32 UTC (permalink / raw)
To: Leon Romanovsky
Cc: Bjorn Helgaas, Logan Gunthorpe, Jason Gunthorpe,
Joerg Roedel (AMD),
Will Deacon, Robin Murphy, Christian König,
Thomas Hellström, linux-pci, linux-kernel, linux-doc, iommu,
Tushar Dave, linux-media, dri-devel, linaro-mm-sig, linux-rdma,
kvm, Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Sumit Semwal
On Thu, Oct 01, 2026 at 02:55:13PM +0300, Leon Romanovsky wrote:
> From: Leon Romanovsky <leonro@nvidia.com>
>
> P2PDMA documentation describes ACS controls as path-wide, although Request
> and Completion controls apply to different transaction directions and only
> affect peer-versus-upstream decisions at the path divergence.
>
> Document the fixed client and provider roles, the divergence port checked
> for each TLP direction, and the conservative handling of unreadable ACS
> state. Clarify which controls disable_acs_redir changes.
>
> Reviewed-by: Logan Gunthorpe <logang@deltatee.com>
> Tested-by: Tushar Dave <tdave@nvidia.com>
> Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
> ---
> Documentation/admin-guide/kernel-parameters.txt | 15 +++++++++------
> Documentation/driver-api/pci/p2pdma.rst | 13 +++++++++++++
> 2 files changed, 22 insertions(+), 6 deletions(-)
>
> diff --git a/Documentation/admin-guide/kernel-parameters.txt b/Documentation/admin-guide/kernel-parameters.txt
> index 68647ff4bdd2..bc83e07dd5fc 100644
> --- a/Documentation/admin-guide/kernel-parameters.txt
> +++ b/Documentation/admin-guide/kernel-parameters.txt
> @@ -5291,12 +5291,15 @@ Kernel parameters
> disable_acs_redir=<pci_dev>[; ...]
> Specify one or more PCI devices (in the format
> specified above) separated by semicolons.
> - Each device specified will have the PCI ACS
> - redirect capabilities forced off which will
> - allow P2P traffic between devices through
> - bridges without forcing it upstream. Note:
> - this removes isolation between devices and
> - may put more devices in an IOMMU group.
> + Each device specified will have the PCI ACS P2P
> + Request Redirect, Completion Redirect, and Egress
> + Control features forced off. This may allow P2P
> + traffic through bridges that would otherwise be
> + redirected upstream. This may allow P2P traffic
> + through bridges that would otherwise be redirected
> + upstream and thus this removes isolation between
> + devices and may cause affected devices to share
> + an IOMMU group.
> config_acs=
> Format:
> <ACS flags>@<pci_dev>[; ...]
> diff --git a/Documentation/driver-api/pci/p2pdma.rst b/Documentation/driver-api/pci/p2pdma.rst
> index 80f8fec9b0e9..42b18610bf7d 100644
> --- a/Documentation/driver-api/pci/p2pdma.rst
> +++ b/Documentation/driver-api/pci/p2pdma.rst
> @@ -15,6 +15,19 @@ then based on the ACS settings the transaction can route entirely within
> the PCIe hierarchy and never reach the root port. The kernel will evaluate
> the PCIe topology and always permit P2P in these well-defined cases.
>
> +The client remains the PCIe requester when it reads or writes provider memory.
> +Where the paths diverge, the kernel therefore evaluates P2P Request Redirect
> +and Egress Control on the client-side port, and P2P Completion Redirect on the
> +provider-side port for completions from a read. An enabled Egress Control is
> +conservatively treated as a Request redirect.
I think "where the paths diverge" means the point where a port decides
whether to route a TLP up through its Upstream Port (or to the RC) or
back down via a sibling Downstream Port?
I'm not sure p2pdma.rst includes the context to interpret "divergence"
yet. I think the code comments in [4/18] might also need a little
more context about divergence.
IIUC, this divergence point is basically a static point where the ACS
controls being applied makes a routing difference.
> +Below the divergence, the route toward the other branch is already upstream,
> +so those P2P redirect controls do not affect it. Redirect controls for the
> +reverse transaction directions do not affect the mapping. P2P DMA is routed
> +through the host bridge when either applicable port redirects. If an ACS
> +Control register cannot be read, P2P DMA is rejected because the kernel cannot
> +establish a usable route.
> +
> This evaluation assumes clients issue strictly ordered Requests carrying an
> Untranslated address. Its result is not defined when clients use Relaxed
> Ordering or issue ATS-translated Requests because those TLP attributes can
>
> --
> 2.55.0
>
^ permalink raw reply [flat|nested] 38+ messages in thread
* Re: [PATCH v9 00/18] PCI/P2PDMA: Route peer-to-peer DMA by TLP class
2026-10-06 21:22 ` Leon Romanovsky
@ 2026-10-06 21:36 ` Bjorn Helgaas
2026-10-07 6:56 ` Bjorn Helgaas
0 siblings, 1 reply; 38+ messages in thread
From: Bjorn Helgaas @ 2026-10-06 21:36 UTC (permalink / raw)
To: Leon Romanovsky
Cc: Bjorn Helgaas, Logan Gunthorpe, Jason Gunthorpe,
Joerg Roedel (AMD),
Will Deacon, Robin Murphy, Christian König,
Thomas Hellström, linux-pci, linux-kernel, linux-doc, iommu,
Tushar Dave, linux-media, dri-devel, linaro-mm-sig, linux-rdma,
kvm, Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Sumit Semwal
On Wed, Oct 07, 2026 at 12:22:27AM +0300, Leon Romanovsky wrote:
> On Tue, Oct 06, 2026 at 02:29:55PM -0500, Bjorn Helgaas wrote:
> > On Thu, Oct 01, 2026 at 02:55:08PM +0300, Leon Romanovsky wrote:
> > > PCI P2PDMA applies Request and Completion Redirect throughout both paths.
> > > This misclassifies asymmetric and nested switches, and reports one answer
> > > for every kind of TLP.
> > >
> > > Three ACS controls act on TLP attributes the client chooses rather than
> > > on the topology: Translation Blocking and Direct Translated P2P act on
> > > a Request's Address Type, and Completion Redirect skips Completions carrying
> > > Relaxed Ordering.
> > >
> > > Evaluate each direction at the path divergence, decide every class from the
> > > one walk, and treat a client with ATS enabled as translating unless its
> > > driver declares per-mapping ATS.
> > >
> > > This completes the P2PDMA side; dma-buf and mlx5 follow separately.
> > > ...
> >
> > > Leon Romanovsky (18):
> > > PCI/P2PDMA: Document the TLP attribute assumptions
> > > PCI/P2PDMA: Derive routing from directional ACS controls
> > > PCI: Reject unreadable ACS controls in isolation checks
> > > PCI/P2PDMA: Evaluate ACS controls at the path divergence
> > > PCI/P2PDMA: Document directional ACS routing
> > > PCI/P2PDMA: Collect the path's ACS controls before deciding
> > > PCI/P2PDMA: Answer routing per TLP class
> > > PCI/P2PDMA: Route Relaxed Ordering Completions directly
> > > PCI/P2PDMA: Reject Translated Requests blocked by Translation Blocking
> > > PCI/P2PDMA: Route Translated Requests under Direct Translated P2P
> > > PCI/P2PDMA: Log detailed ACS routing diagnostics
> > > PCI/P2PDMA: Add KUnit tests for the ACS routing decisions
> > > PCI/P2PDMA: Test the ACS P2P routing walk
> > > PCI: Add KUnit coverage for ACS isolation checks
> > > PCI/P2PDMA: Document TLP-class routing
> > > PCI/P2PDMA: Let a client declare that it selects ATS per mapping
> > > PCI/P2PDMA: Evaluate the ATS path for clients with ATS enabled
> > > PCI/P2PDMA: Test the routing of clients with ATS enabled
> > >
> > > Documentation/admin-guide/kernel-parameters.txt | 15 +-
> > > Documentation/driver-api/pci/p2pdma.rst | 80 +++
> > > drivers/pci/Kconfig | 15 +
> > > drivers/pci/Makefile | 1 +
> > > drivers/pci/p2pdma.c | 719 ++++++++++++++++++--
> > > drivers/pci/pci.c | 7 +-
> > > drivers/pci/pci.h | 58 ++
> > > drivers/pci/pci_acs_test.c | 860 ++++++++++++++++++++++++
> > > drivers/pci/quirks.c | 6 +-
> > > include/linux/pci-p2pdma.h | 12 +-
> > > 10 files changed, 1687 insertions(+), 86 deletions(-)
> > > ---
> > > base-commit: 08dbfad3f5040f5bdb6c529da20d6d4e81fefd72
> > > change-id: 20260821-fix-p2p-acs-v4-0-e72455e3a261
> > > prerequisite-message-id: <20260830-batch-p2p-fixes-v1-0-5044e8dfbe2e@nvidia.com>
> > > prerequisite-patch-id: 6b25c7fcf164cdfc14e9fac5b908d97fcf6509d7
> > > prerequisite-patch-id: 0d083c281001365aae4b35544cf28891a6ab9a96
> > > prerequisite-patch-id: bfd9dabf271f3cc9a3a61f46387d20c20311363d
> > > prerequisite-patch-id: fad0275efc722830fc591509506c0a5e4f581073
> > > prerequisite-patch-id: 0c83bee688fec1f6d1564654df7c630fa6a4a978
> >
> > This seems like material for the PCI tree, but I'm not sure how to
> > apply it. It doesn't apply cleanly on the current pci/p2pdma
> > (https://git.kernel.org/pub/scm/linux/kernel/git/pci/pci.git/log/?h=p2pdma),
> > which does contain your series from
> > 20260830-batch-p2p-fixes-v1-0-5044e8dfbe2e@nvidia.com.
> >
> > I could probably fix the conflicts but I don't know why there should
> > be conflicts, since the only commits on pci/p2pdma other than yours
> > are a few trivial allow-list updates.
>
> So maybe this is related to these trivial updates, my series was based
> on the clean p2pdma branch:
> https://git.kernel.org/pub/scm/linux/kernel/git/leon/linux-rdma.git/log/?h=fix-p2p-acs-v9
Ah, yeah, it must be the allow-list additions that I have but you
don't. I wouldn't have thought the offsets would have been a problem.
^ permalink raw reply [flat|nested] 38+ messages in thread
* Re: [PATCH v9 06/18] PCI/P2PDMA: Collect the path's ACS controls before deciding
2026-10-01 11:55 ` [PATCH v9 06/18] PCI/P2PDMA: Collect the path's ACS controls before deciding Leon Romanovsky
@ 2026-10-06 21:48 ` Bjorn Helgaas
2026-10-07 13:50 ` Leon Romanovsky
0 siblings, 1 reply; 38+ messages in thread
From: Bjorn Helgaas @ 2026-10-06 21:48 UTC (permalink / raw)
To: Leon Romanovsky
Cc: Bjorn Helgaas, Logan Gunthorpe, Jason Gunthorpe,
Joerg Roedel (AMD),
Will Deacon, Robin Murphy, Christian König,
Thomas Hellström, linux-pci, linux-kernel, linux-doc, iommu,
Tushar Dave, linux-media, dri-devel, linaro-mm-sig, linux-rdma,
kvm, Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Sumit Semwal
On Thu, Oct 01, 2026 at 02:55:14PM +0300, Leon Romanovsky wrote:
> From: Leon Romanovsky <leonro@nvidia.com>
>
> calc_map_type_and_dist() reads each divergence port's ACS Control register
> and folds the result into running counters as it goes. Any routing property
> that depends on the kind of TLP being routed would have to be threaded
> through that code, so there is nowhere to put one without reading the
> registers again for each kind.
What is the "one" that there's nowhere to put? I guess the routing
property? So this is an optimization to avoid some config reads?
> Collect the two ports' ACS Control values into struct pci_p2pdma_acs_path
> first, then decide from it. pci_p2pdma_route() applies the same rule as
> before: a path routes directly only when both directions do.
>
> Reviewed-by: Logan Gunthorpe <logang@deltatee.com>
> Tested-by: Tushar Dave <tdave@nvidia.com>
> Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
> ---
> drivers/pci/p2pdma.c | 148 +++++++++++++++++++++++++++++++++------------------
> 1 file changed, 96 insertions(+), 52 deletions(-)
>
> diff --git a/drivers/pci/p2pdma.c b/drivers/pci/p2pdma.c
> index 550e6c7346ef..841c86be31bb 100644
> --- a/drivers/pci/p2pdma.c
> +++ b/drivers/pci/p2pdma.c
> @@ -553,6 +553,80 @@ static void seq_buf_print_bus_devfn(struct seq_buf *buf, struct pci_dev *pdev)
> seq_buf_printf(buf, "%s;", pci_name(pdev));
> }
>
> +/*
> + * What the topology walk found out about one provider/client path. Producing
> + * this costs a walk and one config read per divergence port, none of which
> + * depends on the TLP being routed.
> + *
> + * @req_ctrl: ACS Control of the client-side divergence port. That is the
> + * first port at which a Request can route toward the peer rather
> + * than upstream, so it is where the Request controls apply.
> + * @cpl_ctrl: ACS Control of the provider-side divergence port, likewise for
> + * the Completions travelling back.
> + * @unreadable: First port whose ACS Control could not be read, if any.
> + */
> +struct pci_p2pdma_acs_path {
> + u16 req_ctrl;
> + u16 cpl_ctrl;
> + struct pci_dev *unreadable;
> +};
> +
> +/*
> + * Combine both directions into a mapping type. Only a path that routes the
> + * Request and the Completions it generates directly can be programmed with
> + * the peer's bus addresses.
> + */
> +static enum pci_p2pdma_map_type
> +pci_p2pdma_route(const struct pci_p2pdma_acs_path *path)
> +{
> + if (path->unreadable)
> + return PCI_P2PDMA_MAP_NOT_SUPPORTED;
> +
> + if (pci_acs_p2pdma_request(path->req_ctrl) == PCI_ACS_P2PDMA_DIRECT &&
> + pci_acs_p2pdma_completion(path->cpl_ctrl) == PCI_ACS_P2PDMA_DIRECT)
> + return PCI_P2PDMA_MAP_BUS_ADDR;
> +
> + return PCI_P2PDMA_MAP_THRU_HOST_BRIDGE;
> +}
> +
> +/*
> + * Name the ports that keep this path off a direct route, so that the admin
> + * can hand them to pci=disable_acs_redir=.
> + */
> +static void pci_p2pdma_warn_path(struct pci_dev *client,
> + struct pci_dev *provider,
> + const struct pci_p2pdma_acs_path *path,
> + struct pci_dev *a_child,
> + struct pci_dev *b_child)
> +{
> + struct seq_buf acs_list;
> + char buf[128];
> +
> + if (path->unreadable) {
> + pci_warn(client,
> + "ACS Control is unreadable for provider %s at %s\n",
> + pci_name(provider), pci_name(path->unreadable));
> + return;
> + }
> +
> + seq_buf_init(&acs_list, buf, sizeof(buf));
> + if (pci_acs_p2pdma_completion(path->cpl_ctrl) != PCI_ACS_P2PDMA_DIRECT)
> + seq_buf_print_bus_devfn(&acs_list, a_child);
> + if (pci_acs_p2pdma_request(path->req_ctrl) != PCI_ACS_P2PDMA_DIRECT)
> + seq_buf_print_bus_devfn(&acs_list, b_child);
> +
> + /* Drop the final semicolon; the list is not empty here. */
> + if (!seq_buf_has_overflowed(&acs_list))
> + acs_list.buffer[acs_list.len - 1] = '\0';
> +
> + pci_warn(client,
> + "ACS redirect is set between the client and provider (%s)\n",
> + pci_name(provider));
> + pci_warn(client,
> + "to disable ACS controls for this path, add the kernel parameter: pci=disable_acs_redir=%s\n",
> + seq_buf_str(&acs_list));
> +}
> +
> static bool cpu_supports_p2pdma(void)
> {
> #ifdef CONFIG_X86
> @@ -751,19 +825,13 @@ calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client,
> int *dist, bool verbose)
> {
> enum pci_p2pdma_map_type map_type = PCI_P2PDMA_MAP_THRU_HOST_BRIDGE;
> - enum pci_acs_p2pdma_state state = PCI_ACS_P2PDMA_NOT_SUPPORTED;
> struct pci_dev *a = provider, *b = client, *bb;
> struct pci_dev *a_child = NULL, *b_child = NULL;
> - struct pci_dev *acs_unreadable = NULL;
> + struct pci_p2pdma_acs_path path = {};
> struct pci_p2pdma *p2pdma;
> - struct seq_buf acs_list;
> - int acs_redirect_cnt = 0;
> + bool cpu_p2pdma, host_whitelisted = false;
> int dist_a = 0;
> int dist_b = 0;
> - char buf[128];
> - u16 ctrl;
> -
> - seq_buf_init(&acs_list, buf, sizeof(buf));
>
> /*
> * Note, we don't need to take references to devices returned by
> @@ -806,61 +874,35 @@ calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client,
> * branch is upstream, so redirect controls do not affect the path.
> */
> if (a_child && b_child) {
> - if (pci_acs_p2pdma_ctrl(a_child, &ctrl))
> - state = pci_acs_p2pdma_completion(ctrl);
> - if (state != PCI_ACS_P2PDMA_DIRECT) {
> - seq_buf_print_bus_devfn(&acs_list, a_child);
> - if (state == PCI_ACS_P2PDMA_REDIRECT)
> - acs_redirect_cnt++;
> - else if (!acs_unreadable)
> - acs_unreadable = a_child;
> - }
> -
> - state = PCI_ACS_P2PDMA_NOT_SUPPORTED;
> - if (pci_acs_p2pdma_ctrl(b_child, &ctrl))
> - state = pci_acs_p2pdma_request(ctrl);
> - if (state != PCI_ACS_P2PDMA_DIRECT) {
> - seq_buf_print_bus_devfn(&acs_list, b_child);
> - if (state == PCI_ACS_P2PDMA_REDIRECT)
> - acs_redirect_cnt++;
> - else if (!acs_unreadable)
> - acs_unreadable = b_child;
> - }
> + if (!pci_acs_p2pdma_ctrl(a_child, &path.cpl_ctrl))
> + path.unreadable = a_child;
> + if (!pci_acs_p2pdma_ctrl(b_child, &path.req_ctrl) &&
> + !path.unreadable)
> + path.unreadable = b_child;
> }
>
> /*
> * Below a shared upstream bridge, a path whose divergence ports do not
> * redirect routes the request directly.
> */
> - if (!acs_unreadable && !acs_redirect_cnt) {
> - map_type = PCI_P2PDMA_MAP_BUS_ADDR;
> + map_type = pci_p2pdma_route(&path);
> + if (map_type == PCI_P2PDMA_MAP_BUS_ADDR)
> goto done;
> - }
>
> - if (verbose) {
> - /* Drop the final semicolon; the list is not empty here. */
> - if (!seq_buf_has_overflowed(&acs_list))
> - acs_list.buffer[acs_list.len - 1] = '\0';
> - if (acs_unreadable)
> - pci_warn(client, "ACS Control is unreadable for provider %s at %s\n",
> - pci_name(provider), pci_name(acs_unreadable));
> - else {
> - pci_warn(client, "ACS redirect is set between the client and provider (%s)\n",
> - pci_name(provider));
> - pci_warn(client, "to disable ACS controls for this path, add the kernel parameter: pci=disable_acs_redir=%s\n",
> - seq_buf_str(&acs_list));
> - }
> - }
> + if (verbose)
> + pci_p2pdma_warn_path(client, provider, &path, a_child, b_child);
>
> /* An unreadable control does not establish an upstream redirect. */
> - if (acs_unreadable) {
> - map_type = PCI_P2PDMA_MAP_NOT_SUPPORTED;
> + if (path.unreadable)
> goto done;
> - }
>
> map_through_host_bridge:
> - if (!cpu_supports_p2pdma() &&
> - !host_bridge_whitelist(provider, client, verbose)) {
> + cpu_p2pdma = cpu_supports_p2pdma();
> + if (!cpu_p2pdma)
> + host_whitelisted = host_bridge_whitelist(provider, client,
> + verbose);
> +
> + if (!cpu_p2pdma && !host_whitelisted) {
> if (verbose)
> pci_warn(client, "cannot be used for peer-to-peer DMA as the client and provider (%s) do not share an upstream bridge or whitelisted host bridge\n",
> pci_name(provider));
> @@ -1193,8 +1235,9 @@ enum pci_p2pdma_map_type pci_p2pdma_map_type(struct p2pdma_provider *provider,
> {
> enum pci_p2pdma_map_type type = PCI_P2PDMA_MAP_NOT_SUPPORTED;
> struct pci_dev *pdev = to_pci_dev(provider->owner);
> - struct pci_dev *client;
> struct pci_p2pdma *p2pdma;
> + unsigned long cache_index;
> + struct pci_dev *client;
> int dist;
>
> if (!pdev->p2pdma)
> @@ -1204,13 +1247,14 @@ enum pci_p2pdma_map_type pci_p2pdma_map_type(struct p2pdma_provider *provider,
> return PCI_P2PDMA_MAP_NOT_SUPPORTED;
>
> client = to_pci_dev(dev);
> + cache_index = map_types_idx(client);
>
> rcu_read_lock();
> p2pdma = rcu_dereference(pdev->p2pdma);
>
> if (p2pdma)
> type = xa_to_value(xa_load(&p2pdma->map_types,
> - map_types_idx(client)));
> + cache_index));
> rcu_read_unlock();
>
> if (type == PCI_P2PDMA_MAP_UNKNOWN)
>
> --
> 2.55.0
>
^ permalink raw reply [flat|nested] 38+ messages in thread
* Re: [PATCH v9 08/18] PCI/P2PDMA: Route Relaxed Ordering Completions directly
2026-10-01 11:55 ` [PATCH v9 08/18] PCI/P2PDMA: Route Relaxed Ordering Completions directly Leon Romanovsky
@ 2026-10-06 22:21 ` Bjorn Helgaas
0 siblings, 0 replies; 38+ messages in thread
From: Bjorn Helgaas @ 2026-10-06 22:21 UTC (permalink / raw)
To: Leon Romanovsky
Cc: Bjorn Helgaas, Logan Gunthorpe, Jason Gunthorpe,
Joerg Roedel (AMD),
Will Deacon, Robin Murphy, Christian König,
Thomas Hellström, linux-pci, linux-kernel, linux-doc, iommu,
Tushar Dave, linux-media, dri-devel, linaro-mm-sig, linux-rdma,
kvm, Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Sumit Semwal
On Thu, Oct 01, 2026 at 02:55:16PM +0300, Leon Romanovsky wrote:
> From: Leon Romanovsky <leonro@nvidia.com>
>
> ACS P2P Completion Redirect leaves Completions carrying the Relaxed
> Ordering attribute alone. PCIe r7.0 sec 6.12.1.1 redirects only those "that
> do not have the Relaxed Ordering Attribute bit set", and sec 7.7.12.5
> describes the enable bit as "applicable only to Completions whose Relaxed
> Ordering Attribute is clear". P2PDMA reports one answer for every kind of
> TLP, so a client whose provider returns such Completions is sent through
> the host bridge for a redirect that never happens to it.
I guess pci_acs_p2pdma_completion() returned PCI_ACS_P2PDMA_REDIRECT
even for RO Completions? But that didn't change the hardware
behavior -- maybe the caller *thought* Completions were routed through
the host bridge, but they actually weren't.
I kind of lost the plot here. What does the caller do with this
information? I don't think a device is *required* to set RO even when
it is enabled, and it may set RO on some transactions but not others.
> Add enum pci_p2pdma_tlp_flags and let a caller state that property.
>
> Reviewed-by: Logan Gunthorpe <logang@deltatee.com>
> Tested-by: Tushar Dave <tdave@nvidia.com>
> Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
> ---
> drivers/pci/p2pdma.c | 6 +++++-
> 1 file changed, 5 insertions(+), 1 deletion(-)
>
> diff --git a/drivers/pci/p2pdma.c b/drivers/pci/p2pdma.c
> index 43e225cc5735..1fadef6d0609 100644
> --- a/drivers/pci/p2pdma.c
> +++ b/drivers/pci/p2pdma.c
> @@ -544,11 +544,15 @@ pci_acs_p2pdma_request(u16 ctrl, unsigned int tlp_flags)
> /*
> * Decide how a peer-to-peer Completion at an ACS-capable ingress port routes.
> * PCIe r7.0 sec 6.12.1.1: no ACS control other than P2P Completion Redirect
> - * affects a Completion.
> + * affects a Completion, and that one leaves Completions carrying the Relaxed
> + * Ordering attribute alone.
> */
> static enum pci_acs_p2pdma_state
> pci_acs_p2pdma_completion(u16 ctrl, unsigned int tlp_flags)
> {
> + if (tlp_flags & PCI_P2PDMA_TLP_RELAXED_CPL)
> + return PCI_ACS_P2PDMA_DIRECT;
> +
> return ctrl & PCI_ACS_CR ? PCI_ACS_P2PDMA_REDIRECT :
> PCI_ACS_P2PDMA_DIRECT;
> }
>
> --
> 2.55.0
>
^ permalink raw reply [flat|nested] 38+ messages in thread
* Re: [PATCH v9 10/18] PCI/P2PDMA: Route Translated Requests under Direct Translated P2P
2026-10-01 11:55 ` [PATCH v9 10/18] PCI/P2PDMA: Route Translated Requests under Direct Translated P2P Leon Romanovsky
@ 2026-10-06 22:28 ` Bjorn Helgaas
0 siblings, 0 replies; 38+ messages in thread
From: Bjorn Helgaas @ 2026-10-06 22:28 UTC (permalink / raw)
To: Leon Romanovsky
Cc: Bjorn Helgaas, Logan Gunthorpe, Jason Gunthorpe,
Joerg Roedel (AMD),
Will Deacon, Robin Murphy, Christian König,
Thomas Hellström, linux-pci, linux-kernel, linux-doc, iommu,
Tushar Dave, linux-media, dri-devel, linaro-mm-sig, linux-rdma,
kvm, Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Sumit Semwal
On Thu, Oct 01, 2026 at 02:55:18PM +0300, Leon Romanovsky wrote:
> From: Leon Romanovsky <leonro@nvidia.com>
>
> A Downstream Port with ACS Direct Translated P2P enabled routes a Request
> whose Address Type is Translated "to the peer Egress Port without
> redirection, regardless of ACS P2P Request Redirect and ACS P2P Egress
> Control", per PCIe r7.0 sec 6.12.3. P2PDMA assumes every Request carries an
> Untranslated address, so it sends an ATS client through the host bridge
> even where the fabric would route it straight to the peer.
"sends an ATS client through the host bridge" -- I assume this really
means "we told the caller that Requests would be routed through the
host bridge" when in reality they wouldn't? I don't think this
actually changes any routing in the fabric, does it?
So essentially we told the caller that P2P between A and B was, e.g.,
5 hops when it was really only 2?
> Add PCI_P2PDMA_TLP_TRANSLATED and consult Direct Translated P2P for the
> Requests it describes.
>
> Reviewed-by: Logan Gunthorpe <logang@deltatee.com>
> Tested-by: Tushar Dave <tdave@nvidia.com>
> Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
> ---
> drivers/pci/p2pdma.c | 9 +++++++++
> 1 file changed, 9 insertions(+)
>
> diff --git a/drivers/pci/p2pdma.c b/drivers/pci/p2pdma.c
> index 569a74de3b3a..3fd2cb8d16f0 100644
> --- a/drivers/pci/p2pdma.c
> +++ b/drivers/pci/p2pdma.c
> @@ -548,6 +548,15 @@ pci_acs_p2pdma_request(u16 ctrl, unsigned int tlp_flags)
> */
> if (ctrl & PCI_ACS_TB)
> return PCI_ACS_P2PDMA_BLOCKED;
> +
> + /*
> + * PCIe r7.0 sec 6.12.3: ACS Direct Translated P2P routes a
> + * Request carrying a Translated address to the peer "without
> + * redirection, regardless of ACS P2P Request Redirect and ACS
> + * P2P Egress Control settings".
> + */
> + if (ctrl & PCI_ACS_DT)
> + return PCI_ACS_P2PDMA_DIRECT;
> }
>
> return ctrl & (PCI_ACS_RR | PCI_ACS_EC) ?
>
> --
> 2.55.0
>
^ permalink raw reply [flat|nested] 38+ messages in thread
* Re: [PATCH v9 02/18] PCI/P2PDMA: Derive routing from directional ACS controls
2026-10-06 20:49 ` Bjorn Helgaas
@ 2026-10-07 6:32 ` Leon Romanovsky
2026-10-07 20:40 ` Bjorn Helgaas
0 siblings, 1 reply; 38+ messages in thread
From: Leon Romanovsky @ 2026-10-07 6:32 UTC (permalink / raw)
To: Bjorn Helgaas
Cc: Bjorn Helgaas, Logan Gunthorpe, Jason Gunthorpe,
Joerg Roedel (AMD),
Will Deacon, Robin Murphy, Christian König,
Thomas Hellström, linux-pci, linux-kernel, linux-doc, iommu,
Tushar Dave, linux-media, dri-devel, linaro-mm-sig, linux-rdma,
kvm, Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Sumit Semwal
On Tue, Oct 06, 2026 at 03:49:47PM -0500, Bjorn Helgaas wrote:
> On Thu, Oct 01, 2026 at 02:55:10PM +0300, Leon Romanovsky wrote:
> > From: Leon Romanovsky <leonro@nvidia.com>
> >
> > pci_bridge_has_acs_redir() treats Request and Completion Redirect as
> > interchangeable. On asymmetric fabrics, a control for only the reverse TLP
> > direction can unnecessarily force P2PDMA through the host bridge.
>
> Does "the reverse TLP direction" refer to Completions?
In general, the P2P code treats TLPs flowing from device A to device B the
same as TLPs flowing from device B to device A.
However, in the context of this commit message, yes: completions flow in the
opposite direction from the device's perspective.
>
> > Evaluate Request Redirect for client Requests and Completion Redirect for
> > provider read Completions. Continue treating enabled Egress Control
> > conservatively as a Request redirect.
>
> Completion Redirect is intended to avoid ordering rule violations
> between Completions and Requests when Requests are redirected (PCIe
> r7.0, sec 6.12.1.1). I assume this patch preserves the ordering rule,
> but does the commit log need to say something about that? I don't
> know enough about P2P DMA for it to be obvious to me.
I don't think so, i didn't change anything related to ordering.
>
> Not really a question for this series, but p2pdma.c and p2pdma.rst
> refer to "clients" and "providers", neither of which are mentioned in
> the PCIe spec. In this case it sounds like a client is a Requester
> and a provider is a Completer in spec terms. Is that always the case?
> If so, "client" and "provider" in this paragraph are not adding any
> information.
>
> If "client" is not the same concept as "Requester" and "provider" not
> the same as "Completer", maybe p2pdma.rst could explain the
> difference?
Client vs. provider are actual target vs. initiator. They express the
device role in the flow.
Thanks
>
> > Fixes: 52916982af48 ("PCI/P2PDMA: Support peer-to-peer memory")
> > Reviewed-by: Logan Gunthorpe <logang@deltatee.com>
> > Tested-by: Tushar Dave <tdave@nvidia.com>
> > Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
> > ---
> > drivers/pci/p2pdma.c | 75 ++++++++++++++++++++++++++++++++++++++++------------
> > 1 file changed, 58 insertions(+), 17 deletions(-)
> >
> > diff --git a/drivers/pci/p2pdma.c b/drivers/pci/p2pdma.c
> > index 4e4d2df17a45..12612b82d80d 100644
> > --- a/drivers/pci/p2pdma.c
> > +++ b/drivers/pci/p2pdma.c
> > @@ -21,6 +21,8 @@
> > #include <linux/seq_buf.h>
> > #include <linux/xarray.h>
> >
> > +#include "pci.h"
> > +
> > struct pci_p2pdma {
> > struct gen_pool *pool;
> > bool p2pmem_published;
> > @@ -490,26 +492,56 @@ static struct pci_dev *find_parent_pci_dev(struct device *dev)
> > return NULL;
> > }
> >
> > +enum pci_acs_p2pdma_state {
> > + PCI_ACS_P2PDMA_DIRECT,
> > + PCI_ACS_P2PDMA_REDIRECT,
> > +};
> > +
> > /*
> > - * Check if a PCI bridge has its ACS redirection bits set to redirect P2P
> > - * TLPs upstream via ACS. Returns 1 if the packets will be redirected
> > - * upstream, 0 otherwise.
> > + * Decide how a peer-to-peer Request at an ACS-capable ingress port routes,
> > + * from that port's ACS Control register.
> > + *
> > + * Linux does not read the Egress Control Vector, so Egress Control is treated
> > + * conservatively as a redirect. Per PCIe r7.0 Table 6-11 the outcomes it
> > + * selects are a direct route and an ACS Violation, and neither one lets peer
> > + * bus addressing be assumed.
> > */
> > -static int pci_bridge_has_acs_redir(struct pci_dev *pdev)
> > +static enum pci_acs_p2pdma_state
> > +pci_acs_p2pdma_request(u16 ctrl)
> > {
> > - int pos;
> > - u16 ctrl;
> > + return ctrl & (PCI_ACS_RR | PCI_ACS_EC) ?
> > + PCI_ACS_P2PDMA_REDIRECT : PCI_ACS_P2PDMA_DIRECT;
> > +}
> >
> > - pos = pdev->acs_cap;
> > - if (!pos)
> > - return 0;
> > +/*
> > + * Decide how a peer-to-peer Completion at an ACS-capable ingress port routes.
> > + * PCIe r7.0 sec 6.12.1.1: no ACS control other than P2P Completion Redirect
> > + * affects a Completion.
> > + */
> > +static enum pci_acs_p2pdma_state
> > +pci_acs_p2pdma_completion(u16 ctrl)
> > +{
> > + return ctrl & PCI_ACS_CR ? PCI_ACS_P2PDMA_REDIRECT :
> > + PCI_ACS_P2PDMA_DIRECT;
> > +}
> >
> > - pci_read_config_word(pdev, pos + PCI_ACS_CTRL, &ctrl);
> > +/*
> > + * Read @pdev's ACS Control register. A device without an ACS capability has
> > + * no peer-to-peer controls at all, which routes the same as having them all
> > + * clear. Returns false when the register is present but cannot be read; @ctrl
> > + * is then meaningless.
> > + */
> > +static bool pci_acs_p2pdma_ctrl(struct pci_dev *pdev, u16 *ctrl)
> > +{
> > + int pos;
> >
> > - if (ctrl & (PCI_ACS_RR | PCI_ACS_CR | PCI_ACS_EC))
> > - return 1;
> > + pos = pdev->acs_cap;
> > + if (!pos) {
> > + *ctrl = 0;
> > + return true;
> > + }
> >
> > - return 0;
> > + return !pci_read_config_word(pdev, pos + PCI_ACS_CTRL, ctrl);
> > }
> >
> > static void seq_buf_print_bus_devfn(struct seq_buf *buf, struct pci_dev *pdev)
> > @@ -698,6 +730,10 @@ static unsigned long map_types_idx(struct pci_dev *client)
> > * then to Device B. The mapping type returned depends on the ACS
> > * redirection setting of the ports along the path.
> > *
> > + * The client initiates Requests to provider memory. Check Request Redirect
> > + * on the client path and Completion Redirect for read Completions on the
> > + * provider path.
> > + *
> > * If ACS redirect is set on any port in the path, traffic between the
> > * devices will go through the host bridge, so return
> > * PCI_P2PDMA_MAP_THRU_HOST_BRIDGE; otherwise return
> > @@ -721,6 +757,7 @@ calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client,
> > int dist_a = 0;
> > int dist_b = 0;
> > char buf[128];
> > + u16 ctrl;
> >
> > seq_buf_init(&acs_list, buf, sizeof(buf));
> >
> > @@ -732,7 +769,9 @@ calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client,
> > while (a) {
> > dist_b = 0;
> >
> > - if (pci_bridge_has_acs_redir(a)) {
> > + if (!pci_acs_p2pdma_ctrl(a, &ctrl) ||
> > + pci_acs_p2pdma_completion(ctrl) ==
> > + PCI_ACS_P2PDMA_REDIRECT) {
> > seq_buf_print_bus_devfn(&acs_list, a);
> > acs_cnt++;
> > }
> > @@ -761,7 +800,9 @@ calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client,
> > if (a == bb)
> > break;
> >
> > - if (pci_bridge_has_acs_redir(bb)) {
> > + if (!pci_acs_p2pdma_ctrl(bb, &ctrl) ||
> > + pci_acs_p2pdma_request(ctrl) ==
> > + PCI_ACS_P2PDMA_REDIRECT) {
> > seq_buf_print_bus_devfn(&acs_list, bb);
> > acs_cnt++;
> > }
> > @@ -1109,10 +1150,10 @@ EXPORT_SYMBOL_GPL(pci_p2pmem_publish);
> > /**
> > * pci_p2pdma_map_type - Determine the mapping type for P2PDMA transfers
> > * @provider: P2PDMA provider structure
> > - * @dev: Target device for the transfer
> > + * @dev: Client device that initiates the transfer
> > *
> > * Determines how peer-to-peer DMA transfers should be mapped between
> > - * the provider and the target device. The mapping type indicates whether
> > + * the provider and the client device. The mapping type indicates whether
> > * the transfer can be done directly through PCI switches or must go
> > * through the host bridge.
> > */
> >
> > --
> > 2.55.0
> >
^ permalink raw reply [flat|nested] 38+ messages in thread
* Re: [PATCH v9 00/18] PCI/P2PDMA: Route peer-to-peer DMA by TLP class
2026-10-06 21:36 ` Bjorn Helgaas
@ 2026-10-07 6:56 ` Bjorn Helgaas
2026-10-07 10:54 ` Leon Romanovsky
0 siblings, 1 reply; 38+ messages in thread
From: Bjorn Helgaas @ 2026-10-07 6:56 UTC (permalink / raw)
To: Leon Romanovsky
Cc: Bjorn Helgaas, Logan Gunthorpe, Jason Gunthorpe,
Joerg Roedel (AMD),
Will Deacon, Robin Murphy, Christian König,
Thomas Hellström, linux-pci, linux-kernel, linux-doc, iommu,
Tushar Dave, linux-media, dri-devel, linaro-mm-sig, linux-rdma,
kvm, Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Sumit Semwal
On Tue, Oct 06, 2026 at 04:36:50PM -0500, Bjorn Helgaas wrote:
> On Wed, Oct 07, 2026 at 12:22:27AM +0300, Leon Romanovsky wrote:
> > On Tue, Oct 06, 2026 at 02:29:55PM -0500, Bjorn Helgaas wrote:
> > > On Thu, Oct 01, 2026 at 02:55:08PM +0300, Leon Romanovsky wrote:
> > > > PCI P2PDMA applies Request and Completion Redirect throughout both paths.
> > > > This misclassifies asymmetric and nested switches, and reports one answer
> > > > for every kind of TLP.
> > > >
> > > > Three ACS controls act on TLP attributes the client chooses rather than
> > > > on the topology: Translation Blocking and Direct Translated P2P act on
> > > > a Request's Address Type, and Completion Redirect skips Completions carrying
> > > > Relaxed Ordering.
> > > >
> > > > Evaluate each direction at the path divergence, decide every class from the
> > > > one walk, and treat a client with ATS enabled as translating unless its
> > > > driver declares per-mapping ATS.
> > > >
> > > > This completes the P2PDMA side; dma-buf and mlx5 follow separately.
> > > > ...
> > >
> > > > Leon Romanovsky (18):
> > > > PCI/P2PDMA: Document the TLP attribute assumptions
> > > > PCI/P2PDMA: Derive routing from directional ACS controls
> > > > PCI: Reject unreadable ACS controls in isolation checks
> > > > PCI/P2PDMA: Evaluate ACS controls at the path divergence
> > > > PCI/P2PDMA: Document directional ACS routing
> > > > PCI/P2PDMA: Collect the path's ACS controls before deciding
> > > > PCI/P2PDMA: Answer routing per TLP class
> > > > PCI/P2PDMA: Route Relaxed Ordering Completions directly
> > > > PCI/P2PDMA: Reject Translated Requests blocked by Translation Blocking
> > > > PCI/P2PDMA: Route Translated Requests under Direct Translated P2P
> > > > PCI/P2PDMA: Log detailed ACS routing diagnostics
> > > > PCI/P2PDMA: Add KUnit tests for the ACS routing decisions
> > > > PCI/P2PDMA: Test the ACS P2P routing walk
> > > > PCI: Add KUnit coverage for ACS isolation checks
> > > > PCI/P2PDMA: Document TLP-class routing
> > > > PCI/P2PDMA: Let a client declare that it selects ATS per mapping
> > > > PCI/P2PDMA: Evaluate the ATS path for clients with ATS enabled
> > > > PCI/P2PDMA: Test the routing of clients with ATS enabled
> > > >
> > > > Documentation/admin-guide/kernel-parameters.txt | 15 +-
> > > > Documentation/driver-api/pci/p2pdma.rst | 80 +++
> > > > drivers/pci/Kconfig | 15 +
> > > > drivers/pci/Makefile | 1 +
> > > > drivers/pci/p2pdma.c | 719 ++++++++++++++++++--
> > > > drivers/pci/pci.c | 7 +-
> > > > drivers/pci/pci.h | 58 ++
> > > > drivers/pci/pci_acs_test.c | 860 ++++++++++++++++++++++++
> > > > drivers/pci/quirks.c | 6 +-
> > > > include/linux/pci-p2pdma.h | 12 +-
> > > > 10 files changed, 1687 insertions(+), 86 deletions(-)
> > > > ---
> > > > base-commit: 08dbfad3f5040f5bdb6c529da20d6d4e81fefd72
> > > > change-id: 20260821-fix-p2p-acs-v4-0-e72455e3a261
> > > > prerequisite-message-id: <20260830-batch-p2p-fixes-v1-0-5044e8dfbe2e@nvidia.com>
> > > > prerequisite-patch-id: 6b25c7fcf164cdfc14e9fac5b908d97fcf6509d7
> > > > prerequisite-patch-id: 0d083c281001365aae4b35544cf28891a6ab9a96
> > > > prerequisite-patch-id: bfd9dabf271f3cc9a3a61f46387d20c20311363d
> > > > prerequisite-patch-id: fad0275efc722830fc591509506c0a5e4f581073
> > > > prerequisite-patch-id: 0c83bee688fec1f6d1564654df7c630fa6a4a978
> > >
> > > This seems like material for the PCI tree, but I'm not sure how to
> > > apply it. It doesn't apply cleanly on the current pci/p2pdma
> > > (https://git.kernel.org/pub/scm/linux/kernel/git/pci/pci.git/log/?h=p2pdma),
> > > which does contain your series from
> > > 20260830-batch-p2p-fixes-v1-0-5044e8dfbe2e@nvidia.com.
> > >
> > > I could probably fix the conflicts but I don't know why there should
> > > be conflicts, since the only commits on pci/p2pdma other than yours
> > > are a few trivial allow-list updates.
> >
> > So maybe this is related to these trivial updates, my series was based
> > on the clean p2pdma branch:
> > https://git.kernel.org/pub/scm/linux/kernel/git/leon/linux-rdma.git/log/?h=fix-p2p-acs-v9
>
> Ah, yeah, it must be the allow-list additions that I have but you
> don't. I wouldn't have thought the offsets would have been a problem.
It applies fine if I move the Haswell and Alibaba allowlist patches t
to the end so the base matches fix-p2p-acs-v9, thanks.
^ permalink raw reply [flat|nested] 38+ messages in thread
* Re: [PATCH v9 04/18] PCI/P2PDMA: Evaluate ACS controls at the path divergence
2026-10-06 21:08 ` Bjorn Helgaas
@ 2026-10-07 10:51 ` Leon Romanovsky
0 siblings, 0 replies; 38+ messages in thread
From: Leon Romanovsky @ 2026-10-07 10:51 UTC (permalink / raw)
To: Bjorn Helgaas
Cc: Bjorn Helgaas, Logan Gunthorpe, Jason Gunthorpe,
Joerg Roedel (AMD),
Will Deacon, Robin Murphy, Christian König,
Thomas Hellström, linux-pci, linux-kernel, linux-doc, iommu,
Tushar Dave, linux-media, dri-devel, linaro-mm-sig, linux-rdma,
kvm, Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Sumit Semwal
On Tue, Oct 06, 2026 at 04:08:48PM -0500, Bjorn Helgaas wrote:
> On Thu, Oct 01, 2026 at 02:55:12PM +0300, Leon Romanovsky wrote:
> > From: Leon Romanovsky <leonro@nvidia.com>
> >
> > ACS redirect controls choose between peer and upstream routes only at the
> > path divergence. Applying them below that point rejects valid nested
> > topologies because traffic already has only an upstream route.
>
> I guess the point here is that prior to this patch,
> calc_map_type_and_dist() returned PCI_P2PDMA_MAP_NOT_SUPPORTED in a
> case where it didn't need to? Can you include an example to make this
> concrete?
One possible example I had in mind while preparing the presentation [1] was
ability to connect devices to different switches. Despite devices and switches
support P2P, but from the P2P subsystem's point of view, traffic between them
was not P2P (see slide 4).
This support will become even more important when we reach the point of
enabling P2P inside a VM. It will allow us to hide the underlying PCI
topology from the VM.
Thanks
[1] http://lpc.events/event/20/contributions/2508/
^ permalink raw reply [flat|nested] 38+ messages in thread
* Re: [PATCH v9 00/18] PCI/P2PDMA: Route peer-to-peer DMA by TLP class
2026-10-07 6:56 ` Bjorn Helgaas
@ 2026-10-07 10:54 ` Leon Romanovsky
2026-10-07 11:31 ` Bjorn Helgaas
0 siblings, 1 reply; 38+ messages in thread
From: Leon Romanovsky @ 2026-10-07 10:54 UTC (permalink / raw)
To: Bjorn Helgaas
Cc: Bjorn Helgaas, Logan Gunthorpe, Jason Gunthorpe,
Joerg Roedel (AMD),
Will Deacon, Robin Murphy, Christian König,
Thomas Hellström, linux-pci, linux-kernel, linux-doc, iommu,
Tushar Dave, linux-media, dri-devel, linaro-mm-sig, linux-rdma,
kvm, Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Sumit Semwal
On Wed, Oct 07, 2026 at 01:56:44AM -0500, Bjorn Helgaas wrote:
> On Tue, Oct 06, 2026 at 04:36:50PM -0500, Bjorn Helgaas wrote:
> > On Wed, Oct 07, 2026 at 12:22:27AM +0300, Leon Romanovsky wrote:
> > > On Tue, Oct 06, 2026 at 02:29:55PM -0500, Bjorn Helgaas wrote:
> > > > On Thu, Oct 01, 2026 at 02:55:08PM +0300, Leon Romanovsky wrote:
> > > > > PCI P2PDMA applies Request and Completion Redirect throughout both paths.
> > > > > This misclassifies asymmetric and nested switches, and reports one answer
> > > > > for every kind of TLP.
> > > > >
> > > > > Three ACS controls act on TLP attributes the client chooses rather than
> > > > > on the topology: Translation Blocking and Direct Translated P2P act on
> > > > > a Request's Address Type, and Completion Redirect skips Completions carrying
> > > > > Relaxed Ordering.
> > > > >
> > > > > Evaluate each direction at the path divergence, decide every class from the
> > > > > one walk, and treat a client with ATS enabled as translating unless its
> > > > > driver declares per-mapping ATS.
> > > > >
> > > > > This completes the P2PDMA side; dma-buf and mlx5 follow separately.
> > > > > ...
> > > >
> > > > > Leon Romanovsky (18):
> > > > > PCI/P2PDMA: Document the TLP attribute assumptions
> > > > > PCI/P2PDMA: Derive routing from directional ACS controls
> > > > > PCI: Reject unreadable ACS controls in isolation checks
> > > > > PCI/P2PDMA: Evaluate ACS controls at the path divergence
> > > > > PCI/P2PDMA: Document directional ACS routing
> > > > > PCI/P2PDMA: Collect the path's ACS controls before deciding
> > > > > PCI/P2PDMA: Answer routing per TLP class
> > > > > PCI/P2PDMA: Route Relaxed Ordering Completions directly
> > > > > PCI/P2PDMA: Reject Translated Requests blocked by Translation Blocking
> > > > > PCI/P2PDMA: Route Translated Requests under Direct Translated P2P
> > > > > PCI/P2PDMA: Log detailed ACS routing diagnostics
> > > > > PCI/P2PDMA: Add KUnit tests for the ACS routing decisions
> > > > > PCI/P2PDMA: Test the ACS P2P routing walk
> > > > > PCI: Add KUnit coverage for ACS isolation checks
> > > > > PCI/P2PDMA: Document TLP-class routing
> > > > > PCI/P2PDMA: Let a client declare that it selects ATS per mapping
> > > > > PCI/P2PDMA: Evaluate the ATS path for clients with ATS enabled
> > > > > PCI/P2PDMA: Test the routing of clients with ATS enabled
> > > > >
> > > > > Documentation/admin-guide/kernel-parameters.txt | 15 +-
> > > > > Documentation/driver-api/pci/p2pdma.rst | 80 +++
> > > > > drivers/pci/Kconfig | 15 +
> > > > > drivers/pci/Makefile | 1 +
> > > > > drivers/pci/p2pdma.c | 719 ++++++++++++++++++--
> > > > > drivers/pci/pci.c | 7 +-
> > > > > drivers/pci/pci.h | 58 ++
> > > > > drivers/pci/pci_acs_test.c | 860 ++++++++++++++++++++++++
> > > > > drivers/pci/quirks.c | 6 +-
> > > > > include/linux/pci-p2pdma.h | 12 +-
> > > > > 10 files changed, 1687 insertions(+), 86 deletions(-)
> > > > > ---
> > > > > base-commit: 08dbfad3f5040f5bdb6c529da20d6d4e81fefd72
> > > > > change-id: 20260821-fix-p2p-acs-v4-0-e72455e3a261
> > > > > prerequisite-message-id: <20260830-batch-p2p-fixes-v1-0-5044e8dfbe2e@nvidia.com>
> > > > > prerequisite-patch-id: 6b25c7fcf164cdfc14e9fac5b908d97fcf6509d7
> > > > > prerequisite-patch-id: 0d083c281001365aae4b35544cf28891a6ab9a96
> > > > > prerequisite-patch-id: bfd9dabf271f3cc9a3a61f46387d20c20311363d
> > > > > prerequisite-patch-id: fad0275efc722830fc591509506c0a5e4f581073
> > > > > prerequisite-patch-id: 0c83bee688fec1f6d1564654df7c630fa6a4a978
> > > >
> > > > This seems like material for the PCI tree, but I'm not sure how to
> > > > apply it. It doesn't apply cleanly on the current pci/p2pdma
> > > > (https://git.kernel.org/pub/scm/linux/kernel/git/pci/pci.git/log/?h=p2pdma),
> > > > which does contain your series from
> > > > 20260830-batch-p2p-fixes-v1-0-5044e8dfbe2e@nvidia.com.
> > > >
> > > > I could probably fix the conflicts but I don't know why there should
> > > > be conflicts, since the only commits on pci/p2pdma other than yours
> > > > are a few trivial allow-list updates.
> > >
> > > So maybe this is related to these trivial updates, my series was based
> > > on the clean p2pdma branch:
> > > https://git.kernel.org/pub/scm/linux/kernel/git/leon/linux-rdma.git/log/?h=fix-p2p-acs-v9
> >
> > Ah, yeah, it must be the allow-list additions that I have but you
> > don't. I wouldn't have thought the offsets would have been a problem.
>
> It applies fine if I move the Haswell and Alibaba allowlist patches t
> to the end so the base matches fix-p2p-acs-v9, thanks.
Bjorn,
It is unclear to me if I should to resend.
Thanks
^ permalink raw reply [flat|nested] 38+ messages in thread
* Re: [PATCH v9 00/18] PCI/P2PDMA: Route peer-to-peer DMA by TLP class
2026-10-07 10:54 ` Leon Romanovsky
@ 2026-10-07 11:31 ` Bjorn Helgaas
0 siblings, 0 replies; 38+ messages in thread
From: Bjorn Helgaas @ 2026-10-07 11:31 UTC (permalink / raw)
To: Leon Romanovsky
Cc: Bjorn Helgaas, Logan Gunthorpe, Jason Gunthorpe,
Joerg Roedel (AMD),
Will Deacon, Robin Murphy, Christian König,
Thomas Hellström, linux-pci, linux-kernel, linux-doc, iommu,
Tushar Dave, linux-media, dri-devel, linaro-mm-sig, linux-rdma,
kvm, Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Sumit Semwal
On Wed, Oct 07, 2026 at 01:54:09PM +0300, Leon Romanovsky wrote:
> On Wed, Oct 07, 2026 at 01:56:44AM -0500, Bjorn Helgaas wrote:
> > On Tue, Oct 06, 2026 at 04:36:50PM -0500, Bjorn Helgaas wrote:
> > > On Wed, Oct 07, 2026 at 12:22:27AM +0300, Leon Romanovsky wrote:
> > > > On Tue, Oct 06, 2026 at 02:29:55PM -0500, Bjorn Helgaas wrote:
> > > > > On Thu, Oct 01, 2026 at 02:55:08PM +0300, Leon Romanovsky wrote:
> > > > > > PCI P2PDMA applies Request and Completion Redirect throughout both paths.
> > > > > > This misclassifies asymmetric and nested switches, and reports one answer
> > > > > > for every kind of TLP.
> > > > > >
> > > > > > Three ACS controls act on TLP attributes the client chooses rather than
> > > > > > on the topology: Translation Blocking and Direct Translated P2P act on
> > > > > > a Request's Address Type, and Completion Redirect skips Completions carrying
> > > > > > Relaxed Ordering.
> > > > > >
> > > > > > Evaluate each direction at the path divergence, decide every class from the
> > > > > > one walk, and treat a client with ATS enabled as translating unless its
> > > > > > driver declares per-mapping ATS.
> > > > > >
> > > > > > This completes the P2PDMA side; dma-buf and mlx5 follow separately.
> > > > > > ...
> > > > >
> > > > > > Leon Romanovsky (18):
> > > > > > PCI/P2PDMA: Document the TLP attribute assumptions
> > > > > > PCI/P2PDMA: Derive routing from directional ACS controls
> > > > > > PCI: Reject unreadable ACS controls in isolation checks
> > > > > > PCI/P2PDMA: Evaluate ACS controls at the path divergence
> > > > > > PCI/P2PDMA: Document directional ACS routing
> > > > > > PCI/P2PDMA: Collect the path's ACS controls before deciding
> > > > > > PCI/P2PDMA: Answer routing per TLP class
> > > > > > PCI/P2PDMA: Route Relaxed Ordering Completions directly
> > > > > > PCI/P2PDMA: Reject Translated Requests blocked by Translation Blocking
> > > > > > PCI/P2PDMA: Route Translated Requests under Direct Translated P2P
> > > > > > PCI/P2PDMA: Log detailed ACS routing diagnostics
> > > > > > PCI/P2PDMA: Add KUnit tests for the ACS routing decisions
> > > > > > PCI/P2PDMA: Test the ACS P2P routing walk
> > > > > > PCI: Add KUnit coverage for ACS isolation checks
> > > > > > PCI/P2PDMA: Document TLP-class routing
> > > > > > PCI/P2PDMA: Let a client declare that it selects ATS per mapping
> > > > > > PCI/P2PDMA: Evaluate the ATS path for clients with ATS enabled
> > > > > > PCI/P2PDMA: Test the routing of clients with ATS enabled
> > > > > >
> > > > > > Documentation/admin-guide/kernel-parameters.txt | 15 +-
> > > > > > Documentation/driver-api/pci/p2pdma.rst | 80 +++
> > > > > > drivers/pci/Kconfig | 15 +
> > > > > > drivers/pci/Makefile | 1 +
> > > > > > drivers/pci/p2pdma.c | 719 ++++++++++++++++++--
> > > > > > drivers/pci/pci.c | 7 +-
> > > > > > drivers/pci/pci.h | 58 ++
> > > > > > drivers/pci/pci_acs_test.c | 860 ++++++++++++++++++++++++
> > > > > > drivers/pci/quirks.c | 6 +-
> > > > > > include/linux/pci-p2pdma.h | 12 +-
> > > > > > 10 files changed, 1687 insertions(+), 86 deletions(-)
> > > > > > ---
> > > > > > base-commit: 08dbfad3f5040f5bdb6c529da20d6d4e81fefd72
> > > > > > change-id: 20260821-fix-p2p-acs-v4-0-e72455e3a261
> > > > > > prerequisite-message-id: <20260830-batch-p2p-fixes-v1-0-5044e8dfbe2e@nvidia.com>
> > > > > > prerequisite-patch-id: 6b25c7fcf164cdfc14e9fac5b908d97fcf6509d7
> > > > > > prerequisite-patch-id: 0d083c281001365aae4b35544cf28891a6ab9a96
> > > > > > prerequisite-patch-id: bfd9dabf271f3cc9a3a61f46387d20c20311363d
> > > > > > prerequisite-patch-id: fad0275efc722830fc591509506c0a5e4f581073
> > > > > > prerequisite-patch-id: 0c83bee688fec1f6d1564654df7c630fa6a4a978
> > > > >
> > > > > This seems like material for the PCI tree, but I'm not sure how to
> > > > > apply it. It doesn't apply cleanly on the current pci/p2pdma
> > > > > (https://git.kernel.org/pub/scm/linux/kernel/git/pci/pci.git/log/?h=p2pdma),
> > > > > which does contain your series from
> > > > > 20260830-batch-p2p-fixes-v1-0-5044e8dfbe2e@nvidia.com.
> > > > >
> > > > > I could probably fix the conflicts but I don't know why there should
> > > > > be conflicts, since the only commits on pci/p2pdma other than yours
> > > > > are a few trivial allow-list updates.
> > > >
> > > > So maybe this is related to these trivial updates, my series was based
> > > > on the clean p2pdma branch:
> > > > https://git.kernel.org/pub/scm/linux/kernel/git/leon/linux-rdma.git/log/?h=fix-p2p-acs-v9
> > >
> > > Ah, yeah, it must be the allow-list additions that I have but you
> > > don't. I wouldn't have thought the offsets would have been a problem.
> >
> > It applies fine if I move the Haswell and Alibaba allowlist patches t
> > to the end so the base matches fix-p2p-acs-v9, thanks.
>
> Bjorn,
>
> It is unclear to me if I should to resend.
No need.
^ permalink raw reply [flat|nested] 38+ messages in thread
* Re: [PATCH v9 05/18] PCI/P2PDMA: Document directional ACS routing
2026-10-06 21:32 ` Bjorn Helgaas
@ 2026-10-07 13:46 ` Leon Romanovsky
2026-10-07 20:10 ` Bjorn Helgaas
0 siblings, 1 reply; 38+ messages in thread
From: Leon Romanovsky @ 2026-10-07 13:46 UTC (permalink / raw)
To: Bjorn Helgaas
Cc: Bjorn Helgaas, Logan Gunthorpe, Jason Gunthorpe,
Joerg Roedel (AMD),
Will Deacon, Robin Murphy, Christian König,
Thomas Hellström, linux-pci, linux-kernel, linux-doc, iommu,
Tushar Dave, linux-media, dri-devel, linaro-mm-sig, linux-rdma,
kvm, Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Sumit Semwal
On Tue, Oct 06, 2026 at 04:32:05PM -0500, Bjorn Helgaas wrote:
> On Thu, Oct 01, 2026 at 02:55:13PM +0300, Leon Romanovsky wrote:
> > From: Leon Romanovsky <leonro@nvidia.com>
> >
> > P2PDMA documentation describes ACS controls as path-wide, although Request
> > and Completion controls apply to different transaction directions and only
> > affect peer-versus-upstream decisions at the path divergence.
> >
> > Document the fixed client and provider roles, the divergence port checked
> > for each TLP direction, and the conservative handling of unreadable ACS
> > state. Clarify which controls disable_acs_redir changes.
> >
> > Reviewed-by: Logan Gunthorpe <logang@deltatee.com>
> > Tested-by: Tushar Dave <tdave@nvidia.com>
> > Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
> > ---
> > Documentation/admin-guide/kernel-parameters.txt | 15 +++++++++------
> > Documentation/driver-api/pci/p2pdma.rst | 13 +++++++++++++
> > 2 files changed, 22 insertions(+), 6 deletions(-)
> >
> > diff --git a/Documentation/admin-guide/kernel-parameters.txt b/Documentation/admin-guide/kernel-parameters.txt
> > index 68647ff4bdd2..bc83e07dd5fc 100644
> > --- a/Documentation/admin-guide/kernel-parameters.txt
> > +++ b/Documentation/admin-guide/kernel-parameters.txt
> > @@ -5291,12 +5291,15 @@ Kernel parameters
> > disable_acs_redir=<pci_dev>[; ...]
> > Specify one or more PCI devices (in the format
> > specified above) separated by semicolons.
> > - Each device specified will have the PCI ACS
> > - redirect capabilities forced off which will
> > - allow P2P traffic between devices through
> > - bridges without forcing it upstream. Note:
> > - this removes isolation between devices and
> > - may put more devices in an IOMMU group.
> > + Each device specified will have the PCI ACS P2P
> > + Request Redirect, Completion Redirect, and Egress
> > + Control features forced off. This may allow P2P
> > + traffic through bridges that would otherwise be
> > + redirected upstream. This may allow P2P traffic
> > + through bridges that would otherwise be redirected
> > + upstream and thus this removes isolation between
> > + devices and may cause affected devices to share
> > + an IOMMU group.
> > config_acs=
> > Format:
> > <ACS flags>@<pci_dev>[; ...]
> > diff --git a/Documentation/driver-api/pci/p2pdma.rst b/Documentation/driver-api/pci/p2pdma.rst
> > index 80f8fec9b0e9..42b18610bf7d 100644
> > --- a/Documentation/driver-api/pci/p2pdma.rst
> > +++ b/Documentation/driver-api/pci/p2pdma.rst
> > @@ -15,6 +15,19 @@ then based on the ACS settings the transaction can route entirely within
> > the PCIe hierarchy and never reach the root port. The kernel will evaluate
> > the PCIe topology and always permit P2P in these well-defined cases.
> >
> > +The client remains the PCIe requester when it reads or writes provider memory.
> > +Where the paths diverge, the kernel therefore evaluates P2P Request Redirect
> > +and Egress Control on the client-side port, and P2P Completion Redirect on the
> > +provider-side port for completions from a read. An enabled Egress Control is
> > +conservatively treated as a Request redirect.
>
> I think "where the paths diverge" means the point where a port decides
> whether to route a TLP up through its Upstream Port (or to the RC) or
> back down via a sibling Downstream Port?
Yes.
>
> I'm not sure p2pdma.rst includes the context to interpret "divergence"
> yet. I think the code comments in [4/18] might also need a little
> more context about divergence.
>
> IIUC, this divergence point is basically a static point where the ACS
> controls being applied makes a routing difference.
Right. Previously, we did not account for ACS bits, so any "junction"
automatically caused P2P to be considered unsupported.
Thanks
>
> > +Below the divergence, the route toward the other branch is already upstream,
> > +so those P2P redirect controls do not affect it. Redirect controls for the
> > +reverse transaction directions do not affect the mapping. P2P DMA is routed
> > +through the host bridge when either applicable port redirects. If an ACS
> > +Control register cannot be read, P2P DMA is rejected because the kernel cannot
> > +establish a usable route.
> > +
> > This evaluation assumes clients issue strictly ordered Requests carrying an
> > Untranslated address. Its result is not defined when clients use Relaxed
> > Ordering or issue ATS-translated Requests because those TLP attributes can
> >
> > --
> > 2.55.0
> >
^ permalink raw reply [flat|nested] 38+ messages in thread
* Re: [PATCH v9 06/18] PCI/P2PDMA: Collect the path's ACS controls before deciding
2026-10-06 21:48 ` Bjorn Helgaas
@ 2026-10-07 13:50 ` Leon Romanovsky
0 siblings, 0 replies; 38+ messages in thread
From: Leon Romanovsky @ 2026-10-07 13:50 UTC (permalink / raw)
To: Bjorn Helgaas
Cc: Bjorn Helgaas, Logan Gunthorpe, Jason Gunthorpe,
Joerg Roedel (AMD),
Will Deacon, Robin Murphy, Christian König,
Thomas Hellström, linux-pci, linux-kernel, linux-doc, iommu,
Tushar Dave, linux-media, dri-devel, linaro-mm-sig, linux-rdma,
kvm, Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Sumit Semwal
On Tue, Oct 06, 2026 at 04:48:58PM -0500, Bjorn Helgaas wrote:
> On Thu, Oct 01, 2026 at 02:55:14PM +0300, Leon Romanovsky wrote:
> > From: Leon Romanovsky <leonro@nvidia.com>
> >
> > calc_map_type_and_dist() reads each divergence port's ACS Control register
> > and folds the result into running counters as it goes. Any routing property
> > that depends on the kind of TLP being routed would have to be threaded
> > through that code, so there is nowhere to put one without reading the
> > registers again for each kind.
>
> What is the "one" that there's nowhere to put? I guess the routing
> property? So this is an optimization to avoid some config reads?
Right, cache the ACS bits, since we need to query them for both devices
whenever we calculate routing.
Thanks
^ permalink raw reply [flat|nested] 38+ messages in thread
* Re: [PATCH v9 05/18] PCI/P2PDMA: Document directional ACS routing
2026-10-07 13:46 ` Leon Romanovsky
@ 2026-10-07 20:10 ` Bjorn Helgaas
0 siblings, 0 replies; 38+ messages in thread
From: Bjorn Helgaas @ 2026-10-07 20:10 UTC (permalink / raw)
To: Leon Romanovsky
Cc: Bjorn Helgaas, Logan Gunthorpe, Jason Gunthorpe,
Joerg Roedel (AMD),
Will Deacon, Robin Murphy, Christian König,
Thomas Hellström, linux-pci, linux-kernel, linux-doc, iommu,
Tushar Dave, linux-media, dri-devel, linaro-mm-sig, linux-rdma,
kvm, Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Sumit Semwal
On Wed, Oct 07, 2026 at 04:46:58PM +0300, Leon Romanovsky wrote:
> On Tue, Oct 06, 2026 at 04:32:05PM -0500, Bjorn Helgaas wrote:
> > On Thu, Oct 01, 2026 at 02:55:13PM +0300, Leon Romanovsky wrote:
> > > From: Leon Romanovsky <leonro@nvidia.com>
> > >
> > > P2PDMA documentation describes ACS controls as path-wide, although Request
> > > and Completion controls apply to different transaction directions and only
> > > affect peer-versus-upstream decisions at the path divergence.
> > >
> > > Document the fixed client and provider roles, the divergence port checked
> > > for each TLP direction, and the conservative handling of unreadable ACS
> > > state. Clarify which controls disable_acs_redir changes.
> > >
> > > Reviewed-by: Logan Gunthorpe <logang@deltatee.com>
> > > Tested-by: Tushar Dave <tdave@nvidia.com>
> > > Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
> > > ---
> > > Documentation/admin-guide/kernel-parameters.txt | 15 +++++++++------
> > > Documentation/driver-api/pci/p2pdma.rst | 13 +++++++++++++
> > > 2 files changed, 22 insertions(+), 6 deletions(-)
> > >
> > > diff --git a/Documentation/admin-guide/kernel-parameters.txt b/Documentation/admin-guide/kernel-parameters.txt
> > > index 68647ff4bdd2..bc83e07dd5fc 100644
> > > --- a/Documentation/admin-guide/kernel-parameters.txt
> > > +++ b/Documentation/admin-guide/kernel-parameters.txt
> > > @@ -5291,12 +5291,15 @@ Kernel parameters
> > > disable_acs_redir=<pci_dev>[; ...]
> > > Specify one or more PCI devices (in the format
> > > specified above) separated by semicolons.
> > > - Each device specified will have the PCI ACS
> > > - redirect capabilities forced off which will
> > > - allow P2P traffic between devices through
> > > - bridges without forcing it upstream. Note:
> > > - this removes isolation between devices and
> > > - may put more devices in an IOMMU group.
> > > + Each device specified will have the PCI ACS P2P
> > > + Request Redirect, Completion Redirect, and Egress
> > > + Control features forced off. This may allow P2P
> > > + traffic through bridges that would otherwise be
> > > + redirected upstream. This may allow P2P traffic
> > > + through bridges that would otherwise be redirected
> > > + upstream and thus this removes isolation between
> > > + devices and may cause affected devices to share
> > > + an IOMMU group.
> > > config_acs=
> > > Format:
> > > <ACS flags>@<pci_dev>[; ...]
> > > diff --git a/Documentation/driver-api/pci/p2pdma.rst b/Documentation/driver-api/pci/p2pdma.rst
> > > index 80f8fec9b0e9..42b18610bf7d 100644
> > > --- a/Documentation/driver-api/pci/p2pdma.rst
> > > +++ b/Documentation/driver-api/pci/p2pdma.rst
> > > @@ -15,6 +15,19 @@ then based on the ACS settings the transaction can route entirely within
> > > the PCIe hierarchy and never reach the root port. The kernel will evaluate
> > > the PCIe topology and always permit P2P in these well-defined cases.
> > >
> > > +The client remains the PCIe requester when it reads or writes provider memory.
> > > +Where the paths diverge, the kernel therefore evaluates P2P Request Redirect
> > > +and Egress Control on the client-side port, and P2P Completion Redirect on the
> > > +provider-side port for completions from a read. An enabled Egress Control is
> > > +conservatively treated as a Request redirect.
Does "client remains" imply that the client can also be a Completer in
other circumstances? I'd like to use PCIe terms when possible since
we're talking about PCIe ACS controls.
s/requester/Requester/ since it's defined by PCIe spec.
s/completions/Completions/ similarly.
s/Request redirect/Request Redirect/ to match spec.
s/and Egress Control/and P2P Egress Control/ to match spec and RR and
CR mentions here. The second "Egress Control" matches the "Request
Redirect" so I think that's fine.
> > I think "where the paths diverge" means the point where a port decides
> > whether to route a TLP up through its Upstream Port (or to the RC) or
> > back down via a sibling Downstream Port?
>
> Yes.
Am I right in thinking that there may be several such points e.g., a
Request is routed back downstream by any Downstream Port where P2P
Request Redirect is not enabled?
I think this "paths diverge" needs some background when it's
introduced, e.g. something like this (fix or reword as necessary):
If the Requester and Completer are below the same Root Port, the
path taken by Requests depends on ACS settings of the Downstream
Ports above the Requester. If ACS P2P Request Redirect is enabled
in all of them, the Request is routed upstream; otherwise it will be
routed back downstream via the Downstream Port leading to the
Completer. The path of Completions likewise depends on ACS P2P
Completion Redirect in the Downstream Ports above the Completer.
> > I'm not sure p2pdma.rst includes the context to interpret "divergence"
> > yet. I think the code comments in [4/18] might also need a little
> > more context about divergence.
> >
> > IIUC, this divergence point is basically a static point where the ACS
> > controls being applied makes a routing difference.
>
> Right. Previously, we did not account for ACS bits, so any "junction"
> automatically caused P2P to be considered unsupported.
>
> > > +Below the divergence, the route toward the other branch is already upstream,
> > > +so those P2P redirect controls do not affect it. Redirect controls for the
> > > +reverse transaction directions do not affect the mapping. P2P DMA is routed
> > > +through the host bridge when either applicable port redirects. If an ACS
> > > +Control register cannot be read, P2P DMA is rejected because the kernel cannot
> > > +establish a usable route.
Similarly, I don't think "the divergence" is defined yet here. I
*think* it means the lowest Switch or Root Port shared between
Requester and Completer. That's the first point where a TLP could be
either routed back downstream or upstream.
I think a TLP could also be routed back downstream if a higher Switch
Downstream Port or the Root Port had Request/Completion Redirect
cleared and lower Switch Downstream Ports had it set.
> > > +
> > > This evaluation assumes clients issue strictly ordered Requests carrying an
> > > Untranslated address. Its result is not defined when clients use Relaxed
> > > Ordering or issue ATS-translated Requests because those TLP attributes can
> > >
> > > --
> > > 2.55.0
> > >
^ permalink raw reply [flat|nested] 38+ messages in thread
* Re: [PATCH v9 01/18] PCI/P2PDMA: Document the TLP attribute assumptions
2026-10-01 11:55 ` [PATCH v9 01/18] PCI/P2PDMA: Document the TLP attribute assumptions Leon Romanovsky
@ 2026-10-07 20:12 ` Bjorn Helgaas
0 siblings, 0 replies; 38+ messages in thread
From: Bjorn Helgaas @ 2026-10-07 20:12 UTC (permalink / raw)
To: Leon Romanovsky
Cc: Bjorn Helgaas, Logan Gunthorpe, Jason Gunthorpe,
Joerg Roedel (AMD),
Will Deacon, Robin Murphy, Christian König,
Thomas Hellström, linux-pci, linux-kernel, linux-doc, iommu,
Tushar Dave, linux-media, dri-devel, linaro-mm-sig, linux-rdma,
kvm, Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Sumit Semwal
On Thu, Oct 01, 2026 at 02:55:09PM +0300, Leon Romanovsky wrote:
> From: Leon Romanovsky <leonro@nvidia.com>
>
> P2PDMA selects a mapping without receiving the Request's ordering or
> Address Type attributes. Its ACS handles only strictly ordered Requests
> carrying an Untranslated address.
>
> Document that the result is not defined for Relaxed Ordering or
> ATS-translated Requests because those TLP attributes can select different
> routes through the fabric.
>
> Reviewed-by: Logan Gunthorpe <logang@deltatee.com>
> Tested-by: Tushar Dave <tdave@nvidia.com>
> Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
> ---
> Documentation/driver-api/pci/p2pdma.rst | 7 +++++++
> 1 file changed, 7 insertions(+)
>
> diff --git a/Documentation/driver-api/pci/p2pdma.rst b/Documentation/driver-api/pci/p2pdma.rst
> index 63cff9e4d2c9..80f8fec9b0e9 100644
> --- a/Documentation/driver-api/pci/p2pdma.rst
> +++ b/Documentation/driver-api/pci/p2pdma.rst
> @@ -15,6 +15,13 @@ then based on the ACS settings the transaction can route entirely within
> the PCIe hierarchy and never reach the root port. The kernel will evaluate
> the PCIe topology and always permit P2P in these well-defined cases.
>
> +This evaluation assumes clients issue strictly ordered Requests carrying an
> +Untranslated address. Its result is not defined when clients use Relaxed
> +Ordering or issue ATS-translated Requests because those TLP attributes can
> +select different routes through the fabric. Unless ACS Translation Blocking
> +is enabled, a Port with ACS Direct Translated P2P enabled routes a
> +Translated Request directly to the peer regardless of the redirect controls.
This whole paragraph is removed later in the series, so kudos for
documenting the current state before it changes :)
Unrelated to anything in this series, but this doc mentions
"p2p_provider", which doesn't exist. Should it be "p2pdma_provider"?
^ permalink raw reply [flat|nested] 38+ messages in thread
* Re: [PATCH v9 02/18] PCI/P2PDMA: Derive routing from directional ACS controls
2026-10-07 6:32 ` Leon Romanovsky
@ 2026-10-07 20:40 ` Bjorn Helgaas
0 siblings, 0 replies; 38+ messages in thread
From: Bjorn Helgaas @ 2026-10-07 20:40 UTC (permalink / raw)
To: Leon Romanovsky
Cc: Bjorn Helgaas, Logan Gunthorpe, Jason Gunthorpe,
Joerg Roedel (AMD),
Will Deacon, Robin Murphy, Christian König,
Thomas Hellström, linux-pci, linux-kernel, linux-doc, iommu,
Tushar Dave, linux-media, dri-devel, linaro-mm-sig, linux-rdma,
kvm, Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Sumit Semwal
On Wed, Oct 07, 2026 at 09:32:27AM +0300, Leon Romanovsky wrote:
> On Tue, Oct 06, 2026 at 03:49:47PM -0500, Bjorn Helgaas wrote:
> > On Thu, Oct 01, 2026 at 02:55:10PM +0300, Leon Romanovsky wrote:
> > > From: Leon Romanovsky <leonro@nvidia.com>
> > >
> > > pci_bridge_has_acs_redir() treats Request and Completion Redirect as
> > > interchangeable. On asymmetric fabrics, a control for only the reverse TLP
> > > direction can unnecessarily force P2PDMA through the host bridge.
> >
> > Does "the reverse TLP direction" refer to Completions?
>
> In general, the P2P code treats TLPs flowing from device A to device B the
> same as TLPs flowing from device B to device A.
>
> However, in the context of this commit message, yes: completions flow in the
> opposite direction from the device's perspective.
And I guess asymmetric fabrics must mean fabrics where Request
Redirect and Completion Redirect are not set the same way?
> > > Evaluate Request Redirect for client Requests and Completion Redirect for
> > > provider read Completions. Continue treating enabled Egress Control
> > > conservatively as a Request redirect.
> >
> > Completion Redirect is intended to avoid ordering rule violations
> > between Completions and Requests when Requests are redirected (PCIe
> > r7.0, sec 6.12.1.1). I assume this patch preserves the ordering rule,
> > but does the commit log need to say something about that? I don't
> > know enough about P2P DMA for it to be obvious to me.
>
> I don't think so, i didn't change anything related to ordering.
I don't think there's anything in this whole series that changes any
ACS settings, so I shouldn't have wondered about *preserving* the
ordering rule.
But I asked about ordering because it sounds like this patch expects
to encounter asymmetric fabrics where Request Redirect and Completion
Redirect may not be set the same way, and the spec implies that
asymmetry may result in ordering violations.
> > Not really a question for this series, but p2pdma.c and p2pdma.rst
> > refer to "clients" and "providers", neither of which are mentioned in
> > the PCIe spec. In this case it sounds like a client is a Requester
> > and a provider is a Completer in spec terms. Is that always the case?
> > If so, "client" and "provider" in this paragraph are not adding any
> > information.
> >
> > If "client" is not the same concept as "Requester" and "provider" not
> > the same as "Completer", maybe p2pdma.rst could explain the
> > difference?
>
> Client vs. provider are actual target vs. initiator. They express the
> device role in the flow.
I'd rather use "Requester" and "Completer" when possible because they
have specific meanings in the PCI spec and they correspond to the ACS
control bits. Client, provider, target, initiator are all from the
outer world that makes use of PCIe constructs, but they don't mean
anything inside the PCIe world.
Bjorn
^ permalink raw reply [flat|nested] 38+ messages in thread
end of thread, other threads:[~2026-10-07 20:40 UTC | newest]
Thread overview: 38+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-01 11:55 [PATCH v9 00/18] PCI/P2PDMA: Route peer-to-peer DMA by TLP class Leon Romanovsky
2026-10-01 11:55 ` [PATCH v9 01/18] PCI/P2PDMA: Document the TLP attribute assumptions Leon Romanovsky
2026-10-07 20:12 ` Bjorn Helgaas
2026-10-01 11:55 ` [PATCH v9 02/18] PCI/P2PDMA: Derive routing from directional ACS controls Leon Romanovsky
2026-10-06 20:49 ` Bjorn Helgaas
2026-10-07 6:32 ` Leon Romanovsky
2026-10-07 20:40 ` Bjorn Helgaas
2026-10-01 11:55 ` [PATCH v9 03/18] PCI: Reject unreadable ACS controls in isolation checks Leon Romanovsky
2026-10-01 11:55 ` [PATCH v9 04/18] PCI/P2PDMA: Evaluate ACS controls at the path divergence Leon Romanovsky
2026-10-06 21:08 ` Bjorn Helgaas
2026-10-07 10:51 ` Leon Romanovsky
2026-10-01 11:55 ` [PATCH v9 05/18] PCI/P2PDMA: Document directional ACS routing Leon Romanovsky
2026-10-06 21:32 ` Bjorn Helgaas
2026-10-07 13:46 ` Leon Romanovsky
2026-10-07 20:10 ` Bjorn Helgaas
2026-10-01 11:55 ` [PATCH v9 06/18] PCI/P2PDMA: Collect the path's ACS controls before deciding Leon Romanovsky
2026-10-06 21:48 ` Bjorn Helgaas
2026-10-07 13:50 ` Leon Romanovsky
2026-10-01 11:55 ` [PATCH v9 07/18] PCI/P2PDMA: Answer routing per TLP class Leon Romanovsky
2026-10-01 11:55 ` [PATCH v9 08/18] PCI/P2PDMA: Route Relaxed Ordering Completions directly Leon Romanovsky
2026-10-06 22:21 ` Bjorn Helgaas
2026-10-01 11:55 ` [PATCH v9 09/18] PCI/P2PDMA: Reject Translated Requests blocked by Translation Blocking Leon Romanovsky
2026-10-01 11:55 ` [PATCH v9 10/18] PCI/P2PDMA: Route Translated Requests under Direct Translated P2P Leon Romanovsky
2026-10-06 22:28 ` Bjorn Helgaas
2026-10-01 11:55 ` [PATCH v9 11/18] PCI/P2PDMA: Log detailed ACS routing diagnostics Leon Romanovsky
2026-10-01 11:55 ` [PATCH v9 12/18] PCI/P2PDMA: Add KUnit tests for the ACS routing decisions Leon Romanovsky
2026-10-01 11:55 ` [PATCH v9 13/18] PCI/P2PDMA: Test the ACS P2P routing walk Leon Romanovsky
2026-10-01 11:55 ` [PATCH v9 14/18] PCI: Add KUnit coverage for ACS isolation checks Leon Romanovsky
2026-10-01 11:55 ` [PATCH v9 15/18] PCI/P2PDMA: Document TLP-class routing Leon Romanovsky
2026-10-01 11:55 ` [PATCH v9 16/18] PCI/P2PDMA: Let a client declare that it selects ATS per mapping Leon Romanovsky
2026-10-01 11:55 ` [PATCH v9 17/18] PCI/P2PDMA: Evaluate the ATS path for clients with ATS enabled Leon Romanovsky
2026-10-01 11:55 ` [PATCH v9 18/18] PCI/P2PDMA: Test the routing of " Leon Romanovsky
2026-10-06 19:29 ` [PATCH v9 00/18] PCI/P2PDMA: Route peer-to-peer DMA by TLP class Bjorn Helgaas
2026-10-06 21:22 ` Leon Romanovsky
2026-10-06 21:36 ` Bjorn Helgaas
2026-10-07 6:56 ` Bjorn Helgaas
2026-10-07 10:54 ` Leon Romanovsky
2026-10-07 11:31 ` Bjorn Helgaas
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®