* [RFC PATCH 0/4] PCI: Fix and improve default MPS configuration
@ 2026-09-30 14:50 Niklas Cassel
2026-09-30 14:50 ` [RFC PATCH 1/4] PCI: Update saved Max Payload Size in pcie_set_mps() Niklas Cassel
` (3 more replies)
0 siblings, 4 replies; 5+ messages in thread
From: Niklas Cassel @ 2026-09-30 14:50 UTC (permalink / raw)
To: Bjorn Helgaas, Bjorn Helgaas, Lorenzo Pieralisi,
Krzysztof Wilczyński, Manivannan Sadhasivam, Heiko Stuebner,
Yue Wang, Hans Zhang, Rob Herring, Neil Armstrong, Kevin Hilman,
Jerome Brunet, Martin Blumenstingl, Frank Li, Myron Stowe,
Jon Mason
Cc: Keith Busch, mx2pg, dlemoal, Niklas Cassel, Lukas Wunner,
Mahesh Vaidya, Ricardo Pardini, Shawn Lin, Pali Rohár,
Jingoo Han, linux-pci, linux-kernel, linux-arm-kernel,
linux-amlogic, linux-rockchip, Manivannan Sadhasivam
This RFC reworks Hans's series "PCI: Configure Root Port MPS during host
probing" (v9, link below). Besides configuring Root Ports, which that
series was about, it fixes other problems in how mainline programs the
Maximum Payload Size (MPS), and lists the problems it doesn't fix.
It is an RFC because it changes the default MPS policy, see "Open
questions" below, and because it has only been build-tested so far.
Background
==========
A PCIe device must not send a TLP with a payload larger than the MPS it
is programmed with, and a receiver treats a TLP larger than its own MPS
as Malformed, so all devices on a path need compatible MPS settings.
Unless a "pci=pcie_bus_*" option is given, the kernel uses the default
strategy, PCIE_BUS_DEFAULT: pci_configure_mps() programs each device, as
it is enumerated, to the MPS of its upstream bridge. If the device's MPS
Supported (MPSS) is smaller and the upstream bridge is a Root Port, the
Root Port is lowered instead (9f0e89359775). The other strategies either
configure whole hierarchies in pcie_bus_configure_settings()
(pcie_bus_safe, pcie_bus_perf, pcie_bus_peer2peer) or don't touch MPS at
all (pcie_bus_tune_off).
Problems found in mainline
==========================
Enumeration:
1. The saved MPS goes stale. Fixed by patch 1.
pcie_set_mps() changes Device Control, but not the copy saved by
pci_bus_add_device() and the port driver. The core changes the MPS
of ports whose state has been saved: pci_configure_mps() lowers a
Root Port when a device that supports less is hot-added below it,
and with pci=pcie_bus_safe, pcie_bus_configure_settings() can do the
same after a hot-add. Drivers do it too, e.g. hfi1 on its upstream
bridge.
Since v7.3-rc1 (3fc686d550f6), a Root Port reset through
host->reset_root_port(), e.g. on AER recovery or Link Down with the
qcom and Rockchip DWC drivers, restores the Root Port's saved state
without saving it first. The Root Port goes back to the old, larger
MPS while the device below it is restored to the smaller one, and the
device may receive TLPs that it treats as Malformed.
2. Root Ports keep whatever MPS they came up with. Fixed by patch 3.
Root Ports have no upstream bridge, so pci_configure_mps() never
configures them, and every hierarchy inherits the MPS that firmware or
the hardware reset left in its Root Port. On platforms whose firmware
doesn't configure MPS, which is typical for DT-based systems with
native host controller drivers, every hierarchy runs at the 128-byte
reset value even if the Root Port and all devices below it support
more, at a cost in throughput. Host controller drivers work around
this: the Meson driver programs its Root Port to 256 bytes itself,
regardless of what the devices below it support, which patch 4
removes.
3. Lowering the MPS only works directly below a Root Port. Fixed by
patch 2.
- Below a Switch, the Switch ports keep the larger MPS, and a device
that supports less is left mismatched with a "use
pci=pcie_bus_safe" warning. This happens at boot whenever firmware
or a host controller driver programmed a larger MPS than a device
below a Switch supports. It has been a known limitation since the
default strategy was introduced (27d868b5e6cf).
- For a multi-function device directly below a Root Port, function 0
is programmed to the Root Port's MPS first. If function 1 supports
less, the Root Port is lowered, but function 0 is left at the larger
value and can send TLPs larger than the Root Port now accepts. Only
an informational message about the Root Port is printed. This was
introduced by 9f0e89359775.
Hotplug and rescan:
4. A device hot-added directly below a Root Port. Fixed by patches 1-3.
If the device supports less than the Root Port's MPS, the Root Port
is lowered, with the multi-function problem from 3 and the stale
saved state from 1. If it supports more, it is only matched to the
Root Port's current MPS: a Root Port is never raised again, e.g.
after a device that supported less has been replaced, and without
firmware MPS setup it stays at 128 bytes. Devices found by a rescan
below an empty Root Port slot behave the same.
5. A Switch hot-added directly below a Root Port. Fixed by patch 2.
The Upstream Port can lower the Root Port, but the Downstream Ports
and the devices below them are only matched to their upstream bridge,
and any of them that supports less is left mismatched with a warning.
With patch 2, nothing below the Root Port is in use yet, so the whole
new hierarchy is lowered. Devices hot-added below the new Switch
later are subject to 6.
6. A device hot-added below a Switch. Not addressed.
The Switch ports are in use and can't be lowered, so a device that
supports less than their MPS, e.g. less than what firmware
programmed, is left mismatched with a warning. pcie_bus_safe avoids
this by limiting hierarchies with hotplug-capable Switch ports to 128
bytes. Patch 3 doesn't raise hierarchies with hotplug-capable Switch
ports, so it doesn't make this more likely on hot-add, but a device
that shows up below a Switch without hotplug support, e.g. found by a
rescan, can hit it after patch 3 has raised the hierarchy.
7. A Switch hot-added below another Switch. Not addressed.
Same as 6, for every port and device of the new Switch that supports
less than the Downstream Port it is plugged into, e.g. when chaining
Thunderbolt devices.
8. A device added next to devices already in use below the same Root
Port, e.g. a new function found by a rescan while function 0 is
bound. Only partly addressed.
Mainline lowers the Root Port below the active function, which is
left at the larger MPS. With patch 2, the hierarchy isn't changed,
and the new function is left mismatched with a warning instead.
Other:
9. With pci=pcie_bus_safe or pcie_bus_perf, devices found by a rescan
aren't configured at all: pci_configure_mps() leaves them to
pcie_bus_configure_settings(), which the rescan helpers don't call.
This doesn't affect the default strategy, where pci_configure_mps()
configures each device as it's enumerated. Calling
pcie_bus_configure_settings() from the rescan helpers wouldn't fix
this safely, since those strategies reprogram every device of the
hierarchy, including ones in use. Not addressed.
Approach
========
Two constraints drive the design:
- The MPS of a device that may have a driver bound can't be changed
safely: the device may have DMA in flight, its driver may have cached
the value, and the core's read-modify-write of Device Control isn't
serialized against the driver's own updates of that register.
Changes to a hierarchy are therefore limited to hierarchies in which
no device below the Root Port has been added or made available for
driver binding yet, i.e. during the initial scan, or when devices are
hot-added or rescanned into an empty hierarchy, such as a slot directly
below a Root Port. As today, the Root Port itself may still be
changed.
- A device hot-added below a Switch later can't lower the MPS of a
hierarchy in use, so it only works if it supports the MPS the
hierarchy already runs at (6).
A better solution would be to quiesce the drivers of all devices below
the Root Port, e.g. the way a reset does with ->reset_prepare() and
->reset_done(), change the MPS of the whole hierarchy, and then resume
them. That would allow changing the MPS of a system that is in use,
including when a device that supports less is hot-added below a Switch
(6, 7) or next to devices in use (8), and patch 3 could then also raise
hierarchies with hotplug-capable Switch ports. It requires every
affected driver to support being quiesced and to cope with a changed
MPS, e.g. drivers that cache it, which is beyond this series.
Only PCIE_BUS_DEFAULT changes; the other strategies are unchanged.
Why the Root Port MPS is set from pcie_bus_configure_settings()
===============================================================
v9 raised each Root Port to its maximum MPS when pci_configure_mps() ran
for the Root Port itself, and relied on lowering it again as devices
below it were enumerated. At that point, nothing below the Root Port
has been enumerated yet, so the kernel can't know the smallest MPSS
below it, whether the hierarchy has hotplug-capable Switch ports, or
whether the slot is empty. Raising the Root Port there:
- raises hierarchies with hotplug-capable Switch ports, which can't be
lowered again once a device that supports less is hot-added below the
Switch (6). Undoing the raise for them would mean remembering the
previous MPS of every Root Port;
- raises empty slots, so a Switch hot-added into one later inherits the
raised MPS;
- happens only once, since a Root Port isn't enumerated again on
hot-add, so the Root Port can't be raised again after a device that
supported less has been replaced (4).
pcie_bus_configure_settings() runs after the hierarchy below a Root Port
has been scanned and before its drivers are bound: from pci_host_probe()
and the ACPI and arch-specific root bus scans, and from pciehp, acpiphp
and shpchp after a hot-add. It is where the other strategies already
configure MPS for a whole hierarchy, and pcie_find_smpss() already
computes the smallest MPSS of a hierarchy, returning 128 bytes for
hierarchies with a hotplug bridge other than the Root Port. Since patch
3 only raises, such hierarchies keep their MPS without any extra state.
An empty slot isn't raised until something is hot-added or rescanned
there, and every hot-add directly below a Root Port re-evaluates its
hierarchy. The rescan helpers don't call pcie_bus_configure_settings(),
so patch 3 adds the same raise there.
Configuring the MPS in pci_bus_add_device() instead would also run after
the scan, but once per device rather than once per hierarchy, and it
would move all of the default MPS configuration into the driver binding
path.
Behavior changes
================
- With firmware that programs MPS, e.g. on x86, a hierarchy may now run
at a larger MPS than firmware chose, and different Root Port
hierarchies may end up with different values. This could matter for
peer-to-peer DMA between devices below different Root Ports, unless
the Root Complex splits such TLPs. pci=pcie_bus_peer2peer remains
available for that case.
- A single device with a small MPSS found during enumeration lowers the
MPS of every device below its Root Port.
- A device that shows up later below a Switch without hotplug support,
e.g. found by a rescan, may find its hierarchy raised by patch 3, and
is left mismatched with a warning if it supports less (6).
- Read completion coalescing is now also disabled on Intel 5000/5100 in
the default mode (quirk_intel_mc_errata()), since the default mode can
now raise MPS above what firmware programmed.
- Scan paths that call neither pcie_bus_configure_settings() nor the
rescan helpers, mostly legacy ones or ones without Root Ports, don't
raise anything, just as today.
Open questions
==============
1. Is raising a hierarchy above what firmware programmed acceptable for
PCIE_BUS_DEFAULT, or should the raise be limited, e.g. to Root Ports
that are still at the 128-byte reset value?
2. Hierarchies with hotplug-capable Switch ports keep their current MPS,
so hot-add below a Switch behaves as today (6, 7). The alternative is
to limit such hierarchies to 128 bytes, like pcie_bus_safe does, which
makes every hot-added device work, but costs performance where
firmware programmed more. Which is preferred?
3. Should a device that can't match a hierarchy in use (6, 7, 8) keep
being bound with a warning, as today, or be left without a driver?
Testing
=======
Build-tested only: each patch builds without warnings with W=1 (x86_64
defconfig plus COMPILE_TEST and PCI_MESON), and checkpatch --strict is
clean.
Mahesh, Ricardo, Shawn: patch 3 is a new implementation, so I dropped
your Tested-by tags. A re-test would be much appreciated.
---
Changes compared to v9:
- New patch 1: Update the saved MPS in pcie_set_mps(), so that a Root Port
reset through host->reset_root_port() doesn't restore a stale MPS.
- Patch 2 (was 1): Only lower the hierarchy while no device below the
Root Port has been added or made available for driver binding. v9 also
did it when a device was hot-added or rescanned below a Switch, changing
the MPS of devices that may have drivers bound and DMA in flight, with
an unserialized read-modify-write of Device Control that could race
with their drivers' own updates (reported by sashiko-bot:
https://lore.kernel.org/linux-pci/20260916155239.06ECC1F000FF@smtp.kernel.org/).
- Patch 2: Use pci_warn() with the same message as the rest of
pci_configure_mps(), and reword the comment and commit message.
- Patch 3 (was 2): Reimplemented. Instead of raising each Root Port to
its maximum MPS during enumeration, raise the whole hierarchy after it
has been scanned, from pcie_bus_configure_settings() and the rescan
helpers, and only with PCIE_BUS_DEFAULT. Hierarchies with a hotplug
bridge other than the Root Port, and hierarchies in use, are left alone,
so hot-add below a Switch keeps working as before. v9 also changed the
other strategies and could leave a rescanned Root Port above the MPS of
the devices below it with pci=pcie_bus_safe.
- Patch 3: Apply the Intel 5000/5100 read completion coalescing quirk to
PCIE_BUS_DEFAULT too, since it can now raise MPS above what firmware
programmed.
- Patch 3: New subject and author, with Hans as co-developer, and dropped
the Reviewed-by and Tested-by tags, since the implementation changed.
- Patch 4 (was 3): Use the "PCI: meson:" subject prefix.
- Rebase on v7.3-rc5.
v9: https://lore.kernel.org/linux-pci/20260916153907.60344-1-18255117159@163.com/
Hans Zhang (2):
PCI: Match the hierarchy's MPS to a device's MPSS as necessary
PCI: meson: Remove redundant MPS configuration
Niklas Cassel (2):
PCI: Update saved Max Payload Size in pcie_set_mps()
PCI: Configure Root Port MPS after scanning its hierarchy
drivers/pci/controller/dwc/pci-meson.c | 25 +----
drivers/pci/pci.c | 24 +++-
drivers/pci/probe.c | 149 ++++++++++++++++++++++++-
drivers/pci/quirks.c | 3 +-
4 files changed, 172 insertions(+), 29 deletions(-)
base-commit: 72d3fcf802c45d00b300f25b848a93c3a2bd7c7e
--
2.55.0
^ permalink raw reply [flat|nested] 5+ messages in thread* [RFC PATCH 1/4] PCI: Update saved Max Payload Size in pcie_set_mps()
2026-09-30 14:50 [RFC PATCH 0/4] PCI: Fix and improve default MPS configuration Niklas Cassel
@ 2026-09-30 14:50 ` Niklas Cassel
2026-09-30 14:50 ` [RFC PATCH 2/4] PCI: Match the hierarchy's MPS to a device's MPSS as necessary Niklas Cassel
` (2 subsequent siblings)
3 siblings, 0 replies; 5+ messages in thread
From: Niklas Cassel @ 2026-09-30 14:50 UTC (permalink / raw)
To: Bjorn Helgaas, Bjorn Helgaas, Lorenzo Pieralisi,
Krzysztof Wilczyński, Manivannan Sadhasivam, Heiko Stuebner,
Yue Wang, Hans Zhang, Manivannan Sadhasivam, Frank Li
Cc: Keith Busch, mx2pg, dlemoal, Niklas Cassel, Lukas Wunner,
Mahesh Vaidya, Ricardo Pardini, Shawn Lin, Pali Rohár,
Neil Armstrong, Rob Herring, Jingoo Han, Kevin Hilman,
Jerome Brunet, Martin Blumenstingl, Myron Stowe, Jon Mason,
linux-pci, linux-kernel, linux-arm-kernel, linux-amlogic,
linux-rockchip
pcie_set_mps() changes the Max Payload Size (MPS) in the Device Control
register but leaves the copy saved by pci_save_state() alone, so the next
pci_restore_state() puts back the old value.
The PCI core does change the MPS of devices whose state has already been
saved: pci_configure_mps() reduces a Root Port's MPS when a device with a
smaller MPS Supported is hot-added directly below it (commit 9f0e89359775
("PCI: Match Root Port's MPS to endpoint's MPSS as necessary")), and with
"pci=pcie_bus_safe", pcie_bus_configure_settings() can do the same when
pciehp calls it after a hot-add. The Root Port's state was saved when it
was added and again when portdrv probed it.
Since commit 3fc686d550f6 ("PCI/ERR: Add support for resetting the Root
Ports in a platform-specific way"), pcibios_reset_secondary_bus() restores
that saved state, without saving it first, after host->reset_root_port()
has reset the Root Port, e.g. during AER recovery or on Link Down with the
qcom and Rockchip DWC drivers. The Root Port then goes back to the old,
larger MPS while the device below it is restored to the smaller one, so
the Root Port may send TLPs that the device treats as Malformed.
Update the saved copy of Device Control when pcie_set_mps() changes the
MPS, like commit 909f7bf9b080 ("PCI: Update saved_config_space upon
resource assignment") does for BARs.
Fixes: 3fc686d550f6 ("PCI/ERR: Add support for resetting the Root Ports in a platform-specific way")
Assisted-by: LLM
Signed-off-by: Niklas Cassel <cassel@kernel.org>
---
drivers/pci/pci.c | 24 +++++++++++++++++++++++-
1 file changed, 23 insertions(+), 1 deletion(-)
diff --git a/drivers/pci/pci.c b/drivers/pci/pci.c
index b2879a6be5f8..5b126cd342ca 100644
--- a/drivers/pci/pci.c
+++ b/drivers/pci/pci.c
@@ -1736,6 +1736,24 @@ static void pci_restore_pcie_state(struct pci_dev *dev)
pcie_capability_write_word(dev, PCI_EXP_SLTCTL2, cap[i++]);
}
+/*
+ * Update the saved copy of the Device Control register after changing it, so
+ * that pci_restore_state() doesn't put back a stale value. Device Control is
+ * the first register saved by pci_save_pcie_state().
+ */
+static void pcie_update_saved_devctl(struct pci_dev *dev, u16 clear, u16 set)
+{
+ struct pci_cap_saved_state *save_state;
+ u16 *devctl;
+
+ save_state = pci_find_saved_cap(dev, PCI_CAP_ID_EXP);
+ if (!save_state)
+ return;
+
+ devctl = (u16 *)&save_state->cap.data[0];
+ *devctl = (*devctl & ~clear) | set;
+}
+
static int pci_save_pcix_state(struct pci_dev *dev)
{
int pos;
@@ -5970,8 +5988,12 @@ int pcie_set_mps(struct pci_dev *dev, int mps)
ret = pcie_capability_clear_and_set_word(dev, PCI_EXP_DEVCTL,
PCI_EXP_DEVCTL_PAYLOAD, v);
+ if (ret)
+ return pcibios_err_to_errno(ret);
- return pcibios_err_to_errno(ret);
+ pcie_update_saved_devctl(dev, PCI_EXP_DEVCTL_PAYLOAD, v);
+
+ return 0;
}
EXPORT_SYMBOL(pcie_set_mps);
--
2.55.0
^ permalink raw reply [flat|nested] 5+ messages in thread* [RFC PATCH 2/4] PCI: Match the hierarchy's MPS to a device's MPSS as necessary
2026-09-30 14:50 [RFC PATCH 0/4] PCI: Fix and improve default MPS configuration Niklas Cassel
2026-09-30 14:50 ` [RFC PATCH 1/4] PCI: Update saved Max Payload Size in pcie_set_mps() Niklas Cassel
@ 2026-09-30 14:50 ` Niklas Cassel
2026-09-30 14:50 ` [RFC PATCH 3/4] PCI: Configure Root Port MPS after scanning its hierarchy Niklas Cassel
2026-09-30 14:50 ` [RFC PATCH 4/4] PCI: meson: Remove redundant MPS configuration Niklas Cassel
3 siblings, 0 replies; 5+ messages in thread
From: Niklas Cassel @ 2026-09-30 14:50 UTC (permalink / raw)
To: Bjorn Helgaas, Bjorn Helgaas, Lorenzo Pieralisi,
Krzysztof Wilczyński, Manivannan Sadhasivam, Heiko Stuebner,
Yue Wang, Hans Zhang, Myron Stowe, Jon Mason
Cc: Keith Busch, mx2pg, dlemoal, Lukas Wunner, Frank Li,
Mahesh Vaidya, Ricardo Pardini, Shawn Lin, Pali Rohár,
Neil Armstrong, Rob Herring, Jingoo Han, Kevin Hilman,
Jerome Brunet, Martin Blumenstingl, linux-pci, linux-kernel,
linux-arm-kernel, linux-amlogic, linux-rockchip, Niklas Cassel
From: Hans Zhang <18255117159@163.com>
pci_configure_mps() enumerates top-down and programs each device's Maximum
Payload Size (MPS) to match its upstream bridge. When a device's MPS
Supported (MPSS) is too small to match, commit 9f0e89359775 ("PCI: Match
Root Port's MPS to endpoint's MPSS as necessary") reduces the upstream
bridge instead, but only when that bridge is a Root Port.
That covers an endpoint directly below a Root Port and nothing else. With
a Switch in between, the Switch ports have already inherited the Root
Port's larger MPS, the reduction is skipped because the upstream bridge is
a Switch Downstream Port, and pcie_set_mps() then fails with -EINVAL for
the endpoint. The endpoint is left below a port programmed for a larger
MPS, so any larger TLP it receives is treated as Malformed.
Multi-function devices hit the same hole from the other direction:
reducing the Root Port for a function with a small MPSS leaves the sibling
functions already programmed to the larger value.
Walk the hierarchy from the Root Port down and reduce every device that is
above the new value. Reducing only the ports between the device and the
Root Port is not sufficient, because a Switch does not split TLPs: an
already programmed sibling left at the larger MPS could emit a TLP too
large for its egress port. As a result, a single device with a small MPSS
now lowers the MPS of every device below its Root Port.
Only do this while no device below the Root Port has been added or made
available for driver binding, i.e., during the initial scan, or when
devices are hot-added or rescanned into an empty hierarchy such as a slot
directly below the Root Port. Such devices may have drivers bound and DMA
in flight, so their MPS can't be changed safely. This is the same
constraint that makes PCIE_BUS_SAFE limit fabrics with hotplug bridges
below a Root Port to 128 bytes in pcie_find_smpss(). A device
added next to devices that are already in use is left at its current MPS
and pci_configure_mps() warns and suggests "pci=pcie_bus_safe", which is
what already happens below a Switch today.
This only affects PCIE_BUS_DEFAULT. PCIE_BUS_SAFE, PCIE_BUS_PERFORMANCE
and PCIE_BUS_PEER2PEER program MPS in pcie_bus_configure_settings() and
PCIE_BUS_TUNE_OFF doesn't touch it, so all of them return before this
point.
Fixes: 9f0e89359775 ("PCI: Match Root Port's MPS to endpoint's MPSS as necessary")
Signed-off-by: Hans Zhang <18255117159@163.com>
Assisted-by: LLM
Co-developed-by: Niklas Cassel <cassel@kernel.org>
Signed-off-by: Niklas Cassel <cassel@kernel.org>
---
drivers/pci/probe.c | 66 ++++++++++++++++++++++++++++++++++++++++++---
1 file changed, 62 insertions(+), 4 deletions(-)
diff --git a/drivers/pci/probe.c b/drivers/pci/probe.c
index 27008e2ea5af..d8e58e5ef730 100644
--- a/drivers/pci/probe.c
+++ b/drivers/pci/probe.c
@@ -2200,9 +2200,49 @@ int pci_setup_device(struct pci_dev *dev)
return 0;
}
+static int pcie_reduce_mps(struct pci_dev *dev, void *data)
+{
+ int mps = *(int *)data;
+ int ret;
+
+ /* MPS is of type 'RsvdP' for VFs */
+ if (!pci_is_pcie(dev) || dev->is_virtfn)
+ return 0;
+
+ if (pcie_get_mps(dev) > mps) {
+ ret = pcie_set_mps(dev, mps);
+ if (ret)
+ pci_warn(dev, "can't set Max Payload Size to %d; if necessary, use \"pci=pcie_bus_safe\" and report a bug\n",
+ mps);
+ }
+
+ return 0;
+}
+
+static int pci_dev_check_in_use(struct pci_dev *dev, void *data)
+{
+ bool *in_use = data;
+
+ *in_use = pci_dev_is_added(dev) || !pci_dev_binding_disallowed(dev);
+ return *in_use;
+}
+
+/*
+ * Return true if any device on or below @bus has been added or made available
+ * for driver binding, i.e., may have a driver bound and DMA in flight.
+ */
+static bool pci_bus_in_use(struct pci_bus *bus)
+{
+ bool in_use = false;
+
+ pci_walk_bus(bus, pci_dev_check_in_use, &in_use);
+ return in_use;
+}
+
static void pci_configure_mps(struct pci_dev *dev)
{
struct pci_dev *bridge = pci_upstream_bridge(dev);
+ struct pci_dev *rp;
int mps, mpss, p_mps, rc;
if (!pci_is_pcie(dev))
@@ -2252,10 +2292,28 @@ static void pci_configure_mps(struct pci_dev *dev)
return;
mpss = 128 << dev->pcie_mpss;
- if (mpss < p_mps && pci_pcie_type(bridge) == PCI_EXP_TYPE_ROOT_PORT) {
- pcie_set_mps(bridge, mpss);
- pci_info(dev, "Upstream bridge's Max Payload Size set to %d (was %d, max %d)\n",
- mpss, p_mps, 128 << bridge->pcie_mpss);
+ rp = pcie_find_root_port(bridge);
+ if (mpss < p_mps && rp && !pci_bus_in_use(rp->subordinate)) {
+ /*
+ * dev cannot be programmed to the MPS already in use above
+ * it, so reduce the hierarchy to what dev supports. A Switch
+ * does not split TLPs, so reducing only the upstream bridge
+ * is not enough: every port up to the Root Port has to come
+ * down as well, and so do the devices already programmed
+ * below that Root Port, which would otherwise be left sending
+ * TLPs too large for their egress port.
+ *
+ * Only do this while no device below the Root Port has been
+ * added or made available for driver binding, e.g., during
+ * the initial scan or when hot-adding into a slot directly
+ * below the Root Port. Such devices may have drivers bound
+ * and DMA in flight, so their MPS can't be changed safely
+ * (see pcie_find_smpss()).
+ */
+ pcie_reduce_mps(rp, &mpss);
+ pci_walk_bus(rp->subordinate, pcie_reduce_mps, &mpss);
+ pci_info(dev, "Max Payload Size of %s hierarchy set to %d (was %d)\n",
+ pci_name(rp), mpss, p_mps);
p_mps = pcie_get_mps(bridge);
}
--
2.55.0
^ permalink raw reply [flat|nested] 5+ messages in thread* [RFC PATCH 3/4] PCI: Configure Root Port MPS after scanning its hierarchy
2026-09-30 14:50 [RFC PATCH 0/4] PCI: Fix and improve default MPS configuration Niklas Cassel
2026-09-30 14:50 ` [RFC PATCH 1/4] PCI: Update saved Max Payload Size in pcie_set_mps() Niklas Cassel
2026-09-30 14:50 ` [RFC PATCH 2/4] PCI: Match the hierarchy's MPS to a device's MPSS as necessary Niklas Cassel
@ 2026-09-30 14:50 ` Niklas Cassel
2026-09-30 14:50 ` [RFC PATCH 4/4] PCI: meson: Remove redundant MPS configuration Niklas Cassel
3 siblings, 0 replies; 5+ messages in thread
From: Niklas Cassel @ 2026-09-30 14:50 UTC (permalink / raw)
To: Bjorn Helgaas, Bjorn Helgaas, Lorenzo Pieralisi,
Krzysztof Wilczyński, Manivannan Sadhasivam, Heiko Stuebner,
Yue Wang, Hans Zhang
Cc: Keith Busch, mx2pg, dlemoal, Niklas Cassel, Lukas Wunner,
Frank Li, Mahesh Vaidya, Ricardo Pardini, Shawn Lin,
Pali Rohár, Neil Armstrong, Rob Herring, Jingoo Han,
Kevin Hilman, Jerome Brunet, Martin Blumenstingl, Myron Stowe,
Jon Mason, linux-pci, linux-kernel, linux-arm-kernel,
linux-amlogic, linux-rockchip
With the default MPS strategy (PCIE_BUS_DEFAULT), pci_configure_mps()
only matches each device's Maximum Payload Size (MPS) to its upstream
bridge. Root Ports have no upstream bridge, so their MPS stays at
whatever firmware or the hardware default (128 bytes) left there, and the
hierarchy below them inherits that value even if the Root Port and all
devices below it support more.
Once the hierarchy below a Root Port has been scanned, and before drivers
are bound, raise the Root Port and every device below it to the largest
MPS they all support. Do this from pcie_bus_configure_settings(), which
the host bridge, ACPI and hotplug paths call after scanning, and from
pci_rescan_bus() and pci_rescan_bus_bridge_resize(), which don't.
Leave a hierarchy alone if it contains a hotplug bridge other than the
Root Port, typically a Switch Downstream Port: a device hot-added below it
later can't lower the MPS of a hierarchy in use, so keep the MPS the
hierarchy already had. pcie_find_smpss() already returns the minimum MPS
for such hierarchies, which never raises anything. Don't raise an empty
slot directly below a Root Port either; it is evaluated when a device is
hot-added or rescanned there, which also allows raising the Root Port
again after a device that supported less has been replaced. Hierarchies
with devices that may already have a driver bound are never changed.
Since PCIE_BUS_DEFAULT can now raise MPS above what firmware programmed,
apply the Intel 5000/5100 read completion coalescing quirk to it as well.
The other strategies are unchanged: PCIE_BUS_TUNE_OFF doesn't touch MPS,
and PCIE_BUS_SAFE, PCIE_BUS_PERFORMANCE and PCIE_BUS_PEER2PEER already
configure the hierarchy in pcie_bus_configure_settings().
Suggested-by: Manivannan Sadhasivam <mani@kernel.org>
Co-developed-by: Hans Zhang <18255117159@163.com>
Signed-off-by: Hans Zhang <18255117159@163.com>
Assisted-by: LLM
Signed-off-by: Niklas Cassel <cassel@kernel.org>
---
drivers/pci/probe.c | 83 ++++++++++++++++++++++++++++++++++++++++++++
drivers/pci/quirks.c | 3 +-
2 files changed, 84 insertions(+), 2 deletions(-)
diff --git a/drivers/pci/probe.c b/drivers/pci/probe.c
index d8e58e5ef730..5f37b480b51d 100644
--- a/drivers/pci/probe.c
+++ b/drivers/pci/probe.c
@@ -3083,6 +3083,82 @@ static int pcie_bus_configure_set(struct pci_dev *dev, void *data)
return 0;
}
+static int pcie_raise_mps(struct pci_dev *dev, void *data)
+{
+ int mps = *(int *)data;
+
+ /* MPS is of type 'RsvdP' for VFs */
+ if (!pci_is_pcie(dev) || dev->is_virtfn)
+ return 0;
+
+ if (pcie_get_mps(dev) < mps && pcie_set_mps(dev, mps))
+ pci_err(dev, "can't set Max Payload Size to %d\n", mps);
+
+ return 0;
+}
+
+/*
+ * With PCIE_BUS_DEFAULT, pci_configure_mps() only matches each device to its
+ * upstream bridge, so a hierarchy inherits whatever MPS firmware, or the
+ * hardware default of 128 bytes, left in its Root Port. Once the hierarchy
+ * below a Root Port has been enumerated, raise it to the largest MPS that all
+ * of its devices support.
+ *
+ * Leave hierarchies with a hotplug bridge other than the Root Port alone: a
+ * device hot-added below it can't lower the MPS once the hierarchy is in use.
+ * pcie_find_smpss() returns the minimum MPS for such hierarchies, which never
+ * raises anything. A hotplug slot directly below the Root Port is fine: it
+ * is evaluated when a device is hot-added there, and an empty slot is not
+ * raised, so a Switch hot-added into it doesn't inherit a raised MPS. Also
+ * leave hierarchies alone once any of their devices may have a driver bound.
+ */
+static void pcie_bus_raise_default_mps(struct pci_bus *bus)
+{
+ struct pci_dev *rp = bus->self;
+ u8 smpss = rp->pcie_mpss;
+ int mps, old_mps;
+
+ if (pci_pcie_type(rp) != PCI_EXP_TYPE_ROOT_PORT ||
+ list_empty(&bus->devices) || pci_bus_in_use(bus))
+ return;
+
+ pci_walk_bus(bus, pcie_find_smpss, &smpss);
+ mps = 128 << smpss;
+ old_mps = pcie_get_mps(rp);
+ if (mps <= old_mps)
+ return;
+
+ pcie_raise_mps(rp, &mps);
+ pci_walk_bus(bus, pcie_raise_mps, &mps);
+ pci_info(rp, "Max Payload Size of hierarchy set to %d (was %d)\n",
+ mps, old_mps);
+}
+
+/*
+ * A rescan doesn't go through pcie_bus_configure_settings(), so give the Root
+ * Port hierarchies it may have populated the same chance to be raised: every
+ * hierarchy below a rescanned root bus, or the one @bus belongs to.
+ */
+static void pcie_rescan_raise_default_mps(struct pci_bus *bus)
+{
+ struct pci_bus *child;
+ struct pci_dev *rp;
+
+ if (pcie_bus_config != PCIE_BUS_DEFAULT)
+ return;
+
+ if (pci_is_root_bus(bus)) {
+ list_for_each_entry(child, &bus->children, node)
+ if (child->self && pci_is_pcie(child->self))
+ pcie_bus_raise_default_mps(child);
+ return;
+ }
+
+ rp = pcie_find_root_port(bus->self);
+ if (rp && rp->subordinate)
+ pcie_bus_raise_default_mps(rp->subordinate);
+}
+
/*
* pcie_bus_configure_settings() requires that pci_walk_bus work in a top-down,
* parents then children fashion. If this changes, then this code will not
@@ -3098,6 +3174,11 @@ void pcie_bus_configure_settings(struct pci_bus *bus)
if (!pci_is_pcie(bus->self))
return;
+ if (pcie_bus_config == PCIE_BUS_DEFAULT) {
+ pcie_bus_raise_default_mps(bus);
+ return;
+ }
+
/*
* FIXME - Peer to peer DMA is possible, though the endpoint would need
* to be aware of the MPS of the destination. To work around this,
@@ -3536,6 +3617,7 @@ unsigned int pci_rescan_bus_bridge_resize(struct pci_dev *bridge)
max = pci_scan_child_bus(bus);
pci_assign_unassigned_bridge_resources(bridge);
+ pcie_rescan_raise_default_mps(bus);
pci_bus_add_devices(bus);
@@ -3557,6 +3639,7 @@ unsigned int pci_rescan_bus(struct pci_bus *bus)
max = pci_scan_child_bus(bus);
pci_assign_unassigned_bus_resources(bus);
+ pcie_rescan_raise_default_mps(bus);
pci_bus_add_devices(bus);
return max;
diff --git a/drivers/pci/quirks.c b/drivers/pci/quirks.c
index de9bbccda21f..da171f4babe4 100644
--- a/drivers/pci/quirks.c
+++ b/drivers/pci/quirks.c
@@ -3452,8 +3452,7 @@ static void quirk_intel_mc_errata(struct pci_dev *dev)
int err;
u16 rcc;
- if (pcie_bus_config == PCIE_BUS_TUNE_OFF ||
- pcie_bus_config == PCIE_BUS_DEFAULT)
+ if (pcie_bus_config == PCIE_BUS_TUNE_OFF)
return;
/*
--
2.55.0
^ permalink raw reply [flat|nested] 5+ messages in thread* [RFC PATCH 4/4] PCI: meson: Remove redundant MPS configuration
2026-09-30 14:50 [RFC PATCH 0/4] PCI: Fix and improve default MPS configuration Niklas Cassel
` (2 preceding siblings ...)
2026-09-30 14:50 ` [RFC PATCH 3/4] PCI: Configure Root Port MPS after scanning its hierarchy Niklas Cassel
@ 2026-09-30 14:50 ` Niklas Cassel
3 siblings, 0 replies; 5+ messages in thread
From: Niklas Cassel @ 2026-09-30 14:50 UTC (permalink / raw)
To: Bjorn Helgaas, Bjorn Helgaas, Lorenzo Pieralisi,
Krzysztof Wilczyński, Manivannan Sadhasivam, Heiko Stuebner,
Yue Wang, Hans Zhang, Rob Herring, Neil Armstrong, Kevin Hilman,
Jerome Brunet, Martin Blumenstingl
Cc: Keith Busch, mx2pg, dlemoal, Lukas Wunner, Frank Li,
Mahesh Vaidya, Ricardo Pardini, Shawn Lin, Pali Rohár,
Jingoo Han, Myron Stowe, Jon Mason, linux-pci, linux-kernel,
linux-arm-kernel, linux-amlogic, linux-rockchip, Niklas Cassel
From: Hans Zhang <18255117159@163.com>
The Meson PCIe controller driver manually configures maximum payload
size (MPS) through meson_set_max_payload, duplicating functionality now
centralized in the PCI core. Deprecating redundant code simplifies the
driver and aligns it with the consolidated MPS management strategy,
improving long-term maintainability.
With meson_set_max_payload() gone, meson_size_to_payload() is only used
for MRRS. Rename it to meson_size_to_mrrs(), update the warning message,
and drop the now-unused PCIE_CAP_MAX_PAYLOAD_SIZE and MAX_PAYLOAD_SIZE
macros.
Reviewed-by: Manivannan Sadhasivam <mani@kernel.org>
Signed-off-by: Hans Zhang <18255117159@163.com>
Signed-off-by: Niklas Cassel <cassel@kernel.org>
---
drivers/pci/controller/dwc/pci-meson.c | 25 +++----------------------
1 file changed, 3 insertions(+), 22 deletions(-)
diff --git a/drivers/pci/controller/dwc/pci-meson.c b/drivers/pci/controller/dwc/pci-meson.c
index 8559d132dcde..1d3687566dda 100644
--- a/drivers/pci/controller/dwc/pci-meson.c
+++ b/drivers/pci/controller/dwc/pci-meson.c
@@ -21,7 +21,6 @@
#define to_meson_pcie(x) dev_get_drvdata((x)->dev)
-#define PCIE_CAP_MAX_PAYLOAD_SIZE(x) ((x) << 5)
#define PCIE_CAP_MAX_READ_REQ_SIZE(x) ((x) << 12)
/* PCIe specific config registers */
@@ -37,7 +36,6 @@
#define PM_CURRENT_STATE(x) (((x) >> 7) & 0x1)
#define PORT_CLK_RATE 100000000UL
-#define MAX_PAYLOAD_SIZE 256
#define MAX_READ_REQ_SIZE 256
#define PCIE_RESET_DELAY 500
#define PCIE_SHARED_RESET 1
@@ -256,7 +254,7 @@ static void meson_pcie_ltssm_enable(struct meson_pcie *mp)
meson_cfg_writel(mp, val, PCIE_CFG0);
}
-static int meson_size_to_payload(struct meson_pcie *mp, int size)
+static int meson_size_to_mrrs(struct meson_pcie *mp, int size)
{
struct device *dev = mp->pci.dev;
@@ -266,35 +264,19 @@ static int meson_size_to_payload(struct meson_pcie *mp, int size)
* than 2^12, just set to default size 2^(1+7).
*/
if (!is_power_of_2(size) || size < 128 || size > 4096) {
- dev_warn(dev, "payload size %d, set to default 256\n", size);
+ dev_warn(dev, "MRRS %d, set to default 256\n", size);
return 1;
}
return fls(size) - 8;
}
-static void meson_set_max_payload(struct meson_pcie *mp, int size)
-{
- struct dw_pcie *pci = &mp->pci;
- u32 val;
- u16 offset = dw_pcie_find_capability(pci, PCI_CAP_ID_EXP);
- int max_payload_size = meson_size_to_payload(mp, size);
-
- val = dw_pcie_readl_dbi(pci, offset + PCI_EXP_DEVCTL);
- val &= ~PCI_EXP_DEVCTL_PAYLOAD;
- dw_pcie_writel_dbi(pci, offset + PCI_EXP_DEVCTL, val);
-
- val = dw_pcie_readl_dbi(pci, offset + PCI_EXP_DEVCTL);
- val |= PCIE_CAP_MAX_PAYLOAD_SIZE(max_payload_size);
- dw_pcie_writel_dbi(pci, offset + PCI_EXP_DEVCTL, val);
-}
-
static void meson_set_max_rd_req_size(struct meson_pcie *mp, int size)
{
struct dw_pcie *pci = &mp->pci;
u32 val;
u16 offset = dw_pcie_find_capability(pci, PCI_CAP_ID_EXP);
- int max_rd_req_size = meson_size_to_payload(mp, size);
+ int max_rd_req_size = meson_size_to_mrrs(mp, size);
val = dw_pcie_readl_dbi(pci, offset + PCI_EXP_DEVCTL);
val &= ~PCI_EXP_DEVCTL_READRQ;
@@ -363,7 +345,6 @@ static int meson_pcie_host_init(struct dw_pcie_rp *pp)
pp->bridge->ops = &meson_pci_ops;
- meson_set_max_payload(mp, MAX_PAYLOAD_SIZE);
meson_set_max_rd_req_size(mp, MAX_READ_REQ_SIZE);
return 0;
--
2.55.0
^ permalink raw reply [flat|nested] 5+ messages in thread
end of thread, other threads:[~2026-09-30 14:51 UTC | newest]
Thread overview: 5+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-30 14:50 [RFC PATCH 0/4] PCI: Fix and improve default MPS configuration Niklas Cassel
2026-09-30 14:50 ` [RFC PATCH 1/4] PCI: Update saved Max Payload Size in pcie_set_mps() Niklas Cassel
2026-09-30 14:50 ` [RFC PATCH 2/4] PCI: Match the hierarchy's MPS to a device's MPSS as necessary Niklas Cassel
2026-09-30 14:50 ` [RFC PATCH 3/4] PCI: Configure Root Port MPS after scanning its hierarchy Niklas Cassel
2026-09-30 14:50 ` [RFC PATCH 4/4] PCI: meson: Remove redundant MPS configuration Niklas Cassel
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®