* [PATCH v2 1/3] usb: hcd-pci: Honour pci_save_state() failure
2026-09-30 14:19 [PATCH v2 0/3] PCI/PM: Do not save the config space of an inaccessible device Francisco Beltrán Millalén
@ 2026-09-30 14:19 ` Francisco Beltrán Millalén
2026-09-30 14:19 ` [PATCH v2 2/3] PCI/PM: Do not save the config space of an inaccessible device Francisco Beltrán Millalén
2026-09-30 14:19 ` [PATCH v2 3/3] PCI: Do not mistake an absent device for an active link Francisco Beltrán Millalén
2 siblings, 0 replies; 4+ messages in thread
From: Francisco Beltrán Millalén @ 2026-09-30 14:19 UTC (permalink / raw)
To: Bjorn Helgaas, linux-pci
Cc: Alan Stern, Greg Kroah-Hartman, linux-usb, linux-kernel
hcd_pci_suspend_noirq() ignores the return value of pci_save_state() and
goes on to pci_prepare_to_sleep().
The next patch makes pci_save_state() fail, without marking the state as
saved, when the config space of the device is not accessible. For such
a controller pci_prepare_to_sleep() does not work either:
pci_set_low_power_state() reads the PM control register as all ones and
marks the device D3cold. pci_pm_suspend_noirq() then finds a device
whose power state changed without its state being saved, and warns.
With an earlier version of the next patch on a MacBookPro14,3, whose
Thunderbolt xHCI controllers can be inaccessible at this point:
xhci_hcd 0000:7d:00.0: PCI PM: State of device not saved by hcd_pci_suspend_noirq+0x0/0x180
WARNING: CPU: 7 PID: 96793 at drivers/pci/pci-driver.c:888 pci_pm_suspend_noirq+0x2f4/0x300
Check the return value and, if the state could not be saved, return 0
without putting the controller into a low-power state. The PCI core
then handles it as it handles a driver without a suspend_noirq callback:
it tries to save the state and to put the device into a low-power state
itself, which fails in the same way, but without the warning, as the
power state did not change inside the driver callback.
Move the check that disables wakeup for a dead root hub ahead of it, so
that it is not skipped.
Today pci_save_state() only fails if a capability save buffer is missing
or the VC state cannot be saved. In that case the controller is now left
in D0 instead of being put into a low-power state with an incomplete
saved state.
Assisted-by: LLM
Signed-off-by: Francisco Beltrán Millalén <fbeltranmillalen@gmail.com>
---
v2:
- Rewrote the commit message: the failure handled here comes from the
next patch, not from pci_save_state() as it is today.
- Disable wakeup for a dead root hub before the early return.
- Dropped Alan's Acked-by because of both changes.
drivers/usb/core/hcd-pci.c | 15 +++++++++++++--
1 file changed, 13 insertions(+), 2 deletions(-)
diff --git a/drivers/usb/core/hcd-pci.c b/drivers/usb/core/hcd-pci.c
index cd2234759..7f8a56425 100644
--- a/drivers/usb/core/hcd-pci.c
+++ b/drivers/usb/core/hcd-pci.c
@@ -541,8 +541,6 @@ static int hcd_pci_suspend_noirq(struct device *dev)
if (retval)
return retval;
- pci_save_state(pci_dev);
-
/* If the root hub is dead rather than suspended, disallow remote
* wakeup. usb_hc_died() should ensure that both hosts are marked as
* dying, so we only need to check the primary roothub.
@@ -551,6 +549,19 @@ static int hcd_pci_suspend_noirq(struct device *dev)
device_set_wakeup_enable(dev, 0);
dev_dbg(dev, "wakeup: %d\n", device_may_wakeup(dev));
+ /*
+ * If the state could not be saved, most likely because the controller
+ * is no longer accessible, putting it into a low-power state would
+ * fail as well and leave it marked as D3cold, and the PCI core would
+ * then warn that its state was not saved. Leave it to the PCI core
+ * instead, as for a driver without a suspend_noirq callback.
+ */
+ retval = pci_save_state(pci_dev);
+ if (retval) {
+ dev_dbg(dev, "--> not suspending, state not saved\n");
+ return 0;
+ }
+
/* Possibly enable remote wakeup,
* choose the appropriate low-power state, and go to that state.
*/
^ permalink raw reply [flat|nested] 4+ messages in thread* [PATCH v2 2/3] PCI/PM: Do not save the config space of an inaccessible device
2026-09-30 14:19 [PATCH v2 0/3] PCI/PM: Do not save the config space of an inaccessible device Francisco Beltrán Millalén
2026-09-30 14:19 ` [PATCH v2 1/3] usb: hcd-pci: Honour pci_save_state() failure Francisco Beltrán Millalén
@ 2026-09-30 14:19 ` Francisco Beltrán Millalén
2026-09-30 14:19 ` [PATCH v2 3/3] PCI: Do not mistake an absent device for an active link Francisco Beltrán Millalén
2 siblings, 0 replies; 4+ messages in thread
From: Francisco Beltrán Millalén @ 2026-09-30 14:19 UTC (permalink / raw)
To: Bjorn Helgaas, linux-pci
Cc: Alan Stern, Greg Kroah-Hartman, linux-usb, linux-kernel
pci_save_state() reads the standard header into dev->saved_config_space
and marks it valid without checking that the device answered. If the
device is not accessible, every read returns all ones, the previous
snapshot is overwritten, and pci_restore_state() writes the all-ones
values back once the device answers again. On a bridge that sets every
writable bit of the Bridge Control register, Secondary Bus Reset
included, and sets the primary, secondary and subordinate bus numbers
to 0xff, which cuts off everything below it.
On a MacBookPro14,3 this happens to the upstream bridge of a
Thunderbolt controller that drops off the bus while the system is
suspending: after resume the bridge answers again, but with bus numbers
ff/ff/ff and Secondary Bus Reset asserted, and the xHCI controllers
behind it are removed.
Commit e18d1abc3bff ("PCI: Avoid saving config space state if
inaccessible") added pci_dev_config_accessible() and checks it before a
reset. Check it in pci_save_state() itself, so that system suspend,
where the state is saved by pci_pm_suspend_noirq() or by a driver's
suspend_noirq callback, is covered as well. pci_dev_config_accessible()
reads the Command and Status registers rather than the Vendor and Device
IDs, which always read as all ones on SR-IOV VFs.
Check again after reading, as the device may become inaccessible in the
meantime, and only then replace the previous snapshot of the header and
set state_saved. The capabilities are saved straight into their own
buffers and are not covered by the second check, and both checks are
racy, as pci_dev_config_accessible() notes.
If the device is not accessible, return -EIO and leave state_saved as it
was.
Assisted-by: LLM
Signed-off-by: Francisco Beltrán Millalén <fbeltranmillalen@gmail.com>
---
v2:
- Use pci_dev_config_accessible() instead of reading the Vendor ID,
which always reads as all ones on SR-IOV VFs: v1 refused to save the
state of every VF. In the v1 thread I said I would use
pci_device_is_present(), but for a VF that checks the PF and cannot
tell whether the VF itself answers.
- Do the second check after saving the capabilities, and set
state_saved only if both checks pass.
- The bus numbers are set to 0xff, not cleared.
drivers/pci/pci.c | 53 +++++++++++++++++++++++++++++++++++++----------------
1 file changed, 37 insertions(+), 16 deletions(-)
diff --git a/drivers/pci/pci.c b/drivers/pci/pci.c
index b2879a6be..94de99d31 100644
--- a/drivers/pci/pci.c
+++ b/drivers/pci/pci.c
@@ -1776,31 +1776,52 @@ static void pci_restore_pcix_state(struct pci_dev *dev)
* pci_save_state - save the PCI configuration space of a device before
* suspending
* @dev: PCI device that we're dealing with
+ *
+ * If the config space of @dev is not accessible, nothing is saved and the
+ * previous snapshot of the standard header is kept, as writing back the
+ * all-ones values read from such a device would corrupt it once it is
+ * accessible again.
+ *
+ * Return: 0 on success, -EIO if @dev is not accessible, or another
+ * negative errno if a capability could not be saved.
*/
int pci_save_state(struct pci_dev *dev)
{
- int i;
+ u32 config[16];
+ int i, ret;
+
+ if (!pci_dev_config_accessible(dev, "save state"))
+ return -EIO;
+
/* XXX: 100% dword access ok here? */
for (i = 0; i < 16; i++) {
- pci_read_config_dword(dev, i * 4, &dev->saved_config_space[i]);
- pci_dbg(dev, "save config %#04x: %#010x\n",
- i * 4, dev->saved_config_space[i]);
+ pci_read_config_dword(dev, i * 4, &config[i]);
+ pci_dbg(dev, "save config %#04x: %#010x\n", i * 4, config[i]);
}
- dev->state_saved = true;
- i = pci_save_pcie_state(dev);
- if (i != 0)
- return i;
+ ret = pci_save_pcie_state(dev);
+ if (!ret)
+ ret = pci_save_pcix_state(dev);
+ if (!ret) {
+ pci_save_dpc_state(dev);
+ pci_save_aer_state(dev);
+ pci_save_ptm_state(dev);
+ pci_save_tph_state(dev);
+ ret = pci_save_vc_state(dev);
+ }
+
+ /*
+ * The device may have become inaccessible while it was being read.
+ * Keep the previous header snapshot in that case. The capabilities
+ * are saved directly into their buffers, so they are not protected.
+ */
+ if (!pci_dev_config_accessible(dev, "save state"))
+ return -EIO;
- i = pci_save_pcix_state(dev);
- if (i != 0)
- return i;
+ memcpy(dev->saved_config_space, config, sizeof(config));
+ dev->state_saved = true;
- pci_save_dpc_state(dev);
- pci_save_aer_state(dev);
- pci_save_ptm_state(dev);
- pci_save_tph_state(dev);
- return pci_save_vc_state(dev);
+ return ret;
}
EXPORT_SYMBOL(pci_save_state);
^ permalink raw reply [flat|nested] 4+ messages in thread* [PATCH v2 3/3] PCI: Do not mistake an absent device for an active link
2026-09-30 14:19 [PATCH v2 0/3] PCI/PM: Do not save the config space of an inaccessible device Francisco Beltrán Millalén
2026-09-30 14:19 ` [PATCH v2 1/3] usb: hcd-pci: Honour pci_save_state() failure Francisco Beltrán Millalén
2026-09-30 14:19 ` [PATCH v2 2/3] PCI/PM: Do not save the config space of an inaccessible device Francisco Beltrán Millalén
@ 2026-09-30 14:19 ` Francisco Beltrán Millalén
2 siblings, 0 replies; 4+ messages in thread
From: Francisco Beltrán Millalén @ 2026-09-30 14:19 UTC (permalink / raw)
To: Bjorn Helgaas, linux-pci
Cc: Alan Stern, Greg Kroah-Hartman, linux-usb, linux-kernel
When a bridge is no longer accessible, its Link Status register reads
as 0xffff, which has both PCI_EXP_LNKSTA_DLLLA and PCI_EXP_LNKSTA_LT
set. The code that waits for the link then takes a dead link for an
active one:
- pci_bridge_wait_for_secondary_bus(), for ports up to 5 GT/s, decides
that the link is up and waits PCIE_RESET_READY_POLL_MS for the device
below it;
- pcie_wait_for_link_status(), used for faster ports and by
pcie_retrain_link(), reports success at once when waiting for an
active link, and times out after PCIE_LINK_RETRAIN_TIMEOUT_MS when
waiting for LT to clear.
Treat an all-ones read as the device being gone. All callers of
pcie_wait_for_link_status() and pcie_retrain_link() treat any non-zero
return as failure.
On a MacBookPro14,3 whose Thunderbolt controller does not come back from
S3, the downstream ports of the controller are 2.5 GT/s ports that no
longer answer. Resume spent about 65 seconds waiting for the xHCI
controller behind one of them ("not ready 65535ms after resume"); with
the pci_bridge_wait_for_secondary_bus() change it gives up on that
controller after about one second. That machine never reaches
pcie_wait_for_link_status() in this situation, so that part has not been
tested.
Assisted-by: LLM
Signed-off-by: Francisco Beltrán Millalén <fbeltranmillalen@gmail.com>
---
v2:
- Also check in pcie_wait_for_link_status() (faster ports and
pcie_retrain_link()); not tested.
- Say that it is the bridge that is gone.
drivers/pci/pci.c | 16 +++++++++++-----
1 file changed, 11 insertions(+), 5 deletions(-)
diff --git a/drivers/pci/pci.c b/drivers/pci/pci.c
index 94de99d31..2e31d18b2 100644
--- a/drivers/pci/pci.c
+++ b/drivers/pci/pci.c
@@ -4586,8 +4586,8 @@ static int pci_pm_reset(struct pci_dev *dev, bool probe)
* @use_lt: Use the LT bit if TRUE, or the DLLLA bit if FALSE.
* @active: Waiting for active or inactive?
*
- * Return 0 if successful, or -ETIMEDOUT if status has not changed within
- * PCIE_LINK_RETRAIN_TIMEOUT_MS milliseconds.
+ * Return 0 if successful, -ENODEV if @pdev is not accessible, or -ETIMEDOUT
+ * if status has not changed within PCIE_LINK_RETRAIN_TIMEOUT_MS milliseconds.
*/
static int pcie_wait_for_link_status(struct pci_dev *pdev,
bool use_lt, bool active)
@@ -4602,6 +4602,9 @@ static int pcie_wait_for_link_status(struct pci_dev *pdev,
end_jiffies = jiffies + msecs_to_jiffies(PCIE_LINK_RETRAIN_TIMEOUT_MS);
do {
pcie_capability_read_word(pdev, PCI_EXP_LNKSTA, &lnksta);
+ /* All ones means @pdev is gone; it would look like an active link */
+ if (PCI_POSSIBLE_ERROR(lnksta))
+ return -ENODEV;
if ((lnksta & lnksta_mask) == lnksta_match)
return 0;
msleep(1);
@@ -4624,8 +4627,9 @@ static int pcie_wait_for_link_status(struct pci_dev *pdev,
* according to @use_lt. It is not verified whether the use of the DLLLA
* bit is valid.
*
- * Return 0 if successful, or -ETIMEDOUT if training has not completed
- * within PCIE_LINK_RETRAIN_TIMEOUT_MS milliseconds.
+ * Return 0 if successful, -ENODEV if @pdev is not accessible, or -ETIMEDOUT
+ * if training has not completed within PCIE_LINK_RETRAIN_TIMEOUT_MS
+ * milliseconds.
*/
int pcie_retrain_link(struct pci_dev *pdev, bool use_lt)
{
@@ -4858,8 +4862,10 @@ int pci_bridge_wait_for_secondary_bus(struct pci_dev *dev, char *reset_type)
if (!dev->link_active_reporting)
return -ENOTTY;
+ /* All ones means @dev is gone; it would look like an active link */
pcie_capability_read_word(dev, PCI_EXP_LNKSTA, &status);
- if (!(status & PCI_EXP_LNKSTA_DLLLA))
+ if (PCI_POSSIBLE_ERROR(status) ||
+ !(status & PCI_EXP_LNKSTA_DLLLA))
return -ENOTTY;
return pci_dev_wait(child, reset_type,
^ permalink raw reply [flat|nested] 4+ messages in thread