From: Abhin Parekadan Jose <abhinjoses@gmail.com>
To: Bjorn Helgaas <bhelgaas@google.com>,
Lukas Wunner <lukas@wunner.de>,
"Michael S. Tsirkin" <mst@redhat.com>
Cc: "Ilpo Järvinen" <ilpo.jarvinen@linux.intel.com>,
"Shuai Xue" <xueshuai@linux.alibaba.com>,
"Kees Cook" <kees@kernel.org>,
"Mahesh J Salgaonkar" <mahesh@linux.ibm.com>,
"Oliver O'Halloran" <oohall@gmail.com>,
linux-pci@vger.kernel.org, linuxppc-dev@lists.ozlabs.org,
linux-kernel@vger.kernel.org,
"Abhin Parekadan Jose" <abhinjoses@gmail.com>
Subject: [PATCH RFC v3 1/5] PCI: Report surprise removal event
Date: Sun, 27 Sep 2026 17:51:58 +0000 [thread overview]
Message-ID: <20260927175203.928270-2-abhinjoses@gmail.com> (raw)
In-Reply-To: <20260927175203.928270-1-abhinjoses@gmail.com>
From: "Michael S. Tsirkin" <mst@redhat.com>
At the moment, in case of a surprise removal, the regular remove
callback is invoked, exclusively. This works well, because mostly,
the cleanup would be the same.
However, there's a race: imagine device removal was initiated by a user
action, such as driver unbind, and it in turn initiated some cleanup
and is now waiting for an interrupt from the device. If the device is
now surprise-removed, that never arrives and the remove callback hangs
forever.
For example, this was reported for virtio-blk:
1. the graceful removal is ongoing in the remove() callback,
where disk deletion del_gendisk() is ongoing, which waits
for the requests to complete,
2. Now few requests are yet to complete, and surprise removal
started.
At this point, virtio block driver will not get notified by
the driver core layer, because it is likely serializing
remove() happening by +user/driver unload and PCI hotplug
driver-initiated device removal. So vblk driver doesn't
know that device is removed, block layer is waiting for
requests completions to arrive which it never gets.
So del_gendisk() gets stuck.
Drivers can artificially add timeouts to handle that, but it can be
flaky.
Instead, let's add a way for the driver to be notified about the
disconnect. It can then do any necessary cleanup, knowing that
the device is inactive.
Since cleanups can take a long time, this takes an approach of a work
struct that the driver initiates and enables on probe, and tears down on
remove.
Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
Link: https://lore.kernel.org/all/fba3d235e38c1c6fcef2a30ed083ad9e25b20fa3.1752094439.git.mst@redhat.com/
[Abhin: adapted subject, serialize with a per-device spinlock]
Signed-off-by: Abhin Parekadan Jose <abhinjoses@gmail.com>
---
Changes since RFC v2:
- Protect disconnect_work_enable with a per-device spinlock, held while
testing it and scheduling the work, instead of lockless accesses and
barriers. pci_dev_set_disconnected() could otherwise test the flag,
get preempted, and queue the work while a newly bound driver
re-initializes it. With the lock, cancel_work_sync() is sufficient
again, so go back to it from disable_work_sync(). (Sashiko)
Changes since RFC v1:
- Use disable_work_sync() instead of cancel_work_sync() in
pci_clear_disconnect_work() as schedule_work() on a
disabled work item is a no-op. (Sashiko)
RFC v2: https://lore.kernel.org/all/20260927165459.829900-2-abhinjoses@gmail.com/
Sashiko review of v2: https://lore.kernel.org/all/20260927170707.6F9241F000FF@smtp.kernel.org/
RFC v1: https://lore.kernel.org/all/20260905183905.997833-2-abhinjoses@gmail.com/
Sashiko review of v1: https://lore.kernel.org/all/20260905184649.E8F621F00A3A@smtp.kernel.org/
---
drivers/pci/pci.h | 7 ++++++
drivers/pci/probe.c | 1 +
include/linux/pci.h | 57 +++++++++++++++++++++++++++++++++++++++++++++
3 files changed, 65 insertions(+)
diff --git a/drivers/pci/pci.h b/drivers/pci/pci.h
index ba3c3fddddc23..175273756c183 100644
--- a/drivers/pci/pci.h
+++ b/drivers/pci/pci.h
@@ -802,9 +802,16 @@ static inline bool pci_dev_set_io_state(struct pci_dev *dev,
static inline int pci_dev_set_disconnected(struct pci_dev *dev, void *unused)
{
+ unsigned long flags;
+
pci_dev_set_io_state(dev, pci_channel_io_perm_failure);
pci_doe_disconnected(dev);
+ spin_lock_irqsave(&dev->disconnect_lock, flags);
+ if (dev->disconnect_work_enable)
+ schedule_work(&dev->disconnect_work);
+ spin_unlock_irqrestore(&dev->disconnect_lock, flags);
+
return 0;
}
diff --git a/drivers/pci/probe.c b/drivers/pci/probe.c
index 27008e2ea5afc..3a05ad32bfd54 100644
--- a/drivers/pci/probe.c
+++ b/drivers/pci/probe.c
@@ -2515,6 +2515,7 @@ struct pci_dev *pci_alloc_dev(struct pci_bus *bus)
};
spin_lock_init(&dev->pcie_cap_lock);
+ spin_lock_init(&dev->disconnect_lock);
#ifdef CONFIG_PCI_MSI
raw_spin_lock_init(&dev->msi_lock);
#endif
diff --git a/include/linux/pci.h b/include/linux/pci.h
index d31a8d107b1ef..aed62ae709f59 100644
--- a/include/linux/pci.h
+++ b/include/linux/pci.h
@@ -592,6 +592,10 @@ struct pci_dev {
u8 reset_methods[PCI_NUM_RESET_METHODS]; /* In priority order */
struct gpio_desc *wake; /* WAKE# GPIO */
+ /* Report disconnect events. 0x0 - disable, 0x1 - enable */
+ u8 disconnect_work_enable;
+ spinlock_t disconnect_lock; /* Protects disconnect_work_enable */
+ struct work_struct disconnect_work;
#ifdef CONFIG_PCIE_TPH
u16 tph_cap; /* TPH capability offset */
@@ -2123,6 +2127,59 @@ pci_release_mem_regions(struct pci_dev *pdev)
pci_select_bars(pdev, IORESOURCE_MEM));
}
+/*
+ * Run this first thing after getting a disconnect work, to prevent it from
+ * running multiple times.
+ * Returns: true if disconnect was enabled, proceed. false if disabled, abort.
+ */
+static inline bool pci_test_and_clear_disconnect_enable(struct pci_dev *pdev)
+{
+ unsigned long flags;
+ bool enabled;
+
+ spin_lock_irqsave(&pdev->disconnect_lock, flags);
+ enabled = pdev->disconnect_work_enable;
+ pdev->disconnect_work_enable = 0x0;
+ spin_unlock_irqrestore(&pdev->disconnect_lock, flags);
+
+ return enabled;
+}
+
+/*
+ * Caller must initialize @pdev->disconnect_work before invoking this.
+ * The work function must run and check pci_test_and_clear_disconnect_enable.
+ * Note that device can go away right after this call.
+ */
+static inline void pci_set_disconnect_work(struct pci_dev *pdev)
+{
+ unsigned long flags;
+
+ spin_lock_irqsave(&pdev->disconnect_lock, flags);
+ pdev->disconnect_work_enable = 0x1;
+ spin_unlock_irqrestore(&pdev->disconnect_lock, flags);
+
+ /* check the device did not go away meanwhile. */
+ if (pci_device_is_present(pdev))
+ return;
+
+ spin_lock_irqsave(&pdev->disconnect_lock, flags);
+ if (pdev->disconnect_work_enable)
+ schedule_work(&pdev->disconnect_work);
+ spin_unlock_irqrestore(&pdev->disconnect_lock, flags);
+}
+
+static inline void pci_clear_disconnect_work(struct pci_dev *pdev)
+{
+ unsigned long flags;
+
+ spin_lock_irqsave(&pdev->disconnect_lock, flags);
+ pdev->disconnect_work_enable = 0x0;
+ spin_unlock_irqrestore(&pdev->disconnect_lock, flags);
+
+ /* No one can queue the work any more; wait for a queued or running one */
+ cancel_work_sync(&pdev->disconnect_work);
+}
+
bool pci_suspend_retains_context(struct pci_dev *pdev);
#else /* CONFIG_PCI is not enabled */
--
2.51.1
next prev parent reply other threads:[~2026-09-27 17:52 UTC|newest]
Thread overview: 6+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-27 17:51 [PATCH RFC v3 0/5] PCI: pciehp: Report surprise removal during safe removal Abhin Parekadan Jose
2026-09-27 17:51 ` Abhin Parekadan Jose [this message]
2026-09-27 17:51 ` [PATCH RFC v3 2/5] PCI: pciehp: Add pci_hp_wait_link_change() Abhin Parekadan Jose
2026-09-27 17:52 ` [PATCH RFC v3 3/5] PCI/DPC: Add pci_dpc_wait_recovery() Abhin Parekadan Jose
2026-09-27 17:52 ` [PATCH RFC v3 4/5] PCI: pciehp: Report surprise removal from pciehp_isr() Abhin Parekadan Jose
2026-09-27 17:52 ` [PATCH RFC v3 5/5] misc: Add edu_srpoc surprise removal POC driver Abhin Parekadan Jose
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260927175203.928270-2-abhinjoses@gmail.com \
--to=abhinjoses@gmail.com \
--cc=bhelgaas@google.com \
--cc=ilpo.jarvinen@linux.intel.com \
--cc=kees@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-pci@vger.kernel.org \
--cc=linuxppc-dev@lists.ozlabs.org \
--cc=lukas@wunner.de \
--cc=mahesh@linux.ibm.com \
--cc=mst@redhat.com \
--cc=oohall@gmail.com \
--cc=xueshuai@linux.alibaba.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®