From: Abhin Parekadan Jose <abhinjoses@gmail.com>
To: Bjorn Helgaas <bhelgaas@google.com>,
Lukas Wunner <lukas@wunner.de>,
"Michael S. Tsirkin" <mst@redhat.com>
Cc: "Ilpo Järvinen" <ilpo.jarvinen@linux.intel.com>,
"Shuai Xue" <xueshuai@linux.alibaba.com>,
"Kees Cook" <kees@kernel.org>,
"Mahesh J Salgaonkar" <mahesh@linux.ibm.com>,
"Oliver O'Halloran" <oohall@gmail.com>,
linux-pci@vger.kernel.org, linuxppc-dev@lists.ozlabs.org,
linux-kernel@vger.kernel.org,
"Abhin Parekadan Jose" <abhinjoses@gmail.com>
Subject: [PATCH RFC v4 4/5] PCI: pciehp: Report surprise removal from pciehp_isr()
Date: Sun, 27 Sep 2026 18:20:15 +0000 [thread overview]
Message-ID: <20260927182017.938565-5-abhinjoses@gmail.com> (raw)
In-Reply-To: <20260927182017.938565-1-abhinjoses@gmail.com>
A surprise removal during a safe removal cannot be reported: the
removal blocks waiting on a device interrupt or status read, and the
single-threaded IRQ thread is itself executing that removal, so it
cannot report that the device is gone. The removal hangs.
The hardirq handler pciehp_isr() still runs while the IRQ thread is
blocked, so it can report the disconnect. However, pciehp_ist()
deliberately ignores link and presence changes caused by a Secondary
Bus Reset or Downstream Port Containment, where the device is only
temporarily inaccessible. Distinguishing those normally requires
waiting for the SBR or DPC to conclude, which takes seconds and is not
possible in hardirq context.
Schedule a work item from pciehp_isr() when a PDC or DLLSC event
arrives and neither Presence Detect State nor Data Link Layer Link
Active indicates that a card is present.
This provides us a pathway to wait/block/sleep as we will not be
in pciehp_isr().
In pciehp_disconnect_work(), wait for the DPC recovery or SBR to complete
before checking presence and not consume the link change flags, so that
pciehp_ist() can still see them and ignore the link change if it was
caused by a DPC or SBR. Even after the DPC or SBR has completed, if the
device is really gone then pciehp_disconnect_work() will see that and
report the surprise removal. Like pciehp_ist(), it holds a runtime PM
reference on the port and checks presence under reset_lock, since a
slot reset may make Presence Detect State and Link Active flap.
Link: https://lore.kernel.org/all/aHlZE18kPuHuDtTT@wunner.de/
Signed-off-by: Abhin Parekadan Jose <abhinjoses@gmail.com>
Assisted-by: LLM
---
Changes since RFC v1:
- Drop schedule_notification_work(). pci_dev_set_disconnected() schedules
the driver's disconnect_work again, as in patch 1, for all callers, and
pciehp_disconnect_work() calls it through pci_walk_bus(). (Michael)
- Return early from pciehp_disconnect_work() if pending_events is zero.
pciehp_ist() has then already taken the events and handles them itself.
(Sashiko)
- Treat a read error of the presence check as "not present", both when
scheduling the work in pciehp_isr() and when checking presence in
pciehp_disconnect_work(). (Sashiko)
- Check presence with pciehp_card_present_or_link_active(), as
pciehp_ist() does, so that a port with Presence Detect State hardwired
to zero is not mistaken for an empty slot.
- In pciehp_isr(), check presence before dropping the runtime PM
reference on the port's parent. In pciehp_disconnect_work(), take a
runtime PM reference and check presence under reset_lock.
- Don't call the spurious link change test from pciehp_disconnect_work().
It consumes the one-shot flags PCI_DPC_RECOVERED and PCI_LINK_CHANGED,
so pciehp_ist() could miss them and tear down a device that was only
reset. Await DPC recovery and Secondary Bus Reset with the new
pci_dpc_wait_recovery() and pci_hp_wait_link_change() (patches 2 and 3)
instead, without consuming the flags. pciehp_is_spurious_link_change()
is dropped and pciehp_ist() is unchanged. (Sashiko)
RFC v1: https://lore.kernel.org/all/20260905183905.997833-3-abhinjoses@gmail.com/
Michael's review: https://lore.kernel.org/all/20260912115420-mutt-send-email-mst@kernel.org/
Sashiko review: https://lore.kernel.org/all/20260905185217.E9BC21F00A3A@smtp.kernel.org/
---
drivers/pci/hotplug/pciehp.h | 1 +
drivers/pci/hotplug/pciehp_hpc.c | 67 ++++++++++++++++++++++++++++++++
2 files changed, 68 insertions(+)
diff --git a/drivers/pci/hotplug/pciehp.h b/drivers/pci/hotplug/pciehp.h
index debc79b0adfb2..c8ceb9320e2e9 100644
--- a/drivers/pci/hotplug/pciehp.h
+++ b/drivers/pci/hotplug/pciehp.h
@@ -116,6 +116,7 @@ struct controller {
unsigned int ist_running;
int request_result;
wait_queue_head_t requester;
+ struct work_struct disconnect_work;
};
/**
diff --git a/drivers/pci/hotplug/pciehp_hpc.c b/drivers/pci/hotplug/pciehp_hpc.c
index 4c62140a3cb44..470cc16d4a828 100644
--- a/drivers/pci/hotplug/pciehp_hpc.c
+++ b/drivers/pci/hotplug/pciehp_hpc.c
@@ -620,12 +620,64 @@ static void pciehp_ignore_link_change(struct controller *ctrl,
up_read(&ctrl->reset_lock);
}
+/*
+ * Workaround to not wait in the isr.
+ */
+static void pciehp_disconnect_work(struct work_struct *work)
+{
+ struct pci_bus *bus;
+ struct controller *ctrl = container_of(work, struct controller,
+ disconnect_work);
+ struct pci_dev *pdev = ctrl_dev(ctrl);
+ u32 events;
+ int present;
+
+ events = atomic_read(&ctrl->pending_events);
+
+ /*
+ * Zero means pciehp_ist() has already consumed the events with
+ * atomic_xchg() and is handling them itself, with the full event
+ * mask available for the spurious link change test. Leave it to
+ * the IRQ thread: acting here would override its decision, and a
+ * zero carries no information about the device.
+ *
+ * The events only stay pending for us when the IRQ thread is
+ * blocked and cannot drain them, e.g. in pciehp_unconfigure_device()
+ * during a safe removal. That is the case this work item exists
+ * for.
+ */
+ if (!events)
+ return;
+
+ pci_config_pm_runtime_get(pdev);
+
+ /* Wait for DPC recovery or a SBR to complete */
+ pci_dpc_wait_recovery(pdev);
+ pci_hp_wait_link_change(pdev);
+
+ /* Presence Detect State and Link Active may flap during a slot reset */
+ down_read_nested(&ctrl->reset_lock, ctrl->depth);
+ present = pciehp_card_present_or_link_active(ctrl);
+ up_read(&ctrl->reset_lock);
+
+ pci_config_pm_runtime_put(pdev);
+
+ bus = ctrl->pcie->port->subordinate;
+
+ /* The card may have returned */
+ if (!bus || present > 0)
+ return;
+
+ pci_walk_bus(bus, pci_dev_set_disconnected, NULL);
+}
+
static irqreturn_t pciehp_isr(int irq, void *dev_id)
{
struct controller *ctrl = (struct controller *)dev_id;
struct pci_dev *pdev = ctrl_dev(ctrl);
struct device *parent = pdev->dev.parent;
u16 status, events = 0;
+ int present = 1;
/*
* Interrupts only occur in D3hot or shallower and only if enabled
@@ -697,6 +749,14 @@ static irqreturn_t pciehp_isr(int irq, void *dev_id)
}
ctrl_dbg(ctrl, "pending interrupts %#06x from Slot Status\n", events);
+
+ /*
+ * Check presence while the port is kept accessible. This may race
+ * with a slot reset, so pciehp_disconnect_work() checks again.
+ */
+ if (events & (PCI_EXP_SLTSTA_PDC | PCI_EXP_SLTSTA_DLLSC))
+ present = pciehp_card_present_or_link_active(ctrl);
+
if (parent)
pm_runtime_put(parent);
@@ -722,6 +782,11 @@ static irqreturn_t pciehp_isr(int irq, void *dev_id)
/* Save pending events for consumption by IRQ thread. */
atomic_or(events, &ctrl->pending_events);
+
+ /* The card may be gone, let process context decide */
+ if (present <= 0)
+ schedule_work(&ctrl->disconnect_work);
+
return IRQ_WAKE_THREAD;
}
@@ -1036,6 +1101,7 @@ struct controller *pcie_init(struct pcie_device *dev)
init_waitqueue_head(&ctrl->requester);
init_waitqueue_head(&ctrl->queue);
INIT_DELAYED_WORK(&ctrl->button_work, pciehp_queue_pushbutton_work);
+ INIT_WORK(&ctrl->disconnect_work, pciehp_disconnect_work);
dbg_ctrl(ctrl);
down_read(&pci_bus_sem);
@@ -1096,6 +1162,7 @@ struct controller *pcie_init(struct pcie_device *dev)
void pciehp_release_ctrl(struct controller *ctrl)
{
cancel_delayed_work_sync(&ctrl->button_work);
+ cancel_work_sync(&ctrl->disconnect_work);
kfree(ctrl);
}
--
2.51.1
next prev parent reply other threads:[~2026-09-27 18:20 UTC|newest]
Thread overview: 6+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-27 18:20 [PATCH RFC v4 0/5] PCI: pciehp: Report surprise removal during safe removal Abhin Parekadan Jose
2026-09-27 18:20 ` [PATCH RFC v4 1/5] PCI: Report surprise removal event Abhin Parekadan Jose
2026-09-27 18:20 ` [PATCH RFC v4 2/5] PCI: pciehp: Add pci_hp_wait_link_change() Abhin Parekadan Jose
2026-09-27 18:20 ` [PATCH RFC v4 3/5] PCI/DPC: Add pci_dpc_wait_recovery() Abhin Parekadan Jose
2026-09-27 18:20 ` Abhin Parekadan Jose [this message]
2026-09-27 18:20 ` [PATCH RFC v4 5/5] misc: Add edu_srpoc surprise removal POC driver Abhin Parekadan Jose
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260927182017.938565-5-abhinjoses@gmail.com \
--to=abhinjoses@gmail.com \
--cc=bhelgaas@google.com \
--cc=ilpo.jarvinen@linux.intel.com \
--cc=kees@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-pci@vger.kernel.org \
--cc=linuxppc-dev@lists.ozlabs.org \
--cc=lukas@wunner.de \
--cc=mahesh@linux.ibm.com \
--cc=mst@redhat.com \
--cc=oohall@gmail.com \
--cc=xueshuai@linux.alibaba.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®