mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [RFC PATCH v2 00/16] vfio/pci: Handle PCI error recovery and report state to userspace
@ 2026-09-29 17:32 Shameer Kolothum
  2026-09-29 17:32 ` [RFC PATCH v2 01/16] vfio/pci: Add a device access gate Shameer Kolothum
                   ` (15 more replies)
  0 siblings, 16 replies; 17+ messages in thread
From: Shameer Kolothum @ 2026-09-29 17:32 UTC (permalink / raw)
  To: kvm, linux-pci, linux-kernel
  Cc: alex, jgg, kevin.tian, kbusch, michal.winiarski,
	satyanarayana.k.v.p, sonangp, ankita, nathanc, mochs,
	skolothumtho

Hi,

Currently, vfio-pci signals the error eventfd from error_detected() and
returns PCI_ERS_RESULT_CAN_RECOVER, including for permanent failure. It
has no slot_reset() or resume() callback, so userspace is not notified
when host recovery completes or whether the device was reset.

This series adds opt-in host PCI error recovery for vfio-pci. It blocks
device access during recovery, restores device state after a host reset,
and reports recovery status to userspace. The kernel runs host recovery
and the VMM decides how to recover the guest.

Changes from v1
---------------
 v1: https://lore.kernel.org/kvm/20260901093217.8539-1-skolothumtho@nvidia.com/
  
 - Thanks to Satya and Alex for reviews/feedback.
 - Use SRCU to drain ongoing device accesses and reject new accesses during
   recovery. A mutex protects recovery-state updates. 
 - Return SIGBUS for blocked BAR faults instead of waiting in the fault
   handler. VMM handling still needs target validation.
 - Reject userspace SR-IOV changes while access is blocked.
 - Exclude s390 and all CXL memory devices for now as s390 and CXL RCH
   recovery can omit completion callbacks.
  
Design
------
Userspace opts in through VFIO_DEVICE_FEATURE_PCI_ERROR_RECOVERY. It
registers an eventfd and reads status flags and a sequence number. The
existing error eventfd is still signaled.

The callbacks handle the host recovery sequence:

 - error_detected() blocks new accesses and drains admitted ones,
   quiesces INTx, and revokes BAR mappings and exported DMA-BUFs. On a
   normal channel it also clears bus mastering. It votes CAN_RECOVER
   for a normal channel, NEED_RESET for a frozen channel, and DISCONNECT
   for permanent failure. Failure to save PCI_COMMAND or disable bus
   mastering leaves access blocked and returns NONE so other devices
   can continue recovery.
 - slot_reset() restores the configuration saved at open after a host
   reset, then tears down interrupt configuration. Userspace must set
   up interrupts again.
 - resume() completes restoration, unblocks access and completes INTx
   handling. A restoration failure keeps the device blocked.

Non-fatal errors also block device access because host recovery may
require a reset. Userspace reads the recovery status to determine
whether the device was reset or recovery failed.
  
Without opt-in, the existing error notification behavior is retained.
The open check also rejects a previously recorded host transaction when
recovery is disabled, so closing the device does not discard that state.

Locking and open questions
--------------------------
Several locks are involved in recovery. PCI core holds device_lock
while calling the driver, preventing concurrent driver binding or
unbinding. During the bus walk, it also holds pci_bus_sem for reading
to keep the device list stable. Adding or removing PCI devices requires
the write lock. This semaphore is shared across all PCI buses.
 
VFIO recovery blocks new accesses, waits for existing SRCU readers to
finish, and takes memory_lock. Power transitions also take memory_lock,
but entering D0 can then acquire pci_bus_sem to update PCIe link power
management (ASPM).

Local Sashiko/Claude review reports a lock-order conflict as below:

 1. A guest D0 request holds memory_lock.
 2. PCI core holds pci_bus_sem for reading while VFIO's recovery
    callback waits for memory_lock.
 3. Removal of an unrelated PCI device waits for the bus write lock.
 4. D0's ASPM update requests the bus read lock and waits behind that
    writer.

Each task waits for another, so none can proceed.

One possible solution could be(not implemented):

 - PCI core guarantees that pci_bus_sem is held for reading before
   invoking supported recovery callbacks. VFIO uses
   pci_set_power_state_locked() in those callbacks, avoiding another
   acquisition of the bus lock.
 - A new pci_try_set_power_state() helper tries to acquire pci_bus_sem
   and returns -EAGAIN on contention. Ordinary VFIO D0 paths use this
   helper and handle failure instead of waiting while holding VFIO locks
	  
Not sure there are better ways to handle this or not. Feedbacks appreciated
on this.

Another issue flagged was(I think this is a pre-existing one):
 - Open/close coordination with the whole host recovery transaction
   was already missing. Recovery can start after the open check, and
   close can overlap the physical reset. Restoration of closed devices
   is also not implemented.

Testing
-------
Basic sanity tests performed on a GB200 with an NVIDIA GPU assigned.

Kernel branch:
 https://github.com/shamiali2008/linux/tree/vfio-aer-rfc-v2-ext

QEMU test branch is here(This registers the recovery eventfd and uses
pcie_aer_inject_error() to report the error to Guest)
 https://github.com/shamiali2008/qemu-master/commits/private-master-vfio-aer-test-v2/

Software AER injection was performed using a modified pcieaer_inject
module.

./aer-inject nonfatal.conf

qemu-system-aarch64: info: vfio 0018:06:00.0: host PCI error recovery completed (seq 1)

Guest kernel:

[   59.036988] pcieport 0000:01:00.0: AER: Uncorrectable (Non-Fatal) error message received from 0000:02:00.0
[   59.038326] nvidia 0000:02:00.0: PCIe Bus Error: severity=Uncorrectable (Non-Fatal), type=Transaction Layer, (Completer ID)
[   59.038440] nvidia 0000:02:00.0:   device [10de:2941] error status/mask=00008000/02400000
[   59.038533] nvidia 0000:02:00.0:    [15] CmpltAbrt              (First)
[   59.038705] nvidia 0000:02:00.0: AER:   TLP Header: 0x00000000 0x00000000 0x00000000 0x00000000
...

./aer-inject fatal.conf
qemu-system-aarch64: info: vfio 0018:06:00.0: host PCI error recovery started (seq 2, channel frozen)
qemu-system-aarch64: info: vfio 0018:06:00.0: host PCI error recovery completed (seq 2, device was reset)

Guest kernel:

[  150.909306] pcieport 0000:01:00.0: AER: Uncorrectable (Fatal) error message received from 0000:02:00.0
[  150.909566] nvidia 0000:02:00.0: AER: PCIe Bus Error: severity=Uncorrectable (Fatal), type=Inaccessible, (Unregistered Agent ID)
[  150.909835] nvidia 0000:02:00.0: AER: can't recover (no error_detected callback)
[  150.910062] pcieport 0000:01:00.0: unlocked secondary bus reset via: pciehp_reset_slot+0x54/0x98
[  156.603467] pcieport 0000:01:00.0: AER: Root Port link has been reset (0)
...

Please take a look and let me know your feedback.

Thanks,
Shameer

Shameer Kolothum (16):
  vfio/pci: Add a device access gate
  vfio/pci: Gate config space access
  vfio/pci: Buffer ROM reads before copying to userspace
  vfio/pci: Gate BAR and ROM access
  vfio/pci: Fail BAR faults while access is blocked
  vfio/pci: Gate interrupt configuration
  vfio/pci: Gate function reset and runtime power management
  vfio/pci: Gate device information queries and DMA-BUF export
  vfio/pci: Add PCI error recovery state
  vfio/pci: Quiesce INTx while access is blocked
  vfio/pci: Add INTx recovery start and finish helpers
  vfio/pci: Restore device state from slot_reset()
  vfio/pci: Complete recovery in resume()
  vfio/pci: Block device access during host recovery
  vfio/pci: Add VFIO_DEVICE_FEATURE_PCI_ERROR_RECOVERY
  vfio/pci: Enable host PCI error recovery for vfio-pci

 drivers/vfio/pci/vfio_pci_priv.h   |  10 +
 include/linux/vfio_pci_core.h      |  38 ++
 include/uapi/linux/vfio.h          |  44 ++
 drivers/vfio/pci/vfio_pci.c        |  12 +
 drivers/vfio/pci/vfio_pci_config.c | 104 ++++-
 drivers/vfio/pci/vfio_pci_core.c   | 636 ++++++++++++++++++++++++++++-
 drivers/vfio/pci/vfio_pci_dmabuf.c |  21 +-
 drivers/vfio/pci/vfio_pci_intrs.c  | 131 ++++++
 drivers/vfio/pci/vfio_pci_rdwr.c   | 169 ++++++--
 9 files changed, 1090 insertions(+), 75 deletions(-)

-- 
2.43.0


^ permalink raw reply	[flat|nested] 17+ messages in thread

* [RFC PATCH v2 01/16] vfio/pci: Add a device access gate
  2026-09-29 17:32 [RFC PATCH v2 00/16] vfio/pci: Handle PCI error recovery and report state to userspace Shameer Kolothum
@ 2026-09-29 17:32 ` Shameer Kolothum
  2026-09-29 17:32 ` [RFC PATCH v2 02/16] vfio/pci: Gate config space access Shameer Kolothum
                   ` (14 subsequent siblings)
  15 siblings, 0 replies; 17+ messages in thread
From: Shameer Kolothum @ 2026-09-29 17:32 UTC (permalink / raw)
  To: kvm, linux-pci, linux-kernel
  Cc: alex, jgg, kevin.tian, kbusch, michal.winiarski,
	satyanarayana.k.v.p, sonangp, ankita, nathanc, mochs,
	skolothumtho

PCI error recovery needs to block new device accesses and drain
outstanding accesses before changing device state.

Add an SRCU-based access gate, enabled by pci_recovery_supported.
Serialize open and close state updates with access_lock.

Assisted-by: LLM
Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com>
---
 drivers/vfio/pci/vfio_pci_priv.h |  3 ++
 include/linux/vfio_pci_core.h    | 19 ++++++++++++
 drivers/vfio/pci/vfio_pci_core.c | 53 ++++++++++++++++++++++++++++++++
 3 files changed, 75 insertions(+)

diff --git a/drivers/vfio/pci/vfio_pci_priv.h b/drivers/vfio/pci/vfio_pci_priv.h
index 4e7162234a2e..623e73f379cc 100644
--- a/drivers/vfio/pci/vfio_pci_priv.h
+++ b/drivers/vfio/pci/vfio_pci_priv.h
@@ -73,6 +73,9 @@ u16 vfio_pci_memory_lock_and_enable(struct vfio_pci_core_device *vdev);
 void vfio_pci_memory_unlock_and_restore(struct vfio_pci_core_device *vdev,
 					u16 cmd);
 
+int vfio_pci_core_access_begin(struct vfio_pci_core_device *vdev);
+void vfio_pci_core_access_end(struct vfio_pci_core_device *vdev, int idx);
+
 #ifdef CONFIG_VFIO_PCI_IGD
 bool vfio_pci_is_intel_display(struct pci_dev *pdev);
 int vfio_pci_igd_init(struct vfio_pci_core_device *vdev);
diff --git a/include/linux/vfio_pci_core.h b/include/linux/vfio_pci_core.h
index 9a1674c152aa..de0993280344 100644
--- a/include/linux/vfio_pci_core.h
+++ b/include/linux/vfio_pci_core.h
@@ -13,6 +13,7 @@
 #include <linux/vfio.h>
 #include <linux/irqbypass.h>
 #include <linux/rcupdate.h>
+#include <linux/srcu.h>
 #include <linux/types.h>
 #include <linux/uuid.h>
 #include <linux/notifier.h>
@@ -129,6 +130,7 @@ struct vfio_pci_core_device {
 	bool			disable_idle_d3:1;
 	bool			nointxmask:1;
 	bool			disable_vga:1;
+	bool			pci_recovery_supported:1;
 	/* Flags modified at runtime - dedicated storage unit */
 	bool			needs_reset;
 	bool			pm_intx_masked;
@@ -148,6 +150,23 @@ struct vfio_pci_core_device {
 	struct vfio_pci_core_device	*sriov_pf_core_dev;
 	struct notifier_block	nb;
 	struct rw_semaphore	memory_lock;
+	/*
+	 * Device accesses take access_srcu and check device_open and
+	 * access_blocked. To block access, set access_blocked and wait for
+	 * existing readers with synchronize_srcu(). New accesses fail with
+	 * -EIO while blocked.
+	 *
+	 * Close clears access_blocked before device_open.
+	 *
+	 * INTx paths use irqlock instead of access_srcu to coordinate with
+	 * recovery. They read pci_recovery_enabled and access_blocked with
+	 * READ_ONCE(), since writers do not always hold irqlock.
+	 */
+	struct srcu_struct	access_srcu;
+	bool			access_blocked;
+	/* Set after open completes, cleared before close tears down state. */
+	bool			device_open;
+	struct mutex		access_lock;	/* gate flag writers */
 	struct list_head	dmabufs;
 };
 
diff --git a/drivers/vfio/pci/vfio_pci_core.c b/drivers/vfio/pci/vfio_pci_core.c
index 6757054e9d87..88a78766c178 100644
--- a/drivers/vfio/pci/vfio_pci_core.c
+++ b/drivers/vfio/pci/vfio_pci_core.c
@@ -662,10 +662,18 @@ int vfio_pci_core_enable(struct vfio_pci_core_device *vdev)
 	if (!vfio_vga_disabled(vdev) && vfio_pci_is_vga(pdev))
 		vdev->has_vga = true;
 
+	if (vdev->pci_recovery_supported) {
+		ret = init_srcu_struct(&vdev->access_srcu);
+		if (ret)
+			goto out_free_config;
+	}
+
 	vfio_pci_core_map_bars(vdev);
 
 	return 0;
 
+out_free_config:
+	vfio_config_free(vdev);
 out_free_zdev:
 	vfio_pci_zdev_close_device(vdev);
 out_free_state:
@@ -812,6 +820,9 @@ void vfio_pci_core_disable(struct vfio_pci_core_device *vdev)
 	/* Put the pm-runtime usage counter acquired during enable */
 	if (!vdev->disable_idle_d3)
 		pm_runtime_put(&pdev->dev);
+
+	if (vdev->pci_recovery_supported)
+		cleanup_srcu_struct(&vdev->access_srcu);
 }
 EXPORT_SYMBOL_GPL(vfio_pci_core_disable);
 
@@ -820,6 +831,13 @@ void vfio_pci_core_close_device(struct vfio_device *core_vdev)
 	struct vfio_pci_core_device *vdev =
 		container_of(core_vdev, struct vfio_pci_core_device, vdev);
 
+	if (vdev->pci_recovery_supported) {
+		scoped_guard(mutex, &vdev->access_lock) {
+			WRITE_ONCE(vdev->access_blocked, false);
+			WRITE_ONCE(vdev->device_open, false);
+		}
+	}
+
 	if (vdev->sriov_pf_core_dev) {
 		mutex_lock(&vdev->sriov_pf_core_dev->vf_token->lock);
 		WARN_ON(!vdev->sriov_pf_core_dev->vf_token->users);
@@ -852,6 +870,12 @@ void vfio_pci_core_finish_enable(struct vfio_pci_core_device *vdev)
 		vdev->sriov_pf_core_dev->vf_token->users++;
 		mutex_unlock(&vdev->sriov_pf_core_dev->vf_token->lock);
 	}
+
+	if (vdev->pci_recovery_supported) {
+		guard(mutex)(&vdev->access_lock);
+		WRITE_ONCE(vdev->access_blocked, false);
+		WRITE_ONCE(vdev->device_open, true);
+	}
 }
 EXPORT_SYMBOL_GPL(vfio_pci_core_finish_enable);
 
@@ -1680,6 +1704,33 @@ static ssize_t vfio_pci_rw(struct vfio_pci_core_device *vdev, char __user *buf,
 	return ret;
 }
 
+/*
+ * Return an SRCU index for access_end(), or -EIO if access is blocked
+ * or the device is closed. Do not wait for recovery.
+ */
+int vfio_pci_core_access_begin(struct vfio_pci_core_device *vdev)
+{
+	int idx;
+
+	if (!vdev->pci_recovery_supported)
+		return 0;
+
+	idx = srcu_read_lock(&vdev->access_srcu);
+	if (unlikely(!READ_ONCE(vdev->device_open) ||
+		     READ_ONCE(vdev->access_blocked))) {
+		srcu_read_unlock(&vdev->access_srcu, idx);
+		return -EIO;
+	}
+
+	return idx;
+}
+
+void vfio_pci_core_access_end(struct vfio_pci_core_device *vdev, int idx)
+{
+	if (vdev->pci_recovery_supported)
+		srcu_read_unlock(&vdev->access_srcu, idx);
+}
+
 ssize_t vfio_pci_core_read(struct vfio_device *core_vdev, char __user *buf,
 		size_t count, loff_t *ppos)
 {
@@ -2196,6 +2247,7 @@ int vfio_pci_core_init_dev(struct vfio_device *core_vdev)
 	if (ret && ret != -EOPNOTSUPP)
 		return ret;
 	INIT_LIST_HEAD(&vdev->dmabufs);
+	mutex_init(&vdev->access_lock);
 	init_rwsem(&vdev->memory_lock);
 	xa_init(&vdev->ctx);
 
@@ -2210,6 +2262,7 @@ void vfio_pci_core_release_dev(struct vfio_device *core_vdev)
 
 	mutex_destroy(&vdev->igate);
 	mutex_destroy(&vdev->ioeventfds_lock);
+	mutex_destroy(&vdev->access_lock);
 	kfree(vdev->region);
 	kfree(vdev->pm_save);
 }
-- 
2.43.0


^ permalink raw reply	[flat|nested] 17+ messages in thread

* [RFC PATCH v2 02/16] vfio/pci: Gate config space access
  2026-09-29 17:32 [RFC PATCH v2 00/16] vfio/pci: Handle PCI error recovery and report state to userspace Shameer Kolothum
  2026-09-29 17:32 ` [RFC PATCH v2 01/16] vfio/pci: Add a device access gate Shameer Kolothum
@ 2026-09-29 17:32 ` Shameer Kolothum
  2026-09-29 17:32 ` [RFC PATCH v2 03/16] vfio/pci: Buffer ROM reads before copying to userspace Shameer Kolothum
                   ` (13 subsequent siblings)
  15 siblings, 0 replies; 17+ messages in thread
From: Shameer Kolothum @ 2026-09-29 17:32 UTC (permalink / raw)
  To: kvm, linux-pci, linux-kernel
  Cc: alex, jgg, kevin.tian, kbusch, michal.winiarski,
	satyanarayana.k.v.p, sonangp, ankita, nathanc, mochs,
	skolothumtho

Protect config callbacks with the access gate. Keep user copies outside
SRCU because userfaultfd can stall them indefinitely.

Run D0 transitions after releasing SRCU to avoid waiting for pci_bus_sem
while AER waits for readers to drain. Record blocked D0 requests for
replay in resume(). Return -EIO for blocked D1/D2/D3 requests, which are
not queued for replay. Preserve existing handling of hardware power-state
errors outside recovery.

Add a power_up output parameter to writefn so the PM callback can request
a D0 transition after the caller releases access_srcu.

Assisted-by: LLM
Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com>
---
Note: The memory_lock/pci_bus_sem dependency remains unresolved. See
the cover letter's "Locking and open questions" section.
---
 include/linux/vfio_pci_core.h      |   2 +
 drivers/vfio/pci/vfio_pci_config.c | 104 +++++++++++++++++++++++------
 2 files changed, 84 insertions(+), 22 deletions(-)

diff --git a/include/linux/vfio_pci_core.h b/include/linux/vfio_pci_core.h
index de0993280344..1dc9630dc740 100644
--- a/include/linux/vfio_pci_core.h
+++ b/include/linux/vfio_pci_core.h
@@ -138,6 +138,8 @@ struct vfio_pci_core_device {
 	bool			sriov_active;
 	struct pci_saved_state	*pci_saved_state;
 	struct pci_saved_state	*pm_save;
+	/* Deferred D0 request, protected by memory_lock. */
+	bool			power_up_pending;
 	int			ioeventfds_nr;
 	struct vfio_pci_eventfd __rcu *err_trigger;
 	struct vfio_pci_eventfd __rcu *req_trigger;
diff --git a/drivers/vfio/pci/vfio_pci_config.c b/drivers/vfio/pci/vfio_pci_config.c
index 9914f3ac69ae..e9480b573381 100644
--- a/drivers/vfio/pci/vfio_pci_config.c
+++ b/drivers/vfio/pci/vfio_pci_config.c
@@ -111,8 +111,14 @@ struct perm_bits {
 	u8	*write;		/* writeable bits */
 	int	(*readfn)(struct vfio_pci_core_device *vdev, int pos, int count,
 			  struct perm_bits *perm, int offset, __le32 *val);
+	/*
+	 * Write callbacks run under access_srcu. Defer operations that take
+	 * pci_bus_sem: AER may hold its read lock while waiting for SRCU,
+	 * and a queued writer can block further read acquisitions.
+	 */
 	int	(*writefn)(struct vfio_pci_core_device *vdev, int pos, int count,
-			   struct perm_bits *perm, int offset, __le32 val);
+			   struct perm_bits *perm, int offset, __le32 val,
+			   bool *power_up);
 };
 
 #define	NO_VIRT		0
@@ -200,7 +206,8 @@ static int vfio_default_config_read(struct vfio_pci_core_device *vdev, int pos,
 
 static int vfio_default_config_write(struct vfio_pci_core_device *vdev, int pos,
 				     int count, struct perm_bits *perm,
-				     int offset, __le32 val)
+				     int offset, __le32 val,
+				     bool *power_up)
 {
 	__le32 virt = 0, write = 0;
 
@@ -272,7 +279,8 @@ static int vfio_direct_config_read(struct vfio_pci_core_device *vdev, int pos,
 /* Raw access skips any kind of virtualization */
 static int vfio_raw_config_write(struct vfio_pci_core_device *vdev, int pos,
 				 int count, struct perm_bits *perm,
-				 int offset, __le32 val)
+				 int offset, __le32 val,
+				 bool *power_up)
 {
 	int ret;
 
@@ -299,7 +307,8 @@ static int vfio_raw_config_read(struct vfio_pci_core_device *vdev, int pos,
 /* Virt access uses only virtualization */
 static int vfio_virt_config_write(struct vfio_pci_core_device *vdev, int pos,
 				  int count, struct perm_bits *perm,
-				  int offset, __le32 val)
+				  int offset, __le32 val,
+				  bool *power_up)
 {
 	memcpy(vdev->vconfig + pos, &val, count);
 	return count;
@@ -563,7 +572,8 @@ static bool vfio_need_bar_restore(struct vfio_pci_core_device *vdev)
 
 static int vfio_basic_config_write(struct vfio_pci_core_device *vdev, int pos,
 				   int count, struct perm_bits *perm,
-				   int offset, __le32 val)
+				   int offset, __le32 val,
+				   bool *power_up)
 {
 	struct pci_dev *pdev = vdev->pdev;
 	__le16 *virt_cmd;
@@ -613,7 +623,8 @@ static int vfio_basic_config_write(struct vfio_pci_core_device *vdev, int pos,
 			vfio_bar_restore(vdev);
 	}
 
-	count = vfio_default_config_write(vdev, pos, count, perm, offset, val);
+	count = vfio_default_config_write(vdev, pos, count, perm, offset, val,
+					  power_up);
 	if (count < 0) {
 		if (offset == PCI_COMMAND)
 			up_write(&vdev->memory_lock);
@@ -709,8 +720,8 @@ static int __init init_pci_cap_basic_perm(struct perm_bits *perm)
  * It takes all the required locks to protect the access of power related
  * variables and then invokes vfio_pci_set_power_state().
  */
-static void vfio_lock_and_set_power_state(struct vfio_pci_core_device *vdev,
-					  pci_power_t state)
+static int vfio_lock_and_set_power_state(struct vfio_pci_core_device *vdev,
+					 pci_power_t state)
 {
 	if (state >= PCI_D3hot) {
 		vfio_pci_zap_and_down_write_memory_lock(vdev);
@@ -719,27 +730,51 @@ static void vfio_lock_and_set_power_state(struct vfio_pci_core_device *vdev,
 		down_write(&vdev->memory_lock);
 	}
 
+	/*
+	 * Defer D0 until recovery completes if access is already blocked.
+	 *
+	 * This check does not prevent a block starting during the transition.
+	 * A lock dependency remains: D0 takes pci_bus_sem under memory_lock,
+	 * while AER takes memory_lock under pci_bus_sem. A queued bus writer
+	 * can block the D0 reader and deadlock both paths.
+	 */
+	if (vdev->pci_recovery_supported && READ_ONCE(vdev->access_blocked)) {
+		if (state == PCI_D0)
+			vdev->power_up_pending = true;
+		up_write(&vdev->memory_lock);
+		return state == PCI_D0 ? 0 : -EIO;
+	}
+
 	vfio_pci_set_power_state(vdev, state);
 	if (__vfio_pci_memory_enabled(vdev))
 		vfio_pci_dma_buf_move(vdev, false);
 	up_write(&vdev->memory_lock);
+
+	return 0;
 }
 
 static int vfio_pm_config_write(struct vfio_pci_core_device *vdev, int pos,
 				int count, struct perm_bits *perm,
-				int offset, __le32 val)
+				int offset, __le32 val,
+				bool *power_up)
 {
-	count = vfio_default_config_write(vdev, pos, count, perm, offset, val);
+	count = vfio_default_config_write(vdev, pos, count, perm, offset, val,
+					  power_up);
 	if (count < 0)
 		return count;
 
 	if (offset == PCI_PM_CTRL) {
 		pci_power_t state;
+		int ret;
 
 		switch (le32_to_cpu(val) & PCI_PM_CTRL_STATE_MASK) {
 		case 0:
-			state = PCI_D0;
-			break;
+			/*
+			 * Request D0 after dropping access_srcu; the ASPM update can
+			 * otherwise deadlock with AER waiting for SRCU readers.
+			 */
+			*power_up = true;
+			return count;
 		case 1:
 			state = PCI_D1;
 			break;
@@ -751,7 +786,9 @@ static int vfio_pm_config_write(struct vfio_pci_core_device *vdev, int pos,
 			break;
 		}
 
-		vfio_lock_and_set_power_state(vdev, state);
+		ret = vfio_lock_and_set_power_state(vdev, state);
+		if (ret)
+			return ret;
 	}
 
 	return count;
@@ -799,7 +836,8 @@ static int __init init_pci_cap_pm_perm(struct perm_bits *perm)
 
 static int vfio_vpd_config_write(struct vfio_pci_core_device *vdev, int pos,
 				 int count, struct perm_bits *perm,
-				 int offset, __le32 val)
+				 int offset, __le32 val,
+				 bool *power_up)
 {
 	struct pci_dev *pdev = vdev->pdev;
 	__le16 *paddr = (__le16 *)(vdev->vconfig + pos - offset + PCI_VPD_ADDR);
@@ -812,7 +850,8 @@ static int vfio_vpd_config_write(struct vfio_pci_core_device *vdev, int pos,
 	 * of PCI_VPD_ADDR, then the PCI_VPD_ADDR_F bit is written and we
 	 * have work to do.
 	 */
-	count = vfio_default_config_write(vdev, pos, count, perm, offset, val);
+	count = vfio_default_config_write(vdev, pos, count, perm, offset, val,
+					  power_up);
 	if (count < 0 || offset > PCI_VPD_ADDR + 1 ||
 	    offset + count <= PCI_VPD_ADDR + 1)
 		return count;
@@ -881,13 +920,15 @@ static int __init init_pci_cap_pcix_perm(struct perm_bits *perm)
 
 static int vfio_exp_config_write(struct vfio_pci_core_device *vdev, int pos,
 				 int count, struct perm_bits *perm,
-				 int offset, __le32 val)
+				 int offset, __le32 val,
+				 bool *power_up)
 {
 	__le16 *ctrl = (__le16 *)(vdev->vconfig + pos -
 				  offset + PCI_EXP_DEVCTL);
 	int readrq = le16_to_cpu(*ctrl) & PCI_EXP_DEVCTL_READRQ;
 
-	count = vfio_default_config_write(vdev, pos, count, perm, offset, val);
+	count = vfio_default_config_write(vdev, pos, count, perm, offset, val,
+					  power_up);
 	if (count < 0)
 		return count;
 
@@ -968,11 +1009,13 @@ static int __init init_pci_cap_exp_perm(struct perm_bits *perm)
 
 static int vfio_af_config_write(struct vfio_pci_core_device *vdev, int pos,
 				int count, struct perm_bits *perm,
-				int offset, __le32 val)
+				int offset, __le32 val,
+				bool *power_up)
 {
 	u8 *ctrl = vdev->vconfig + pos - offset + PCI_AF_CTRL;
 
-	count = vfio_default_config_write(vdev, pos, count, perm, offset, val);
+	count = vfio_default_config_write(vdev, pos, count, perm, offset, val,
+					  power_up);
 	if (count < 0)
 		return count;
 
@@ -1168,9 +1211,11 @@ static int vfio_msi_config_read(struct vfio_pci_core_device *vdev, int pos,
 
 static int vfio_msi_config_write(struct vfio_pci_core_device *vdev, int pos,
 				 int count, struct perm_bits *perm,
-				 int offset, __le32 val)
+				 int offset, __le32 val,
+				 bool *power_up)
 {
-	count = vfio_default_config_write(vdev, pos, count, perm, offset, val);
+	count = vfio_default_config_write(vdev, pos, count, perm, offset, val,
+					  power_up);
 	if (count < 0)
 		return count;
 
@@ -1889,6 +1934,8 @@ ssize_t vfio_pci_config_rw_single(struct vfio_pci_core_device *vdev,
 	struct perm_bits *perm;
 	__le32 val = 0;
 	int cap_start = 0, offset;
+	int idx;
+	bool power_up = false;
 	u8 cap_id;
 	ssize_t ret;
 
@@ -1957,11 +2004,24 @@ ssize_t vfio_pci_config_rw_single(struct vfio_pci_core_device *vdev,
 		if (copy_from_user(&val, buf, count))
 			return -EFAULT;
 
-		ret = perm->writefn(vdev, *ppos, count, perm, offset, val);
+		idx = vfio_pci_core_access_begin(vdev);
+		if (idx < 0)
+			return idx;
+		ret = perm->writefn(vdev, *ppos, count, perm, offset, val,
+				    &power_up);
+		vfio_pci_core_access_end(vdev, idx);
+		if (ret < 0)
+			return ret;
+		if (power_up)
+			vfio_lock_and_set_power_state(vdev, PCI_D0);
 	} else {
 		if (perm->readfn) {
+			idx = vfio_pci_core_access_begin(vdev);
+			if (idx < 0)
+				return idx;
 			ret = perm->readfn(vdev, *ppos, count,
 					   perm, offset, &val);
+			vfio_pci_core_access_end(vdev, idx);
 			if (ret < 0)
 				return ret;
 		}
-- 
2.43.0


^ permalink raw reply	[flat|nested] 17+ messages in thread

* [RFC PATCH v2 03/16] vfio/pci: Buffer ROM reads before copying to userspace
  2026-09-29 17:32 [RFC PATCH v2 00/16] vfio/pci: Handle PCI error recovery and report state to userspace Shameer Kolothum
  2026-09-29 17:32 ` [RFC PATCH v2 01/16] vfio/pci: Add a device access gate Shameer Kolothum
  2026-09-29 17:32 ` [RFC PATCH v2 02/16] vfio/pci: Gate config space access Shameer Kolothum
@ 2026-09-29 17:32 ` Shameer Kolothum
  2026-09-29 17:32 ` [RFC PATCH v2 04/16] vfio/pci: Gate BAR and ROM access Shameer Kolothum
                   ` (12 subsequent siblings)
  15 siblings, 0 replies; 17+ messages in thread
From: Shameer Kolothum @ 2026-09-29 17:32 UTC (permalink / raw)
  To: kvm, linux-pci, linux-kernel
  Cc: alex, jgg, kevin.tian, kbusch, michal.winiarski,
	satyanarayana.k.v.p, sonangp, ankita, nathanc, mochs,
	skolothumtho

Move ROM reads into a separate helper because vfio_pci_core_do_io_rw()
copies directly to userspace. Buffer the requested ROM data and unmap
the ROM before copying it to userspace.

The next patch adds recovery gating around ROM mapping, reading and
unmapping. Buffering keeps userspace faults outside that gate, so they
cannot stall recovery or leave ROM reads to resume after recovery has
disabled decoding.

Preserve aligned byte, word and dword reads and 0xff padding. The
temporary allocation can fail with -ENOMEM.

Assisted-by: LLM
Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com>
---
 drivers/vfio/pci/vfio_pci_rdwr.c | 147 ++++++++++++++++++++++---------
 1 file changed, 104 insertions(+), 43 deletions(-)

diff --git a/drivers/vfio/pci/vfio_pci_rdwr.c b/drivers/vfio/pci/vfio_pci_rdwr.c
index 7f14dd46de17..b4f2c9d967dd 100644
--- a/drivers/vfio/pci/vfio_pci_rdwr.c
+++ b/drivers/vfio/pci/vfio_pci_rdwr.c
@@ -10,8 +10,10 @@
  * Author: Tom Lyon, pugs@cisco.com
  */
 
+#include <linux/align.h>
 #include <linux/fs.h>
 #include <linux/pci.h>
+#include <linux/slab.h>
 #include <linux/uaccess.h>
 #include <linux/io.h>
 #include <linux/vfio.h>
@@ -198,6 +200,96 @@ ssize_t vfio_pci_core_do_io_rw(struct vfio_pci_core_device *vdev, bool test_mem,
 }
 EXPORT_SYMBOL_GPL(vfio_pci_core_do_io_rw);
 
+static ssize_t vfio_pci_rom_read(struct vfio_pci_core_device *vdev,
+				 char __user *buf, size_t count, loff_t pos)
+{
+	struct pci_dev *pdev = vdev->pdev;
+	bool rom_bar = pci_resource_start(pdev, PCI_ROM_RESOURCE);
+	size_t size, length = 0, done = 0;
+	void *data = NULL;
+	void __iomem *io;
+	ssize_t ret;
+
+	/* Serialize ROM decoding with power-state and other ROM accesses. */
+	down_write(&vdev->memory_lock);
+	if (rom_bar) {
+		io = pci_map_rom(pdev, &size);
+	} else {
+		io = ioremap(pdev->rom, pdev->romlen);
+		size = pdev->romlen;
+	}
+	if (!io) {
+		ret = -ENOMEM;
+		goto out_unlock;
+	}
+
+	/* Buffer only ROM data, not unused space in a large ROM BAR. */
+	if (pos < size)
+		length = min(count, size - (size_t)pos);
+	if (length) {
+		if ((pci_resource_flags(pdev, PCI_ROM_RESOURCE) & IORESOURCE_MEM) &&
+		    !__vfio_pci_memory_enabled(vdev)) {
+			ret = -EIO;
+			goto out_unmap;
+		}
+		data = kvmalloc(length, GFP_KERNEL_ACCOUNT);
+		if (!data) {
+			ret = -ENOMEM;
+			goto out_unmap;
+		}
+	}
+
+	/*
+	 * Certain devices (e.g. Intel X710) don't support qword
+	 * access to the ROM bar. Otherwise PCI AER errors might be
+	 * triggered.
+	 *
+	 * Disable qword access to the ROM bar universally, which
+	 * worked reliably for years before qword access is enabled.
+	 */
+	while (done < length) {
+		if (length - done >= 4 && IS_ALIGNED(pos + done, 4)) {
+			u32 val = vfio_ioread32(io + pos + done);
+
+			memcpy(data + done, &val, sizeof(val));
+			done += sizeof(val);
+		} else if (length - done >= 2 && IS_ALIGNED(pos + done, 2)) {
+			u16 val = vfio_ioread16(io + pos + done);
+
+			memcpy(data + done, &val, sizeof(val));
+			done += sizeof(val);
+		} else {
+			((u8 *)data)[done] = vfio_ioread8(io + pos + done);
+			done++;
+		}
+	}
+	ret = count;
+
+out_unmap:
+	if (rom_bar)
+		pci_unmap_rom(pdev, io);
+	else
+		iounmap(io);
+out_unlock:
+	up_write(&vdev->memory_lock);
+	if (ret < 0)
+		goto out_free;
+
+	if (length && copy_to_user(buf, data, length)) {
+		ret = -EFAULT;
+		goto out_free;
+	}
+	for (done = length; done < count; done++) {
+		if (put_user((u8)0xff, buf + done)) {
+			ret = -EFAULT;
+			break;
+		}
+	}
+out_free:
+	kvfree(data);
+	return ret;
+}
+
 ssize_t vfio_pci_bar_rw(struct vfio_pci_core_device *vdev, char __user *buf,
 			size_t count, loff_t *ppos, bool iswrite)
 {
@@ -209,7 +301,6 @@ ssize_t vfio_pci_bar_rw(struct vfio_pci_core_device *vdev, char __user *buf,
 	void __iomem *io;
 	struct resource *res = &vdev->pdev->resource[bar];
 	ssize_t done;
-	enum vfio_pci_io_width max_width = VFIO_PCI_IO_WIDTH_8;
 
 	if (pci_resource_start(pdev, bar))
 		end = pci_resource_len(pdev, bar);
@@ -224,57 +315,27 @@ ssize_t vfio_pci_bar_rw(struct vfio_pci_core_device *vdev, char __user *buf,
 	count = min(count, (size_t)(end - pos));
 
 	if (bar == PCI_ROM_RESOURCE) {
-		/*
-		 * The ROM can fill less space than the BAR, so we start the
-		 * excluded range at the end of the actual ROM.  This makes
-		 * filling large ROM BARs much faster.
-		 */
-		if (pci_resource_start(pdev, bar)) {
-			io = pci_map_rom(pdev, &x_start);
-		} else {
-			io = ioremap(pdev->rom, pdev->romlen);
-			x_start = pdev->romlen;
-		}
-		if (!io)
-			return -ENOMEM;
-		x_end = end;
-
-		/*
-		 * Certain devices (e.g. Intel X710) don't support qword
-		 * access to the ROM bar. Otherwise PCI AER errors might be
-		 * triggered.
-		 *
-		 * Disable qword access to the ROM bar universally, which
-		 * worked reliably for years before qword access is enabled.
-		 */
-		max_width = VFIO_PCI_IO_WIDTH_4;
+		if (iswrite)
+			return -EINVAL;
+		done = vfio_pci_rom_read(vdev, buf, count, pos);
 	} else {
 		io = vfio_pci_core_get_iomap(vdev, bar);
-		if (IS_ERR(io)) {
-			done = PTR_ERR(io);
-			goto out;
+		if (IS_ERR(io))
+			return PTR_ERR(io);
+
+		if (bar == vdev->msix_bar) {
+			x_start = vdev->msix_offset;
+			x_end = vdev->msix_offset + vdev->msix_size;
 		}
-	}
 
-	if (bar == vdev->msix_bar) {
-		x_start = vdev->msix_offset;
-		x_end = vdev->msix_offset + vdev->msix_size;
+		done = vfio_pci_core_do_io_rw(vdev, res->flags & IORESOURCE_MEM,
+					      io, buf, pos, count, x_start, x_end,
+					      iswrite, VFIO_PCI_IO_WIDTH_8);
 	}
 
-	done = vfio_pci_core_do_io_rw(vdev, res->flags & IORESOURCE_MEM, io, buf, pos,
-				      count, x_start, x_end, iswrite, max_width);
-
 	if (done >= 0)
 		*ppos += done;
 
-	if (bar == PCI_ROM_RESOURCE) {
-		if (pci_resource_start(pdev, bar))
-			pci_unmap_rom(pdev, io);
-		else
-			iounmap(io);
-	}
-
-out:
 	return done;
 }
 
-- 
2.43.0


^ permalink raw reply	[flat|nested] 17+ messages in thread

* [RFC PATCH v2 04/16] vfio/pci: Gate BAR and ROM access
  2026-09-29 17:32 [RFC PATCH v2 00/16] vfio/pci: Handle PCI error recovery and report state to userspace Shameer Kolothum
                   ` (2 preceding siblings ...)
  2026-09-29 17:32 ` [RFC PATCH v2 03/16] vfio/pci: Buffer ROM reads before copying to userspace Shameer Kolothum
@ 2026-09-29 17:32 ` Shameer Kolothum
  2026-09-29 17:32 ` [RFC PATCH v2 05/16] vfio/pci: Fail BAR faults while access is blocked Shameer Kolothum
                   ` (11 subsequent siblings)
  15 siblings, 0 replies; 17+ messages in thread
From: Shameer Kolothum @ 2026-09-29 17:32 UTC (permalink / raw)
  To: kvm, linux-pci, linux-kernel
  Cc: alex, jgg, kevin.tian, kbusch, michal.winiarski,
	satyanarayana.k.v.p, sonangp, ankita, nathanc, mochs,
	skolothumtho

Protect the common I/O accessors used for BAR, VGA, port I/O and ioeventfd
accesses with access_srcu. Recovery can then block new accesses and wait
for admitted accesses to finish. Keep user copies outside SRCU so a
userspace fault cannot stall recovery.

Return -EIO before runtime resume if region access is already blocked.
This check avoids waking the device unnecessarily; it does not serialize
runtime resume against recovery. The I/O accessors check the gate again
under SRCU before touching hardware.

Hold one SRCU section across ROM mapping, the buffered hardware read and
unmapping. Copy the snapshot to userspace afterward, so recovery cannot
disable ROM decoding between read chunks or wait for a userspace fault.

Assisted-by: LLM
Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com>
---
 drivers/vfio/pci/vfio_pci_core.c |  4 ++++
 drivers/vfio/pci/vfio_pci_rdwr.c | 22 ++++++++++++++++++++++
 2 files changed, 26 insertions(+)

diff --git a/drivers/vfio/pci/vfio_pci_core.c b/drivers/vfio/pci/vfio_pci_core.c
index 88a78766c178..323038cd0ca5 100644
--- a/drivers/vfio/pci/vfio_pci_core.c
+++ b/drivers/vfio/pci/vfio_pci_core.c
@@ -1666,6 +1666,10 @@ static ssize_t vfio_pci_rw(struct vfio_pci_core_device *vdev, char __user *buf,
 	if (index >= VFIO_PCI_NUM_REGIONS + vdev->num_regions)
 		return -EINVAL;
 
+	/* Avoid runtime resume when blocked; region handlers check access again. */
+	if (vdev->pci_recovery_supported && READ_ONCE(vdev->access_blocked))
+		return -EIO;
+
 	ret = pm_runtime_resume_and_get(&vdev->pdev->dev);
 	if (ret) {
 		pci_info_ratelimited(vdev->pdev, "runtime resume failed %d\n",
diff --git a/drivers/vfio/pci/vfio_pci_rdwr.c b/drivers/vfio/pci/vfio_pci_rdwr.c
index b4f2c9d967dd..40c9dd02c3df 100644
--- a/drivers/vfio/pci/vfio_pci_rdwr.c
+++ b/drivers/vfio/pci/vfio_pci_rdwr.c
@@ -44,10 +44,17 @@
 int vfio_pci_core_iowrite##size(struct vfio_pci_core_device *vdev,	\
 			bool test_mem, u##size val, void __iomem *io)	\
 {									\
+	int idx;							\
+									\
+	idx = vfio_pci_core_access_begin(vdev);				\
+	if (idx < 0)							\
+		return idx;						\
+									\
 	if (test_mem) {							\
 		down_read(&vdev->memory_lock);				\
 		if (!__vfio_pci_memory_enabled(vdev)) {			\
 			up_read(&vdev->memory_lock);			\
+			vfio_pci_core_access_end(vdev, idx);		\
 			return -EIO;					\
 		}							\
 	}								\
@@ -56,6 +63,7 @@ int vfio_pci_core_iowrite##size(struct vfio_pci_core_device *vdev,	\
 									\
 	if (test_mem)							\
 		up_read(&vdev->memory_lock);				\
+	vfio_pci_core_access_end(vdev, idx);				\
 									\
 	return 0;							\
 }									\
@@ -70,10 +78,17 @@ VFIO_IOWRITE(64)
 int vfio_pci_core_ioread##size(struct vfio_pci_core_device *vdev,	\
 			bool test_mem, u##size *val, void __iomem *io)	\
 {									\
+	int idx;							\
+									\
+	idx = vfio_pci_core_access_begin(vdev);				\
+	if (idx < 0)							\
+		return idx;						\
+									\
 	if (test_mem) {							\
 		down_read(&vdev->memory_lock);				\
 		if (!__vfio_pci_memory_enabled(vdev)) {			\
 			up_read(&vdev->memory_lock);			\
+			vfio_pci_core_access_end(vdev, idx);		\
 			return -EIO;					\
 		}							\
 	}								\
@@ -82,6 +97,7 @@ int vfio_pci_core_ioread##size(struct vfio_pci_core_device *vdev,	\
 									\
 	if (test_mem)							\
 		up_read(&vdev->memory_lock);				\
+	vfio_pci_core_access_end(vdev, idx);				\
 									\
 	return 0;							\
 }									\
@@ -209,6 +225,11 @@ static ssize_t vfio_pci_rom_read(struct vfio_pci_core_device *vdev,
 	void *data = NULL;
 	void __iomem *io;
 	ssize_t ret;
+	int idx;
+
+	idx = vfio_pci_core_access_begin(vdev);
+	if (idx < 0)
+		return idx;
 
 	/* Serialize ROM decoding with power-state and other ROM accesses. */
 	down_write(&vdev->memory_lock);
@@ -272,6 +293,7 @@ static ssize_t vfio_pci_rom_read(struct vfio_pci_core_device *vdev,
 		iounmap(io);
 out_unlock:
 	up_write(&vdev->memory_lock);
+	vfio_pci_core_access_end(vdev, idx);
 	if (ret < 0)
 		goto out_free;
 
-- 
2.43.0


^ permalink raw reply	[flat|nested] 17+ messages in thread

* [RFC PATCH v2 05/16] vfio/pci: Fail BAR faults while access is blocked
  2026-09-29 17:32 [RFC PATCH v2 00/16] vfio/pci: Handle PCI error recovery and report state to userspace Shameer Kolothum
                   ` (3 preceding siblings ...)
  2026-09-29 17:32 ` [RFC PATCH v2 04/16] vfio/pci: Gate BAR and ROM access Shameer Kolothum
@ 2026-09-29 17:32 ` Shameer Kolothum
  2026-09-29 17:32 ` [RFC PATCH v2 06/16] vfio/pci: Gate interrupt configuration Shameer Kolothum
                   ` (10 subsequent siblings)
  15 siblings, 0 replies; 17+ messages in thread
From: Shameer Kolothum @ 2026-09-29 17:32 UTC (permalink / raw)
  To: kvm, linux-pci, linux-kernel
  Cc: alex, jgg, kevin.tian, kbusch, michal.winiarski,
	satyanarayana.k.v.p, sonangp, ankita, nathanc, mochs,
	skolothumtho

Hold access_srcu across PFN insertion in the BAR fault handler. This
ensures recovery waits for an insertion before revoking the mapping.

Return VM_FAULT_SIGBUS rather than wait for recovery in the fault
handler when access is blocked. Whether userspace can handle the
resulting failure depends on the architecture and KVM fault-reporting
path.

Assisted-by: LLM
Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com>
---
 drivers/vfio/pci/vfio_pci_core.c | 11 +++++++++++
 1 file changed, 11 insertions(+)

diff --git a/drivers/vfio/pci/vfio_pci_core.c b/drivers/vfio/pci/vfio_pci_core.c
index 323038cd0ca5..1c8e39f3f509 100644
--- a/drivers/vfio/pci/vfio_pci_core.c
+++ b/drivers/vfio/pci/vfio_pci_core.c
@@ -1842,8 +1842,19 @@ static vm_fault_t vfio_pci_mmap_huge_fault(struct vm_fault *vmf,
 	vm_fault_t ret = VM_FAULT_FALLBACK;
 
 	if (is_aligned_for_order(vma, addr, pfn, order)) {
+		int idx;
+
+		/*
+		 * Hold access_srcu across PFN insertion so recovery cannot
+		 * finish revoking mappings before this insertion completes.
+		 */
+		idx = vfio_pci_core_access_begin(vdev);
+		if (idx < 0)
+			return VM_FAULT_SIGBUS;
+
 		scoped_guard(rwsem_read, &vdev->memory_lock)
 			ret = vfio_pci_vmf_insert_pfn(vdev, vmf, pfn, order);
+		vfio_pci_core_access_end(vdev, idx);
 	}
 
 	dev_dbg_ratelimited(&vdev->pdev->dev,
-- 
2.43.0


^ permalink raw reply	[flat|nested] 17+ messages in thread

* [RFC PATCH v2 06/16] vfio/pci: Gate interrupt configuration
  2026-09-29 17:32 [RFC PATCH v2 00/16] vfio/pci: Handle PCI error recovery and report state to userspace Shameer Kolothum
                   ` (4 preceding siblings ...)
  2026-09-29 17:32 ` [RFC PATCH v2 05/16] vfio/pci: Fail BAR faults while access is blocked Shameer Kolothum
@ 2026-09-29 17:32 ` Shameer Kolothum
  2026-09-29 17:32 ` [RFC PATCH v2 07/16] vfio/pci: Gate function reset and runtime power management Shameer Kolothum
                   ` (9 subsequent siblings)
  15 siblings, 0 replies; 17+ messages in thread
From: Shameer Kolothum @ 2026-09-29 17:32 UTC (permalink / raw)
  To: kvm, linux-pci, linux-kernel
  Cc: alex, jgg, kevin.tian, kbusch, michal.winiarski,
	satyanarayana.k.v.p, sonangp, ankita, nathanc, mochs,
	skolothumtho

Protect INTx, MSI and MSI-X hardware accesses in GET_IRQ_INFO and
SET_IRQS with access_srcu. ERR and REQ only update eventfds, so leave
them available while device access is blocked.

SET_IRQS uses separate read sections for validation and configuration,
with the user copy between them. Recovery must also wait for IRQ teardown
to finish flushing the global virqfd workqueue.

Assisted-by: LLM
Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com>
---
 drivers/vfio/pci/vfio_pci_core.c | 54 ++++++++++++++++++++++++++++++--
 1 file changed, 51 insertions(+), 3 deletions(-)

diff --git a/drivers/vfio/pci/vfio_pci_core.c b/drivers/vfio/pci/vfio_pci_core.c
index 1c8e39f3f509..91c92e8f3daa 100644
--- a/drivers/vfio/pci/vfio_pci_core.c
+++ b/drivers/vfio/pci/vfio_pci_core.c
@@ -1266,6 +1266,19 @@ int vfio_pci_ioctl_get_region_info(struct vfio_device *core_vdev,
 }
 EXPORT_SYMBOL_GPL(vfio_pci_ioctl_get_region_info);
 
+/* ERR and REQ do not access hardware. */
+static bool vfio_pci_irq_index_is_device(u32 index)
+{
+	switch (index) {
+	case VFIO_PCI_INTX_IRQ_INDEX:
+	case VFIO_PCI_MSI_IRQ_INDEX:
+	case VFIO_PCI_MSIX_IRQ_INDEX:
+		return true;
+	default:
+		return false;
+	}
+}
+
 static int vfio_pci_ioctl_get_irq_info(struct vfio_pci_core_device *vdev,
 				       struct vfio_irq_info __user *arg)
 {
@@ -1289,7 +1302,16 @@ static int vfio_pci_ioctl_get_irq_info(struct vfio_pci_core_device *vdev,
 
 	info.flags = VFIO_IRQ_INFO_EVENTFD;
 
-	info.count = vfio_pci_get_irq_count(vdev, info.index);
+	if (vfio_pci_irq_index_is_device(info.index)) {
+		int idx = vfio_pci_core_access_begin(vdev);
+
+		if (idx < 0)
+			return idx;
+		info.count = vfio_pci_get_irq_count(vdev, info.index);
+		vfio_pci_core_access_end(vdev, idx);
+	} else {
+		info.count = vfio_pci_get_irq_count(vdev, info.index);
+	}
 
 	if (info.index == VFIO_PCI_INTX_IRQ_INDEX)
 		info.flags |=
@@ -1306,13 +1328,24 @@ static int vfio_pci_ioctl_set_irqs(struct vfio_pci_core_device *vdev,
 	unsigned long minsz = offsetofend(struct vfio_irq_set, count);
 	struct vfio_irq_set hdr;
 	u8 *data = NULL;
-	int max, ret = 0;
+	bool device_irq;
+	int max, ret = 0, idx = 0;
 	size_t data_size = 0;
 
 	if (copy_from_user(&hdr, arg, minsz))
 		return -EFAULT;
 
-	max = vfio_pci_get_irq_count(vdev, hdr.index);
+	/* Use separate SRCU sections to keep user copies outside access_srcu. */
+	device_irq = vfio_pci_irq_index_is_device(hdr.index);
+	if (device_irq) {
+		idx = vfio_pci_core_access_begin(vdev);
+		if (idx < 0)
+			return idx;
+		max = vfio_pci_get_irq_count(vdev, hdr.index);
+		vfio_pci_core_access_end(vdev, idx);
+	} else {
+		max = vfio_pci_get_irq_count(vdev, hdr.index);
+	}
 
 	ret = vfio_set_irqs_validate_and_prepare(&hdr, max, VFIO_PCI_NUM_IRQS,
 						 &data_size);
@@ -1325,12 +1358,27 @@ static int vfio_pci_ioctl_set_irqs(struct vfio_pci_core_device *vdev,
 			return PTR_ERR(data);
 	}
 
+	/*
+	 * IRQ teardown flushes the global virqfd workqueue. Recovery must
+	 * wait for that flush, including pending work for other devices.
+	 */
+	if (device_irq) {
+		idx = vfio_pci_core_access_begin(vdev);
+		if (idx < 0) {
+			ret = idx;
+			goto out_free;
+		}
+	}
+	/* SRCU drains accesses for recovery; igate serializes IRQ changes. */
 	mutex_lock(&vdev->igate);
 
 	ret = vfio_pci_set_irqs_ioctl(vdev, hdr.flags, hdr.index, hdr.start,
 				      hdr.count, data);
 
 	mutex_unlock(&vdev->igate);
+	if (device_irq)
+		vfio_pci_core_access_end(vdev, idx);
+out_free:
 	kfree(data);
 
 	return ret;
-- 
2.43.0


^ permalink raw reply	[flat|nested] 17+ messages in thread

* [RFC PATCH v2 07/16] vfio/pci: Gate function reset and runtime power management
  2026-09-29 17:32 [RFC PATCH v2 00/16] vfio/pci: Handle PCI error recovery and report state to userspace Shameer Kolothum
                   ` (5 preceding siblings ...)
  2026-09-29 17:32 ` [RFC PATCH v2 06/16] vfio/pci: Gate interrupt configuration Shameer Kolothum
@ 2026-09-29 17:32 ` Shameer Kolothum
  2026-09-29 17:32 ` [RFC PATCH v2 08/16] vfio/pci: Gate device information queries and DMA-BUF export Shameer Kolothum
                   ` (8 subsequent siblings)
  15 siblings, 0 replies; 17+ messages in thread
From: Shameer Kolothum @ 2026-09-29 17:32 UTC (permalink / raw)
  To: kvm, linux-pci, linux-kernel
  Cc: alex, jgg, kevin.tian, kbusch, michal.winiarski,
	satyanarayana.k.v.p, sonangp, ankita, nathanc, mochs,
	skolothumtho

Reject function reset and low-power entry or exit while device access is
blocked. Hold access_srcu across reset and runtime power-management
updates so recovery waits for admitted operations to finish.

Perform the reset ioctl's D0 transition before taking SRCU, preserving
VFIO's pm_save handling. Keep memory_lock held across D0 preparation and
reset so a concurrent guest D3 write cannot intervene. Check the blocked
state before D0 and check SRCU admission afterward. A rejected reset does
not queue a deferred D0 request.

The reset uses pci_try_reset_function(), which returns -EAGAIN if AER
holds device_lock.

Assisted-by: LLM
Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com>
---
Note: D0 remains outside SRCU, but this series introduces a recovery
locking dependency involving memory_lock and pci_bus_sem through ASPM.
This remains unresolved. See the cover letter's "Locking and open
questions" section.
---
 drivers/vfio/pci/vfio_pci_priv.h   |  3 +++
 drivers/vfio/pci/vfio_pci_config.c |  4 ++--
 drivers/vfio/pci/vfio_pci_core.c   | 32 ++++++++++++++++++++++++++++--
 3 files changed, 35 insertions(+), 4 deletions(-)

diff --git a/drivers/vfio/pci/vfio_pci_priv.h b/drivers/vfio/pci/vfio_pci_priv.h
index 623e73f379cc..f757e5407ce4 100644
--- a/drivers/vfio/pci/vfio_pci_priv.h
+++ b/drivers/vfio/pci/vfio_pci_priv.h
@@ -41,6 +41,9 @@ ssize_t vfio_pci_config_rw_single(struct vfio_pci_core_device *vdev,
 				  char __user *buf, size_t count, loff_t *ppos,
 				  bool iswrite);
 
+int vfio_lock_and_set_power_state(struct vfio_pci_core_device *vdev,
+				 pci_power_t state);
+
 ssize_t vfio_pci_bar_rw(struct vfio_pci_core_device *vdev, char __user *buf,
 			size_t count, loff_t *ppos, bool iswrite);
 
diff --git a/drivers/vfio/pci/vfio_pci_config.c b/drivers/vfio/pci/vfio_pci_config.c
index e9480b573381..f3366df4af68 100644
--- a/drivers/vfio/pci/vfio_pci_config.c
+++ b/drivers/vfio/pci/vfio_pci_config.c
@@ -720,8 +720,8 @@ static int __init init_pci_cap_basic_perm(struct perm_bits *perm)
  * It takes all the required locks to protect the access of power related
  * variables and then invokes vfio_pci_set_power_state().
  */
-static int vfio_lock_and_set_power_state(struct vfio_pci_core_device *vdev,
-					 pci_power_t state)
+int vfio_lock_and_set_power_state(struct vfio_pci_core_device *vdev,
+				 pci_power_t state)
 {
 	if (state >= PCI_D3hot) {
 		vfio_pci_zap_and_down_write_memory_lock(vdev);
diff --git a/drivers/vfio/pci/vfio_pci_core.c b/drivers/vfio/pci/vfio_pci_core.c
index 91c92e8f3daa..38551bdedc89 100644
--- a/drivers/vfio/pci/vfio_pci_core.c
+++ b/drivers/vfio/pci/vfio_pci_core.c
@@ -372,6 +372,12 @@ int vfio_pci_set_power_state(struct vfio_pci_core_device *vdev, pci_power_t stat
 static int vfio_pci_runtime_pm_entry(struct vfio_pci_core_device *vdev,
 				     struct eventfd_ctx *efdctx)
 {
+	int idx;
+
+	idx = vfio_pci_core_access_begin(vdev);
+	if (idx < 0)
+		return idx;
+
 	/*
 	 * The vdev power related flags are protected with 'memory_lock'
 	 * semaphore.
@@ -381,6 +387,7 @@ static int vfio_pci_runtime_pm_entry(struct vfio_pci_core_device *vdev,
 
 	if (vdev->pm_runtime_engaged) {
 		up_write(&vdev->memory_lock);
+		vfio_pci_core_access_end(vdev, idx);
 		return -EINVAL;
 	}
 
@@ -388,6 +395,7 @@ static int vfio_pci_runtime_pm_entry(struct vfio_pci_core_device *vdev,
 	vdev->pm_wake_eventfd_ctx = efdctx;
 	pm_runtime_put_noidle(&vdev->pdev->dev);
 	up_write(&vdev->memory_lock);
+	vfio_pci_core_access_end(vdev, idx);
 
 	return 0;
 }
@@ -470,7 +478,7 @@ static void vfio_pci_runtime_pm_exit(struct vfio_pci_core_device *vdev)
 static int vfio_pci_core_pm_exit(struct vfio_pci_core_device *vdev, u32 flags,
 				 void __user *arg, size_t argsz)
 {
-	int ret;
+	int ret, idx;
 
 	ret = vfio_check_feature(flags, argsz, VFIO_DEVICE_FEATURE_SET, 0);
 	if (ret != 1)
@@ -483,7 +491,11 @@ static int vfio_pci_core_pm_exit(struct vfio_pci_core_device *vdev, u32 flags,
 	 * already signaled the eventfd and exited low power mode itself.
 	 * pm_runtime_engaged protects the redundant call here.
 	 */
+	idx = vfio_pci_core_access_begin(vdev);
+	if (idx < 0)
+		return idx;
 	vfio_pci_runtime_pm_exit(vdev);
+	vfio_pci_core_access_end(vdev, idx);
 	return 0;
 }
 
@@ -1387,12 +1399,16 @@ static int vfio_pci_ioctl_set_irqs(struct vfio_pci_core_device *vdev,
 static int vfio_pci_ioctl_reset(struct vfio_pci_core_device *vdev,
 				void __user *arg)
 {
-	int ret;
+	int ret, idx;
 
 	if (!vdev->reset_works)
 		return -EINVAL;
 
 	vfio_pci_zap_and_down_write_memory_lock(vdev);
+	if (vdev->pci_recovery_supported && READ_ONCE(vdev->access_blocked)) {
+		up_write(&vdev->memory_lock);
+		return -EIO;
+	}
 
 	/*
 	 * This function can be invoked while the power state is non-D0. If
@@ -1402,15 +1418,27 @@ static int vfio_pci_ioctl_reset(struct vfio_pci_core_device *vdev,
 	 * have NoSoftRst-, the reset function can cause the PCI config space
 	 * reset without restoring the original state (saved locally in
 	 * 'vdev->pm_save').
+	 *
+	 * Do the D0 transition outside access_srcu because it can acquire
+	 * pci_bus_sem through ASPM. Keep memory_lock held through reset to
+	 * prevent a concurrent power-state write from putting the device in D3.
 	 */
 	vfio_pci_set_power_state(vdev, PCI_D0);
 
+	idx = vfio_pci_core_access_begin(vdev);
+	if (idx < 0) {
+		up_write(&vdev->memory_lock);
+		return idx;
+	}
+
 	vfio_pci_dma_buf_move(vdev, true);
 	ret = pci_try_reset_function(vdev->pdev);
 	if (__vfio_pci_memory_enabled(vdev))
 		vfio_pci_dma_buf_move(vdev, false);
 	up_write(&vdev->memory_lock);
 
+	vfio_pci_core_access_end(vdev, idx);
+
 	return ret;
 }
 
-- 
2.43.0


^ permalink raw reply	[flat|nested] 17+ messages in thread

* [RFC PATCH v2 08/16] vfio/pci: Gate device information queries and DMA-BUF export
  2026-09-29 17:32 [RFC PATCH v2 00/16] vfio/pci: Handle PCI error recovery and report state to userspace Shameer Kolothum
                   ` (6 preceding siblings ...)
  2026-09-29 17:32 ` [RFC PATCH v2 07/16] vfio/pci: Gate function reset and runtime power management Shameer Kolothum
@ 2026-09-29 17:32 ` Shameer Kolothum
  2026-09-29 17:32 ` [RFC PATCH v2 09/16] vfio/pci: Add PCI error recovery state Shameer Kolothum
                   ` (7 subsequent siblings)
  15 siblings, 0 replies; 17+ messages in thread
From: Shameer Kolothum @ 2026-09-29 17:32 UTC (permalink / raw)
  To: kvm, linux-pci, linux-kernel
  Cc: alex, jgg, kevin.tian, kbusch, michal.winiarski,
	satyanarayana.k.v.p, sonangp, ankita, nathanc, mochs,
	skolothumtho

GET_INFO accesses PCI configuration to report AtomicOps support, and
GET_REGION_INFO maps the ROM to validate its contents. Protect these
hardware accesses with access_srcu.

Hold SRCU through DMA-BUF creation and list insertion so recovery cannot
finish revoking buffers before a concurrent export becomes visible.

Assisted-by: LLM
Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com>
---
 drivers/vfio/pci/vfio_pci_core.c   | 13 ++++++++++++-
 drivers/vfio/pci/vfio_pci_dmabuf.c | 16 +++++++++++++---
 2 files changed, 25 insertions(+), 4 deletions(-)

diff --git a/drivers/vfio/pci/vfio_pci_core.c b/drivers/vfio/pci/vfio_pci_core.c
index 38551bdedc89..92497224f471 100644
--- a/drivers/vfio/pci/vfio_pci_core.c
+++ b/drivers/vfio/pci/vfio_pci_core.c
@@ -1112,7 +1112,7 @@ static int vfio_pci_ioctl_get_info(struct vfio_pci_core_device *vdev,
 	unsigned long minsz = offsetofend(struct vfio_device_info, num_irqs);
 	struct vfio_device_info info = {};
 	struct vfio_info_cap caps = { .buf = NULL, .size = 0 };
-	int ret;
+	int ret, idx;
 
 	if (copy_from_user(&info, arg, minsz))
 		return -EFAULT;
@@ -1137,7 +1137,13 @@ static int vfio_pci_ioctl_get_info(struct vfio_pci_core_device *vdev,
 		return ret;
 	}
 
+	idx = vfio_pci_core_access_begin(vdev);
+	if (idx < 0) {
+		kfree(caps.buf);
+		return idx;
+	}
 	ret = vfio_pci_info_atomic_cap(vdev, &caps);
+	vfio_pci_core_access_end(vdev, idx);
 	if (ret && ret != -ENODEV) {
 		pci_warn(vdev->pdev,
 			 "Failed to setup AtomicOps info capability\n");
@@ -1213,6 +1219,10 @@ int vfio_pci_ioctl_get_region_info(struct vfio_device *core_vdev,
 			 * Check ROM content is valid. Need to enable memory
 			 * decode for ROM access in pci_map_rom().
 			 */
+			int idx = vfio_pci_core_access_begin(vdev);
+
+			if (idx < 0)
+				return idx;
 			cmd = vfio_pci_memory_lock_and_enable(vdev);
 			io = pci_map_rom(pdev, &size);
 			if (io) {
@@ -1223,6 +1233,7 @@ int vfio_pci_ioctl_get_region_info(struct vfio_device *core_vdev,
 				pci_unmap_rom(pdev, io);
 			}
 			vfio_pci_memory_unlock_and_restore(vdev, cmd);
+			vfio_pci_core_access_end(vdev, idx);
 		} else if (pdev->rom && pdev->romlen) {
 			info->flags = VFIO_REGION_INFO_FLAG_READ;
 			/* Report BAR size as power of two. */
diff --git a/drivers/vfio/pci/vfio_pci_dmabuf.c b/drivers/vfio/pci/vfio_pci_dmabuf.c
index c16f460c01d6..c1a0250af680 100644
--- a/drivers/vfio/pci/vfio_pci_dmabuf.c
+++ b/drivers/vfio/pci/vfio_pci_dmabuf.c
@@ -227,7 +227,7 @@ int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,
 	DEFINE_DMA_BUF_EXPORT_INFO(exp_info);
 	struct vfio_pci_dma_buf *priv;
 	size_t length;
-	int ret;
+	int ret, idx;
 
 	if (!vdev->pci_ops || !vdev->pci_ops->get_dmabuf_phys)
 		return -EOPNOTSUPP;
@@ -274,19 +274,26 @@ int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,
 	priv->vdev = vdev;
 	priv->nr_ranges = get_dma_buf.nr_ranges;
 	priv->size = length;
+
+	idx = vfio_pci_core_access_begin(vdev);
+	if (idx < 0) {
+		ret = idx;
+		goto err_free_phys;
+	}
+
 	ret = vdev->pci_ops->get_dmabuf_phys(vdev, &priv->provider,
 					     get_dma_buf.region_index,
 					     priv->phys_vec, dma_ranges,
 					     priv->nr_ranges);
 	if (ret)
-		goto err_free_phys;
+		goto err_access;
 
 	kfree(dma_ranges);
 	dma_ranges = NULL;
 
 	if (!vfio_device_try_get_registration(&vdev->vdev)) {
 		ret = -ENODEV;
-		goto err_free_phys;
+		goto err_access;
 	}
 
 	exp_info.ops = &vfio_pci_dmabuf_ops;
@@ -311,6 +318,7 @@ int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,
 	list_add_tail(&priv->dmabufs_elm, &vdev->dmabufs);
 	dma_resv_unlock(priv->dmabuf->resv);
 	up_write(&vdev->memory_lock);
+	vfio_pci_core_access_end(vdev, idx);
 
 	/*
 	 * dma_buf_fd() consumes the reference, when the file closes the dmabuf
@@ -324,6 +332,8 @@ int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,
 
 err_dev_put:
 	vfio_device_put_registration(&vdev->vdev);
+err_access:
+	vfio_pci_core_access_end(vdev, idx);
 err_free_phys:
 	kfree(priv->phys_vec);
 err_free_priv:
-- 
2.43.0


^ permalink raw reply	[flat|nested] 17+ messages in thread

* [RFC PATCH v2 09/16] vfio/pci: Add PCI error recovery state
  2026-09-29 17:32 [RFC PATCH v2 00/16] vfio/pci: Handle PCI error recovery and report state to userspace Shameer Kolothum
                   ` (7 preceding siblings ...)
  2026-09-29 17:32 ` [RFC PATCH v2 08/16] vfio/pci: Gate device information queries and DMA-BUF export Shameer Kolothum
@ 2026-09-29 17:32 ` Shameer Kolothum
  2026-09-29 17:32 ` [RFC PATCH v2 10/16] vfio/pci: Quiesce INTx while access is blocked Shameer Kolothum
                   ` (6 subsequent siblings)
  15 siblings, 0 replies; 17+ messages in thread
From: Shameer Kolothum @ 2026-09-29 17:32 UTC (permalink / raw)
  To: kvm, linux-pci, linux-kernel
  Cc: alex, jgg, kevin.tian, kbusch, michal.winiarski,
	satyanarayana.k.v.p, sonangp, ankita, nathanc, mochs,
	skolothumtho

Add recovery flags, a sequence number, a saved PCI_COMMAND value and an
eventfd for recovery notifications. Initialize the per-open state at
open and clear it at close.

Track the host transaction separately from the userspace session. The
host-active flag survives close so a later patch can reject reopen until
the outstanding recovery completes.

Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com>
---
 include/linux/vfio_pci_core.h    | 19 ++++++++++++++++++-
 drivers/vfio/pci/vfio_pci_core.c | 11 +++++++++++
 2 files changed, 29 insertions(+), 1 deletion(-)

diff --git a/include/linux/vfio_pci_core.h b/include/linux/vfio_pci_core.h
index 1dc9630dc740..61d21c3f8089 100644
--- a/include/linux/vfio_pci_core.h
+++ b/include/linux/vfio_pci_core.h
@@ -96,6 +96,11 @@ static inline int vfio_pci_core_get_dmabuf_phys(
 }
 #endif
 
+#define VFIO_PCI_RECOVERY_IN_PROGRESS	BIT(0)
+#define VFIO_PCI_RECOVERY_FROZEN	BIT(1)
+#define VFIO_PCI_RECOVERY_RESET		BIT(2)
+#define VFIO_PCI_RECOVERY_FAILED	BIT(3)
+
 struct vfio_pci_core_device {
 	struct vfio_device	vdev;
 	struct pci_dev		*pdev;
@@ -143,6 +148,7 @@ struct vfio_pci_core_device {
 	int			ioeventfds_nr;
 	struct vfio_pci_eventfd __rcu *err_trigger;
 	struct vfio_pci_eventfd __rcu *req_trigger;
+	struct vfio_pci_eventfd __rcu *pci_recovery_trigger;
 	struct eventfd_ctx	*pm_wake_eventfd_ctx;
 	struct list_head	dummy_resources_list;
 	struct mutex		ioeventfds_lock;
@@ -168,7 +174,18 @@ struct vfio_pci_core_device {
 	bool			access_blocked;
 	/* Set after open completes, cleared before close tears down state. */
 	bool			device_open;
-	struct mutex		access_lock;	/* gate flag writers */
+	struct mutex		access_lock;	/* gate and recovery state writers */
+	u32			pci_recovery_flags;
+	u64			pci_recovery_sequence;
+	/* Saved PCI_COMMAND, valid when pci_recovery_command_valid is set. */
+	u16			pci_recovery_command;
+	bool			pci_recovery_enabled;
+	bool			pci_recovery_command_valid;
+	/*
+	 * Host recovery in progress, protected by access_lock. Unlike the
+	 * per-open recovery flags, this is preserved across close.
+	 */
+	bool			pci_recovery_host_active;
 	struct list_head	dmabufs;
 };
 
diff --git a/drivers/vfio/pci/vfio_pci_core.c b/drivers/vfio/pci/vfio_pci_core.c
index 92497224f471..667c5813f6c7 100644
--- a/drivers/vfio/pci/vfio_pci_core.c
+++ b/drivers/vfio/pci/vfio_pci_core.c
@@ -845,8 +845,11 @@ void vfio_pci_core_close_device(struct vfio_device *core_vdev)
 
 	if (vdev->pci_recovery_supported) {
 		scoped_guard(mutex, &vdev->access_lock) {
+			WRITE_ONCE(vdev->pci_recovery_enabled, false);
+			vdev->pci_recovery_command_valid = false;
 			WRITE_ONCE(vdev->access_blocked, false);
 			WRITE_ONCE(vdev->device_open, false);
+			WRITE_ONCE(vdev->pci_recovery_flags, 0);
 		}
 	}
 
@@ -866,6 +869,10 @@ void vfio_pci_core_close_device(struct vfio_device *core_vdev)
 	mutex_lock(&vdev->igate);
 	vfio_pci_eventfd_replace_locked(vdev, &vdev->err_trigger, NULL);
 	vfio_pci_eventfd_replace_locked(vdev, &vdev->req_trigger, NULL);
+	if (vdev->pci_recovery_supported)
+		vfio_pci_eventfd_replace_locked(vdev,
+						&vdev->pci_recovery_trigger,
+						NULL);
 	mutex_unlock(&vdev->igate);
 }
 EXPORT_SYMBOL_GPL(vfio_pci_core_close_device);
@@ -885,6 +892,10 @@ void vfio_pci_core_finish_enable(struct vfio_pci_core_device *vdev)
 
 	if (vdev->pci_recovery_supported) {
 		guard(mutex)(&vdev->access_lock);
+		WRITE_ONCE(vdev->pci_recovery_flags, 0);
+		vdev->pci_recovery_sequence = 0;
+		vdev->pci_recovery_command_valid = false;
+		WRITE_ONCE(vdev->pci_recovery_enabled, false);
 		WRITE_ONCE(vdev->access_blocked, false);
 		WRITE_ONCE(vdev->device_open, true);
 	}
-- 
2.43.0


^ permalink raw reply	[flat|nested] 17+ messages in thread

* [RFC PATCH v2 10/16] vfio/pci: Quiesce INTx while access is blocked
  2026-09-29 17:32 [RFC PATCH v2 00/16] vfio/pci: Handle PCI error recovery and report state to userspace Shameer Kolothum
                   ` (8 preceding siblings ...)
  2026-09-29 17:32 ` [RFC PATCH v2 09/16] vfio/pci: Add PCI error recovery state Shameer Kolothum
@ 2026-09-29 17:32 ` Shameer Kolothum
  2026-09-29 17:33 ` [RFC PATCH v2 11/16] vfio/pci: Add INTx recovery start and finish helpers Shameer Kolothum
                   ` (5 subsequent siblings)
  15 siblings, 0 replies; 17+ messages in thread
From: Shameer Kolothum @ 2026-09-29 17:32 UTC (permalink / raw)
  To: kvm, linux-pci, linux-kernel
  Cc: alex, jgg, kevin.tian, kbusch, michal.winiarski,
	satyanarayana.k.v.p, sonangp, ankita, nathanc, mochs,
	skolothumtho

The INTx handler and unmask virqfd can access PCI_COMMAND without taking
access_srcu. Check the recovery state in these paths under irqlock.

Mask pending INTx while access is blocked to prevent an interrupt storm.
For shared lines, claim only interrupts identified by
pci_check_and_mask_intx(). Record unmask requests for replay when recovery
completes.

Assisted-by: LLM
Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com>
---
 drivers/vfio/pci/vfio_pci_intrs.c | 45 +++++++++++++++++++++++++++++++
 1 file changed, 45 insertions(+)

diff --git a/drivers/vfio/pci/vfio_pci_intrs.c b/drivers/vfio/pci/vfio_pci_intrs.c
index 64f80f64ff57..ff29377c1ae8 100644
--- a/drivers/vfio/pci/vfio_pci_intrs.c
+++ b/drivers/vfio/pci/vfio_pci_intrs.c
@@ -29,6 +29,7 @@ struct vfio_pci_irq_ctx {
 	struct virqfd			*mask;
 	char				*name;
 	bool				masked;
+	bool				unmask_pending;
 	struct irq_bypass_producer	producer;
 };
 
@@ -49,6 +50,15 @@ static bool is_irq_none(struct vfio_pci_core_device *vdev)
 		 vdev->irq_type == VFIO_PCI_MSIX_IRQ_INDEX);
 }
 
+static bool vfio_pci_recovery_blocks_irq(struct vfio_pci_core_device *vdev)
+{
+	if (!vdev->pci_recovery_supported)
+		return false;
+
+	return READ_ONCE(vdev->pci_recovery_enabled) &&
+	       READ_ONCE(vdev->access_blocked);
+}
+
 static
 struct vfio_pci_irq_ctx *vfio_irq_ctx_get(struct vfio_pci_core_device *vdev,
 					  unsigned long index)
@@ -171,6 +181,12 @@ static int vfio_pci_intx_unmask_handler(void *opaque, void *data)
 	int ret = 0;
 
 	spin_lock_irqsave(&vdev->irqlock, flags);
+	/* Defer unmasking until recovery completes. */
+	if (unlikely(vfio_pci_recovery_blocks_irq(vdev))) {
+		if (is_intx(vdev))
+			ctx->unmask_pending = true;
+		goto out_unlock;
+	}
 
 	/*
 	 * Unmasking comes from ioctl or config, so again, have the
@@ -182,6 +198,8 @@ static int vfio_pci_intx_unmask_handler(void *opaque, void *data)
 		goto out_unlock;
 	}
 
+	ctx->unmask_pending = false;
+
 	if (ctx->masked && !vdev->virq_disabled) {
 		/*
 		 * A pending interrupt here would immediately trigger,
@@ -220,6 +238,24 @@ void vfio_pci_intx_unmask(struct vfio_pci_core_device *vdev)
 	mutex_unlock(&vdev->igate);
 }
 
+/* Mask INTx for a blocked device. Returns true if this call masked it. */
+static bool vfio_pci_intx_mask_for_recovery(struct vfio_pci_core_device *vdev,
+					    struct vfio_pci_irq_ctx *ctx)
+{
+	lockdep_assert_held(&vdev->irqlock);
+
+	if (ctx->masked)
+		return false;
+
+	if (!vdev->pci_2_3)
+		disable_irq_nosync(vdev->pdev->irq);
+	else if (!pci_check_and_mask_intx(vdev->pdev))
+		return false;
+
+	ctx->masked = true;
+	return true;
+}
+
 static irqreturn_t vfio_intx_handler(int irq, void *dev_id)
 {
 	struct vfio_pci_irq_ctx *ctx = dev_id;
@@ -228,6 +264,14 @@ static irqreturn_t vfio_intx_handler(int irq, void *dev_id)
 	int ret = IRQ_NONE;
 
 	spin_lock_irqsave(&vdev->irqlock, flags);
+	if (unlikely(vfio_pci_recovery_blocks_irq(vdev))) {
+		/* Mask pending INTx to prevent an interrupt storm. */
+		if (vfio_pci_intx_mask_for_recovery(vdev, ctx))
+			ret = IRQ_HANDLED;
+		else if (ctx->masked && !vdev->pci_2_3)
+			ret = IRQ_HANDLED;
+		goto out_unlock;
+	}
 
 	if (!vdev->pci_2_3) {
 		disable_irq_nosync(vdev->pdev->irq);
@@ -239,6 +283,7 @@ static irqreturn_t vfio_intx_handler(int irq, void *dev_id)
 		ret = IRQ_HANDLED;
 	}
 
+out_unlock:
 	spin_unlock_irqrestore(&vdev->irqlock, flags);
 
 	if (ret == IRQ_HANDLED)
-- 
2.43.0


^ permalink raw reply	[flat|nested] 17+ messages in thread

* [RFC PATCH v2 11/16] vfio/pci: Add INTx recovery start and finish helpers
  2026-09-29 17:32 [RFC PATCH v2 00/16] vfio/pci: Handle PCI error recovery and report state to userspace Shameer Kolothum
                   ` (9 preceding siblings ...)
  2026-09-29 17:32 ` [RFC PATCH v2 10/16] vfio/pci: Quiesce INTx while access is blocked Shameer Kolothum
@ 2026-09-29 17:33 ` Shameer Kolothum
  2026-09-29 17:33 ` [RFC PATCH v2 12/16] vfio/pci: Restore device state from slot_reset() Shameer Kolothum
                   ` (4 subsequent siblings)
  15 siblings, 0 replies; 17+ messages in thread
From: Shameer Kolothum @ 2026-09-29 17:33 UTC (permalink / raw)
  To: kvm, linux-pci, linux-kernel
  Cc: alex, jgg, kevin.tian, kbusch, michal.winiarski,
	satyanarayana.k.v.p, sonangp, ankita, nathanc, mochs,
	skolothumtho

Add helpers to mask INTx at recovery entry and restore it on completion.
Mask INTx at recovery entry even when no interrupt is pending. Record
whether recovery changed the mask so the completion helper can undo it
and apply deferred unmask requests.

Preserve the current INTx mask when restoring PCI_COMMAND. Only update
INTX_DISABLE for devices supporting that bit; other devices use the IRQ
controller mask. Run the completion helper after access is unblocked.

An explicit INTx mask cancels pending recovery unmasking, even if the
interrupt is already masked. This preserves a mask request made after
access reopens but before the recovery completion helper runs.

Assisted-by: LLM
Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com>
---
 drivers/vfio/pci/vfio_pci_priv.h  |  4 ++
 drivers/vfio/pci/vfio_pci_intrs.c | 92 ++++++++++++++++++++++++++++++-
 2 files changed, 93 insertions(+), 3 deletions(-)

diff --git a/drivers/vfio/pci/vfio_pci_priv.h b/drivers/vfio/pci/vfio_pci_priv.h
index f757e5407ce4..387a831f23c0 100644
--- a/drivers/vfio/pci/vfio_pci_priv.h
+++ b/drivers/vfio/pci/vfio_pci_priv.h
@@ -25,6 +25,10 @@ struct vfio_pci_ioeventfd {
 
 bool vfio_pci_intx_mask(struct vfio_pci_core_device *vdev);
 void vfio_pci_intx_unmask(struct vfio_pci_core_device *vdev);
+void vfio_pci_intx_recovery_start(struct vfio_pci_core_device *vdev);
+void vfio_pci_intx_recovery_finish(struct vfio_pci_core_device *vdev);
+u16 vfio_pci_intx_recovery_update_command(struct vfio_pci_core_device *vdev,
+					  u16 command);
 
 int vfio_pci_eventfd_replace_locked(struct vfio_pci_core_device *vdev,
 				    struct vfio_pci_eventfd __rcu **peventfd,
diff --git a/drivers/vfio/pci/vfio_pci_intrs.c b/drivers/vfio/pci/vfio_pci_intrs.c
index ff29377c1ae8..d78260776ac1 100644
--- a/drivers/vfio/pci/vfio_pci_intrs.c
+++ b/drivers/vfio/pci/vfio_pci_intrs.c
@@ -29,6 +29,7 @@ struct vfio_pci_irq_ctx {
 	struct virqfd			*mask;
 	char				*name;
 	bool				masked;
+	bool				recovery_masked;
 	bool				unmask_pending;
 	struct irq_bypass_producer	producer;
 };
@@ -136,6 +137,10 @@ static bool __vfio_pci_intx_mask(struct vfio_pci_core_device *vdev)
 	if (WARN_ON_ONCE(!ctx))
 		goto out_unlock;
 
+	/* An explicit mask supersedes any pending recovery unmask. */
+	ctx->recovery_masked = false;
+	ctx->unmask_pending = false;
+
 	if (!ctx->masked) {
 		/*
 		 * Can't use check_and_mask here because we always want to
@@ -199,6 +204,7 @@ static int vfio_pci_intx_unmask_handler(void *opaque, void *data)
 	}
 
 	ctx->unmask_pending = false;
+	ctx->recovery_masked = false;
 
 	if (ctx->masked && !vdev->virq_disabled) {
 		/*
@@ -238,9 +244,14 @@ void vfio_pci_intx_unmask(struct vfio_pci_core_device *vdev)
 	mutex_unlock(&vdev->igate);
 }
 
-/* Mask INTx for a blocked device. Returns true if this call masked it. */
+/*
+ * Recovery masks INTx unconditionally and records whether to unmask it.
+ * The interrupt handler only masks a pending interrupt, since another
+ * device may share the line.
+ */
 static bool vfio_pci_intx_mask_for_recovery(struct vfio_pci_core_device *vdev,
-					    struct vfio_pci_irq_ctx *ctx)
+					    struct vfio_pci_irq_ctx *ctx,
+					    bool quiesce)
 {
 	lockdep_assert_held(&vdev->irqlock);
 
@@ -249,10 +260,14 @@ static bool vfio_pci_intx_mask_for_recovery(struct vfio_pci_core_device *vdev,
 
 	if (!vdev->pci_2_3)
 		disable_irq_nosync(vdev->pdev->irq);
+	else if (quiesce)
+		pci_intx(vdev->pdev, 0);
 	else if (!pci_check_and_mask_intx(vdev->pdev))
 		return false;
 
 	ctx->masked = true;
+	if (quiesce)
+		ctx->recovery_masked = true;
 	return true;
 }
 
@@ -266,7 +281,7 @@ static irqreturn_t vfio_intx_handler(int irq, void *dev_id)
 	spin_lock_irqsave(&vdev->irqlock, flags);
 	if (unlikely(vfio_pci_recovery_blocks_irq(vdev))) {
 		/* Mask pending INTx to prevent an interrupt storm. */
-		if (vfio_pci_intx_mask_for_recovery(vdev, ctx))
+		if (vfio_pci_intx_mask_for_recovery(vdev, ctx, false))
 			ret = IRQ_HANDLED;
 		else if (ctx->masked && !vdev->pci_2_3)
 			ret = IRQ_HANDLED;
@@ -292,6 +307,77 @@ static irqreturn_t vfio_intx_handler(int irq, void *dev_id)
 	return ret;
 }
 
+void vfio_pci_intx_recovery_start(struct vfio_pci_core_device *vdev)
+{
+	struct vfio_pci_irq_ctx *ctx;
+	unsigned long flags;
+
+	lockdep_assert_held(&vdev->access_lock);
+
+	spin_lock_irqsave(&vdev->irqlock, flags);
+	if (!is_intx(vdev))
+		goto out_unlock;
+
+	ctx = vfio_irq_ctx_get(vdev, 0);
+	if (WARN_ON_ONCE(!ctx))
+		goto out_unlock;
+
+	vfio_pci_intx_mask_for_recovery(vdev, ctx, true);
+
+out_unlock:
+	spin_unlock_irqrestore(&vdev->irqlock, flags);
+}
+
+/* Preserve the current INTx mask when restoring PCI_COMMAND. */
+u16 vfio_pci_intx_recovery_update_command(struct vfio_pci_core_device *vdev,
+					  u16 command)
+{
+	struct vfio_pci_irq_ctx *ctx;
+
+	lockdep_assert_held(&vdev->irqlock);
+
+	if (!vdev->pci_2_3 || !is_intx(vdev))
+		return command;
+
+	ctx = vfio_irq_ctx_get(vdev, 0);
+	if (ctx && ctx->masked)
+		command |= PCI_COMMAND_INTX_DISABLE;
+
+	return command;
+}
+
+/*
+ * Undo the recovery mask and apply deferred unmask requests.
+ * Call after clearing access_blocked so unmasking is not deferred again.
+ */
+void vfio_pci_intx_recovery_finish(struct vfio_pci_core_device *vdev)
+{
+	struct vfio_pci_irq_ctx *ctx;
+	unsigned long flags;
+	bool replay = false;
+
+	lockdep_assert_held(&vdev->access_lock);
+
+	mutex_lock(&vdev->igate);
+	spin_lock_irqsave(&vdev->irqlock, flags);
+	if (!is_intx(vdev))
+		goto out_unlock;
+
+	ctx = vfio_irq_ctx_get(vdev, 0);
+	if (WARN_ON_ONCE(!ctx))
+		goto out_unlock;
+
+	replay = ctx->recovery_masked || ctx->unmask_pending;
+	ctx->recovery_masked = false;
+	ctx->unmask_pending = false;
+
+out_unlock:
+	spin_unlock_irqrestore(&vdev->irqlock, flags);
+	if (replay)
+		__vfio_pci_intx_unmask(vdev);
+	mutex_unlock(&vdev->igate);
+}
+
 static int vfio_intx_enable(struct vfio_pci_core_device *vdev,
 			    struct eventfd_ctx *trigger)
 {
-- 
2.43.0


^ permalink raw reply	[flat|nested] 17+ messages in thread

* [RFC PATCH v2 12/16] vfio/pci: Restore device state from slot_reset()
  2026-09-29 17:32 [RFC PATCH v2 00/16] vfio/pci: Handle PCI error recovery and report state to userspace Shameer Kolothum
                   ` (10 preceding siblings ...)
  2026-09-29 17:33 ` [RFC PATCH v2 11/16] vfio/pci: Add INTx recovery start and finish helpers Shameer Kolothum
@ 2026-09-29 17:33 ` Shameer Kolothum
  2026-09-29 17:33 ` [RFC PATCH v2 13/16] vfio/pci: Complete recovery in resume() Shameer Kolothum
                   ` (3 subsequent siblings)
  15 siblings, 0 replies; 17+ messages in thread
From: Shameer Kolothum @ 2026-09-29 17:33 UTC (permalink / raw)
  To: kvm, linux-pci, linux-kernel
  Cc: alex, jgg, kevin.tian, kbusch, michal.winiarski,
	satyanarayana.k.v.p, sonangp, ankita, nathanc, mochs,
	skolothumtho

A host reset clears device configuration. Add slot_reset() to restore
the state saved at open before tearing down interrupts. MSI-X shutdown
requires the restored BARs to access its table.

Load vdev->pci_saved_state before entering D0, since the power transition
may restore BARs. PCI core's saved copy may have been overwritten by a
user reset or D3 entry. Discard the stale pre-reset pm_save and update
the power state without restoring it. Keep access blocked and report
failure if the snapshot is missing or restoration fails.

Keep active INTx masked in the loaded snapshot so restoration cannot
re-enable it before IRQ teardown.

Use pci_set_power_state() as existing recovery callbacks do.

Hold access_lock to exclude close's interrupt teardown. Return NONE on
failure to allow recovery of other devices under the bridge to continue.

Assisted-by: LLM
Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com>
---
Note: This series adds a VFIO recovery callback that can reacquire
pci_bus_sem through the D0 ASPM update while AER already holds it. This
can deadlock behind a queued writer. Similar callback patterns in other
drivers do not resolve this new VFIO locking issue; topology does not
establish lock ownership. See the cover letter's "Locking and open
questions" section.
---
 drivers/vfio/pci/vfio_pci_core.c | 101 +++++++++++++++++++++++++++++++
 1 file changed, 101 insertions(+)

diff --git a/drivers/vfio/pci/vfio_pci_core.c b/drivers/vfio/pci/vfio_pci_core.c
index 667c5813f6c7..a7b7499e071c 100644
--- a/drivers/vfio/pci/vfio_pci_core.c
+++ b/drivers/vfio/pci/vfio_pci_core.c
@@ -311,6 +311,15 @@ static void vfio_pci_probe_power_state(struct vfio_pci_core_device *vdev)
 	vdev->needs_pm_restore = !(pmcsr & PCI_PM_CTRL_NO_SOFT_RESET);
 }
 
+/* Keep recovery's INTx mask when restoring a loaded PCI state snapshot. */
+static void vfio_pci_recovery_mask_saved_intx(struct vfio_pci_core_device *vdev)
+{
+	if (READ_ONCE(vdev->access_blocked) && vdev->pci_2_3 &&
+	    READ_ONCE(vdev->irq_type) == VFIO_PCI_INTX_IRQ_INDEX)
+		vdev->pdev->saved_config_space[PCI_COMMAND / 4] |=
+			PCI_COMMAND_INTX_DISABLE;
+}
+
 /*
  * pci_set_power_state() wrapper handling devices which perform a soft reset on
  * D3->D0 transition.  Save state prior to D0/1/2->D3, stash it on the vdev,
@@ -2499,6 +2508,18 @@ void vfio_pci_core_unregister_device(struct vfio_pci_core_device *vdev)
 }
 EXPORT_SYMBOL_GPL(vfio_pci_core_unregister_device);
 
+static void
+vfio_pci_signal_recovery_event(struct vfio_pci_core_device *vdev)
+{
+	struct vfio_pci_eventfd *eventfd;
+
+	rcu_read_lock();
+	eventfd = rcu_dereference(vdev->pci_recovery_trigger);
+	if (eventfd)
+		eventfd_signal(eventfd->ctx);
+	rcu_read_unlock();
+}
+
 pci_ers_result_t vfio_pci_core_aer_err_detected(struct pci_dev *pdev,
 						pci_channel_state_t state)
 {
@@ -2515,6 +2536,85 @@ pci_ers_result_t vfio_pci_core_aer_err_detected(struct pci_dev *pdev,
 }
 EXPORT_SYMBOL_GPL(vfio_pci_core_aer_err_detected);
 
+/* Caller holds memory_lock. Discard the pre-reset PM snapshot. */
+static int vfio_pci_recovery_restore_state(struct vfio_pci_core_device *vdev)
+{
+	struct pci_dev *pdev = vdev->pdev;
+	int ret;
+
+	if (!vdev->pci_saved_state)
+		return -ENODATA;
+	ret = pci_load_saved_state(pdev, vdev->pci_saved_state);
+	if (ret)
+		return ret;
+
+	kfree(vdev->pm_save);
+	vdev->pm_save = NULL;
+	ret = pci_set_power_state(pdev, PCI_D0);
+	if (ret)
+		return ret;
+
+	vfio_pci_recovery_mask_saved_intx(vdev);
+	pci_restore_state(pdev);
+	return 0;
+}
+
+static pci_ers_result_t vfio_pci_core_aer_slot_reset(struct pci_dev *pdev)
+{
+	struct vfio_pci_core_device *vdev = dev_get_drvdata(&pdev->dev);
+	int ret = 0;
+
+	mutex_lock(&vdev->access_lock);
+	if (!(vdev->pci_recovery_flags & VFIO_PCI_RECOVERY_IN_PROGRESS) ||
+	    !vdev->device_open) {
+		mutex_unlock(&vdev->access_lock);
+		return PCI_ERS_RESULT_NONE;
+	}
+
+	/*
+	 * Load the open snapshot before D0, which may restore BARs from
+	 * saved state. Restore BARs before IRQ teardown so MSI-X shutdown
+	 * can access the table after the host reset.
+	 */
+	down_write(&vdev->memory_lock);
+	ret = vfio_pci_recovery_restore_state(vdev);
+	up_write(&vdev->memory_lock);
+	if (ret)
+		goto out_failed;
+
+	/*
+	 * Close clears device_open under access_lock before IRQ teardown,
+	 * excluding this callback from the teardown path.
+	 */
+	mutex_lock(&vdev->igate);
+	if (vdev->irq_type < VFIO_PCI_NUM_IRQS)
+		ret = vfio_pci_set_irqs_ioctl(vdev,
+					      VFIO_IRQ_SET_DATA_NONE |
+					      VFIO_IRQ_SET_ACTION_TRIGGER,
+					      vdev->irq_type, 0, 0, NULL);
+	mutex_unlock(&vdev->igate);
+	if (ret)
+		goto out_failed;
+
+	WRITE_ONCE(vdev->pci_recovery_flags,
+		   vdev->pci_recovery_flags | VFIO_PCI_RECOVERY_RESET);
+	mutex_unlock(&vdev->access_lock);
+
+	return PCI_ERS_RESULT_RECOVERED;
+
+out_failed:
+	WRITE_ONCE(vdev->pci_recovery_flags,
+		   (vdev->pci_recovery_flags | VFIO_PCI_RECOVERY_FAILED) &
+		   ~VFIO_PCI_RECOVERY_IN_PROGRESS);
+	vdev->pci_recovery_command_valid = false;
+	mutex_unlock(&vdev->access_lock);
+	/* Report failure here; resume() skips completed transactions. */
+	vfio_pci_signal_recovery_event(vdev);
+
+	/* Allow recovery of other devices under the bridge to continue. */
+	return PCI_ERS_RESULT_NONE;
+}
+
 int vfio_pci_core_sriov_configure(struct vfio_pci_core_device *vdev,
 				  int nr_virtfn)
 {
@@ -2587,6 +2687,7 @@ EXPORT_SYMBOL_GPL(vfio_pci_core_sriov_configure);
 
 const struct pci_error_handlers vfio_pci_core_err_handlers = {
 	.error_detected = vfio_pci_core_aer_err_detected,
+	.slot_reset = vfio_pci_core_aer_slot_reset,
 };
 EXPORT_SYMBOL_GPL(vfio_pci_core_err_handlers);
 
-- 
2.43.0


^ permalink raw reply	[flat|nested] 17+ messages in thread

* [RFC PATCH v2 13/16] vfio/pci: Complete recovery in resume()
  2026-09-29 17:32 [RFC PATCH v2 00/16] vfio/pci: Handle PCI error recovery and report state to userspace Shameer Kolothum
                   ` (11 preceding siblings ...)
  2026-09-29 17:33 ` [RFC PATCH v2 12/16] vfio/pci: Restore device state from slot_reset() Shameer Kolothum
@ 2026-09-29 17:33 ` Shameer Kolothum
  2026-09-29 17:33 ` [RFC PATCH v2 14/16] vfio/pci: Block device access during host recovery Shameer Kolothum
                   ` (2 subsequent siblings)
  15 siblings, 0 replies; 17+ messages in thread
From: Shameer Kolothum @ 2026-09-29 17:33 UTC (permalink / raw)
  To: kvm, linux-pci, linux-kernel
  Cc: alex, jgg, kevin.tian, kbusch, michal.winiarski,
	satyanarayana.k.v.p, sonangp, ankita, nathanc, mochs,
	skolothumtho

Add resume() to complete the host transaction and notify userspace.
For recovery without reset, apply a deferred D0 request before restoring
PCI_COMMAND. After a reset, slot_reset() has already established D0 and
restored the open snapshot. Use the existing VFIO power helper to retain
pm_save handling. Keep active INTx masked when restoring pm_save while
access is blocked, until recovery completes.

After restoring device state, disable ROM decoding unless the ROM
resource is marked to remain enabled. Then unblock access, restore
DMA-BUF availability and complete INTx unmasking. Protect PCI_COMMAND
and the unblock operation with irqlock. A power or command restore
failure leaves access blocked and sets FAILED.

Clear the host-active flag even if close or a local failure has already
cleared IN_PROGRESS. In that case, skip the per-open completion work.

Assisted-by: LLM
Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com>
---
Note: As in slot_reset(), this series introduces a VFIO recovery path
that can reacquire pci_bus_sem through the ordinary PCI power helper's
ASPM update. This remains unresolved. See the cover letter's "Locking
and open questions" section.
---
 drivers/vfio/pci/vfio_pci_core.c | 83 ++++++++++++++++++++++++++++++++
 1 file changed, 83 insertions(+)

diff --git a/drivers/vfio/pci/vfio_pci_core.c b/drivers/vfio/pci/vfio_pci_core.c
index a7b7499e071c..d6cc34240b25 100644
--- a/drivers/vfio/pci/vfio_pci_core.c
+++ b/drivers/vfio/pci/vfio_pci_core.c
@@ -371,6 +371,7 @@ int vfio_pci_set_power_state(struct vfio_pci_core_device *vdev, pci_power_t stat
 			vdev->pm_save = pci_store_saved_state(pdev);
 		} else if (needs_restore) {
 			pci_load_and_free_saved_state(pdev, &vdev->pm_save);
+			vfio_pci_recovery_mask_saved_intx(vdev);
 			pci_restore_state(pdev);
 		}
 	}
@@ -2615,6 +2616,87 @@ static pci_ers_result_t vfio_pci_core_aer_slot_reset(struct pci_dev *pdev)
 	return PCI_ERS_RESULT_NONE;
 }
 
+static void vfio_pci_core_aer_resume(struct pci_dev *pdev)
+{
+	struct vfio_pci_core_device *vdev = dev_get_drvdata(&pdev->dev);
+	unsigned long irq_flags;
+	bool notify_recovery = false;
+	bool power_up;
+	u32 flags;
+	int ret = 0;
+
+	mutex_lock(&vdev->access_lock);
+	vdev->pci_recovery_host_active = false;
+	if (!(vdev->pci_recovery_flags & VFIO_PCI_RECOVERY_IN_PROGRESS))
+		goto out_unlock;
+
+	notify_recovery = true;
+	down_write(&vdev->memory_lock);
+	/*
+	 * Apply deferred D0 requests before restoring PCI_COMMAND.
+	 * slot_reset() has already established D0 after a reset.
+	 */
+	power_up = vdev->power_up_pending;
+	vdev->power_up_pending = false;
+	if (power_up && !(vdev->pci_recovery_flags & VFIO_PCI_RECOVERY_RESET)) {
+		ret = vfio_pci_set_power_state(vdev, PCI_D0);
+		if (ret)
+			goto out_memory;
+	}
+
+	/*
+	 * Serialize with INTx updates to PCI_COMMAND. Keep masked INTx
+	 * disabled until vfio_pci_intx_recovery_finish().
+	 */
+	spin_lock_irqsave(&vdev->irqlock, irq_flags);
+	if (!(vdev->pci_recovery_flags & VFIO_PCI_RECOVERY_RESET) &&
+	    vdev->pci_recovery_command_valid) {
+		u16 cmd = vdev->pci_recovery_command;
+
+		cmd = vfio_pci_intx_recovery_update_command(vdev, cmd);
+		ret = pci_write_config_word(pdev, PCI_COMMAND, cmd);
+	}
+	spin_unlock_irqrestore(&vdev->irqlock, irq_flags);
+	if (ret)
+		goto out_memory;
+
+	/*
+	 * A saved-state restore may leave ROM decode enabled. Disable it
+	 * before allowing access to shared decoders.
+	 */
+	if (pci_resource_start(pdev, PCI_ROM_RESOURCE) &&
+	    !(pdev->resource[PCI_ROM_RESOURCE].flags & IORESOURCE_ROM_ENABLE))
+		pci_disable_rom(pdev);
+
+	/*
+	 * Finish restoring hardware state before allowing accesses.
+	 * Serialize the flag update with the INTx handler.
+	 */
+	spin_lock_irqsave(&vdev->irqlock, irq_flags);
+	WRITE_ONCE(vdev->access_blocked, false);
+	spin_unlock_irqrestore(&vdev->irqlock, irq_flags);
+
+	if (__vfio_pci_memory_enabled(vdev))
+		vfio_pci_dma_buf_move(vdev, false);
+
+out_memory:
+	up_write(&vdev->memory_lock);
+	vdev->pci_recovery_command_valid = false;
+	flags = vdev->pci_recovery_flags & ~VFIO_PCI_RECOVERY_IN_PROGRESS;
+	if (ret) {
+		WRITE_ONCE(vdev->pci_recovery_flags,
+			   flags | VFIO_PCI_RECOVERY_FAILED);
+		goto out_unlock;
+	}
+	WRITE_ONCE(vdev->pci_recovery_flags, flags);
+	vfio_pci_intx_recovery_finish(vdev);
+
+out_unlock:
+	mutex_unlock(&vdev->access_lock);
+	if (notify_recovery)
+		vfio_pci_signal_recovery_event(vdev);
+}
+
 int vfio_pci_core_sriov_configure(struct vfio_pci_core_device *vdev,
 				  int nr_virtfn)
 {
@@ -2688,6 +2770,7 @@ EXPORT_SYMBOL_GPL(vfio_pci_core_sriov_configure);
 const struct pci_error_handlers vfio_pci_core_err_handlers = {
 	.error_detected = vfio_pci_core_aer_err_detected,
 	.slot_reset = vfio_pci_core_aer_slot_reset,
+	.resume = vfio_pci_core_aer_resume,
 };
 EXPORT_SYMBOL_GPL(vfio_pci_core_err_handlers);
 
-- 
2.43.0


^ permalink raw reply	[flat|nested] 17+ messages in thread

* [RFC PATCH v2 14/16] vfio/pci: Block device access during host recovery
  2026-09-29 17:32 [RFC PATCH v2 00/16] vfio/pci: Handle PCI error recovery and report state to userspace Shameer Kolothum
                   ` (12 preceding siblings ...)
  2026-09-29 17:33 ` [RFC PATCH v2 13/16] vfio/pci: Complete recovery in resume() Shameer Kolothum
@ 2026-09-29 17:33 ` Shameer Kolothum
  2026-09-29 17:33 ` [RFC PATCH v2 15/16] vfio/pci: Add VFIO_DEVICE_FEATURE_PCI_ERROR_RECOVERY Shameer Kolothum
  2026-09-29 17:33 ` [RFC PATCH v2 16/16] vfio/pci: Enable host PCI error recovery for vfio-pci Shameer Kolothum
  15 siblings, 0 replies; 17+ messages in thread
From: Shameer Kolothum @ 2026-09-29 17:33 UTC (permalink / raw)
  To: kvm, linux-pci, linux-kernel
  Cc: alex, jgg, kevin.tian, kbusch, michal.winiarski,
	satyanarayana.k.v.p, sonangp, ankita, nathanc, mochs,
	skolothumtho

For userspace that enables recovery, extend error_detected() to block
device access and drain existing SRCU readers. Mask INTx, revoke BAR
mappings and DMA-BUFs, and disable bus mastering on a normal channel.

Return CAN_RECOVER for a normal channel, NEED_RESET for a frozen channel
and DISCONNECT for permanent failure. If saving PCI_COMMAND or disabling
bus mastering fails, return NONE and leave access blocked. Preserve the
original command and sequence number across nested events.

Track host recovery across close and reopen, including without userspace
opt-in. Reject open during a recorded transaction or after disconnection.
Clear deferred power requests at close and keep DMA-BUFs revoked while
access is blocked.

Reject hot reset for devices in ongoing recovery, but allow failed
devices so healthy siblings can be reset.

Reject userspace SR-IOV configuration while access is blocked. 

Assisted-by: LLM
Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com>
---
Note: The existing open/close paths are not serialized against the full
host recovery transaction. This series blocks reopening during an
already recorded transaction, but recovery can still start after the
open check, and close can overlap the physical reset between callbacks.
See the cover letter's "Locking and open questions" section.
---
 drivers/vfio/pci/vfio_pci.c        |   3 +
 drivers/vfio/pci/vfio_pci_core.c   | 166 ++++++++++++++++++++++++++++-
 drivers/vfio/pci/vfio_pci_dmabuf.c |   5 +
 3 files changed, 173 insertions(+), 1 deletion(-)

diff --git a/drivers/vfio/pci/vfio_pci.c b/drivers/vfio/pci/vfio_pci.c
index 830369ff878d..4887dd062c77 100644
--- a/drivers/vfio/pci/vfio_pci.c
+++ b/drivers/vfio/pci/vfio_pci.c
@@ -213,6 +213,9 @@ static int vfio_pci_sriov_configure(struct pci_dev *pdev, int nr_virtfn)
 	if (!enable_sriov)
 		return -ENOENT;
 
+	if (vdev->pci_recovery_supported && READ_ONCE(vdev->access_blocked))
+		return -EIO;
+
 	return vfio_pci_core_sriov_configure(vdev, nr_virtfn);
 }
 
diff --git a/drivers/vfio/pci/vfio_pci_core.c b/drivers/vfio/pci/vfio_pci_core.c
index d6cc34240b25..2ec8022e4a69 100644
--- a/drivers/vfio/pci/vfio_pci_core.c
+++ b/drivers/vfio/pci/vfio_pci_core.c
@@ -617,6 +617,15 @@ int vfio_pci_core_enable(struct vfio_pci_core_device *vdev)
 	u16 cmd;
 	u8 msix_pos;
 
+	if (vdev->pci_recovery_supported) {
+		if (pci_dev_is_disconnected(pdev))
+			return -ENODEV;
+		scoped_guard(mutex, &vdev->access_lock) {
+			if (vdev->pci_recovery_host_active)
+				return -EBUSY;
+		}
+	}
+
 	if (!vdev->disable_idle_d3) {
 		ret = pm_runtime_resume_and_get(&pdev->dev);
 		if (ret < 0)
@@ -860,6 +869,10 @@ void vfio_pci_core_close_device(struct vfio_device *core_vdev)
 			WRITE_ONCE(vdev->access_blocked, false);
 			WRITE_ONCE(vdev->device_open, false);
 			WRITE_ONCE(vdev->pci_recovery_flags, 0);
+			/* Discard pending D0 so it cannot be replayed after reopen. */
+			down_write(&vdev->memory_lock);
+			vdev->power_up_pending = false;
+			up_write(&vdev->memory_lock);
 		}
 	}
 
@@ -1843,6 +1856,15 @@ void vfio_pci_core_access_end(struct vfio_pci_core_device *vdev, int idx)
 		srcu_read_unlock(&vdev->access_srcu, idx);
 }
 
+/* Reject new accesses and wait for existing SRCU readers. */
+static void vfio_pci_core_block_access(struct vfio_pci_core_device *vdev)
+{
+	lockdep_assert_held(&vdev->access_lock);
+
+	WRITE_ONCE(vdev->access_blocked, true);
+	synchronize_srcu(&vdev->access_srcu);
+}
+
 ssize_t vfio_pci_core_read(struct vfio_device *core_vdev, char __user *buf,
 		size_t count, loff_t *ppos)
 {
@@ -2521,19 +2543,147 @@ vfio_pci_signal_recovery_event(struct vfio_pci_core_device *vdev)
 	rcu_read_unlock();
 }
 
+/*
+ * Save PCI_COMMAND, clear bus mastering and set the requested bits.
+ * Reject an all-ones read from an inaccessible device.
+ */
+static int
+vfio_pci_recovery_save_and_clear_master(struct vfio_pci_core_device *vdev, u16 set)
+{
+	struct pci_dev *pdev = vdev->pdev;
+	u16 command;
+	int ret;
+
+	ret = pci_read_config_word(pdev, PCI_COMMAND, &vdev->pci_recovery_command);
+	if (ret)
+		return ret;
+	if (PCI_POSSIBLE_ERROR(vdev->pci_recovery_command))
+		return -EIO;
+
+	command = (vdev->pci_recovery_command & ~PCI_COMMAND_MASTER) | set;
+	ret = pci_write_config_word(pdev, PCI_COMMAND, command);
+	if (ret)
+		return ret;
+
+	vdev->pci_recovery_command_valid = true;
+	return 0;
+}
+
 pci_ers_result_t vfio_pci_core_aer_err_detected(struct pci_dev *pdev,
 						pci_channel_state_t state)
 {
 	struct vfio_pci_core_device *vdev = dev_get_drvdata(&pdev->dev);
 	struct vfio_pci_eventfd *eventfd;
+	pci_ers_result_t result = PCI_ERS_RESULT_CAN_RECOVER;
+	unsigned long irq_flags;
+	bool notify_recovery = false;
+	bool nested;
+	u32 flags;
+	int ret = 0;
+
+	if (!vdev->pci_recovery_supported)
+		goto out;
+
+	mutex_lock(&vdev->access_lock);
+	/*
+	 * Track the host transaction even when userspace has not enabled
+	 * recovery. Reopen must wait for resume() or permanent failure.
+	 */
+	vdev->pci_recovery_host_active =
+		state != pci_channel_io_perm_failure;
+	if (!vdev->pci_recovery_enabled)
+		goto out_unlock;
 
+	/* Keep failed devices blocked until close. */
+	if (vdev->pci_recovery_flags & VFIO_PCI_RECOVERY_FAILED) {
+		result = PCI_ERS_RESULT_NONE;
+		goto out_unlock;
+	}
+
+	if (!vdev->device_open) {
+		result = PCI_ERS_RESULT_NONE;
+		goto out_unlock;
+	}
+
+	notify_recovery = true;
+	vfio_pci_core_block_access(vdev);
+	/*
+	 * Preserve the saved command for nested events. The current value
+	 * may already have bus mastering disabled by recovery.
+	 */
+	nested = vdev->pci_recovery_flags & VFIO_PCI_RECOVERY_IN_PROGRESS;
+	if (!nested)
+		vdev->pci_recovery_command_valid = false;
+	vfio_pci_intx_recovery_start(vdev);
+	vfio_pci_zap_and_down_write_memory_lock(vdev);
+	/*
+	 * Serialize PCI_COMMAND updates with the INTx handler, which does
+	 * not take access_srcu.
+	 */
+	spin_lock_irqsave(&vdev->irqlock, irq_flags);
+	if (state == pci_channel_io_normal && vdev->pci_2_3 && !nested)
+		ret = vfio_pci_recovery_save_and_clear_master(
+			vdev, PCI_COMMAND_INTX_DISABLE);
+	spin_unlock_irqrestore(&vdev->irqlock, irq_flags);
+	vfio_pci_dma_buf_move(vdev, true);
+
+	/*
+	 * Start a new sequence unless this event belongs to an active
+	 * transaction. Publish the flags in one store to avoid exposing
+	 * a transient cleared state to lockless readers.
+	 */
+	if (nested) {
+		flags = vdev->pci_recovery_flags;
+	} else {
+		if (++vdev->pci_recovery_sequence == 0)
+			vdev->pci_recovery_sequence++;
+		flags = 0;
+	}
+
+	if (state == pci_channel_io_perm_failure) {
+		WRITE_ONCE(vdev->pci_recovery_flags,
+			   (flags | VFIO_PCI_RECOVERY_FAILED) &
+			   ~VFIO_PCI_RECOVERY_IN_PROGRESS);
+		vdev->pci_recovery_command_valid = false;
+		result = PCI_ERS_RESULT_DISCONNECT;
+		goto out_memory;
+	}
+
+	if (state == pci_channel_io_frozen) {
+		WRITE_ONCE(vdev->pci_recovery_flags,
+			   flags | VFIO_PCI_RECOVERY_IN_PROGRESS |
+			   VFIO_PCI_RECOVERY_FROZEN);
+		result = PCI_ERS_RESULT_NEED_RESET;
+		goto out_memory;
+	}
+
+	if (!ret && !vdev->pci_2_3 && !nested)
+		ret = vfio_pci_recovery_save_and_clear_master(vdev, 0);
+
+	if (ret) {
+		flags |= VFIO_PCI_RECOVERY_FAILED;
+		flags &= ~VFIO_PCI_RECOVERY_IN_PROGRESS;
+		result = PCI_ERS_RESULT_NONE;
+	} else {
+		flags |= VFIO_PCI_RECOVERY_IN_PROGRESS;
+	}
+	WRITE_ONCE(vdev->pci_recovery_flags, flags);
+
+out_memory:
+	up_write(&vdev->memory_lock);
+out_unlock:
+	mutex_unlock(&vdev->access_lock);
+
+out:
 	rcu_read_lock();
 	eventfd = rcu_dereference(vdev->err_trigger);
 	if (eventfd)
 		eventfd_signal(eventfd->ctx);
 	rcu_read_unlock();
+	if (notify_recovery)
+		vfio_pci_signal_recovery_event(vdev);
 
-	return PCI_ERS_RESULT_CAN_RECOVER;
+	return result;
 }
 EXPORT_SYMBOL_GPL(vfio_pci_core_aer_err_detected);
 
@@ -2926,6 +3076,20 @@ static int vfio_pci_dev_set_hot_reset(struct vfio_device_set *dev_set,
 			break;
 		}
 
+		/*
+		 * Recovery releases memory_lock between callbacks. Check the
+		 * access block after taking the lock as well. Allow failed
+		 * devices so they do not prevent resetting healthy siblings.
+		 */
+		if (vdev->pci_recovery_supported &&
+		    READ_ONCE(vdev->access_blocked) &&
+		    !(READ_ONCE(vdev->pci_recovery_flags) &
+		      VFIO_PCI_RECOVERY_FAILED)) {
+			up_write(&vdev->memory_lock);
+			ret = -EBUSY;
+			break;
+		}
+
 		vfio_pci_dma_buf_move(vdev, true);
 		vfio_pci_zap_bars(vdev);
 	}
diff --git a/drivers/vfio/pci/vfio_pci_dmabuf.c b/drivers/vfio/pci/vfio_pci_dmabuf.c
index c1a0250af680..69c41c3146f8 100644
--- a/drivers/vfio/pci/vfio_pci_dmabuf.c
+++ b/drivers/vfio/pci/vfio_pci_dmabuf.c
@@ -350,6 +350,11 @@ void vfio_pci_dma_buf_move(struct vfio_pci_core_device *vdev, bool revoked)
 
 	lockdep_assert_held_write(&vdev->memory_lock);
 
+	/* Reset and power-state cleanup must not undo recovery revocation. */
+	if (!revoked && vdev->pci_recovery_supported &&
+	    READ_ONCE(vdev->access_blocked))
+		return;
+
 	list_for_each_entry_safe(priv, tmp, &vdev->dmabufs, dmabufs_elm) {
 		if (!get_file_active(&priv->dmabuf->file))
 			continue;
-- 
2.43.0


^ permalink raw reply	[flat|nested] 17+ messages in thread

* [RFC PATCH v2 15/16] vfio/pci: Add VFIO_DEVICE_FEATURE_PCI_ERROR_RECOVERY
  2026-09-29 17:32 [RFC PATCH v2 00/16] vfio/pci: Handle PCI error recovery and report state to userspace Shameer Kolothum
                   ` (13 preceding siblings ...)
  2026-09-29 17:33 ` [RFC PATCH v2 14/16] vfio/pci: Block device access during host recovery Shameer Kolothum
@ 2026-09-29 17:33 ` Shameer Kolothum
  2026-09-29 17:33 ` [RFC PATCH v2 16/16] vfio/pci: Enable host PCI error recovery for vfio-pci Shameer Kolothum
  15 siblings, 0 replies; 17+ messages in thread
From: Shameer Kolothum @ 2026-09-29 17:33 UTC (permalink / raw)
  To: kvm, linux-pci, linux-kernel
  Cc: alex, jgg, kevin.tian, kbusch, michal.winiarski,
	satyanarayana.k.v.p, sonangp, ankita, nathanc, mochs,
	skolothumtho

Add a device feature to enable host PCI error recovery and query its
status. SET registers an eventfd for start and completion notifications;
eventfd -1 disables recovery. GET returns the status flags and sequence
number, including while recovery is active.

Reject SET while access is blocked or recovery has failed. Reset the
per-open status and sequence when updating the eventfd. Existing
VFIO_PCI_ERR_IRQ_INDEX notifications are unchanged.

Assisted-by: LLM
Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com>
---
 include/uapi/linux/vfio.h        |  44 +++++++++++++
 drivers/vfio/pci/vfio_pci_core.c | 108 +++++++++++++++++++++++++++++++
 2 files changed, 152 insertions(+)

diff --git a/include/uapi/linux/vfio.h b/include/uapi/linux/vfio.h
index e41437fa17ad..b755849fd39b 100644
--- a/include/uapi/linux/vfio.h
+++ b/include/uapi/linux/vfio.h
@@ -1555,6 +1555,50 @@ struct vfio_device_feature_zpci_err {
 
 #define VFIO_DEVICE_FEATURE_ZPCI_ERROR 13
 
+/*
+ * Report host PCI error recovery state for this device.
+ *
+ * VFIO_DEVICE_FEATURE_SET with a valid eventfd enables recovery for this
+ * open and registers notifications for the start and end of each event.
+ * SET with eventfd -1 disables recovery. flags and sequence must be zero.
+ * SET returns -EBUSY during recovery, after recovery failure, or while
+ * device access is blocked, and -ENODEV after close has started.
+ * VFIO_PCI_ERR_IRQ_INDEX notifications are unchanged.
+ *
+ * VFIO_DEVICE_FEATURE_GET returns the current state and eventfd = -1.
+ * GET is allowed during recovery, but may wait for a running callback.
+ *
+ * The sequence number increments for each event and resets to zero when
+ * recovery is enabled. Use it to detect coalesced notifications. Device
+ * accesses return -EIO while IN_PROGRESS is set. Recovery may complete
+ * before userspace reads the notification, so userspace must check the
+ * sequence and status rather than wait to observe IN_PROGRESS.
+ *
+ * CHANNEL_FROZEN records a frozen channel. DEVICE_RESET records a host
+ * reset; userspace must reconfigure interrupts with VFIO_DEVICE_SET_IRQS.
+ * FAILED indicates that recovery failed and access remains blocked until
+ * close and reopen. Status bits persist until the next event or SET.
+ * ENABLED indicates that recovery is enabled.
+ *
+ * Enabling recovery during an existing event does not provide status or
+ * notifications for that event.
+ *
+ * Open returns -EBUSY while a previously recorded host recovery is active,
+ * including events reported while recovery was disabled.
+ */
+struct vfio_device_pci_error_recovery {
+	__u32 flags;
+#define VFIO_PCI_ERROR_RECOVERY_IN_PROGRESS	(1U << 0)
+#define VFIO_PCI_ERROR_RECOVERY_CHANNEL_FROZEN	(1U << 1)
+#define VFIO_PCI_ERROR_RECOVERY_DEVICE_RESET	(1U << 2)
+#define VFIO_PCI_ERROR_RECOVERY_FAILED		(1U << 3)
+#define VFIO_PCI_ERROR_RECOVERY_ENABLED		(1U << 4)
+	__s32 eventfd;
+	__aligned_u64 sequence;
+};
+
+#define VFIO_DEVICE_FEATURE_PCI_ERROR_RECOVERY 14
+
 /* -------- API for Type1 VFIO IOMMU -------- */
 
 /**
diff --git a/drivers/vfio/pci/vfio_pci_core.c b/drivers/vfio/pci/vfio_pci_core.c
index 2ec8022e4a69..314d77868803 100644
--- a/drivers/vfio/pci/vfio_pci_core.c
+++ b/drivers/vfio/pci/vfio_pci_core.c
@@ -1752,6 +1752,111 @@ static int vfio_pci_core_feature_token(struct vfio_pci_core_device *vdev,
 	return 0;
 }
 
+/* Consumes the eventfd reference on both success and failure. */
+static int vfio_pci_recovery_set(struct vfio_pci_core_device *vdev,
+				 struct eventfd_ctx *ctx, bool enable)
+{
+	int ret;
+
+	guard(mutex)(&vdev->access_lock);
+	if (!vdev->device_open) {
+		ret = -ENODEV;
+		goto out_put;
+	}
+	if (vdev->access_blocked) {
+		ret = -EBUSY;
+		goto out_put;
+	}
+	if (!enable &&
+	    (vdev->pci_recovery_flags &
+	     (VFIO_PCI_RECOVERY_IN_PROGRESS | VFIO_PCI_RECOVERY_FAILED))) {
+		ret = -EBUSY;
+		goto out_put;
+	}
+
+	mutex_lock(&vdev->igate);
+	ret = vfio_pci_eventfd_replace_locked(vdev, &vdev->pci_recovery_trigger,
+					      ctx);
+	mutex_unlock(&vdev->igate);
+	if (ret)
+		goto out_put;
+
+	WRITE_ONCE(vdev->pci_recovery_enabled, enable);
+	/*
+	 * Reset the per-open status when updating the recovery eventfd.
+	 * The access block and host transaction state are unchanged.
+	 */
+	WRITE_ONCE(vdev->pci_recovery_flags, 0);
+	vdev->pci_recovery_sequence = 0;
+	vdev->pci_recovery_command_valid = false;
+
+	return 0;
+
+out_put:
+	if (ctx)
+		eventfd_ctx_put(ctx);
+	return ret;
+}
+
+static int
+vfio_pci_core_feature_error_recovery(struct vfio_pci_core_device *vdev, u32 flags,
+				     struct vfio_device_pci_error_recovery __user *arg,
+				     size_t argsz)
+{
+	struct vfio_device_pci_error_recovery state = { .eventfd = -1 };
+	struct eventfd_ctx *ctx = NULL;
+	bool enable;
+	int ret;
+
+	if (!vdev->pci_recovery_supported)
+		return -ENOTTY;
+
+	ret = vfio_check_feature(flags, argsz,
+				 VFIO_DEVICE_FEATURE_GET |
+				 VFIO_DEVICE_FEATURE_SET, sizeof(state));
+	if (ret != 1)
+		return ret;
+
+	if (flags & VFIO_DEVICE_FEATURE_GET) {
+		scoped_guard(mutex, &vdev->access_lock) {
+			u32 rflags = vdev->pci_recovery_flags;
+
+			if (vdev->pci_recovery_enabled)
+				state.flags |= VFIO_PCI_ERROR_RECOVERY_ENABLED;
+			if (rflags & VFIO_PCI_RECOVERY_IN_PROGRESS)
+				state.flags |=
+					VFIO_PCI_ERROR_RECOVERY_IN_PROGRESS;
+			if (rflags & VFIO_PCI_RECOVERY_FROZEN)
+				state.flags |=
+					VFIO_PCI_ERROR_RECOVERY_CHANNEL_FROZEN;
+			if (rflags & VFIO_PCI_RECOVERY_RESET)
+				state.flags |=
+					VFIO_PCI_ERROR_RECOVERY_DEVICE_RESET;
+			if (rflags & VFIO_PCI_RECOVERY_FAILED)
+				state.flags |= VFIO_PCI_ERROR_RECOVERY_FAILED;
+			state.sequence = vdev->pci_recovery_sequence;
+		}
+
+		if (copy_to_user(arg, &state, sizeof(state)))
+			return -EFAULT;
+		return 0;
+	}
+
+	if (copy_from_user(&state, arg, sizeof(state)))
+		return -EFAULT;
+	if (state.flags || state.sequence || state.eventfd < -1)
+		return -EINVAL;
+
+	enable = state.eventfd >= 0;
+	if (enable) {
+		ctx = eventfd_ctx_fdget(state.eventfd);
+		if (IS_ERR(ctx))
+			return PTR_ERR(ctx);
+	}
+
+	return vfio_pci_recovery_set(vdev, ctx, enable);
+}
+
 int vfio_pci_core_ioctl_feature(struct vfio_device *device, u32 flags,
 				void __user *arg, size_t argsz)
 {
@@ -1772,6 +1877,9 @@ int vfio_pci_core_ioctl_feature(struct vfio_device *device, u32 flags,
 		return vfio_pci_core_feature_dma_buf(vdev, flags, arg, argsz);
 	case VFIO_DEVICE_FEATURE_ZPCI_ERROR:
 		return vfio_pci_zdev_feature_err(device, flags, arg, argsz);
+
+	case VFIO_DEVICE_FEATURE_PCI_ERROR_RECOVERY:
+		return vfio_pci_core_feature_error_recovery(vdev, flags, arg, argsz);
 	default:
 		return -ENOTTY;
 	}
-- 
2.43.0


^ permalink raw reply	[flat|nested] 17+ messages in thread

* [RFC PATCH v2 16/16] vfio/pci: Enable host PCI error recovery for vfio-pci
  2026-09-29 17:32 [RFC PATCH v2 00/16] vfio/pci: Handle PCI error recovery and report state to userspace Shameer Kolothum
                   ` (14 preceding siblings ...)
  2026-09-29 17:33 ` [RFC PATCH v2 15/16] vfio/pci: Add VFIO_DEVICE_FEATURE_PCI_ERROR_RECOVERY Shameer Kolothum
@ 2026-09-29 17:33 ` Shameer Kolothum
  15 siblings, 0 replies; 17+ messages in thread
From: Shameer Kolothum @ 2026-09-29 17:33 UTC (permalink / raw)
  To: kvm, linux-pci, linux-kernel
  Cc: alex, jgg, kevin.tian, kbusch, michal.winiarski,
	satyanarayana.k.v.p, sonangp, ankita, nathanc, mochs,
	skolothumtho

Enable recovery support for generic vfio-pci. Userspace opts in through
VFIO_DEVICE_FEATURE_PCI_ERROR_RECOVERY. Open rejects an active host
recovery even without opt-in. Leave variant drivers disabled until their
access paths are covered.

Exclude CXL memory devices and s390: Restricted CXL Host (RCH) and s390
mediated recovery can omit completion callbacks. On s390, permanent-failure
notification also blocks config access, so draining SRCU can deadlock
with config readers.

Assisted-by: LLM
Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com>
---
 drivers/vfio/pci/vfio_pci.c | 9 +++++++++
 1 file changed, 9 insertions(+)

diff --git a/drivers/vfio/pci/vfio_pci.c b/drivers/vfio/pci/vfio_pci.c
index 4887dd062c77..043371c2882c 100644
--- a/drivers/vfio/pci/vfio_pci.c
+++ b/drivers/vfio/pci/vfio_pci.c
@@ -129,6 +129,7 @@ static int vfio_pci_init_dev(struct vfio_device *core_vdev)
 {
 	struct vfio_pci_core_device *vdev =
 		container_of(core_vdev, struct vfio_pci_core_device, vdev);
+	struct pci_dev *pdev = to_pci_dev(core_vdev->dev);
 
 	/*
 	 * These behaviors originated in vfio-pci and moved into
@@ -139,6 +140,14 @@ static int vfio_pci_init_dev(struct vfio_device *core_vdev)
 	 */
 	vdev->nointxmask = nointxmask;
 	vdev->disable_idle_d3 = disable_idle_d3;
+	/*
+	 * CXL RCH and s390 mediated recovery can omit completion callbacks.
+	 * s390 also blocks config access across failure notification, which
+	 * prevents draining SRCU readers waiting for config access.
+	 */
+	if (!IS_ENABLED(CONFIG_S390) &&
+	    (pdev->class >> 8) != PCI_CLASS_MEMORY_CXL)
+		vdev->pci_recovery_supported = true;
 #ifdef CONFIG_VFIO_PCI_VGA
 	vdev->disable_vga = disable_vga;
 #endif
-- 
2.43.0


^ permalink raw reply	[flat|nested] 17+ messages in thread

end of thread, other threads:[~2026-09-29 17:35 UTC | newest]

Thread overview: 17+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-29 17:32 [RFC PATCH v2 00/16] vfio/pci: Handle PCI error recovery and report state to userspace Shameer Kolothum
2026-09-29 17:32 ` [RFC PATCH v2 01/16] vfio/pci: Add a device access gate Shameer Kolothum
2026-09-29 17:32 ` [RFC PATCH v2 02/16] vfio/pci: Gate config space access Shameer Kolothum
2026-09-29 17:32 ` [RFC PATCH v2 03/16] vfio/pci: Buffer ROM reads before copying to userspace Shameer Kolothum
2026-09-29 17:32 ` [RFC PATCH v2 04/16] vfio/pci: Gate BAR and ROM access Shameer Kolothum
2026-09-29 17:32 ` [RFC PATCH v2 05/16] vfio/pci: Fail BAR faults while access is blocked Shameer Kolothum
2026-09-29 17:32 ` [RFC PATCH v2 06/16] vfio/pci: Gate interrupt configuration Shameer Kolothum
2026-09-29 17:32 ` [RFC PATCH v2 07/16] vfio/pci: Gate function reset and runtime power management Shameer Kolothum
2026-09-29 17:32 ` [RFC PATCH v2 08/16] vfio/pci: Gate device information queries and DMA-BUF export Shameer Kolothum
2026-09-29 17:32 ` [RFC PATCH v2 09/16] vfio/pci: Add PCI error recovery state Shameer Kolothum
2026-09-29 17:32 ` [RFC PATCH v2 10/16] vfio/pci: Quiesce INTx while access is blocked Shameer Kolothum
2026-09-29 17:33 ` [RFC PATCH v2 11/16] vfio/pci: Add INTx recovery start and finish helpers Shameer Kolothum
2026-09-29 17:33 ` [RFC PATCH v2 12/16] vfio/pci: Restore device state from slot_reset() Shameer Kolothum
2026-09-29 17:33 ` [RFC PATCH v2 13/16] vfio/pci: Complete recovery in resume() Shameer Kolothum
2026-09-29 17:33 ` [RFC PATCH v2 14/16] vfio/pci: Block device access during host recovery Shameer Kolothum
2026-09-29 17:33 ` [RFC PATCH v2 15/16] vfio/pci: Add VFIO_DEVICE_FEATURE_PCI_ERROR_RECOVERY Shameer Kolothum
2026-09-29 17:33 ` [RFC PATCH v2 16/16] vfio/pci: Enable host PCI error recovery for vfio-pci Shameer Kolothum

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®