mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [RFC PATCH v2 00/16] vfio/pci: Handle PCI error recovery and report state to userspace
@ 2026-09-29 17:32 Shameer Kolothum
  2026-09-29 17:32 ` [RFC PATCH v2 01/16] vfio/pci: Add a device access gate Shameer Kolothum
                   ` (15 more replies)
  0 siblings, 16 replies; 17+ messages in thread
From: Shameer Kolothum @ 2026-09-29 17:32 UTC (permalink / raw)
  To: kvm, linux-pci, linux-kernel
  Cc: alex, jgg, kevin.tian, kbusch, michal.winiarski,
	satyanarayana.k.v.p, sonangp, ankita, nathanc, mochs,
	skolothumtho

Hi,

Currently, vfio-pci signals the error eventfd from error_detected() and
returns PCI_ERS_RESULT_CAN_RECOVER, including for permanent failure. It
has no slot_reset() or resume() callback, so userspace is not notified
when host recovery completes or whether the device was reset.

This series adds opt-in host PCI error recovery for vfio-pci. It blocks
device access during recovery, restores device state after a host reset,
and reports recovery status to userspace. The kernel runs host recovery
and the VMM decides how to recover the guest.

Changes from v1
---------------
 v1: https://lore.kernel.org/kvm/20260901093217.8539-1-skolothumtho@nvidia.com/
  
 - Thanks to Satya and Alex for reviews/feedback.
 - Use SRCU to drain ongoing device accesses and reject new accesses during
   recovery. A mutex protects recovery-state updates. 
 - Return SIGBUS for blocked BAR faults instead of waiting in the fault
   handler. VMM handling still needs target validation.
 - Reject userspace SR-IOV changes while access is blocked.
 - Exclude s390 and all CXL memory devices for now as s390 and CXL RCH
   recovery can omit completion callbacks.
  
Design
------
Userspace opts in through VFIO_DEVICE_FEATURE_PCI_ERROR_RECOVERY. It
registers an eventfd and reads status flags and a sequence number. The
existing error eventfd is still signaled.

The callbacks handle the host recovery sequence:

 - error_detected() blocks new accesses and drains admitted ones,
   quiesces INTx, and revokes BAR mappings and exported DMA-BUFs. On a
   normal channel it also clears bus mastering. It votes CAN_RECOVER
   for a normal channel, NEED_RESET for a frozen channel, and DISCONNECT
   for permanent failure. Failure to save PCI_COMMAND or disable bus
   mastering leaves access blocked and returns NONE so other devices
   can continue recovery.
 - slot_reset() restores the configuration saved at open after a host
   reset, then tears down interrupt configuration. Userspace must set
   up interrupts again.
 - resume() completes restoration, unblocks access and completes INTx
   handling. A restoration failure keeps the device blocked.

Non-fatal errors also block device access because host recovery may
require a reset. Userspace reads the recovery status to determine
whether the device was reset or recovery failed.
  
Without opt-in, the existing error notification behavior is retained.
The open check also rejects a previously recorded host transaction when
recovery is disabled, so closing the device does not discard that state.

Locking and open questions
--------------------------
Several locks are involved in recovery. PCI core holds device_lock
while calling the driver, preventing concurrent driver binding or
unbinding. During the bus walk, it also holds pci_bus_sem for reading
to keep the device list stable. Adding or removing PCI devices requires
the write lock. This semaphore is shared across all PCI buses.
 
VFIO recovery blocks new accesses, waits for existing SRCU readers to
finish, and takes memory_lock. Power transitions also take memory_lock,
but entering D0 can then acquire pci_bus_sem to update PCIe link power
management (ASPM).

Local Sashiko/Claude review reports a lock-order conflict as below:

 1. A guest D0 request holds memory_lock.
 2. PCI core holds pci_bus_sem for reading while VFIO's recovery
    callback waits for memory_lock.
 3. Removal of an unrelated PCI device waits for the bus write lock.
 4. D0's ASPM update requests the bus read lock and waits behind that
    writer.

Each task waits for another, so none can proceed.

One possible solution could be(not implemented):

 - PCI core guarantees that pci_bus_sem is held for reading before
   invoking supported recovery callbacks. VFIO uses
   pci_set_power_state_locked() in those callbacks, avoiding another
   acquisition of the bus lock.
 - A new pci_try_set_power_state() helper tries to acquire pci_bus_sem
   and returns -EAGAIN on contention. Ordinary VFIO D0 paths use this
   helper and handle failure instead of waiting while holding VFIO locks
	  
Not sure there are better ways to handle this or not. Feedbacks appreciated
on this.

Another issue flagged was(I think this is a pre-existing one):
 - Open/close coordination with the whole host recovery transaction
   was already missing. Recovery can start after the open check, and
   close can overlap the physical reset. Restoration of closed devices
   is also not implemented.

Testing
-------
Basic sanity tests performed on a GB200 with an NVIDIA GPU assigned.

Kernel branch:
 https://github.com/shamiali2008/linux/tree/vfio-aer-rfc-v2-ext

QEMU test branch is here(This registers the recovery eventfd and uses
pcie_aer_inject_error() to report the error to Guest)
 https://github.com/shamiali2008/qemu-master/commits/private-master-vfio-aer-test-v2/

Software AER injection was performed using a modified pcieaer_inject
module.

./aer-inject nonfatal.conf

qemu-system-aarch64: info: vfio 0018:06:00.0: host PCI error recovery completed (seq 1)

Guest kernel:

[   59.036988] pcieport 0000:01:00.0: AER: Uncorrectable (Non-Fatal) error message received from 0000:02:00.0
[   59.038326] nvidia 0000:02:00.0: PCIe Bus Error: severity=Uncorrectable (Non-Fatal), type=Transaction Layer, (Completer ID)
[   59.038440] nvidia 0000:02:00.0:   device [10de:2941] error status/mask=00008000/02400000
[   59.038533] nvidia 0000:02:00.0:    [15] CmpltAbrt              (First)
[   59.038705] nvidia 0000:02:00.0: AER:   TLP Header: 0x00000000 0x00000000 0x00000000 0x00000000
...

./aer-inject fatal.conf
qemu-system-aarch64: info: vfio 0018:06:00.0: host PCI error recovery started (seq 2, channel frozen)
qemu-system-aarch64: info: vfio 0018:06:00.0: host PCI error recovery completed (seq 2, device was reset)

Guest kernel:

[  150.909306] pcieport 0000:01:00.0: AER: Uncorrectable (Fatal) error message received from 0000:02:00.0
[  150.909566] nvidia 0000:02:00.0: AER: PCIe Bus Error: severity=Uncorrectable (Fatal), type=Inaccessible, (Unregistered Agent ID)
[  150.909835] nvidia 0000:02:00.0: AER: can't recover (no error_detected callback)
[  150.910062] pcieport 0000:01:00.0: unlocked secondary bus reset via: pciehp_reset_slot+0x54/0x98
[  156.603467] pcieport 0000:01:00.0: AER: Root Port link has been reset (0)
...

Please take a look and let me know your feedback.

Thanks,
Shameer

Shameer Kolothum (16):
  vfio/pci: Add a device access gate
  vfio/pci: Gate config space access
  vfio/pci: Buffer ROM reads before copying to userspace
  vfio/pci: Gate BAR and ROM access
  vfio/pci: Fail BAR faults while access is blocked
  vfio/pci: Gate interrupt configuration
  vfio/pci: Gate function reset and runtime power management
  vfio/pci: Gate device information queries and DMA-BUF export
  vfio/pci: Add PCI error recovery state
  vfio/pci: Quiesce INTx while access is blocked
  vfio/pci: Add INTx recovery start and finish helpers
  vfio/pci: Restore device state from slot_reset()
  vfio/pci: Complete recovery in resume()
  vfio/pci: Block device access during host recovery
  vfio/pci: Add VFIO_DEVICE_FEATURE_PCI_ERROR_RECOVERY
  vfio/pci: Enable host PCI error recovery for vfio-pci

 drivers/vfio/pci/vfio_pci_priv.h   |  10 +
 include/linux/vfio_pci_core.h      |  38 ++
 include/uapi/linux/vfio.h          |  44 ++
 drivers/vfio/pci/vfio_pci.c        |  12 +
 drivers/vfio/pci/vfio_pci_config.c | 104 ++++-
 drivers/vfio/pci/vfio_pci_core.c   | 636 ++++++++++++++++++++++++++++-
 drivers/vfio/pci/vfio_pci_dmabuf.c |  21 +-
 drivers/vfio/pci/vfio_pci_intrs.c  | 131 ++++++
 drivers/vfio/pci/vfio_pci_rdwr.c   | 169 ++++++--
 9 files changed, 1090 insertions(+), 75 deletions(-)

-- 
2.43.0


^ permalink raw reply	[flat|nested] 17+ messages in thread

end of thread, other threads:[~2026-09-29 17:35 UTC | newest]

Thread overview: 17+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-29 17:32 [RFC PATCH v2 00/16] vfio/pci: Handle PCI error recovery and report state to userspace Shameer Kolothum
2026-09-29 17:32 ` [RFC PATCH v2 01/16] vfio/pci: Add a device access gate Shameer Kolothum
2026-09-29 17:32 ` [RFC PATCH v2 02/16] vfio/pci: Gate config space access Shameer Kolothum
2026-09-29 17:32 ` [RFC PATCH v2 03/16] vfio/pci: Buffer ROM reads before copying to userspace Shameer Kolothum
2026-09-29 17:32 ` [RFC PATCH v2 04/16] vfio/pci: Gate BAR and ROM access Shameer Kolothum
2026-09-29 17:32 ` [RFC PATCH v2 05/16] vfio/pci: Fail BAR faults while access is blocked Shameer Kolothum
2026-09-29 17:32 ` [RFC PATCH v2 06/16] vfio/pci: Gate interrupt configuration Shameer Kolothum
2026-09-29 17:32 ` [RFC PATCH v2 07/16] vfio/pci: Gate function reset and runtime power management Shameer Kolothum
2026-09-29 17:32 ` [RFC PATCH v2 08/16] vfio/pci: Gate device information queries and DMA-BUF export Shameer Kolothum
2026-09-29 17:32 ` [RFC PATCH v2 09/16] vfio/pci: Add PCI error recovery state Shameer Kolothum
2026-09-29 17:32 ` [RFC PATCH v2 10/16] vfio/pci: Quiesce INTx while access is blocked Shameer Kolothum
2026-09-29 17:33 ` [RFC PATCH v2 11/16] vfio/pci: Add INTx recovery start and finish helpers Shameer Kolothum
2026-09-29 17:33 ` [RFC PATCH v2 12/16] vfio/pci: Restore device state from slot_reset() Shameer Kolothum
2026-09-29 17:33 ` [RFC PATCH v2 13/16] vfio/pci: Complete recovery in resume() Shameer Kolothum
2026-09-29 17:33 ` [RFC PATCH v2 14/16] vfio/pci: Block device access during host recovery Shameer Kolothum
2026-09-29 17:33 ` [RFC PATCH v2 15/16] vfio/pci: Add VFIO_DEVICE_FEATURE_PCI_ERROR_RECOVERY Shameer Kolothum
2026-09-29 17:33 ` [RFC PATCH v2 16/16] vfio/pci: Enable host PCI error recovery for vfio-pci Shameer Kolothum

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®