From: Shameer Kolothum <skolothumtho@nvidia.com>
To: <kvm@vger.kernel.org>, <linux-pci@vger.kernel.org>,
<linux-kernel@vger.kernel.org>
Cc: <alex@shazbot.org>, <jgg@ziepe.ca>, <kevin.tian@intel.com>,
<kbusch@meta.com>, <michal.winiarski@intel.com>,
<satyanarayana.k.v.p@intel.com>, <sonangp@nvidia.com>,
<ankita@nvidia.com>, <nathanc@nvidia.com>, <mochs@nvidia.com>,
<skolothumtho@nvidia.com>
Subject: [RFC PATCH v2 00/16] vfio/pci: Handle PCI error recovery and report state to userspace
Date: Tue, 29 Sep 2026 18:32:49 +0100 [thread overview]
Message-ID: <20260929173305.204856-1-skolothumtho@nvidia.com> (raw)
Hi,
Currently, vfio-pci signals the error eventfd from error_detected() and
returns PCI_ERS_RESULT_CAN_RECOVER, including for permanent failure. It
has no slot_reset() or resume() callback, so userspace is not notified
when host recovery completes or whether the device was reset.
This series adds opt-in host PCI error recovery for vfio-pci. It blocks
device access during recovery, restores device state after a host reset,
and reports recovery status to userspace. The kernel runs host recovery
and the VMM decides how to recover the guest.
Changes from v1
---------------
v1: https://lore.kernel.org/kvm/20260901093217.8539-1-skolothumtho@nvidia.com/
- Thanks to Satya and Alex for reviews/feedback.
- Use SRCU to drain ongoing device accesses and reject new accesses during
recovery. A mutex protects recovery-state updates.
- Return SIGBUS for blocked BAR faults instead of waiting in the fault
handler. VMM handling still needs target validation.
- Reject userspace SR-IOV changes while access is blocked.
- Exclude s390 and all CXL memory devices for now as s390 and CXL RCH
recovery can omit completion callbacks.
Design
------
Userspace opts in through VFIO_DEVICE_FEATURE_PCI_ERROR_RECOVERY. It
registers an eventfd and reads status flags and a sequence number. The
existing error eventfd is still signaled.
The callbacks handle the host recovery sequence:
- error_detected() blocks new accesses and drains admitted ones,
quiesces INTx, and revokes BAR mappings and exported DMA-BUFs. On a
normal channel it also clears bus mastering. It votes CAN_RECOVER
for a normal channel, NEED_RESET for a frozen channel, and DISCONNECT
for permanent failure. Failure to save PCI_COMMAND or disable bus
mastering leaves access blocked and returns NONE so other devices
can continue recovery.
- slot_reset() restores the configuration saved at open after a host
reset, then tears down interrupt configuration. Userspace must set
up interrupts again.
- resume() completes restoration, unblocks access and completes INTx
handling. A restoration failure keeps the device blocked.
Non-fatal errors also block device access because host recovery may
require a reset. Userspace reads the recovery status to determine
whether the device was reset or recovery failed.
Without opt-in, the existing error notification behavior is retained.
The open check also rejects a previously recorded host transaction when
recovery is disabled, so closing the device does not discard that state.
Locking and open questions
--------------------------
Several locks are involved in recovery. PCI core holds device_lock
while calling the driver, preventing concurrent driver binding or
unbinding. During the bus walk, it also holds pci_bus_sem for reading
to keep the device list stable. Adding or removing PCI devices requires
the write lock. This semaphore is shared across all PCI buses.
VFIO recovery blocks new accesses, waits for existing SRCU readers to
finish, and takes memory_lock. Power transitions also take memory_lock,
but entering D0 can then acquire pci_bus_sem to update PCIe link power
management (ASPM).
Local Sashiko/Claude review reports a lock-order conflict as below:
1. A guest D0 request holds memory_lock.
2. PCI core holds pci_bus_sem for reading while VFIO's recovery
callback waits for memory_lock.
3. Removal of an unrelated PCI device waits for the bus write lock.
4. D0's ASPM update requests the bus read lock and waits behind that
writer.
Each task waits for another, so none can proceed.
One possible solution could be(not implemented):
- PCI core guarantees that pci_bus_sem is held for reading before
invoking supported recovery callbacks. VFIO uses
pci_set_power_state_locked() in those callbacks, avoiding another
acquisition of the bus lock.
- A new pci_try_set_power_state() helper tries to acquire pci_bus_sem
and returns -EAGAIN on contention. Ordinary VFIO D0 paths use this
helper and handle failure instead of waiting while holding VFIO locks
Not sure there are better ways to handle this or not. Feedbacks appreciated
on this.
Another issue flagged was(I think this is a pre-existing one):
- Open/close coordination with the whole host recovery transaction
was already missing. Recovery can start after the open check, and
close can overlap the physical reset. Restoration of closed devices
is also not implemented.
Testing
-------
Basic sanity tests performed on a GB200 with an NVIDIA GPU assigned.
Kernel branch:
https://github.com/shamiali2008/linux/tree/vfio-aer-rfc-v2-ext
QEMU test branch is here(This registers the recovery eventfd and uses
pcie_aer_inject_error() to report the error to Guest)
https://github.com/shamiali2008/qemu-master/commits/private-master-vfio-aer-test-v2/
Software AER injection was performed using a modified pcieaer_inject
module.
./aer-inject nonfatal.conf
qemu-system-aarch64: info: vfio 0018:06:00.0: host PCI error recovery completed (seq 1)
Guest kernel:
[ 59.036988] pcieport 0000:01:00.0: AER: Uncorrectable (Non-Fatal) error message received from 0000:02:00.0
[ 59.038326] nvidia 0000:02:00.0: PCIe Bus Error: severity=Uncorrectable (Non-Fatal), type=Transaction Layer, (Completer ID)
[ 59.038440] nvidia 0000:02:00.0: device [10de:2941] error status/mask=00008000/02400000
[ 59.038533] nvidia 0000:02:00.0: [15] CmpltAbrt (First)
[ 59.038705] nvidia 0000:02:00.0: AER: TLP Header: 0x00000000 0x00000000 0x00000000 0x00000000
...
./aer-inject fatal.conf
qemu-system-aarch64: info: vfio 0018:06:00.0: host PCI error recovery started (seq 2, channel frozen)
qemu-system-aarch64: info: vfio 0018:06:00.0: host PCI error recovery completed (seq 2, device was reset)
Guest kernel:
[ 150.909306] pcieport 0000:01:00.0: AER: Uncorrectable (Fatal) error message received from 0000:02:00.0
[ 150.909566] nvidia 0000:02:00.0: AER: PCIe Bus Error: severity=Uncorrectable (Fatal), type=Inaccessible, (Unregistered Agent ID)
[ 150.909835] nvidia 0000:02:00.0: AER: can't recover (no error_detected callback)
[ 150.910062] pcieport 0000:01:00.0: unlocked secondary bus reset via: pciehp_reset_slot+0x54/0x98
[ 156.603467] pcieport 0000:01:00.0: AER: Root Port link has been reset (0)
...
Please take a look and let me know your feedback.
Thanks,
Shameer
Shameer Kolothum (16):
vfio/pci: Add a device access gate
vfio/pci: Gate config space access
vfio/pci: Buffer ROM reads before copying to userspace
vfio/pci: Gate BAR and ROM access
vfio/pci: Fail BAR faults while access is blocked
vfio/pci: Gate interrupt configuration
vfio/pci: Gate function reset and runtime power management
vfio/pci: Gate device information queries and DMA-BUF export
vfio/pci: Add PCI error recovery state
vfio/pci: Quiesce INTx while access is blocked
vfio/pci: Add INTx recovery start and finish helpers
vfio/pci: Restore device state from slot_reset()
vfio/pci: Complete recovery in resume()
vfio/pci: Block device access during host recovery
vfio/pci: Add VFIO_DEVICE_FEATURE_PCI_ERROR_RECOVERY
vfio/pci: Enable host PCI error recovery for vfio-pci
drivers/vfio/pci/vfio_pci_priv.h | 10 +
include/linux/vfio_pci_core.h | 38 ++
include/uapi/linux/vfio.h | 44 ++
drivers/vfio/pci/vfio_pci.c | 12 +
drivers/vfio/pci/vfio_pci_config.c | 104 ++++-
drivers/vfio/pci/vfio_pci_core.c | 636 ++++++++++++++++++++++++++++-
drivers/vfio/pci/vfio_pci_dmabuf.c | 21 +-
drivers/vfio/pci/vfio_pci_intrs.c | 131 ++++++
drivers/vfio/pci/vfio_pci_rdwr.c | 169 ++++++--
9 files changed, 1090 insertions(+), 75 deletions(-)
--
2.43.0
next reply other threads:[~2026-09-29 17:33 UTC|newest]
Thread overview: 17+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-29 17:32 Shameer Kolothum [this message]
2026-09-29 17:32 ` [RFC PATCH v2 01/16] vfio/pci: Add a device access gate Shameer Kolothum
2026-09-29 17:32 ` [RFC PATCH v2 02/16] vfio/pci: Gate config space access Shameer Kolothum
2026-09-29 17:32 ` [RFC PATCH v2 03/16] vfio/pci: Buffer ROM reads before copying to userspace Shameer Kolothum
2026-09-29 17:32 ` [RFC PATCH v2 04/16] vfio/pci: Gate BAR and ROM access Shameer Kolothum
2026-09-29 17:32 ` [RFC PATCH v2 05/16] vfio/pci: Fail BAR faults while access is blocked Shameer Kolothum
2026-09-29 17:32 ` [RFC PATCH v2 06/16] vfio/pci: Gate interrupt configuration Shameer Kolothum
2026-09-29 17:32 ` [RFC PATCH v2 07/16] vfio/pci: Gate function reset and runtime power management Shameer Kolothum
2026-09-29 17:32 ` [RFC PATCH v2 08/16] vfio/pci: Gate device information queries and DMA-BUF export Shameer Kolothum
2026-09-29 17:32 ` [RFC PATCH v2 09/16] vfio/pci: Add PCI error recovery state Shameer Kolothum
2026-09-29 17:32 ` [RFC PATCH v2 10/16] vfio/pci: Quiesce INTx while access is blocked Shameer Kolothum
2026-09-29 17:33 ` [RFC PATCH v2 11/16] vfio/pci: Add INTx recovery start and finish helpers Shameer Kolothum
2026-09-29 17:33 ` [RFC PATCH v2 12/16] vfio/pci: Restore device state from slot_reset() Shameer Kolothum
2026-09-29 17:33 ` [RFC PATCH v2 13/16] vfio/pci: Complete recovery in resume() Shameer Kolothum
2026-09-29 17:33 ` [RFC PATCH v2 14/16] vfio/pci: Block device access during host recovery Shameer Kolothum
2026-09-29 17:33 ` [RFC PATCH v2 15/16] vfio/pci: Add VFIO_DEVICE_FEATURE_PCI_ERROR_RECOVERY Shameer Kolothum
2026-09-29 17:33 ` [RFC PATCH v2 16/16] vfio/pci: Enable host PCI error recovery for vfio-pci Shameer Kolothum
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260929173305.204856-1-skolothumtho@nvidia.com \
--to=skolothumtho@nvidia.com \
--cc=alex@shazbot.org \
--cc=ankita@nvidia.com \
--cc=jgg@ziepe.ca \
--cc=kbusch@meta.com \
--cc=kevin.tian@intel.com \
--cc=kvm@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-pci@vger.kernel.org \
--cc=michal.winiarski@intel.com \
--cc=mochs@nvidia.com \
--cc=nathanc@nvidia.com \
--cc=satyanarayana.k.v.p@intel.com \
--cc=sonangp@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®