* [PATCH v5 0/2] fix two issues in hisi_acc_vfio_pci driver
@ 2026-09-23 8:25 Longfang Liu
2026-09-23 8:25 ` [PATCH v5 1/2] hisi_acc_vfio_pci: fix NULL dereference in reset_prepare on PF passthrough Longfang Liu
2026-09-23 8:25 ` [PATCH v5 2/2] hisi_acc_vfio_pci: reject live migration on 64KB page with QM_HW_V3 hardware Longfang Liu
0 siblings, 2 replies; 4+ messages in thread
From: Longfang Liu @ 2026-09-23 8:25 UTC (permalink / raw)
To: alex.williamson, jgg; +Cc: kvm, linux-kernel, liulongfang
This series fixes two issues in the HiSilicon ACC VFIO PCI driver:
1. A NULL pointer dereference in reset_prepare() when a PF is bound
to the driver and passed through to a VM.
2. A 64KB page incompatibility on QM_HW_V3 hardware that breaks
migration and the guest kernel.
Changes since v4:
- Removed dev_err() log from reset_prepare().
- Removed redundant pdev->is_virtfn check from probe().
Changes since v3:
- Merged the set_reset_flag fix into patch 1 (was a separate patch in v3)
- Replaced pf_qm guards with a mig_ops check at function entry.
- Moved the 64KB page check from open_device to probe, selecting generic
ops instead of rejecting the device.
Changes since v2:
- Remove redundant qm->io_base check.
- Drop PAGE_ALIGN for BAR2 configuration region.
Longfang Liu (2):
hisi_acc_vfio_pci: fix NULL dereference in reset_prepare on PF
passthrough
hisi_acc_vfio_pci: reject live migration on 64KB page with QM_HW_V3
hardware
drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c | 43 ++++++++++++-------
1 file changed, 28 insertions(+), 15 deletions(-)
--
2.43.0
^ permalink raw reply [flat|nested] 4+ messages in thread* [PATCH v5 1/2] hisi_acc_vfio_pci: fix NULL dereference in reset_prepare on PF passthrough 2026-09-23 8:25 [PATCH v5 0/2] fix two issues in hisi_acc_vfio_pci driver Longfang Liu @ 2026-09-23 8:25 ` Longfang Liu 2026-09-29 21:51 ` Alex Williamson 2026-09-23 8:25 ` [PATCH v5 2/2] hisi_acc_vfio_pci: reject live migration on 64KB page with QM_HW_V3 hardware Longfang Liu 1 sibling, 1 reply; 4+ messages in thread From: Longfang Liu @ 2026-09-23 8:25 UTC (permalink / raw) To: alex.williamson, jgg; +Cc: kvm, linux-kernel, liulongfang When a PF is bound to the driver via driver_override and passed through to a VM, its pf_qm stays NULL. The PCI error handler reset_prepare() runs during open_device through pci_try_reset_function(), before the mig_ops gate, and dereferences the NULL pf_qm for the timeout log, crashing the kernel. Move the mig_ops check to the entry of reset_prepare() and aer_reset_done() so non-migration devices skip the QM_RESETTING coordination. Also clear set_reset_flag together with QM_RESETTING in aer_reset_done(); the flag was never cleared before, so a later timed-out reset could release a foreign lock. Fixes: b0eed085903e ("hisi_acc_vfio_pci: Add support for VFIO live migration") Fixes: a22099ed7936f ("hisi_acc_vfio_pci: fix VF reset timeout issue") Signed-off-by: Longfang Liu <liulongfang@huawei.com> --- drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c | 11 ++++++++--- 1 file changed, 8 insertions(+), 3 deletions(-) diff --git a/drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c b/drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c index 86362ec424a5..6a09252258b9 100644 --- a/drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c +++ b/drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c @@ -1157,6 +1157,9 @@ static void hisi_acc_vf_pci_reset_prepare(struct pci_dev *pdev) struct device *dev = &qm->pdev->dev; u32 delay = 0; + if (!hisi_acc_vdev->core_device.vdev.mig_ops) + return; + /* All reset requests need to be queued for processing */ while (test_and_set_bit(QM_RESETTING, &qm->misc_ctl)) { msleep(1); @@ -1174,12 +1177,14 @@ static void hisi_acc_vf_pci_aer_reset_done(struct pci_dev *pdev) struct hisi_acc_vf_core_device *hisi_acc_vdev = hisi_acc_drvdata(pdev); struct hisi_qm *qm = hisi_acc_vdev->pf_qm; - if (hisi_acc_vdev->set_reset_flag) - clear_bit(QM_RESETTING, &qm->misc_ctl); - if (!hisi_acc_vdev->core_device.vdev.mig_ops) return; + if (hisi_acc_vdev->set_reset_flag) { + clear_bit(QM_RESETTING, &qm->misc_ctl); + hisi_acc_vdev->set_reset_flag = false; + } + mutex_lock(&hisi_acc_vdev->state_mutex); hisi_acc_vf_reset(hisi_acc_vdev); mutex_unlock(&hisi_acc_vdev->state_mutex); -- 2.43.0 ^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [PATCH v5 1/2] hisi_acc_vfio_pci: fix NULL dereference in reset_prepare on PF passthrough 2026-09-23 8:25 ` [PATCH v5 1/2] hisi_acc_vfio_pci: fix NULL dereference in reset_prepare on PF passthrough Longfang Liu @ 2026-09-29 21:51 ` Alex Williamson 0 siblings, 0 replies; 4+ messages in thread From: Alex Williamson @ 2026-09-29 21:51 UTC (permalink / raw) To: Longfang Liu; +Cc: alex.williamson, jgg, kvm, linux-kernel, alex On Wed, 23 Sep 2026 16:25:50 +0800 Longfang Liu <liulongfang@huawei.com> wrote: > When a PF is bound to the driver via driver_override and passed > through to a VM, its pf_qm stays NULL. The PCI error handler > reset_prepare() runs during open_device through > pci_try_reset_function(), before the mig_ops gate, and dereferences > the NULL pf_qm for the timeout log, crashing the kernel. > > Move the mig_ops check to the entry of reset_prepare() and > aer_reset_done() so non-migration devices skip the QM_RESETTING > coordination. Also clear set_reset_flag together with QM_RESETTING > in aer_reset_done(); the flag was never cleared before, so a later > timed-out reset could release a foreign lock. > > Fixes: b0eed085903e ("hisi_acc_vfio_pci: Add support for VFIO live migration") > Fixes: a22099ed7936f ("hisi_acc_vfio_pci: fix VF reset timeout issue") > Signed-off-by: Longfang Liu <liulongfang@huawei.com> > --- > drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c | 11 ++++++++--- > 1 file changed, 8 insertions(+), 3 deletions(-) > > diff --git a/drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c b/drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c > index 86362ec424a5..6a09252258b9 100644 > --- a/drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c > +++ b/drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c > @@ -1157,6 +1157,9 @@ static void hisi_acc_vf_pci_reset_prepare(struct pci_dev *pdev) Rolling back to one more line of context in this function: struct hisi_qm *qm = hisi_acc_vdev->pf_qm; As per the commit log, we're trying to fix a problem where pf_qm is NULL and dereferenced, such as in this next line: > struct device *dev = &qm->pdev->dev; > u32 delay = 0; > > + if (!hisi_acc_vdev->core_device.vdev.mig_ops) > + return; > + > /* All reset requests need to be queued for processing */ > while (test_and_set_bit(QM_RESETTING, &qm->misc_ctl)) { > msleep(1); Further context: if (++delay > QM_RESET_WAIT_TIMEOUT) { dev_err(dev, "reset prepare failed\n"); return; } So the added test exits before the inline dereference testing QM_RESETTING, but you're relying on the compiler to move the resolution of dev into the branch here. The dereference bug is not fully resolved, it's only masked by the compiler optimization. The easy solution is to either move the struct device deref into the branch where it's needed, or (better) just open code &qm->pdev->dev on the dev_err() call. Thanks, Alex > @@ -1174,12 +1177,14 @@ static void hisi_acc_vf_pci_aer_reset_done(struct pci_dev *pdev) > struct hisi_acc_vf_core_device *hisi_acc_vdev = hisi_acc_drvdata(pdev); > struct hisi_qm *qm = hisi_acc_vdev->pf_qm; > > - if (hisi_acc_vdev->set_reset_flag) > - clear_bit(QM_RESETTING, &qm->misc_ctl); > - > if (!hisi_acc_vdev->core_device.vdev.mig_ops) > return; > > + if (hisi_acc_vdev->set_reset_flag) { > + clear_bit(QM_RESETTING, &qm->misc_ctl); > + hisi_acc_vdev->set_reset_flag = false; > + } > + > mutex_lock(&hisi_acc_vdev->state_mutex); > hisi_acc_vf_reset(hisi_acc_vdev); > mutex_unlock(&hisi_acc_vdev->state_mutex); ^ permalink raw reply [flat|nested] 4+ messages in thread
* [PATCH v5 2/2] hisi_acc_vfio_pci: reject live migration on 64KB page with QM_HW_V3 hardware 2026-09-23 8:25 [PATCH v5 0/2] fix two issues in hisi_acc_vfio_pci driver Longfang Liu 2026-09-23 8:25 ` [PATCH v5 1/2] hisi_acc_vfio_pci: fix NULL dereference in reset_prepare on PF passthrough Longfang Liu @ 2026-09-23 8:25 ` Longfang Liu 1 sibling, 0 replies; 4+ messages in thread From: Longfang Liu @ 2026-09-23 8:25 UTC (permalink / raw) To: alex.williamson, jgg; +Cc: kvm, linux-kernel, liulongfang On QM_HW_V3 (VF_CTRL mode) hardware with 64KB pages, the functional region and migration registers share one BAR2 physical page with no isolation, which leaks migration registers to the guest and causes guest kernel calltrace under the KVM emulated device scheme. The previous fix rejected the device in hisi_acc_vf_qm_init() during open_device, returning -EINVAL and breaking passthrough entirely. Move the drv_mode check into a helper called from probe() so that a VF_CTRL device on a 64KB page is bound to the generic ops, keeping passthrough working while disabling live migration. Fixes: b0eed085903e ("hisi_acc_vfio_pci: Add support for VFIO live migration") Signed-off-by: Longfang Liu <liulongfang@huawei.com> --- .../vfio/pci/hisilicon/hisi_acc_vfio_pci.c | 32 ++++++++++++------- 1 file changed, 20 insertions(+), 12 deletions(-) diff --git a/drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c b/drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c index 6a09252258b9..3b014f248078 100644 --- a/drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c +++ b/drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c @@ -1196,13 +1196,6 @@ static int hisi_acc_vf_qm_init(struct hisi_acc_vf_core_device *hisi_acc_vdev) struct hisi_qm *vf_qm = &hisi_acc_vdev->vf_qm; struct hisi_qm *pf_qm = hisi_acc_vdev->pf_qm; struct pci_dev *vf_dev = vdev->pdev; - u32 val; - - val = readl(pf_qm->io_base + QM_MIG_REGION_SEL); - if (pf_qm->ver > QM_HW_V3 && (val & QM_MIG_REGION_EN)) - hisi_acc_vdev->drv_mode = HW_ACC_MIG_PF_CTRL; - else - hisi_acc_vdev->drv_mode = HW_ACC_MIG_VF_CTRL; if (hisi_acc_vdev->drv_mode == HW_ACC_MIG_PF_CTRL) { /* @@ -1670,21 +1663,34 @@ static void hisi_acc_vf_debugfs_exit(struct hisi_acc_vf_core_device *hisi_acc_vd hisi_acc_vdev->debug_migf = NULL; } +static enum hw_drv_mode hisi_acc_vf_get_drv_mode(struct hisi_qm *pf_qm) +{ + u32 val; + + val = readl(pf_qm->io_base + QM_MIG_REGION_SEL); + if (pf_qm->ver > QM_HW_V3 && (val & QM_MIG_REGION_EN)) + return HW_ACC_MIG_PF_CTRL; + + return HW_ACC_MIG_VF_CTRL; +} + static int hisi_acc_vfio_pci_probe(struct pci_dev *pdev, const struct pci_device_id *id) { struct hisi_acc_vf_core_device *hisi_acc_vdev; const struct vfio_device_ops *ops = &hisi_acc_vfio_pci_ops; + enum hw_drv_mode drv_mode = HW_ACC_MIG_VF_CTRL; + resource_size_t func_len; struct hisi_qm *pf_qm; - int vf_id; int ret; pf_qm = hisi_acc_get_pf_qm(pdev); if (pf_qm && pf_qm->ver >= QM_HW_V3) { - vf_id = pci_iov_vf_id(pdev); - if (vf_id >= 0) - ops = &hisi_acc_vfio_pci_migrn_ops; + func_len = pci_resource_len(pdev, VFIO_PCI_BAR2_REGION_INDEX) >> 1; + drv_mode = hisi_acc_vf_get_drv_mode(pf_qm); + if (drv_mode == HW_ACC_MIG_VF_CTRL && func_len < PAGE_SIZE) + pci_warn(pdev, "migration not supported on 64KB pages with QM_HW_V3\n"); else - pci_warn(pdev, "migration support failed, continue with generic interface\n"); + ops = &hisi_acc_vfio_pci_migrn_ops; } hisi_acc_vdev = vfio_alloc_device(hisi_acc_vf_core_device, @@ -1692,6 +1698,8 @@ static int hisi_acc_vfio_pci_probe(struct pci_dev *pdev, const struct pci_device if (IS_ERR(hisi_acc_vdev)) return PTR_ERR(hisi_acc_vdev); + hisi_acc_vdev->drv_mode = drv_mode; + dev_set_drvdata(&pdev->dev, &hisi_acc_vdev->core_device); ret = vfio_pci_core_register_device(&hisi_acc_vdev->core_device); if (ret) -- 2.43.0 ^ permalink raw reply [flat|nested] 4+ messages in thread
end of thread, other threads:[~2026-09-29 21:51 UTC | newest] Thread overview: 4+ messages (download: mbox.gz / follow: Atom feed) -- links below jump to the message on this page -- 2026-09-23 8:25 [PATCH v5 0/2] fix two issues in hisi_acc_vfio_pci driver Longfang Liu 2026-09-23 8:25 ` [PATCH v5 1/2] hisi_acc_vfio_pci: fix NULL dereference in reset_prepare on PF passthrough Longfang Liu 2026-09-29 21:51 ` Alex Williamson 2026-09-23 8:25 ` [PATCH v5 2/2] hisi_acc_vfio_pci: reject live migration on 64KB page with QM_HW_V3 hardware Longfang Liu
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®