mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v2] powerpc/eeh: Fix recursive locking on devices without EEH sensitive driver
@ 2026-07-14 17:16 Shivaprasad G Bhat
  2026-07-15 10:16 ` Amit Machhiwal
  0 siblings, 1 reply; 2+ messages in thread
From: Shivaprasad G Bhat @ 2026-07-14 17:16 UTC (permalink / raw)
  To: maddy, linuxppc-dev
  Cc: harshpb, mpe, npiggin, chleroy, sbhat, ritesh.list, linux-kernel

The commit 1010b4c012b0 ("powerpc/eeh: Make EEH driver device hotplug
safe") refactored the EEH code such that the pci_rescan_remove_lock is
held at the beginning of eeh_handle_normal_event() and the
eeh_reset_device() is called with that lock being held. Looks like the
commit missed to remove the existing lock/unlock inside eeh_rmv_device()
which is no longer necessary. This is causing the eehd to hang on the
lock which it actually holds when that code path is taken.

[<0>] 0xc00000011c78f870
[<0>] __switch_to+0xfc/0x1a0
[<0>] pci_lock_rescan_remove+0x30/0x44
[<0>] eeh_rmv_device+0x290/0x2e0
[<0>] eeh_pe_dev_traverse+0x80/0x130
[<0>] eeh_reset_device+0xcc/0x23c
[<0>] eeh_handle_normal_event+0x830/0xa80
[<0>] eeh_event_handler+0xf8/0x190
[<0>] kthread+0x194/0x1b0
[<0>] start_kernel_thread+0x14/0x18

The issue is seen for cases where the errors are detected on the PHB
directly AND|OR for devices where the driver error_detected() returns
PCI_ERS_RESULT_NEED_RESET, and driver being not EEH sensitive(i.e no
error handlers like slot_reset(), resume() etc defined).

Fixes: 1010b4c012b0 ("powerpc/eeh: Make EEH driver device hotplug safe")
Cc: stable <stable@kernel.org>
Reviewed-by: Ritesh Harjani (IBM) <ritesh.list@gmail.com>
Signed-off-by: Shivaprasad G Bhat <sbhat@linux.ibm.com>
---
 arch/powerpc/kernel/eeh_driver.c |    2 --
 1 file changed, 2 deletions(-)

diff --git a/arch/powerpc/kernel/eeh_driver.c b/arch/powerpc/kernel/eeh_driver.c
index 028f69158532..d64cce17a4e0 100644
--- a/arch/powerpc/kernel/eeh_driver.c
+++ b/arch/powerpc/kernel/eeh_driver.c
@@ -533,9 +533,7 @@ static void eeh_rmv_device(struct eeh_dev *edev, void *userdata)
 		if (rmv_data)
 			list_add(&edev->rmv_entry, &rmv_data->removed_vf_list);
 	} else {
-		pci_lock_rescan_remove();
 		pci_stop_and_remove_bus_device(dev);
-		pci_unlock_rescan_remove();
 	}
 }
 



^ permalink raw reply	[flat|nested] 2+ messages in thread

* Re: [PATCH v2] powerpc/eeh: Fix recursive locking on devices without EEH sensitive driver
  2026-07-14 17:16 [PATCH v2] powerpc/eeh: Fix recursive locking on devices without EEH sensitive driver Shivaprasad G Bhat
@ 2026-07-15 10:16 ` Amit Machhiwal
  0 siblings, 0 replies; 2+ messages in thread
From: Amit Machhiwal @ 2026-07-15 10:16 UTC (permalink / raw)
  To: Shivaprasad G Bhat
  Cc: maddy, linuxppc-dev, harshpb, mpe, npiggin, chleroy, ritesh.list,
	linux-kernel

On 2026/07/14 05:16 PM, Shivaprasad G Bhat wrote:
> The commit 1010b4c012b0 ("powerpc/eeh: Make EEH driver device hotplug
> safe") refactored the EEH code such that the pci_rescan_remove_lock is
> held at the beginning of eeh_handle_normal_event() and the
> eeh_reset_device() is called with that lock being held. Looks like the
> commit missed to remove the existing lock/unlock inside eeh_rmv_device()
> which is no longer necessary. This is causing the eehd to hang on the
> lock which it actually holds when that code path is taken.
> 
> [<0>] 0xc00000011c78f870
> [<0>] __switch_to+0xfc/0x1a0
> [<0>] pci_lock_rescan_remove+0x30/0x44
> [<0>] eeh_rmv_device+0x290/0x2e0
> [<0>] eeh_pe_dev_traverse+0x80/0x130
> [<0>] eeh_reset_device+0xcc/0x23c
> [<0>] eeh_handle_normal_event+0x830/0xa80
> [<0>] eeh_event_handler+0xf8/0x190
> [<0>] kthread+0x194/0x1b0
> [<0>] start_kernel_thread+0x14/0x18
> 
> The issue is seen for cases where the errors are detected on the PHB
> directly AND|OR for devices where the driver error_detected() returns
> PCI_ERS_RESULT_NEED_RESET, and driver being not EEH sensitive(i.e no
> error handlers like slot_reset(), resume() etc defined).
> 
> Fixes: 1010b4c012b0 ("powerpc/eeh: Make EEH driver device hotplug safe")
> Cc: stable <stable@kernel.org>
> Reviewed-by: Ritesh Harjani (IBM) <ritesh.list@gmail.com>
> Signed-off-by: Shivaprasad G Bhat <sbhat@linux.ibm.com>
> ---
>  arch/powerpc/kernel/eeh_driver.c |    2 --
>  1 file changed, 2 deletions(-)

The change makes sense. eeh_handle_normal_event() already holds the
rescan/remove lock across eeh_reset_device(), so taking it again in
eeh_rmv_device() can deadlock exactly as described.

Reviewed-by: Amit Machhiwal <amachhiw@linux.ibm.com>

Thanks,
Amit

> 
> diff --git a/arch/powerpc/kernel/eeh_driver.c b/arch/powerpc/kernel/eeh_driver.c
> index 028f69158532..d64cce17a4e0 100644
> --- a/arch/powerpc/kernel/eeh_driver.c
> +++ b/arch/powerpc/kernel/eeh_driver.c
> @@ -533,9 +533,7 @@ static void eeh_rmv_device(struct eeh_dev *edev, void *userdata)
>  		if (rmv_data)
>  			list_add(&edev->rmv_entry, &rmv_data->removed_vf_list);
>  	} else {
> -		pci_lock_rescan_remove();
>  		pci_stop_and_remove_bus_device(dev);
> -		pci_unlock_rescan_remove();
>  	}
>  }
>  
> 
> 

^ permalink raw reply	[flat|nested] 2+ messages in thread

end of thread, other threads:[~2026-07-15 10:16 UTC | newest]

Thread overview: 2+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-07-14 17:16 [PATCH v2] powerpc/eeh: Fix recursive locking on devices without EEH sensitive driver Shivaprasad G Bhat
2026-07-15 10:16 ` Amit Machhiwal

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

Powered by JetHome