mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH net v3 0/2] net: ngbe: fix error handling in resume and open paths
@ 2026-09-30  9:47 Zhang Yunfei
  2026-09-30  9:47 ` [PATCH net v3 1/2] net: ngbe: propagate resume errors to the PM core Zhang Yunfei
                   ` (2 more replies)
  0 siblings, 3 replies; 6+ messages in thread
From: Zhang Yunfei @ 2026-09-30  9:47 UTC (permalink / raw)
  To: netdev
  Cc: jiawenwu, mengyuanlou, andrew+netdev, davem, edumazet, kuba,
	pabeni, aleksandr.loktionov, u.kleine-koenig, weirongguang,
	zhangyunfei1, linux-kernel, stable, leitao

Two error-handling fixes for the ngbe PM/open paths.

ngbe_resume() declared err as u32 and returned 0 unconditionally, so a
failed ngbe_reset_hw(), wx_init_interrupt_scheme() or ngbe_open() left
the device in netif_device_detach() state with a broken interrupt
scheme while the PM core was told the resume succeeded; the reset task
bails out on the missing netif_device_present() check, so the device
cannot self-heal. Patch 1 fixes the type and propagates all of these
errors, making the whole tail of the resume path consistent with the
pci_enable_device_mem() failure path at the top, which already reports
its error.

ngbe_open() sets the WX_CFG_PORT_CTL_DRV_LOAD bit to tell the
management firmware the host has taken over the port, but no error
path cleared it, leaving the firmware owning a port whose rings and
IRQs are gone. Patch 2 rolls the bit back on all open error paths,
matching ngbe_close() and ngbe_dev_shutdown().

---
Changes in v3:
- patch 1: set WX_STATE_RES_FREED on every failing return of
  ngbe_resume() (the pci_enable_device_mem() failure, the hardware
  reset failure early return and the wx_init_interrupt_scheme()/
  ngbe_open() failures), so that a later ngbe_close() (ndo_stop or
  unregister_netdev()) skips re-running the teardown on the
  already-freed post-suspend state instead of freeing IRQs that are no
  longer requested (Sashiko review of v2);
- patch 2: correct the Fixes tag to a1cf597b99a7, the commit that
  introduced the bug (the first ngbe_open() failure path after the
  DRV_LOAD bit is set; v2 pointed at e7956139a6cf, which added more
  failing returns but not the first one); no code change.

Changes in v2:
- also propagate the ngbe_reset_hw() failure, so the whole tail of
  ngbe_resume() reports errors to the PM core (Sashiko review of v1);
- drop the inaccurate "device can be re-probed" claim: the PM core
  records and logs the failure, there is no re-probe (Sashiko review
  of v1);
- patch 2/2 unchanged.

Link: https://lore.kernel.org/netdev/20260922100836.1147718-1-zhangyunfei1@kylinos.cn/T/#u/


Zhang Yunfei (2):
  net: ngbe: propagate resume errors to the PM core
  net: ngbe: clear DRV_LOAD bit when ngbe_open() fails

 drivers/net/ethernet/wangxun/ngbe/ngbe_main.c | 18 ++++++++++++++----
 1 file changed, 14 insertions(+), 4 deletions(-)


base-commit: 93f51579e7df248780214094418f205253383cc5
-- 
2.25.1


^ permalink raw reply	[flat|nested] 6+ messages in thread

end of thread, other threads:[~2026-09-30 20:01 UTC | newest]

Thread overview: 6+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-30  9:47 [PATCH net v3 0/2] net: ngbe: fix error handling in resume and open paths Zhang Yunfei
2026-09-30  9:47 ` [PATCH net v3 1/2] net: ngbe: propagate resume errors to the PM core Zhang Yunfei
2026-09-30 20:00   ` Joe Damato
2026-09-30  9:47 ` [PATCH net v3 2/2] net: ngbe: clear DRV_LOAD bit when ngbe_open() fails Zhang Yunfei
2026-09-30 19:58   ` Joe Damato
2026-09-30 20:01 ` [PATCH net v3 0/2] net: ngbe: fix error handling in resume and open paths Joe Damato

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®