From: Joe Damato <joe@dama.to>
To: Zhang Yunfei <zhangyunfei1@kylinos.cn>
Cc: netdev@vger.kernel.org, jiawenwu@trustnetic.com,
mengyuanlou@net-swift.com, andrew+netdev@lunn.ch,
davem@davemloft.net, edumazet@google.com, kuba@kernel.org,
pabeni@redhat.com, aleksandr.loktionov@intel.com,
u.kleine-koenig@baylibre.com, weirongguang@kylinos.cn,
linux-kernel@vger.kernel.org, stable@vger.kernel.org,
leitao@debian.org, Sashiko <sashiko-bot@kernel.org>
Subject: Re: [PATCH net v3 1/2] net: ngbe: propagate resume errors to the PM core
Date: Wed, 30 Sep 2026 13:00:03 -0700 [thread overview]
Message-ID: <ar1qQ0z41yRMPflz@devvm20253.cco0.facebook.com> (raw)
In-Reply-To: <20260930094748.1198085-2-zhangyunfei1@kylinos.cn>
On Wed, Sep 30, 2026 at 05:47:47PM +0800, Zhang Yunfei wrote:
> After a failed resume from suspend, the device stays detached from
> the networking stack: every attempt to bring the interface up fails
> the netif_device_present() check in __dev_open() with -ENODEV,
> while the PM core is told that resume succeeded.
>
> ngbe_resume() declares err as u32 and unconditionally returns 0:
> the return value of ngbe_reset_hw() is ignored entirely, and
> failures of wx_init_interrupt_scheme() and ngbe_open() are silently
> swallowed. The interrupt scheme torn down at suspend is never
> rebuilt, so the device cannot self-heal.
>
> Fix the type to int and propagate the errors, making the whole tail
> of the resume path consistent with the pci_enable_device_mem()
> failure path at the top. If the hardware reset fails, return early:
> the remaining resume steps cannot succeed. The suspend path may
> already have torn the interface down (ngbe_close() and
> wx_clear_interrupt_scheme()), leaving freed rings and IRQs behind
> while netif_running() still reports true, so set WX_STATE_RES_FREED
> on every failing return, the same mark ngbe_down_suspend() uses for
> the PCI error recovery path: a later ngbe_close() skips the
> teardown of the already-freed state, and ngbe_up_complete() clears
> the bit again so later opens are unaffected. This is safe on the
> wx_init_interrupt_scheme() failure path too, as it cleans up after
> itself.
>
> Found by manual code inspection of the PM error paths; the missing
> ngbe_reset_hw() propagation was reported by Sashiko in its review
> of v1. The failure paths are unreachable without fault injection:
> a loadable test module injects failures into ngbe_reset_hw(),
> wx_init_interrupt_scheme() and ngbe_open(). In a QEMU VM the
> unfixed driver hits the kernel "Trying to free already-free IRQ"
> warning on the post-suspend close; with the fix, the error is
> reported and no warning appears. No physical ngbe device is
> involved.
>
> Fixes: 6963e463256e ("net: ngbe: add Wake on Lan support")
> Reported-by: Sashiko <sashiko-bot@kernel.org>
> Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260917090050.1927999-1-zhangyunfei1%40kylinos.cn
> Cc: stable@vger.kernel.org
> Signed-off-by: Zhang Yunfei <zhangyunfei1@kylinos.cn>
> ---
> Changes in v3:
> - set WX_STATE_RES_FREED on every failing return of ngbe_resume()
> (the pci_enable_device_mem() failure, the reset failure early return
> and the wx_init_interrupt_scheme()/ngbe_open() failures), so that a
> later ngbe_close() skips re-running the teardown on the already-freed
> post-suspend state (Sashiko review of v2);
>
> Changes in v2:
> - also propagate the ngbe_reset_hw() failure, so the whole tail of
> ngbe_resume() reports errors to the PM core;
> - drop the inaccurate "device can be re-probed" claim: the PM core
> records and logs the failure, there is no re-probe.
>
> drivers/net/ethernet/wangxun/ngbe/ngbe_main.c | 14 +++++++++++---
> 1 file changed, 11 insertions(+), 3 deletions(-)
propagating the error through seems sensible, so:
Reviewed-by: Joe Damato <joe@dama.to>
next prev parent reply other threads:[~2026-09-30 20:00 UTC|newest]
Thread overview: 6+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-30 9:47 [PATCH net v3 0/2] net: ngbe: fix error handling in resume and open paths Zhang Yunfei
2026-09-30 9:47 ` [PATCH net v3 1/2] net: ngbe: propagate resume errors to the PM core Zhang Yunfei
2026-09-30 20:00 ` Joe Damato [this message]
2026-09-30 9:47 ` [PATCH net v3 2/2] net: ngbe: clear DRV_LOAD bit when ngbe_open() fails Zhang Yunfei
2026-09-30 19:58 ` Joe Damato
2026-09-30 20:01 ` [PATCH net v3 0/2] net: ngbe: fix error handling in resume and open paths Joe Damato
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=ar1qQ0z41yRMPflz@devvm20253.cco0.facebook.com \
--to=joe@dama.to \
--cc=aleksandr.loktionov@intel.com \
--cc=andrew+netdev@lunn.ch \
--cc=davem@davemloft.net \
--cc=edumazet@google.com \
--cc=jiawenwu@trustnetic.com \
--cc=kuba@kernel.org \
--cc=leitao@debian.org \
--cc=linux-kernel@vger.kernel.org \
--cc=mengyuanlou@net-swift.com \
--cc=netdev@vger.kernel.org \
--cc=pabeni@redhat.com \
--cc=sashiko-bot@kernel.org \
--cc=stable@vger.kernel.org \
--cc=u.kleine-koenig@baylibre.com \
--cc=weirongguang@kylinos.cn \
--cc=zhangyunfei1@kylinos.cn \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®