mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Joe Damato <joe@dama.to>
To: Zhang Yunfei <zhangyunfei1@kylinos.cn>
Cc: netdev@vger.kernel.org, jiawenwu@trustnetic.com,
	mengyuanlou@net-swift.com, andrew+netdev@lunn.ch,
	davem@davemloft.net, edumazet@google.com, kuba@kernel.org,
	pabeni@redhat.com, aleksandr.loktionov@intel.com,
	u.kleine-koenig@baylibre.com, weirongguang@kylinos.cn,
	linux-kernel@vger.kernel.org, stable@vger.kernel.org,
	leitao@debian.org, Sashiko <sashiko-bot@kernel.org>
Subject: Re: [PATCH net v3 1/2] net: ngbe: propagate resume errors to the PM core
Date: Wed, 30 Sep 2026 13:00:03 -0700	[thread overview]
Message-ID: <ar1qQ0z41yRMPflz@devvm20253.cco0.facebook.com> (raw)
In-Reply-To: <20260930094748.1198085-2-zhangyunfei1@kylinos.cn>

On Wed, Sep 30, 2026 at 05:47:47PM +0800, Zhang Yunfei wrote:
> After a failed resume from suspend, the device stays detached from
> the networking stack: every attempt to bring the interface up fails
> the netif_device_present() check in __dev_open() with -ENODEV,
> while the PM core is told that resume succeeded.
> 
> ngbe_resume() declares err as u32 and unconditionally returns 0:
> the return value of ngbe_reset_hw() is ignored entirely, and
> failures of wx_init_interrupt_scheme() and ngbe_open() are silently
> swallowed. The interrupt scheme torn down at suspend is never
> rebuilt, so the device cannot self-heal.
> 
> Fix the type to int and propagate the errors, making the whole tail
> of the resume path consistent with the pci_enable_device_mem()
> failure path at the top. If the hardware reset fails, return early:
> the remaining resume steps cannot succeed. The suspend path may
> already have torn the interface down (ngbe_close() and
> wx_clear_interrupt_scheme()), leaving freed rings and IRQs behind
> while netif_running() still reports true, so set WX_STATE_RES_FREED
> on every failing return, the same mark ngbe_down_suspend() uses for
> the PCI error recovery path: a later ngbe_close() skips the
> teardown of the already-freed state, and ngbe_up_complete() clears
> the bit again so later opens are unaffected. This is safe on the
> wx_init_interrupt_scheme() failure path too, as it cleans up after
> itself.
> 
> Found by manual code inspection of the PM error paths; the missing
> ngbe_reset_hw() propagation was reported by Sashiko in its review
> of v1. The failure paths are unreachable without fault injection:
> a loadable test module injects failures into ngbe_reset_hw(),
> wx_init_interrupt_scheme() and ngbe_open(). In a QEMU VM the
> unfixed driver hits the kernel "Trying to free already-free IRQ"
> warning on the post-suspend close; with the fix, the error is
> reported and no warning appears. No physical ngbe device is
> involved.
> 
> Fixes: 6963e463256e ("net: ngbe: add Wake on Lan support")
> Reported-by: Sashiko <sashiko-bot@kernel.org>
> Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260917090050.1927999-1-zhangyunfei1%40kylinos.cn
> Cc: stable@vger.kernel.org
> Signed-off-by: Zhang Yunfei <zhangyunfei1@kylinos.cn>
> ---
> Changes in v3:
> - set WX_STATE_RES_FREED on every failing return of ngbe_resume()
>   (the pci_enable_device_mem() failure, the reset failure early return
>   and the wx_init_interrupt_scheme()/ngbe_open() failures), so that a
>   later ngbe_close() skips re-running the teardown on the already-freed
>   post-suspend state (Sashiko review of v2);
> 
> Changes in v2:
> - also propagate the ngbe_reset_hw() failure, so the whole tail of
>   ngbe_resume() reports errors to the PM core;
> - drop the inaccurate "device can be re-probed" claim: the PM core
>   records and logs the failure, there is no re-probe.
> 
>  drivers/net/ethernet/wangxun/ngbe/ngbe_main.c | 14 +++++++++++---
>  1 file changed, 11 insertions(+), 3 deletions(-)

propagating the error through seems sensible, so:

Reviewed-by: Joe Damato <joe@dama.to>

  reply	other threads:[~2026-09-30 20:00 UTC|newest]

Thread overview: 6+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-30  9:47 [PATCH net v3 0/2] net: ngbe: fix error handling in resume and open paths Zhang Yunfei
2026-09-30  9:47 ` [PATCH net v3 1/2] net: ngbe: propagate resume errors to the PM core Zhang Yunfei
2026-09-30 20:00   ` Joe Damato [this message]
2026-09-30  9:47 ` [PATCH net v3 2/2] net: ngbe: clear DRV_LOAD bit when ngbe_open() fails Zhang Yunfei
2026-09-30 19:58   ` Joe Damato
2026-09-30 20:01 ` [PATCH net v3 0/2] net: ngbe: fix error handling in resume and open paths Joe Damato

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=ar1qQ0z41yRMPflz@devvm20253.cco0.facebook.com \
    --to=joe@dama.to \
    --cc=aleksandr.loktionov@intel.com \
    --cc=andrew+netdev@lunn.ch \
    --cc=davem@davemloft.net \
    --cc=edumazet@google.com \
    --cc=jiawenwu@trustnetic.com \
    --cc=kuba@kernel.org \
    --cc=leitao@debian.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=mengyuanlou@net-swift.com \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=sashiko-bot@kernel.org \
    --cc=stable@vger.kernel.org \
    --cc=u.kleine-koenig@baylibre.com \
    --cc=weirongguang@kylinos.cn \
    --cc=zhangyunfei1@kylinos.cn \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®