mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* r8152: RX stops until rebind after -EPROTO on the bulk-in endpoint
       [not found] <CAM6_QD_PoTdwdM7DvzzgHRs-G5EvQu-c5oNGve=GuKV4wFYAvQ@mail.gmail.com>
@ 2026-09-27 22:01 ` Dane Linssen
  2026-09-28 10:03   ` Michal Pecio
  2026-09-28  1:33 ` Andrew Lunn
  1 sibling, 1 reply; 4+ messages in thread
From: Dane Linssen @ 2026-09-27 22:01 UTC (permalink / raw)
  To: linux-usb, netdev
  Cc: Mathias Nyman, Alan Stern, Greg Kroah-Hartman, Andrew Lunn,
	David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
	linux-kernel

Hi,

An RTL8156B on r8152 stops receiving and stays that way until the
driver is unbound and rebound. TX keeps working and nothing is logged.
This has happened 5 times since 2026-09-23 on one host. The two on
2026-09-27 were captured; both followed several minutes of about
10k rx / 6k tx packets/s.

  Kernel: 7.2.6-1-default (openSUSE MicroOS)
  Device: 0bda:8156, bcdDevice 31.04, fw rtl8156b-2 v3 10/20/23
  Host:   HP ENVY 13-ah0xxx (i5-8250U), xHCI 8086:9d2f, 5 Gbps
  Link:   1000baseT/Full, EEE inactive, pause off

During the stall at 18:08, read twice 5 s apart:

  ethtool -S  rx_packets  117909777 -> 117909777
              rx_missed        9214 ->      9504
              tx_packets  111602790 -> 111603644
  host        rx_packets  117835758 -> 117835758

  xhci debugfs devices/02/ep-context, bulk-in endpoint:
    State running mult 1 max P. Streams 0 interval 125 us
    max ESIT payload 0 CErr 3 Type Bulk IN burst 3 maxp 1024
    deq 000000010bd3f960 avg trb len 0, virt_state:0x40

virt_state reads 0x0 when healthy, both here and on a second host
with the same adapter. The stall at 18:59 looked the same.

0x40 is EP_HARD_CLEAR_TOGGLE. xhci sets it when it hard-resets a
halted endpoint, which includes a bulk transaction error that
outlasted MAX_SOFT_RETRY, and clears it in xhci_endpoint_reset() once
the class driver calls usb_clear_halt(). r8152 never calls
usb_clear_halt(). On -EPROTO, read_bulk_callback() re-queues the
buffer without logging, so as far as I can tell the endpoint stays
out of sync until a rebind resets it. No "Rx status" warning was
ever logged, so it wasn't a stall (-EPIPE).

usbnet doesn't clear the halt on -EPROTO either, so I'm not sure
whether the fix belongs in r8152 or on the xhci side. It looks close
to what Mathias's RFC "fix xhci endpoint restart at EPROTO" (March
2026) discusses [1].

The host now keeps an xhci-hcd trace instance with only the reset,
stop, set-dequeue and configure events, plus dynamic debug for the
"Transfer error" and "-reset ep" messages. I'll send those from the
next stall, or anything else that would help.

[1] https://lore.kernel.org/linux-usb/?q=s%3A%22fix+xhci+endpoint+restart+at+EPROTO%22

Thanks,
Dane Linssen

^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: r8152: RX stops until rebind after -EPROTO on the bulk-in endpoint
       [not found] <CAM6_QD_PoTdwdM7DvzzgHRs-G5EvQu-c5oNGve=GuKV4wFYAvQ@mail.gmail.com>
  2026-09-27 22:01 ` r8152: RX stops until rebind after -EPROTO on the bulk-in endpoint Dane Linssen
@ 2026-09-28  1:33 ` Andrew Lunn
  1 sibling, 0 replies; 4+ messages in thread
From: Andrew Lunn @ 2026-09-28  1:33 UTC (permalink / raw)
  To: Dane Linssen
  Cc: linux-usb, netdev, Mathias Nyman, Alan Stern, Greg Kroah-Hartman,
	Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
	Paolo Abeni, linux-kernel

On Sun, Sep 27, 2026 at 11:55:25PM +0200, Dane Linssen wrote:
> Hi,
> 
> An RTL8156B on r8152 stops receiving and stays that way until the
> driver is unbound and rebound. TX keeps working and nothing is logged.

Is this new, after a kernel upgrade? If so, can you do i git bisect to
find the change which broke it?

	Andrew

^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: r8152: RX stops until rebind after -EPROTO on the bulk-in endpoint
  2026-09-27 22:01 ` r8152: RX stops until rebind after -EPROTO on the bulk-in endpoint Dane Linssen
@ 2026-09-28 10:03   ` Michal Pecio
  2026-09-28 12:48     ` Dane Linssen
  0 siblings, 1 reply; 4+ messages in thread
From: Michal Pecio @ 2026-09-28 10:03 UTC (permalink / raw)
  To: Dane Linssen
  Cc: linux-usb, netdev, Mathias Nyman, Alan Stern, Greg Kroah-Hartman,
	Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
	Paolo Abeni, linux-kernel

On Mon, 28 Sep 2026 00:01:31 +0200, Dane Linssen wrote:
>   xhci debugfs devices/02/ep-context, bulk-in endpoint:
>     State running mult 1 max P. Streams 0 interval 125 us
>     max ESIT payload 0 CErr 3 Type Bulk IN burst 3 maxp 1024
>     deq 000000010bd3f960 avg trb len 0, virt_state:0x40
> 
> virt_state reads 0x0 when healthy, both here and on a second host
> with the same adapter. The stall at 18:59 looked the same.
> 
> 0x40 is EP_HARD_CLEAR_TOGGLE. xhci sets it when it hard-resets a
> halted endpoint, which includes a bulk transaction error that
> outlasted MAX_SOFT_RETRY, and clears it in xhci_endpoint_reset()
> once the class driver calls usb_clear_halt(). r8152 never calls
> usb_clear_halt(). On -EPROTO, read_bulk_callback() re-queues the
> buffer without logging, so as far as I can tell the endpoint stays
> out of sync until a rebind resets it. No "Rx status" warning was
> ever logged, so it wasn't a stall (-EPIPE).

Oh great, this stuff again. Yes, it seems you are getting xHCI
"USB Transaction Error" aka -EPROTO and then xhci-hcd resets the
host endpoint and clears it sequence state, but device endpoint
remains in the former state. In USB 3.x, sequence number mismatch
causes all future URBs to complete with -EPROTO again. (USB 2.0
could lose one USB packet but then it "recovers").

This should be visible with usbmon.

Alternatively, on ASMedia controllers, future URBs never complete;
AFAIU it's a HW bug. Out of curiosity, what's your xHCI chip?

As a bandaid, you could try increasing MAX_SOFT_RETRY or this:
https://lore.kernel.org/linux-usb/20260905101837.4b7849c5.michal.pecio@gmail.com/

You mentioned using dynamic debug. Do you see "Transfer error"
messages randomly during operation, or is it only one burst right
before the failure? Is your controller ASM4242 by any chance?

> usbnet doesn't clear the halt on -EPROTO either, so I'm not sure
> whether the fix belongs in r8152 or on the xhci side. It looks close
> to what Mathias's RFC "fix xhci endpoint restart at EPROTO" (March
> 2026) discusses [1].

This xHCI patch only tried to prevent URB execution before the driver
calls usb_clear_halt(). This seems a good policy for -EPIPE and maybe
also for -EPROTO on USB 3.x devices, since sequence mismatch renders
them unusable anyway. We've been reluctant to touch USB 2.0.

But somebody still needs to call usb_clear_halt(). Alan Stern thought
it could be USB core, but these patches haven't materialized and TBH
there is nothing wrong with drivers like r8152 calling it. In fact,
some class specs (like mass storage) seem to imply this, by wanting
other class-specific operations to happen before clear halt.

Is it doable for r8152 to call this before continuing operation?
xHCI side can be fixed and things might work, at least for r8152.

Regards,
Michal

^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: r8152: RX stops until rebind after -EPROTO on the bulk-in endpoint
  2026-09-28 10:03   ` Michal Pecio
@ 2026-09-28 12:48     ` Dane Linssen
  0 siblings, 0 replies; 4+ messages in thread
From: Dane Linssen @ 2026-09-28 12:48 UTC (permalink / raw)
  To: Michal Pecio
  Cc: linux-usb, netdev, Mathias Nyman, Alan Stern, Greg Kroah-Hartman,
	Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
	Paolo Abeni, linux-kernel

Thanks, everyone, for the quick replies.

Andrew:
As far as I can tell it isn't a regression. The adapter is new
(2026-09-17) and has only run 7.2.5 and 7.2.6, which don't differ in
r8152 or xhci. e765ab012f73 ("usb: xhci: Improve Soft Retries after
short transfers", v7.1) should make it rarer, not more common.
Bisecting isn't practically possible, as the stalls are days apart.

Michal:
> Out of curiosity, what's your xHCI chip?

Intel Sunrise Point-LP, 8086:9d2f (i5-8250U), so not ASMedia. The
other host with the same adapter, which hasn't stalled, has an Intel
Comet Lake-LP, 8086:02ed.

> As a bandaid, you could try increasing MAX_SOFT_RETRY or this:
> https://lore.kernel.org/linux-usb/20260905101837.4b7849c5.michal.pecio@gmail.com/

Thank you. Would you like me to try this to gather more data? If not,
I already have a userspace watchdog that rebinds r8152 when the LAN is
unreachable and the bulk-in endpoint shows virt_state 0x40.

> You mentioned using dynamic debug. Do you see "Transfer error"
> messages randomly during operation, or is it only one burst right
> before the failure?

I only enabled it after the two stalls I reported. In the 19 hours
since, including 71 minutes above 5k rx packets/s (peak about 29k),
there hasn't been a single "Transfer error" message, and no stall. So
nothing random so far. The next stall will show whether it's a burst.

> This should be visible with usbmon.

usbmon's debugfs files are blocked by lockdown here, but /dev/usbmonN
works, so if I have another stall I can send a 2 s usbmon capture of
the bus as well.

Regards,
Dane

^ permalink raw reply	[flat|nested] 4+ messages in thread

end of thread, other threads:[~2026-09-28 12:48 UTC | newest]

Thread overview: 4+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
     [not found] <CAM6_QD_PoTdwdM7DvzzgHRs-G5EvQu-c5oNGve=GuKV4wFYAvQ@mail.gmail.com>
2026-09-27 22:01 ` r8152: RX stops until rebind after -EPROTO on the bulk-in endpoint Dane Linssen
2026-09-28 10:03   ` Michal Pecio
2026-09-28 12:48     ` Dane Linssen
2026-09-28  1:33 ` Andrew Lunn

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®