* r8152: RX stops until rebind after -EPROTO on the bulk-in endpoint [not found] <CAM6_QD_PoTdwdM7DvzzgHRs-G5EvQu-c5oNGve=GuKV4wFYAvQ@mail.gmail.com> @ 2026-09-27 22:01 ` Dane Linssen 2026-09-28 10:03 ` Michal Pecio 2026-09-28 1:33 ` Andrew Lunn 1 sibling, 1 reply; 4+ messages in thread From: Dane Linssen @ 2026-09-27 22:01 UTC (permalink / raw) To: linux-usb, netdev Cc: Mathias Nyman, Alan Stern, Greg Kroah-Hartman, Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni, linux-kernel Hi, An RTL8156B on r8152 stops receiving and stays that way until the driver is unbound and rebound. TX keeps working and nothing is logged. This has happened 5 times since 2026-09-23 on one host. The two on 2026-09-27 were captured; both followed several minutes of about 10k rx / 6k tx packets/s. Kernel: 7.2.6-1-default (openSUSE MicroOS) Device: 0bda:8156, bcdDevice 31.04, fw rtl8156b-2 v3 10/20/23 Host: HP ENVY 13-ah0xxx (i5-8250U), xHCI 8086:9d2f, 5 Gbps Link: 1000baseT/Full, EEE inactive, pause off During the stall at 18:08, read twice 5 s apart: ethtool -S rx_packets 117909777 -> 117909777 rx_missed 9214 -> 9504 tx_packets 111602790 -> 111603644 host rx_packets 117835758 -> 117835758 xhci debugfs devices/02/ep-context, bulk-in endpoint: State running mult 1 max P. Streams 0 interval 125 us max ESIT payload 0 CErr 3 Type Bulk IN burst 3 maxp 1024 deq 000000010bd3f960 avg trb len 0, virt_state:0x40 virt_state reads 0x0 when healthy, both here and on a second host with the same adapter. The stall at 18:59 looked the same. 0x40 is EP_HARD_CLEAR_TOGGLE. xhci sets it when it hard-resets a halted endpoint, which includes a bulk transaction error that outlasted MAX_SOFT_RETRY, and clears it in xhci_endpoint_reset() once the class driver calls usb_clear_halt(). r8152 never calls usb_clear_halt(). On -EPROTO, read_bulk_callback() re-queues the buffer without logging, so as far as I can tell the endpoint stays out of sync until a rebind resets it. No "Rx status" warning was ever logged, so it wasn't a stall (-EPIPE). usbnet doesn't clear the halt on -EPROTO either, so I'm not sure whether the fix belongs in r8152 or on the xhci side. It looks close to what Mathias's RFC "fix xhci endpoint restart at EPROTO" (March 2026) discusses [1]. The host now keeps an xhci-hcd trace instance with only the reset, stop, set-dequeue and configure events, plus dynamic debug for the "Transfer error" and "-reset ep" messages. I'll send those from the next stall, or anything else that would help. [1] https://lore.kernel.org/linux-usb/?q=s%3A%22fix+xhci+endpoint+restart+at+EPROTO%22 Thanks, Dane Linssen ^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: r8152: RX stops until rebind after -EPROTO on the bulk-in endpoint 2026-09-27 22:01 ` r8152: RX stops until rebind after -EPROTO on the bulk-in endpoint Dane Linssen @ 2026-09-28 10:03 ` Michal Pecio 2026-09-28 12:48 ` Dane Linssen 0 siblings, 1 reply; 4+ messages in thread From: Michal Pecio @ 2026-09-28 10:03 UTC (permalink / raw) To: Dane Linssen Cc: linux-usb, netdev, Mathias Nyman, Alan Stern, Greg Kroah-Hartman, Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni, linux-kernel On Mon, 28 Sep 2026 00:01:31 +0200, Dane Linssen wrote: > xhci debugfs devices/02/ep-context, bulk-in endpoint: > State running mult 1 max P. Streams 0 interval 125 us > max ESIT payload 0 CErr 3 Type Bulk IN burst 3 maxp 1024 > deq 000000010bd3f960 avg trb len 0, virt_state:0x40 > > virt_state reads 0x0 when healthy, both here and on a second host > with the same adapter. The stall at 18:59 looked the same. > > 0x40 is EP_HARD_CLEAR_TOGGLE. xhci sets it when it hard-resets a > halted endpoint, which includes a bulk transaction error that > outlasted MAX_SOFT_RETRY, and clears it in xhci_endpoint_reset() > once the class driver calls usb_clear_halt(). r8152 never calls > usb_clear_halt(). On -EPROTO, read_bulk_callback() re-queues the > buffer without logging, so as far as I can tell the endpoint stays > out of sync until a rebind resets it. No "Rx status" warning was > ever logged, so it wasn't a stall (-EPIPE). Oh great, this stuff again. Yes, it seems you are getting xHCI "USB Transaction Error" aka -EPROTO and then xhci-hcd resets the host endpoint and clears it sequence state, but device endpoint remains in the former state. In USB 3.x, sequence number mismatch causes all future URBs to complete with -EPROTO again. (USB 2.0 could lose one USB packet but then it "recovers"). This should be visible with usbmon. Alternatively, on ASMedia controllers, future URBs never complete; AFAIU it's a HW bug. Out of curiosity, what's your xHCI chip? As a bandaid, you could try increasing MAX_SOFT_RETRY or this: https://lore.kernel.org/linux-usb/20260905101837.4b7849c5.michal.pecio@gmail.com/ You mentioned using dynamic debug. Do you see "Transfer error" messages randomly during operation, or is it only one burst right before the failure? Is your controller ASM4242 by any chance? > usbnet doesn't clear the halt on -EPROTO either, so I'm not sure > whether the fix belongs in r8152 or on the xhci side. It looks close > to what Mathias's RFC "fix xhci endpoint restart at EPROTO" (March > 2026) discusses [1]. This xHCI patch only tried to prevent URB execution before the driver calls usb_clear_halt(). This seems a good policy for -EPIPE and maybe also for -EPROTO on USB 3.x devices, since sequence mismatch renders them unusable anyway. We've been reluctant to touch USB 2.0. But somebody still needs to call usb_clear_halt(). Alan Stern thought it could be USB core, but these patches haven't materialized and TBH there is nothing wrong with drivers like r8152 calling it. In fact, some class specs (like mass storage) seem to imply this, by wanting other class-specific operations to happen before clear halt. Is it doable for r8152 to call this before continuing operation? xHCI side can be fixed and things might work, at least for r8152. Regards, Michal ^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: r8152: RX stops until rebind after -EPROTO on the bulk-in endpoint 2026-09-28 10:03 ` Michal Pecio @ 2026-09-28 12:48 ` Dane Linssen 0 siblings, 0 replies; 4+ messages in thread From: Dane Linssen @ 2026-09-28 12:48 UTC (permalink / raw) To: Michal Pecio Cc: linux-usb, netdev, Mathias Nyman, Alan Stern, Greg Kroah-Hartman, Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni, linux-kernel Thanks, everyone, for the quick replies. Andrew: As far as I can tell it isn't a regression. The adapter is new (2026-09-17) and has only run 7.2.5 and 7.2.6, which don't differ in r8152 or xhci. e765ab012f73 ("usb: xhci: Improve Soft Retries after short transfers", v7.1) should make it rarer, not more common. Bisecting isn't practically possible, as the stalls are days apart. Michal: > Out of curiosity, what's your xHCI chip? Intel Sunrise Point-LP, 8086:9d2f (i5-8250U), so not ASMedia. The other host with the same adapter, which hasn't stalled, has an Intel Comet Lake-LP, 8086:02ed. > As a bandaid, you could try increasing MAX_SOFT_RETRY or this: > https://lore.kernel.org/linux-usb/20260905101837.4b7849c5.michal.pecio@gmail.com/ Thank you. Would you like me to try this to gather more data? If not, I already have a userspace watchdog that rebinds r8152 when the LAN is unreachable and the bulk-in endpoint shows virt_state 0x40. > You mentioned using dynamic debug. Do you see "Transfer error" > messages randomly during operation, or is it only one burst right > before the failure? I only enabled it after the two stalls I reported. In the 19 hours since, including 71 minutes above 5k rx packets/s (peak about 29k), there hasn't been a single "Transfer error" message, and no stall. So nothing random so far. The next stall will show whether it's a burst. > This should be visible with usbmon. usbmon's debugfs files are blocked by lockdown here, but /dev/usbmonN works, so if I have another stall I can send a 2 s usbmon capture of the bus as well. Regards, Dane ^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: r8152: RX stops until rebind after -EPROTO on the bulk-in endpoint [not found] <CAM6_QD_PoTdwdM7DvzzgHRs-G5EvQu-c5oNGve=GuKV4wFYAvQ@mail.gmail.com> 2026-09-27 22:01 ` r8152: RX stops until rebind after -EPROTO on the bulk-in endpoint Dane Linssen @ 2026-09-28 1:33 ` Andrew Lunn 1 sibling, 0 replies; 4+ messages in thread From: Andrew Lunn @ 2026-09-28 1:33 UTC (permalink / raw) To: Dane Linssen Cc: linux-usb, netdev, Mathias Nyman, Alan Stern, Greg Kroah-Hartman, Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni, linux-kernel On Sun, Sep 27, 2026 at 11:55:25PM +0200, Dane Linssen wrote: > Hi, > > An RTL8156B on r8152 stops receiving and stays that way until the > driver is unbound and rebound. TX keeps working and nothing is logged. Is this new, after a kernel upgrade? If so, can you do i git bisect to find the change which broke it? Andrew ^ permalink raw reply [flat|nested] 4+ messages in thread
end of thread, other threads:[~2026-09-28 12:48 UTC | newest]
Thread overview: 4+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
[not found] <CAM6_QD_PoTdwdM7DvzzgHRs-G5EvQu-c5oNGve=GuKV4wFYAvQ@mail.gmail.com>
2026-09-27 22:01 ` r8152: RX stops until rebind after -EPROTO on the bulk-in endpoint Dane Linssen
2026-09-28 10:03 ` Michal Pecio
2026-09-28 12:48 ` Dane Linssen
2026-09-28 1:33 ` Andrew Lunn
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®