Re: r8152: RX stops until rebind after -EPROTO on the bulk-in endpoint
From: Michal Pecio
Date: Mon Sep 28 2026 - 06:04:05 EST
On Mon, 28 Sep 2026 00:01:31 +0200, Dane Linssen wrote:
> xhci debugfs devices/02/ep-context, bulk-in endpoint:
> State running mult 1 max P. Streams 0 interval 125 us
> max ESIT payload 0 CErr 3 Type Bulk IN burst 3 maxp 1024
> deq 000000010bd3f960 avg trb len 0, virt_state:0x40
>
> virt_state reads 0x0 when healthy, both here and on a second host
> with the same adapter. The stall at 18:59 looked the same.
>
> 0x40 is EP_HARD_CLEAR_TOGGLE. xhci sets it when it hard-resets a
> halted endpoint, which includes a bulk transaction error that
> outlasted MAX_SOFT_RETRY, and clears it in xhci_endpoint_reset()
> once the class driver calls usb_clear_halt(). r8152 never calls
> usb_clear_halt(). On -EPROTO, read_bulk_callback() re-queues the
> buffer without logging, so as far as I can tell the endpoint stays
> out of sync until a rebind resets it. No "Rx status" warning was
> ever logged, so it wasn't a stall (-EPIPE).
Oh great, this stuff again. Yes, it seems you are getting xHCI
"USB Transaction Error" aka -EPROTO and then xhci-hcd resets the
host endpoint and clears it sequence state, but device endpoint
remains in the former state. In USB 3.x, sequence number mismatch
causes all future URBs to complete with -EPROTO again. (USB 2.0
could lose one USB packet but then it "recovers").
This should be visible with usbmon.
Alternatively, on ASMedia controllers, future URBs never complete;
AFAIU it's a HW bug. Out of curiosity, what's your xHCI chip?
As a bandaid, you could try increasing MAX_SOFT_RETRY or this:
https://lore.kernel.org/linux-usb/20260905101837.4b7849c5.michal.pecio@xxxxxxxxx/
You mentioned using dynamic debug. Do you see "Transfer error"
messages randomly during operation, or is it only one burst right
before the failure? Is your controller ASM4242 by any chance?
> usbnet doesn't clear the halt on -EPROTO either, so I'm not sure
> whether the fix belongs in r8152 or on the xhci side. It looks close
> to what Mathias's RFC "fix xhci endpoint restart at EPROTO" (March
> 2026) discusses [1].
This xHCI patch only tried to prevent URB execution before the driver
calls usb_clear_halt(). This seems a good policy for -EPIPE and maybe
also for -EPROTO on USB 3.x devices, since sequence mismatch renders
them unusable anyway. We've been reluctant to touch USB 2.0.
But somebody still needs to call usb_clear_halt(). Alan Stern thought
it could be USB core, but these patches haven't materialized and TBH
there is nothing wrong with drivers like r8152 calling it. In fact,
some class specs (like mass storage) seem to imply this, by wanting
other class-specific operations to happen before clear halt.
Is it doable for r8152 to call this before continuing operation?
xHCI side can be fixed and things might work, at least for r8152.
Regards,
Michal