* [PATCH v3] usb: dwc3: core: Fix RAM interface getting stuck during enumeration @ 2026-09-29 5:19 ` Krishna Kurapati 2026-09-29 6:45 ` Selvarasu Ganesan 0 siblings, 1 reply; 5+ messages in thread From: Krishna Kurapati @ 2026-09-29 5:19 UTC (permalink / raw) To: Thinh Nguyen, Greg Kroah-Hartman Cc: linux-usb, linux-kernel, Krishna Kurapati During plug-in/plug-out test cases, it is sometimes seen that no events are generated by the controller and all CSR register reads give "0" and CSR_Timeout bit gets set indicating that CSR reads/writes are timing out or timed out. The issue comes up on different instnaces of enumeration on different platforms. On SM8550, the debug log is as follows: Prepared a TRB on ep0out and did start transfer to get set address request from host: <...>-7191 [000] D..1. 66.421006: dwc3_gadget_ep_cmd: ep0out: cmd 'Start Transfer' [406] params 00000000 efffa000 00000000 --> status: Successful <...>-7191 [000] D..1. 66.421196: dwc3_event: event (0000c040): ep0out: Transfer Complete (sIL) [Setup Phase] <...>-7191 [000] D..1. 66.421197: dwc3_ctrl_req: Set Address(Addr = 01) An XFER NRDY is received on ep0in for zero length status phase and a Start Transfer was done on ep0in with 0-length packet in 2 Stage status phase: <...>-7191 [000] D..1. 66.421249: dwc3_event: event (000020c2): ep0in: Transfer Not Ready [00000000] (Not Active) [Status Phase] <...>-7191 [000] D..1. 66.421266: dwc3_prepare_trb: ep0in: trb ffffffc00fcfd000 (E0:D0) buf 00000000efffa000 size 0 ctrl 00000c33 sofn 00000000 (HLcs:SC:status2) <...>-7191 [000] D..1. 66.421387: dwc3_gadget_ep_cmd: ep0in: cmd 'Start Transfer' [406] params 00000000 efffa000 00000000 -->status: Successful A bus reset was then received directly after 500 msec. Software never got the cmd complete for the start transfer done in status phase. Here the RAM interface is stuck. So host issues a bus reset as link is idle for 500 msec: <...>-7191 [000] D..1. 66.935603: dwc3_event: event (00000101): Reset [U0] Then software sees that it is in status phase and we issue an ENDXFER on ep0in and it gets timedout waiting for the CMDACT to go '0': <...>-7191 [000] D..1. 66.958249: dwc3_gadget_ep_cmd: ep0in: cmd 'End Transfer' [10508] params 00000000 00000000 00000000 --> status: Timed Out Upon debug with Synopsys, the root cause is as follows: During any transfer, if the data is not successfully transmitted, then a Done (with failure) handshake is returned, so that the BMU can re-attempt the same data again by rewinding its data pointers. But, if the USB IN is a 0-length payload (which is what is happening in this case - 2 stage status phase of set_address), then there is no need to rewind the pointers and the Done (with failure) handshake is not returned for failure case. This keeps the Request-Done interface busy till the next Done handshake. The MAC sends the 0-length payload again when the host requests. If the transmission is successful this time, the Done (with success) handshake is provided back. Otherwise, it repeats the same steps again. If the cable is disconnected or if the Host aborts the transfer on 3 consecutive failed attempts, the Request-Done handshake is not complete. This keeps the interface busy. The subsequent RAM access cannot proceed until the above pending transfer is complete. This results in failure of any access to RAM address locations. Many of the EndPoint commands need to access the RAM and they would fail to complete successfully. Furthermore when cable removal happens, this would not generate a disconnect event and the "connected" flag remains true always blockin suspend. Synopsys confirmed that the issue is present on all USB3 devices and as a workaround, suggested to re-initialize device mode. Signed-off-by: Krishna Kurapati <krishna.kurapati@oss.qualcomm.com> --- This series has only been compile tested. The issue was reproduced easily with a certain kind of cable and CDP port of AMD based Lenovo laptop. I don't have access to the cable currently and hence only compile testing the fix for now. But the issue has popped up on OEM testing as well. Also, didn't add locking while calling error recovery work in gadget_ep_cmd since the caller is supposed to handle it. Changes to v3: - Using error receovery mechanism from [1]. Link to v2: https://lore.kernel.org/all/20260806-ram-interface-code-v2-1-fe4a0de31d42@oss.qualcomm.com/ Changes in v2: - Implemented gadget recovery mechanism during gadget_ep_cmd instead of handling this issue only during disconnect. Link to RFC: https://lore.kernel.org/all/20231011100214.25720-1-quic_kriskura@quicinc.com/ [1]: https://lore.kernel.org/all/20260915110637.17658-1-jiazi.liu1984@gmail.com/ --- drivers/usb/dwc3/gadget.c | 19 +++++++++++++++++++ 1 file changed, 19 insertions(+) diff --git a/drivers/usb/dwc3/gadget.c b/drivers/usb/dwc3/gadget.c index ee837235630a..a7c4514cf0e2 100644 --- a/drivers/usb/dwc3/gadget.c +++ b/drivers/usb/dwc3/gadget.c @@ -283,6 +283,8 @@ int dwc3_send_gadget_generic_command(struct dwc3 *dwc, unsigned int cmd, return ret; } +static void dwc3_schedule_err_recovery(struct dwc3 *dwc); + /** * dwc3_send_gadget_ep_cmd - issue an endpoint command * @dep: the endpoint to which the command is going to be issued @@ -432,6 +434,23 @@ int dwc3_send_gadget_ep_cmd(struct dwc3_ep *dep, unsigned int cmd, cmd_status = -ETIMEDOUT; } + /* + * STAR 5001544 - In some situations, like the cable is + * disconnected or if the Host aborts the transfer on 3 + * consecutive failed attempts, the Request-Done handshake is not + * complete. This keeps the RAM interface busy. + * + * The subsequent RAM access cannot proceed until the pending + * transfer is complete. This results in failure of any access + * to RAM address locations. Many of the EndPoint commands need to + * access the RAM and they would fail to complete successfully. + * + * If the depcmd doesn't match the actual command, trigger controller + * recovery. + */ + if (DWC3_DEPCMD_CMD(reg) != DWC3_DEPCMD_CMD(cmd)) + dwc3_schedule_err_recovery(dwc); + skip_status: trace_dwc3_gadget_ep_cmd(dep, cmd, params, cmd_status); --- base-commit: ab29ca7714b82485ebb31d66835ccd5364221a76 change-id: 20260929-ram-interface-stuck-v3-aaf0c4b13b98 Best regards, -- Krishna Kurapati <krishna.kurapati@oss.qualcomm.com> ^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH v3] usb: dwc3: core: Fix RAM interface getting stuck during enumeration 2026-09-29 5:19 ` [PATCH v3] usb: dwc3: core: Fix RAM interface getting stuck during enumeration Krishna Kurapati @ 2026-09-29 6:45 ` Selvarasu Ganesan 2026-09-29 8:35 ` Krishna Kurapati 0 siblings, 1 reply; 5+ messages in thread From: Selvarasu Ganesan @ 2026-09-29 6:45 UTC (permalink / raw) To: Krishna Kurapati, Thinh Nguyen, Greg Kroah-Hartman Cc: linux-usb, linux-kernel On 9/29/2026 10:49 AM, Krishna Kurapati wrote: > During plug-in/plug-out test cases, it is sometimes seen that no events > are generated by the controller and all CSR register reads give "0" and > CSR_Timeout bit gets set indicating that CSR reads/writes are timing out > or timed out. Hi Krishna, As per our discussion with synopsys we got a below feedback for CSR timeout in different issue. "Note: CSR timeout is not expected to occur during functional mode. Once timeout is reported by controller, s/w should treat it as a fatal error and should identify the root cause of issuing the parallel access during soft reset and fix it" As per our understanding, there is no recovery from CSR timeout. Are you able to recover from CSR timeout in your case even after trigger controller recovery?. > > The issue comes up on different instnaces of enumeration on different > platforms. On SM8550, the debug log is as follows: > > Prepared a TRB on ep0out and did start transfer to get set > address request from host: > > <...>-7191 [000] D..1. 66.421006: dwc3_gadget_ep_cmd: ep0out: > cmd 'Start Transfer' [406] params 00000000 efffa000 00000000 --> > status: Successful > > <...>-7191 [000] D..1. 66.421196: dwc3_event: event (0000c040): > ep0out: Transfer Complete (sIL) [Setup Phase] > > <...>-7191 [000] D..1. 66.421197: dwc3_ctrl_req: Set > Address(Addr = 01) > > An XFER NRDY is received on ep0in for zero length status phase and > a Start Transfer was done on ep0in with 0-length packet in 2 Stage > status phase: > > <...>-7191 [000] D..1. 66.421249: dwc3_event: event (000020c2): > ep0in: Transfer Not Ready [00000000] (Not Active) [Status Phase] > > <...>-7191 [000] D..1. 66.421266: dwc3_prepare_trb: ep0in: trb > ffffffc00fcfd000 (E0:D0) buf 00000000efffa000 size 0 ctrl 00000c33 > sofn 00000000 (HLcs:SC:status2) > > <...>-7191 [000] D..1. 66.421387: dwc3_gadget_ep_cmd: ep0in: cmd > 'Start Transfer' [406] params 00000000 efffa000 00000000 -->status: > Successful > > A bus reset was then received directly after 500 msec. Software never > got the cmd complete for the start transfer done in status phase. Here > the RAM interface is stuck. So host issues a bus reset as link is > idle for 500 msec: > > <...>-7191 [000] D..1. 66.935603: dwc3_event: event (00000101): > Reset [U0] > > Then software sees that it is in status phase and we issue an ENDXFER > on ep0in and it gets timedout waiting for the CMDACT to go '0': > > <...>-7191 [000] D..1. 66.958249: dwc3_gadget_ep_cmd: ep0in: cmd > 'End Transfer' [10508] params 00000000 00000000 00000000 --> status: > Timed Out > > Upon debug with Synopsys, the root cause is as follows: > > During any transfer, if the data is not successfully transmitted, > then a Done (with failure) handshake is returned, so that the BMU > can re-attempt the same data again by rewinding its data pointers. > > But, if the USB IN is a 0-length payload (which is what is happening > in this case - 2 stage status phase of set_address), then there is no > need to rewind the pointers and the Done (with failure) handshake is > not returned for failure case. This keeps the Request-Done interface > busy till the next Done handshake. The MAC sends the 0-length payload > again when the host requests. If the transmission is successful this > time, the Done (with success) handshake is provided back. Otherwise, > it repeats the same steps again. > > If the cable is disconnected or if the Host aborts the transfer on 3 > consecutive failed attempts, the Request-Done handshake is not > complete. This keeps the interface busy. > > The subsequent RAM access cannot proceed until the above pending > transfer is complete. This results in failure of any access to RAM > address locations. Many of the EndPoint commands need to access the > RAM and they would fail to complete successfully. > > Furthermore when cable removal happens, this would not generate a > disconnect event and the "connected" flag remains true always blockin > suspend. > > Synopsys confirmed that the issue is present on all USB3 devices and > as a workaround, suggested to re-initialize device mode. > > Signed-off-by: Krishna Kurapati <krishna.kurapati@oss.qualcomm.com> > --- > This series has only been compile tested. The issue was reproduced > easily with a certain kind of cable and CDP port of AMD based Lenovo > laptop. I don't have access to the cable currently and hence only > compile testing the fix for now. But the issue has popped up on OEM > testing as well. > > Also, didn't add locking while calling error recovery work in > gadget_ep_cmd since the caller is supposed to handle it. > > Changes to v3: > - Using error receovery mechanism from [1]. > > Link to v2: > https://lore.kernel.org/all/20260806-ram-interface-code-v2-1-fe4a0de31d42@oss.qualcomm.com/ > > Changes in v2: > - Implemented gadget recovery mechanism during gadget_ep_cmd instead of > handling this issue only during disconnect. > > Link to RFC: > https://lore.kernel.org/all/20231011100214.25720-1-quic_kriskura@quicinc.com/ > > [1]: https://lore.kernel.org/all/20260915110637.17658-1-jiazi.liu1984@gmail.com/ > --- > drivers/usb/dwc3/gadget.c | 19 +++++++++++++++++++ > 1 file changed, 19 insertions(+) > > diff --git a/drivers/usb/dwc3/gadget.c b/drivers/usb/dwc3/gadget.c > index ee837235630a..a7c4514cf0e2 100644 > --- a/drivers/usb/dwc3/gadget.c > +++ b/drivers/usb/dwc3/gadget.c > @@ -283,6 +283,8 @@ int dwc3_send_gadget_generic_command(struct dwc3 *dwc, unsigned int cmd, > return ret; > } > > +static void dwc3_schedule_err_recovery(struct dwc3 *dwc); Where is the function implementation for this function?. Could you please point me on what you are doing in this function? Thanks, Selva > + > /** > * dwc3_send_gadget_ep_cmd - issue an endpoint command > * @dep: the endpoint to which the command is going to be issued > @@ -432,6 +434,23 @@ int dwc3_send_gadget_ep_cmd(struct dwc3_ep *dep, unsigned int cmd, > cmd_status = -ETIMEDOUT; > } > > + /* > + * STAR 5001544 - In some situations, like the cable is > + * disconnected or if the Host aborts the transfer on 3 > + * consecutive failed attempts, the Request-Done handshake is not > + * complete. This keeps the RAM interface busy. > + * > + * The subsequent RAM access cannot proceed until the pending > + * transfer is complete. This results in failure of any access > + * to RAM address locations. Many of the EndPoint commands need to > + * access the RAM and they would fail to complete successfully. > + * > + * If the depcmd doesn't match the actual command, trigger controller > + * recovery. > + */ > + if (DWC3_DEPCMD_CMD(reg) != DWC3_DEPCMD_CMD(cmd)) > + dwc3_schedule_err_recovery(dwc); > + > skip_status: > trace_dwc3_gadget_ep_cmd(dep, cmd, params, cmd_status); > > > --- > base-commit: ab29ca7714b82485ebb31d66835ccd5364221a76 > change-id: 20260929-ram-interface-stuck-v3-aaf0c4b13b98 > > Best regards, > -- > Krishna Kurapati <krishna.kurapati@oss.qualcomm.com> ^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH v3] usb: dwc3: core: Fix RAM interface getting stuck during enumeration 2026-09-29 6:45 ` Selvarasu Ganesan @ 2026-09-29 8:35 ` Krishna Kurapati 2026-09-29 11:02 ` Selvarasu Ganesan 0 siblings, 1 reply; 5+ messages in thread From: Krishna Kurapati @ 2026-09-29 8:35 UTC (permalink / raw) To: Selvarasu Ganesan, Thinh Nguyen Cc: linux-usb, linux-kernel, Greg Kroah-Hartman On 9/29/2026 12:15 PM, Selvarasu Ganesan wrote: > > On 9/29/2026 10:49 AM, Krishna Kurapati wrote: >> During plug-in/plug-out test cases, it is sometimes seen that no events >> are generated by the controller and all CSR register reads give "0" and >> CSR_Timeout bit gets set indicating that CSR reads/writes are timing out >> or timed out. > Hi Krishna, > > As per our discussion with synopsys we got a below feedback for CSR > timeout in different issue. > > "Note: CSR timeout is not expected to occur during functional mode. Once > timeout is reported by controller, s/w should treat it as a fatal error > and should identify the root cause of issuing the parallel access during > soft reset and fix it" > > As per our understanding, there is no recovery from CSR timeout. > > Are you able to recover from CSR timeout in your case even after trigger > controller recovery?. > Disabling and re-enabling gadget mode is the recommended fix. I did not test this particular patch as mentioned below, but this error recovery mechanism does the same thing as a soft_disconnect followed by soft_connect. Regards, Krishna, > >> >> The issue comes up on different instnaces of enumeration on different >> platforms. On SM8550, the debug log is as follows: >> >> Prepared a TRB on ep0out and did start transfer to get set >> address request from host: >> >> <...>-7191 [000] D..1. 66.421006: dwc3_gadget_ep_cmd: ep0out: >> cmd 'Start Transfer' [406] params 00000000 efffa000 00000000 --> >> status: Successful >> >> <...>-7191 [000] D..1. 66.421196: dwc3_event: event (0000c040): >> ep0out: Transfer Complete (sIL) [Setup Phase] >> >> <...>-7191 [000] D..1. 66.421197: dwc3_ctrl_req: Set >> Address(Addr = 01) >> >> An XFER NRDY is received on ep0in for zero length status phase and >> a Start Transfer was done on ep0in with 0-length packet in 2 Stage >> status phase: >> >> <...>-7191 [000] D..1. 66.421249: dwc3_event: event (000020c2): >> ep0in: Transfer Not Ready [00000000] (Not Active) [Status Phase] >> >> <...>-7191 [000] D..1. 66.421266: dwc3_prepare_trb: ep0in: trb >> ffffffc00fcfd000 (E0:D0) buf 00000000efffa000 size 0 ctrl 00000c33 >> sofn 00000000 (HLcs:SC:status2) >> >> <...>-7191 [000] D..1. 66.421387: dwc3_gadget_ep_cmd: ep0in: cmd >> 'Start Transfer' [406] params 00000000 efffa000 00000000 -->status: >> Successful >> >> A bus reset was then received directly after 500 msec. Software never >> got the cmd complete for the start transfer done in status phase. Here >> the RAM interface is stuck. So host issues a bus reset as link is >> idle for 500 msec: >> >> <...>-7191 [000] D..1. 66.935603: dwc3_event: event (00000101): >> Reset [U0] >> >> Then software sees that it is in status phase and we issue an ENDXFER >> on ep0in and it gets timedout waiting for the CMDACT to go '0': >> >> <...>-7191 [000] D..1. 66.958249: dwc3_gadget_ep_cmd: ep0in: cmd >> 'End Transfer' [10508] params 00000000 00000000 00000000 --> status: >> Timed Out >> >> Upon debug with Synopsys, the root cause is as follows: >> >> During any transfer, if the data is not successfully transmitted, >> then a Done (with failure) handshake is returned, so that the BMU >> can re-attempt the same data again by rewinding its data pointers. >> >> But, if the USB IN is a 0-length payload (which is what is happening >> in this case - 2 stage status phase of set_address), then there is no >> need to rewind the pointers and the Done (with failure) handshake is >> not returned for failure case. This keeps the Request-Done interface >> busy till the next Done handshake. The MAC sends the 0-length payload >> again when the host requests. If the transmission is successful this >> time, the Done (with success) handshake is provided back. Otherwise, >> it repeats the same steps again. >> >> If the cable is disconnected or if the Host aborts the transfer on 3 >> consecutive failed attempts, the Request-Done handshake is not >> complete. This keeps the interface busy. >> >> The subsequent RAM access cannot proceed until the above pending >> transfer is complete. This results in failure of any access to RAM >> address locations. Many of the EndPoint commands need to access the >> RAM and they would fail to complete successfully. >> >> Furthermore when cable removal happens, this would not generate a >> disconnect event and the "connected" flag remains true always blockin >> suspend. >> >> Synopsys confirmed that the issue is present on all USB3 devices and >> as a workaround, suggested to re-initialize device mode. >> >> Signed-off-by: Krishna Kurapati <krishna.kurapati@oss.qualcomm.com> >> --- >> This series has only been compile tested. The issue was reproduced >> easily with a certain kind of cable and CDP port of AMD based Lenovo >> laptop. I don't have access to the cable currently and hence only >> compile testing the fix for now. But the issue has popped up on OEM >> testing as well. >> >> Also, didn't add locking while calling error recovery work in >> gadget_ep_cmd since the caller is supposed to handle it. >> >> Changes to v3: >> - Using error receovery mechanism from [1]. >> >> Link to v2: >> https://lore.kernel.org/all/20260806-ram-interface-code-v2-1-fe4a0de31d42@oss.qualcomm.com/ >> >> Changes in v2: >> - Implemented gadget recovery mechanism during gadget_ep_cmd instead of >> handling this issue only during disconnect. >> >> Link to RFC: >> https://lore.kernel.org/all/20231011100214.25720-1-quic_kriskura@quicinc.com/ >> >> [1]: https://lore.kernel.org/all/20260915110637.17658-1-jiazi.liu1984@gmail.com/ >> --- >> drivers/usb/dwc3/gadget.c | 19 +++++++++++++++++++ >> 1 file changed, 19 insertions(+) >> >> diff --git a/drivers/usb/dwc3/gadget.c b/drivers/usb/dwc3/gadget.c >> index ee837235630a..a7c4514cf0e2 100644 >> --- a/drivers/usb/dwc3/gadget.c >> +++ b/drivers/usb/dwc3/gadget.c >> @@ -283,6 +283,8 @@ int dwc3_send_gadget_generic_command(struct dwc3 *dwc, unsigned int cmd, >> return ret; >> } >> >> +static void dwc3_schedule_err_recovery(struct dwc3 *dwc); > > Where is the function implementation for this function?. Could you > please point me on what you are doing in this function? > > > Thanks, > Selva > > >> + >> /** >> * dwc3_send_gadget_ep_cmd - issue an endpoint command >> * @dep: the endpoint to which the command is going to be issued >> @@ -432,6 +434,23 @@ int dwc3_send_gadget_ep_cmd(struct dwc3_ep *dep, unsigned int cmd, >> cmd_status = -ETIMEDOUT; >> } >> >> + /* >> + * STAR 5001544 - In some situations, like the cable is >> + * disconnected or if the Host aborts the transfer on 3 >> + * consecutive failed attempts, the Request-Done handshake is not >> + * complete. This keeps the RAM interface busy. >> + * >> + * The subsequent RAM access cannot proceed until the pending >> + * transfer is complete. This results in failure of any access >> + * to RAM address locations. Many of the EndPoint commands need to >> + * access the RAM and they would fail to complete successfully. >> + * >> + * If the depcmd doesn't match the actual command, trigger controller >> + * recovery. >> + */ >> + if (DWC3_DEPCMD_CMD(reg) != DWC3_DEPCMD_CMD(cmd)) >> + dwc3_schedule_err_recovery(dwc); >> + >> skip_status: >> trace_dwc3_gadget_ep_cmd(dep, cmd, params, cmd_status); >> >> >> --- >> base-commit: ab29ca7714b82485ebb31d66835ccd5364221a76 >> change-id: 20260929-ram-interface-stuck-v3-aaf0c4b13b98 >> >> Best regards, >> -- >> Krishna Kurapati <krishna.kurapati@oss.qualcomm.com> ^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH v3] usb: dwc3: core: Fix RAM interface getting stuck during enumeration 2026-09-29 8:35 ` Krishna Kurapati @ 2026-09-29 11:02 ` Selvarasu Ganesan 2026-09-29 17:50 ` Krishna Kurapati 0 siblings, 1 reply; 5+ messages in thread From: Selvarasu Ganesan @ 2026-09-29 11:02 UTC (permalink / raw) To: Krishna Kurapati, Thinh Nguyen Cc: linux-usb, linux-kernel, Greg Kroah-Hartman On 9/29/2026 2:05 PM, Krishna Kurapati wrote: > > > On 9/29/2026 12:15 PM, Selvarasu Ganesan wrote: >> >> On 9/29/2026 10:49 AM, Krishna Kurapati wrote: >>> During plug-in/plug-out test cases, it is sometimes seen that no events >>> are generated by the controller and all CSR register reads give "0" and >>> CSR_Timeout bit gets set indicating that CSR reads/writes are timing >>> out >>> or timed out. >> Hi Krishna, >> >> As per our discussion with synopsys we got a below feedback for CSR >> timeout in different issue. >> >> "Note: CSR timeout is not expected to occur during functional mode. Once >> timeout is reported by controller, s/w should treat it as a fatal error >> and should identify the root cause of issuing the parallel access during >> soft reset and fix it" >> >> As per our understanding, there is no recovery from CSR timeout. >> >> Are you able to recover from CSR timeout in your case even after trigger >> controller recovery?. >> > > Disabling and re-enabling gadget mode is the recommended fix. I did > not test this particular patch as mentioned below, but this error > recovery mechanism does the same thing as a soft_disconnect followed > by soft_connect. Hi Krishna, Thanks for your update. We tried to recover from the CSR timeout, but it was unsuccessful. The sequence below is provided for your reference. It would be best to identify the root cause of why the CSR timeout is occurring in your case. Issue Sequence: 1. Start core soft reset 2. Parallel access of other registers 3. End core soft reset 4. Getting CSR timeout Try to recovery from CSR timeout: 5. Clear CSR timeout bit by RW1C same bit 6. Start core soft reset 7. End core softt reset 8. Access other registers 9. Getting CSR timeout Thanks, Selva > > Regards, > Krishna, > >> >>> >>> The issue comes up on different instnaces of enumeration on different >>> platforms. On SM8550, the debug log is as follows: >>> >>> Prepared a TRB on ep0out and did start transfer to get set >>> address request from host: >>> >>> <...>-7191 [000] D..1. 66.421006: dwc3_gadget_ep_cmd: ep0out: >>> cmd 'Start Transfer' [406] params 00000000 efffa000 00000000 --> >>> status: Successful >>> >>> <...>-7191 [000] D..1. 66.421196: dwc3_event: event (0000c040): >>> ep0out: Transfer Complete (sIL) [Setup Phase] >>> >>> <...>-7191 [000] D..1. 66.421197: dwc3_ctrl_req: Set >>> Address(Addr = 01) >>> >>> An XFER NRDY is received on ep0in for zero length status phase and >>> a Start Transfer was done on ep0in with 0-length packet in 2 Stage >>> status phase: >>> >>> <...>-7191 [000] D..1. 66.421249: dwc3_event: event (000020c2): >>> ep0in: Transfer Not Ready [00000000] (Not Active) [Status Phase] >>> >>> <...>-7191 [000] D..1. 66.421266: dwc3_prepare_trb: ep0in: trb >>> ffffffc00fcfd000 (E0:D0) buf 00000000efffa000 size 0 ctrl 00000c33 >>> sofn 00000000 (HLcs:SC:status2) >>> >>> <...>-7191 [000] D..1. 66.421387: dwc3_gadget_ep_cmd: ep0in: cmd >>> 'Start Transfer' [406] params 00000000 efffa000 00000000 -->status: >>> Successful >>> >>> A bus reset was then received directly after 500 msec. Software never >>> got the cmd complete for the start transfer done in status phase. Here >>> the RAM interface is stuck. So host issues a bus reset as link is >>> idle for 500 msec: >>> >>> <...>-7191 [000] D..1. 66.935603: dwc3_event: event (00000101): >>> Reset [U0] >>> >>> Then software sees that it is in status phase and we issue an ENDXFER >>> on ep0in and it gets timedout waiting for the CMDACT to go '0': >>> >>> <...>-7191 [000] D..1. 66.958249: dwc3_gadget_ep_cmd: ep0in: cmd >>> 'End Transfer' [10508] params 00000000 00000000 00000000 --> status: >>> Timed Out >>> >>> Upon debug with Synopsys, the root cause is as follows: >>> >>> During any transfer, if the data is not successfully transmitted, >>> then a Done (with failure) handshake is returned, so that the BMU >>> can re-attempt the same data again by rewinding its data pointers. >>> >>> But, if the USB IN is a 0-length payload (which is what is happening >>> in this case - 2 stage status phase of set_address), then there is no >>> need to rewind the pointers and the Done (with failure) handshake is >>> not returned for failure case. This keeps the Request-Done interface >>> busy till the next Done handshake. The MAC sends the 0-length payload >>> again when the host requests. If the transmission is successful this >>> time, the Done (with success) handshake is provided back. Otherwise, >>> it repeats the same steps again. >>> >>> If the cable is disconnected or if the Host aborts the transfer on 3 >>> consecutive failed attempts, the Request-Done handshake is not >>> complete. This keeps the interface busy. >>> >>> The subsequent RAM access cannot proceed until the above pending >>> transfer is complete. This results in failure of any access to RAM >>> address locations. Many of the EndPoint commands need to access the >>> RAM and they would fail to complete successfully. >>> >>> Furthermore when cable removal happens, this would not generate a >>> disconnect event and the "connected" flag remains true always blockin >>> suspend. >>> >>> Synopsys confirmed that the issue is present on all USB3 devices and >>> as a workaround, suggested to re-initialize device mode. >>> >>> Signed-off-by: Krishna Kurapati <krishna.kurapati@oss.qualcomm.com> >>> --- >>> This series has only been compile tested. The issue was reproduced >>> easily with a certain kind of cable and CDP port of AMD based Lenovo >>> laptop. I don't have access to the cable currently and hence only >>> compile testing the fix for now. But the issue has popped up on OEM >>> testing as well. >>> >>> Also, didn't add locking while calling error recovery work in >>> gadget_ep_cmd since the caller is supposed to handle it. >>> >>> Changes to v3: >>> - Using error receovery mechanism from [1]. >>> >>> Link to v2: >>> https://lore.kernel.org/all/20260806-ram-interface-code-v2-1-fe4a0de31d42@oss.qualcomm.com/ >>> >>> >>> Changes in v2: >>> - Implemented gadget recovery mechanism during gadget_ep_cmd instead of >>> handling this issue only during disconnect. >>> >>> Link to RFC: >>> https://lore.kernel.org/all/20231011100214.25720-1-quic_kriskura@quicinc.com/ >>> >>> >>> [1]: >>> https://lore.kernel.org/all/20260915110637.17658-1-jiazi.liu1984@gmail.com/ >>> --- >>> drivers/usb/dwc3/gadget.c | 19 +++++++++++++++++++ >>> 1 file changed, 19 insertions(+) >>> >>> diff --git a/drivers/usb/dwc3/gadget.c b/drivers/usb/dwc3/gadget.c >>> index ee837235630a..a7c4514cf0e2 100644 >>> --- a/drivers/usb/dwc3/gadget.c >>> +++ b/drivers/usb/dwc3/gadget.c >>> @@ -283,6 +283,8 @@ int dwc3_send_gadget_generic_command(struct dwc3 >>> *dwc, unsigned int cmd, >>> return ret; >>> } >>> +static void dwc3_schedule_err_recovery(struct dwc3 *dwc); >> >> Where is the function implementation for this function?. Could you >> please point me on what you are doing in this function? >> >> >> Thanks, >> Selva >> >> >>> + >>> /** >>> * dwc3_send_gadget_ep_cmd - issue an endpoint command >>> * @dep: the endpoint to which the command is going to be issued >>> @@ -432,6 +434,23 @@ int dwc3_send_gadget_ep_cmd(struct dwc3_ep >>> *dep, unsigned int cmd, >>> cmd_status = -ETIMEDOUT; >>> } >>> + /* >>> + * STAR 5001544 - In some situations, like the cable is >>> + * disconnected or if the Host aborts the transfer on 3 >>> + * consecutive failed attempts, the Request-Done handshake is not >>> + * complete. This keeps the RAM interface busy. >>> + * >>> + * The subsequent RAM access cannot proceed until the pending >>> + * transfer is complete. This results in failure of any access >>> + * to RAM address locations. Many of the EndPoint commands need to >>> + * access the RAM and they would fail to complete successfully. >>> + * >>> + * If the depcmd doesn't match the actual command, trigger >>> controller >>> + * recovery. >>> + */ >>> + if (DWC3_DEPCMD_CMD(reg) != DWC3_DEPCMD_CMD(cmd)) >>> + dwc3_schedule_err_recovery(dwc); >>> + >>> skip_status: >>> trace_dwc3_gadget_ep_cmd(dep, cmd, params, cmd_status); >>> >>> --- >>> base-commit: ab29ca7714b82485ebb31d66835ccd5364221a76 >>> change-id: 20260929-ram-interface-stuck-v3-aaf0c4b13b98 >>> >>> Best regards, >>> -- >>> Krishna Kurapati <krishna.kurapati@oss.qualcomm.com> > ^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH v3] usb: dwc3: core: Fix RAM interface getting stuck during enumeration 2026-09-29 11:02 ` Selvarasu Ganesan @ 2026-09-29 17:50 ` Krishna Kurapati 0 siblings, 0 replies; 5+ messages in thread From: Krishna Kurapati @ 2026-09-29 17:50 UTC (permalink / raw) To: Selvarasu Ganesan, Thinh Nguyen Cc: linux-usb, linux-kernel, Greg Kroah-Hartman On 9/29/2026 4:32 PM, Selvarasu Ganesan wrote: > > On 9/29/2026 2:05 PM, Krishna Kurapati wrote: >> >> >> On 9/29/2026 12:15 PM, Selvarasu Ganesan wrote: >>> >>> On 9/29/2026 10:49 AM, Krishna Kurapati wrote: >>>> During plug-in/plug-out test cases, it is sometimes seen that no events >>>> are generated by the controller and all CSR register reads give "0" and >>>> CSR_Timeout bit gets set indicating that CSR reads/writes are timing >>>> out >>>> or timed out. >>> Hi Krishna, >>> >>> As per our discussion with synopsys we got a below feedback for CSR >>> timeout in different issue. >>> >>> "Note: CSR timeout is not expected to occur during functional mode. Once >>> timeout is reported by controller, s/w should treat it as a fatal error >>> and should identify the root cause of issuing the parallel access during >>> soft reset and fix it" >>> >>> As per our understanding, there is no recovery from CSR timeout. >>> >>> Are you able to recover from CSR timeout in your case even after trigger >>> controller recovery?. >>> >> >> Disabling and re-enabling gadget mode is the recommended fix. I did >> not test this particular patch as mentioned below, but this error >> recovery mechanism does the same thing as a soft_disconnect followed >> by soft_connect. > > Hi Krishna, > > Thanks for your update. > > We tried to recover from the CSR timeout, but it was unsuccessful. The > sequence below is provided for your reference. > > It would be best to identify the root cause of why the CSR timeout is > occurring in your case. > > Issue Sequence: > 1. Start core soft reset > 2. Parallel access of other registers > 3. End core soft reset > 4. Getting CSR timeout > > > Try to recovery from CSR timeout: > 5. Clear CSR timeout bit by RW1C same bit > 6. Start core soft reset > 7. End core softt reset > 8. Access other registers > 9. Getting CSR timeout > Hi Selva, As per official issue description we got from SNPS: --- Problem Description: If a USB disconnect happens during 0-length IN 2.0 transfer or if the host aborts the 0-length IN 2.0 transfer on a Transaction Error, the driver application fails to access the RAM data subsequently. The read to the RAM registers returns all-0. Some EndPoint commands hang without completion as they need to access the RAM locations. --- Root cause: For normal transfers, the AXI BMU (Bus Master Unit) module reads the data from the memory and issues a transmit Command Request to the MAC to transmit the data to the USB Host. • Once the data is successfully transmitted to the Host, the Done (with success) handshake is provided back to the BMU from the MAC. • If the data is not successfully transmitted, then a Done (with failure) handshake is returned, so that the BMU can re-attempt the same data again by rewinding its data pointers.– But, if the USB IN is a 0-length payload, then there is no need to rewind the pointers and the Done (with failure) handshake is not returned for failure case. This keeps the Request-Done interface busy till the next Done handshake. The MAC sends the 0-length payload again when the host requests. If the transmission is successful this time, the Done (with success) handshake is provided back. Otherwise, it repeats the same steps again.– If the cable is disconnected or if the Host aborts the transfer on 3 consecutive failed attempts, the Request-Done handshake is not complete. This keeps the interface busy. • The subsequent RAM access cannot proceed until the above pending transfer is complete. This results in failure of any access to RAM address locations. Many of the EndPoint commands need to access the RAM and they would fail to complete successfully. --- Workaround • Issue the DCTL[30] Soft Reset and re-initialize the device controller -- Error recovery work does the exact same thing and hence we are invoking that in gadget_ep_cmd when registers read all zeroes. Regards, Krishna, > > Thanks, > Selva >> >> Regards, >> Krishna, >> >>> >>>> >>>> The issue comes up on different instnaces of enumeration on different >>>> platforms. On SM8550, the debug log is as follows: >>>> >>>> Prepared a TRB on ep0out and did start transfer to get set >>>> address request from host: >>>> >>>> <...>-7191 [000] D..1. 66.421006: dwc3_gadget_ep_cmd: ep0out: >>>> cmd 'Start Transfer' [406] params 00000000 efffa000 00000000 --> >>>> status: Successful >>>> >>>> <...>-7191 [000] D..1. 66.421196: dwc3_event: event (0000c040): >>>> ep0out: Transfer Complete (sIL) [Setup Phase] >>>> >>>> <...>-7191 [000] D..1. 66.421197: dwc3_ctrl_req: Set >>>> Address(Addr = 01) >>>> >>>> An XFER NRDY is received on ep0in for zero length status phase and >>>> a Start Transfer was done on ep0in with 0-length packet in 2 Stage >>>> status phase: >>>> >>>> <...>-7191 [000] D..1. 66.421249: dwc3_event: event (000020c2): >>>> ep0in: Transfer Not Ready [00000000] (Not Active) [Status Phase] >>>> >>>> <...>-7191 [000] D..1. 66.421266: dwc3_prepare_trb: ep0in: trb >>>> ffffffc00fcfd000 (E0:D0) buf 00000000efffa000 size 0 ctrl 00000c33 >>>> sofn 00000000 (HLcs:SC:status2) >>>> >>>> <...>-7191 [000] D..1. 66.421387: dwc3_gadget_ep_cmd: ep0in: cmd >>>> 'Start Transfer' [406] params 00000000 efffa000 00000000 -->status: >>>> Successful >>>> >>>> A bus reset was then received directly after 500 msec. Software never >>>> got the cmd complete for the start transfer done in status phase. Here >>>> the RAM interface is stuck. So host issues a bus reset as link is >>>> idle for 500 msec: >>>> >>>> <...>-7191 [000] D..1. 66.935603: dwc3_event: event (00000101): >>>> Reset [U0] >>>> >>>> Then software sees that it is in status phase and we issue an ENDXFER >>>> on ep0in and it gets timedout waiting for the CMDACT to go '0': >>>> >>>> <...>-7191 [000] D..1. 66.958249: dwc3_gadget_ep_cmd: ep0in: cmd >>>> 'End Transfer' [10508] params 00000000 00000000 00000000 --> status: >>>> Timed Out >>>> >>>> Upon debug with Synopsys, the root cause is as follows: >>>> >>>> During any transfer, if the data is not successfully transmitted, >>>> then a Done (with failure) handshake is returned, so that the BMU >>>> can re-attempt the same data again by rewinding its data pointers. >>>> >>>> But, if the USB IN is a 0-length payload (which is what is happening >>>> in this case - 2 stage status phase of set_address), then there is no >>>> need to rewind the pointers and the Done (with failure) handshake is >>>> not returned for failure case. This keeps the Request-Done interface >>>> busy till the next Done handshake. The MAC sends the 0-length payload >>>> again when the host requests. If the transmission is successful this >>>> time, the Done (with success) handshake is provided back. Otherwise, >>>> it repeats the same steps again. >>>> >>>> If the cable is disconnected or if the Host aborts the transfer on 3 >>>> consecutive failed attempts, the Request-Done handshake is not >>>> complete. This keeps the interface busy. >>>> >>>> The subsequent RAM access cannot proceed until the above pending >>>> transfer is complete. This results in failure of any access to RAM >>>> address locations. Many of the EndPoint commands need to access the >>>> RAM and they would fail to complete successfully. >>>> >>>> Furthermore when cable removal happens, this would not generate a >>>> disconnect event and the "connected" flag remains true always blockin >>>> suspend. >>>> >>>> Synopsys confirmed that the issue is present on all USB3 devices and >>>> as a workaround, suggested to re-initialize device mode. >>>> >>>> Signed-off-by: Krishna Kurapati <krishna.kurapati@oss.qualcomm.com> >>>> --- >>>> This series has only been compile tested. The issue was reproduced >>>> easily with a certain kind of cable and CDP port of AMD based Lenovo >>>> laptop. I don't have access to the cable currently and hence only >>>> compile testing the fix for now. But the issue has popped up on OEM >>>> testing as well. >>>> >>>> Also, didn't add locking while calling error recovery work in >>>> gadget_ep_cmd since the caller is supposed to handle it. >>>> >>>> Changes to v3: >>>> - Using error receovery mechanism from [1]. >>>> >>>> Link to v2: >>>> https://lore.kernel.org/all/20260806-ram-interface-code-v2-1-fe4a0de31d42@oss.qualcomm.com/ >>>> >>>> >>>> Changes in v2: >>>> - Implemented gadget recovery mechanism during gadget_ep_cmd instead of >>>> handling this issue only during disconnect. >>>> >>>> Link to RFC: >>>> https://lore.kernel.org/all/20231011100214.25720-1-quic_kriskura@quicinc.com/ >>>> >>>> >>>> [1]: >>>> https://lore.kernel.org/all/20260915110637.17658-1-jiazi.liu1984@gmail.com/ >>>> --- >>>> drivers/usb/dwc3/gadget.c | 19 +++++++++++++++++++ >>>> 1 file changed, 19 insertions(+) >>>> >>>> diff --git a/drivers/usb/dwc3/gadget.c b/drivers/usb/dwc3/gadget.c >>>> index ee837235630a..a7c4514cf0e2 100644 >>>> --- a/drivers/usb/dwc3/gadget.c >>>> +++ b/drivers/usb/dwc3/gadget.c >>>> @@ -283,6 +283,8 @@ int dwc3_send_gadget_generic_command(struct dwc3 >>>> *dwc, unsigned int cmd, >>>> return ret; >>>> } >>>> +static void dwc3_schedule_err_recovery(struct dwc3 *dwc); >>> >>> Where is the function implementation for this function?. Could you >>> please point me on what you are doing in this function? >>> >>> >>> Thanks, >>> Selva >>> >>> >>>> + >>>> /** >>>> * dwc3_send_gadget_ep_cmd - issue an endpoint command >>>> * @dep: the endpoint to which the command is going to be issued >>>> @@ -432,6 +434,23 @@ int dwc3_send_gadget_ep_cmd(struct dwc3_ep >>>> *dep, unsigned int cmd, >>>> cmd_status = -ETIMEDOUT; >>>> } >>>> + /* >>>> + * STAR 5001544 - In some situations, like the cable is >>>> + * disconnected or if the Host aborts the transfer on 3 >>>> + * consecutive failed attempts, the Request-Done handshake is not >>>> + * complete. This keeps the RAM interface busy. >>>> + * >>>> + * The subsequent RAM access cannot proceed until the pending >>>> + * transfer is complete. This results in failure of any access >>>> + * to RAM address locations. Many of the EndPoint commands need to >>>> + * access the RAM and they would fail to complete successfully. >>>> + * >>>> + * If the depcmd doesn't match the actual command, trigger >>>> controller >>>> + * recovery. >>>> + */ >>>> + if (DWC3_DEPCMD_CMD(reg) != DWC3_DEPCMD_CMD(cmd)) >>>> + dwc3_schedule_err_recovery(dwc); >>>> + >>>> skip_status: >>>> trace_dwc3_gadget_ep_cmd(dep, cmd, params, cmd_status); >>>> >>>> --- >>>> base-commit: ab29ca7714b82485ebb31d66835ccd5364221a76 >>>> change-id: 20260929-ram-interface-stuck-v3-aaf0c4b13b98 >>>> >>>> Best regards, >>>> -- >>>> Krishna Kurapati <krishna.kurapati@oss.qualcomm.com> >> ^ permalink raw reply [flat|nested] 5+ messages in thread
end of thread, other threads:[~2026-09-29 17:50 UTC | newest]
Thread overview: 5+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
[not found] <CGME20260929051941epcas5p27a5c69151937b9105660941776cda819@epcas5p2.samsung.com>
2026-09-29 5:19 ` [PATCH v3] usb: dwc3: core: Fix RAM interface getting stuck during enumeration Krishna Kurapati
2026-09-29 6:45 ` Selvarasu Ganesan
2026-09-29 8:35 ` Krishna Kurapati
2026-09-29 11:02 ` Selvarasu Ganesan
2026-09-29 17:50 ` Krishna Kurapati
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®