* WQ_UNBOUND workqueue warnings from multiple drivers @ 2024-03-18 22:33 Kamaljit Singh 2024-03-20 9:11 ` Sagi Grimberg 0 siblings, 1 reply; 9+ messages in thread From: Kamaljit Singh @ 2024-03-18 22:33 UTC (permalink / raw) To: linux-nvme, linux-kernel, kbusch, Sagi Grimberg Hello, After switching from Kernel v6.6.2 to v6.6.21 we're now seeing these workqueue warnings. I found a discussion thread about the the Intel drm driver here https://lore.kernel.org/lkml/ZO-BkaGuVCgdr3wc@slm.duckdns.org/T/ and this related bug report https://gitlab.freedesktop.org/drm/intel/-/issues/9245 but that that drm fix isn't merged into v6.6.21. It appears that we may need the same WQ_UNBOUND change to the nvme host tcp driver among others. [Fri Mar 15 22:30:06 2024] workqueue: nvme_tcp_io_work [nvme_tcp] hogged CPU for >10000us 4 times, consider switching to WQ_UNBOUND [Fri Mar 15 23:44:58 2024] workqueue: drain_vmap_area_work hogged CPU for >10000us 4 times, consider switching to WQ_UNBOUND [Sat Mar 16 09:55:27 2024] workqueue: drain_vmap_area_work hogged CPU for >10000us 8 times, consider switching to WQ_UNBOUND [Sat Mar 16 17:51:18 2024] workqueue: nvme_tcp_io_work [nvme_tcp] hogged CPU for >10000us 8 times, consider switching to WQ_UNBOUND [Sat Mar 16 23:04:14 2024] workqueue: nvme_tcp_io_work [nvme_tcp] hogged CPU for >10000us 16 times, consider switching to WQ_UNBOUND [Sun Mar 17 21:35:46 2024] perf: interrupt took too long (2707 > 2500), lowering kernel.perf_event_max_sample_rate to 73750 [Sun Mar 17 21:49:34 2024] workqueue: drain_vmap_area_work hogged CPU for >10000us 16 times, consider switching to WQ_UNBOUND ... workqueue: drm_fb_helper_damage_work [drm_kms_helper] hogged CPU for >10000us 32 times, consider switching to WQ_UNBOUND Thanks, Kamaljit Singh ^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: WQ_UNBOUND workqueue warnings from multiple drivers 2024-03-18 22:33 WQ_UNBOUND workqueue warnings from multiple drivers Kamaljit Singh @ 2024-03-20 9:11 ` Sagi Grimberg 2024-03-21 17:36 ` Chaitanya Kulkarni 0 siblings, 1 reply; 9+ messages in thread From: Sagi Grimberg @ 2024-03-20 9:11 UTC (permalink / raw) To: Kamaljit Singh, linux-nvme, linux-kernel, kbusch On 19/03/2024 0:33, Kamaljit Singh wrote: > Hello, > > After switching from Kernel v6.6.2 to v6.6.21 we're now seeing these workqueue > warnings. I found a discussion thread about the the Intel drm driver here > https://lore.kernel.org/lkml/ZO-BkaGuVCgdr3wc@slm.duckdns.org/T/ > > and this related bug report https://gitlab.freedesktop.org/drm/intel/-/issues/9245 > but that that drm fix isn't merged into v6.6.21. It appears that we may need the same > WQ_UNBOUND change to the nvme host tcp driver among others. > > [Fri Mar 15 22:30:06 2024] workqueue: nvme_tcp_io_work [nvme_tcp] hogged CPU for >10000us 4 times, consider switching to WQ_UNBOUND > [Fri Mar 15 23:44:58 2024] workqueue: drain_vmap_area_work hogged CPU for >10000us 4 times, consider switching to WQ_UNBOUND > [Sat Mar 16 09:55:27 2024] workqueue: drain_vmap_area_work hogged CPU for >10000us 8 times, consider switching to WQ_UNBOUND > [Sat Mar 16 17:51:18 2024] workqueue: nvme_tcp_io_work [nvme_tcp] hogged CPU for >10000us 8 times, consider switching to WQ_UNBOUND > [Sat Mar 16 23:04:14 2024] workqueue: nvme_tcp_io_work [nvme_tcp] hogged CPU for >10000us 16 times, consider switching to WQ_UNBOUND > [Sun Mar 17 21:35:46 2024] perf: interrupt took too long (2707 > 2500), lowering kernel.perf_event_max_sample_rate to 73750 > [Sun Mar 17 21:49:34 2024] workqueue: drain_vmap_area_work hogged CPU for >10000us 16 times, consider switching to WQ_UNBOUND > ... > workqueue: drm_fb_helper_damage_work [drm_kms_helper] hogged CPU for >10000us 32 times, consider switching to WQ_UNBOUND Hey Kamaljit, Its interesting that this happens because nvme_tcp_io_work is bound to 1 jiffie. Although in theory we do not stop receiving from a socket once we started, so I guess this can happen in some extreme cases. Was the test you were running read-heavy? I was thinking that we may want to optionally move the recv path to softirq instead to get some latency improvements, although I don't know if that would improve the situation if we end up spending a lot of time in soft-irq... > > > Thanks, > Kamaljit Singh ^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: WQ_UNBOUND workqueue warnings from multiple drivers 2024-03-20 9:11 ` Sagi Grimberg @ 2024-03-21 17:36 ` Chaitanya Kulkarni 2024-04-02 23:50 ` Kamaljit Singh 0 siblings, 1 reply; 9+ messages in thread From: Chaitanya Kulkarni @ 2024-03-21 17:36 UTC (permalink / raw) To: Sagi Grimberg, Kamaljit Singh; +Cc: kbusch, linux-kernel, linux-nvme On 3/20/24 02:11, Sagi Grimberg wrote: > > > On 19/03/2024 0:33, Kamaljit Singh wrote: >> Hello, >> >> After switching from Kernel v6.6.2 to v6.6.21 we're now seeing these >> workqueue >> warnings. I found a discussion thread about the the Intel drm driver >> here >> https://lore.kernel.org/lkml/ZO-BkaGuVCgdr3wc@slm.duckdns.org/T/ >> >> and this related bug report >> https://gitlab.freedesktop.org/drm/intel/-/issues/9245 >> but that that drm fix isn't merged into v6.6.21. It appears that we >> may need the same >> WQ_UNBOUND change to the nvme host tcp driver among others. >> [Fri Mar 15 22:30:06 2024] workqueue: nvme_tcp_io_work [nvme_tcp] >> hogged CPU for >10000us 4 times, consider switching to WQ_UNBOUND >> [Fri Mar 15 23:44:58 2024] workqueue: drain_vmap_area_work hogged CPU >> for >10000us 4 times, consider switching to WQ_UNBOUND >> [Sat Mar 16 09:55:27 2024] workqueue: drain_vmap_area_work hogged CPU >> for >10000us 8 times, consider switching to WQ_UNBOUND >> [Sat Mar 16 17:51:18 2024] workqueue: nvme_tcp_io_work [nvme_tcp] >> hogged CPU for >10000us 8 times, consider switching to WQ_UNBOUND >> [Sat Mar 16 23:04:14 2024] workqueue: nvme_tcp_io_work [nvme_tcp] >> hogged CPU for >10000us 16 times, consider switching to WQ_UNBOUND >> [Sun Mar 17 21:35:46 2024] perf: interrupt took too long (2707 > >> 2500), lowering kernel.perf_event_max_sample_rate to 73750 >> [Sun Mar 17 21:49:34 2024] workqueue: drain_vmap_area_work hogged CPU >> for >10000us 16 times, consider switching to WQ_UNBOUND >> ... >> workqueue: drm_fb_helper_damage_work [drm_kms_helper] hogged CPU for >> >10000us 32 times, consider switching to WQ_UNBOUND > > Hey Kamaljit, > > Its interesting that this happens because nvme_tcp_io_work is bound to > 1 jiffie. > Although in theory we do not stop receiving from a socket once we > started, so > I guess this can happen in some extreme cases. Was the test you were > running > read-heavy? > > I was thinking that we may want to optionally move the recv path to > softirq instead to > get some latency improvements, although I don't know if that would > improve the situation > if we end up spending a lot of time in soft-irq... > >> Thanks, >> Kamaljit Singh > > we need a regular test for this in blktests as it doesn't look like we caught this in regular testing ... Kamaljit, can you please provide details of the tests you are running so we can reproduce ? -ck ^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: WQ_UNBOUND workqueue warnings from multiple drivers 2024-03-21 17:36 ` Chaitanya Kulkarni @ 2024-04-02 23:50 ` Kamaljit Singh 2024-04-07 20:08 ` Sagi Grimberg 0 siblings, 1 reply; 9+ messages in thread From: Kamaljit Singh @ 2024-04-02 23:50 UTC (permalink / raw) To: Chaitanya Kulkarni, Sagi Grimberg; +Cc: kbusch, linux-kernel, linux-nvme Sagi, Chaitanya, Sorry for the delay, found your replies in the junk folder :( > Was the test you were running read-heavy? No, most of the failing fio tests were doing heavy writes. All were with 8 Controllers and 32 NS each. io-specs are below. [1] bs=16k, iodepth=16, rwmixread=0, numjobs=16 Failed in ~1 min Some others were: [2] bs=8k, iodepth=16, rwmixread=5, numjobs=16 [3] bs=8k, iodepth=16, rwmixread=50, numjobs=16 Thanks, Kamaljit From: Chaitanya Kulkarni <chaitanyak@nvidia.com> Date: Thursday, March 21, 2024 at 10:36 To: Sagi Grimberg <sagi@grimberg.me>, Kamaljit Singh <Kamaljit.Singh1@wdc.com> Cc: kbusch@kernel.org <kbusch@kernel.org>, linux-kernel@vger.kernel.org <linux-kernel@vger.kernel.org>, linux-nvme@lists.infradead.org <linux-nvme@lists.infradead.org> Subject: Re: WQ_UNBOUND workqueue warnings from multiple drivers CAUTION: This email originated from outside of Western Digital. Do not click on links or open attachments unless you recognize the sender and know that the content is safe. On 3/20/24 02:11, Sagi Grimberg wrote: > > > On 19/03/2024 0:33, Kamaljit Singh wrote: >> Hello, >> >> After switching from Kernel v6.6.2 to v6.6.21 we're now seeing these >> workqueue >> warnings. I found a discussion thread about the the Intel drm driver >> here >> https://lore.kernel.org/lkml/ZO-BkaGuVCgdr3wc@slm.duckdns.org/T/ >> >> and this related bug report >> https://gitlab.freedesktop.org/drm/intel/-/issues/9245 >> but that that drm fix isn't merged into v6.6.21. It appears that we >> may need the same >> WQ_UNBOUND change to the nvme host tcp driver among others. >> [Fri Mar 15 22:30:06 2024] workqueue: nvme_tcp_io_work [nvme_tcp] >> hogged CPU for >10000us 4 times, consider switching to WQ_UNBOUND >> [Fri Mar 15 23:44:58 2024] workqueue: drain_vmap_area_work hogged CPU >> for >10000us 4 times, consider switching to WQ_UNBOUND >> [Sat Mar 16 09:55:27 2024] workqueue: drain_vmap_area_work hogged CPU >> for >10000us 8 times, consider switching to WQ_UNBOUND >> [Sat Mar 16 17:51:18 2024] workqueue: nvme_tcp_io_work [nvme_tcp] >> hogged CPU for >10000us 8 times, consider switching to WQ_UNBOUND >> [Sat Mar 16 23:04:14 2024] workqueue: nvme_tcp_io_work [nvme_tcp] >> hogged CPU for >10000us 16 times, consider switching to WQ_UNBOUND >> [Sun Mar 17 21:35:46 2024] perf: interrupt took too long (2707 > >> 2500), lowering kernel.perf_event_max_sample_rate to 73750 >> [Sun Mar 17 21:49:34 2024] workqueue: drain_vmap_area_work hogged CPU >> for >10000us 16 times, consider switching to WQ_UNBOUND >> ... >> workqueue: drm_fb_helper_damage_work [drm_kms_helper] hogged CPU for >> >10000us 32 times, consider switching to WQ_UNBOUND > > Hey Kamaljit, > > Its interesting that this happens because nvme_tcp_io_work is bound to > 1 jiffie. > Although in theory we do not stop receiving from a socket once we > started, so > I guess this can happen in some extreme cases. Was the test you were > running > read-heavy? > > I was thinking that we may want to optionally move the recv path to > softirq instead to > get some latency improvements, although I don't know if that would > improve the situation > if we end up spending a lot of time in soft-irq... > >> Thanks, >> Kamaljit Singh > > we need a regular test for this in blktests as it doesn't look like we caught this in regular testing ... Kamaljit, can you please provide details of the tests you are running so we can reproduce ? -ck ^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: WQ_UNBOUND workqueue warnings from multiple drivers 2024-04-02 23:50 ` Kamaljit Singh @ 2024-04-07 20:08 ` Sagi Grimberg 2024-05-08 23:16 ` Kamaljit Singh 0 siblings, 1 reply; 9+ messages in thread From: Sagi Grimberg @ 2024-04-07 20:08 UTC (permalink / raw) To: Kamaljit Singh, Chaitanya Kulkarni; +Cc: kbusch, linux-kernel, linux-nvme On 03/04/2024 2:50, Kamaljit Singh wrote: > Sagi, Chaitanya, > > Sorry for the delay, found your replies in the junk folder :( > >> Was the test you were running read-heavy? > No, most of the failing fio tests were doing heavy writes. All were with 8 Controllers and 32 NS each. io-specs are below. > > [1] bs=16k, iodepth=16, rwmixread=0, numjobs=16 > Failed in ~1 min > > Some others were: > [2] bs=8k, iodepth=16, rwmixread=5, numjobs=16 > [3] bs=8k, iodepth=16, rwmixread=50, numjobs=16 Interesting, that is the opposite of what I would suspect (I thought that the workload would be read-only or read-mostly). Does this happen with a 90-%100% read workload? If we look at nvme_tcp_io_work() it is essentially looping doing send() and recv() and every iteration checks if a 1ms deadline elapsed. The fact that it happens on a 100% write workload leads me to conclude that the only way this can happen if sending a single 16K request to a controller on its own takes more than 10ms, which is unexpected... Question, are you working with a Linux controller? what is the ctrl ioccsz? ^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: WQ_UNBOUND workqueue warnings from multiple drivers 2024-04-07 20:08 ` Sagi Grimberg @ 2024-05-08 23:16 ` Kamaljit Singh 2024-05-09 6:36 ` Sagi Grimberg 0 siblings, 1 reply; 9+ messages in thread From: Kamaljit Singh @ 2024-05-08 23:16 UTC (permalink / raw) To: Sagi Grimberg, Chaitanya Kulkarni; +Cc: kbusch, linux-kernel, linux-nvme Sagi, >Does this happen with a 90-%100% read workload? Yes, we’ve now seen it with 100% reads as well. Here’s the Medusa cmd we used. I’ve removed the devices for brevity. sudo /opt/medusa_labs/test_tools/bin/maim 20g -b8K -Q128 -Y1 -M30 --full-device -B3 -r -d900000 <device_list> We saw the original issue with the upstream kernel v6.6.21. But now we’re also seeing it with Ubuntu 24.04 (kernel 6.8.0-31-generic), where IOs are timing out and forcing connection drops. >Question, are you working with a Linux controller? No, with our ASIC (NVMe Fabrics bridge). >what is the ctrl ioccsz? ioccsz : 4 Thanks, Kamaljit From: Sagi Grimberg <sagi@grimberg.me> Date: Sunday, April 7, 2024 at 13:08 To: Kamaljit Singh <Kamaljit.Singh1@wdc.com>, Chaitanya Kulkarni <chaitanyak@nvidia.com> Cc: kbusch@kernel.org <kbusch@kernel.org>, linux-kernel@vger.kernel.org <linux-kernel@vger.kernel.org>, linux-nvme@lists.infradead.org <linux-nvme@lists.infradead.org> Subject: Re: WQ_UNBOUND workqueue warnings from multiple drivers CAUTION: This email originated from outside of Western Digital. Do not click on links or open attachments unless you recognize the sender and know that the content is safe. On 03/04/2024 2:50, Kamaljit Singh wrote: > Sagi, Chaitanya, > > Sorry for the delay, found your replies in the junk folder :( > >> Was the test you were running read-heavy? > No, most of the failing fio tests were doing heavy writes. All were with 8 Controllers and 32 NS each. io-specs are below. > > [1] bs=16k, iodepth=16, rwmixread=0, numjobs=16 > Failed in ~1 min > > Some others were: > [2] bs=8k, iodepth=16, rwmixread=5, numjobs=16 > [3] bs=8k, iodepth=16, rwmixread=50, numjobs=16 Interesting, that is the opposite of what I would suspect (I thought that the workload would be read-only or read-mostly). Does this happen with a 90-%100% read workload? If we look at nvme_tcp_io_work() it is essentially looping doing send() and recv() and every iteration checks if a 1ms deadline elapsed. The fact that it happens on a 100% write workload leads me to conclude that the only way this can happen if sending a single 16K request to a controller on its own takes more than 10ms, which is unexpected... Question, are you working with a Linux controller? what is the ctrl ioccsz? ^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: WQ_UNBOUND workqueue warnings from multiple drivers 2024-05-08 23:16 ` Kamaljit Singh @ 2024-05-09 6:36 ` Sagi Grimberg [not found] ` <BYAPR04MB41514F3EA640750D6C68AF56BCF12@BYAPR04MB4151.namprd04.prod.outlook.com> 0 siblings, 1 reply; 9+ messages in thread From: Sagi Grimberg @ 2024-05-09 6:36 UTC (permalink / raw) To: Kamaljit Singh, Chaitanya Kulkarni; +Cc: kbusch, linux-kernel, linux-nvme On 09/05/2024 2:16, Kamaljit Singh wrote: > Sagi, > >> Does this happen with a 90-%100% read workload? > Yes, we’ve now seen it with 100% reads as well. Here’s the Medusa cmd we used. I’ve removed the devices for brevity. > sudo /opt/medusa_labs/test_tools/bin/maim 20g -b8K -Q128 -Y1 -M30 --full-device -B3 -r -d900000 <device_list> > > We saw the original issue with the upstream kernel v6.6.21. But now we’re also seeing it with Ubuntu 24.04 (kernel 6.8.0-31-generic), where IOs are timing out and forcing connection drops. Thanks for the info. > > >> Question, are you working with a Linux controller? > No, with our ASIC (NVMe Fabrics bridge). > >> what is the ctrl ioccsz? > ioccsz : 4 I see, what is the mdts? and what are the r2t lengths the controller is sending to the host? ^ permalink raw reply [flat|nested] 9+ messages in thread
[parent not found: <BYAPR04MB41514F3EA640750D6C68AF56BCF12@BYAPR04MB4151.namprd04.prod.outlook.com>]
* Re: WQ_UNBOUND workqueue warnings from multiple drivers [not found] ` <BYAPR04MB41514F3EA640750D6C68AF56BCF12@BYAPR04MB4151.namprd04.prod.outlook.com> @ 2024-05-31 4:27 ` Sagi Grimberg 2024-06-04 0:06 ` Kamaljit Singh 0 siblings, 1 reply; 9+ messages in thread From: Sagi Grimberg @ 2024-05-31 4:27 UTC (permalink / raw) To: Kamaljit Singh, Chaitanya Kulkarni; +Cc: kbusch, linux-kernel, linux-nvme >> I see, what is the mdts? > MDTS=3, i.e. 32K > >> and what are the r2t lengths the controller is sending to the host? > r2t lengths are 24 bytes or 28 bytes. > Umm, so your controller requests 24 or 28 bytes from the host at a time?? That's unlikely. by "r2t lengths" I meant r2t_length field in the r2t, which tells the host how much to send in a h2cdata... ^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: WQ_UNBOUND workqueue warnings from multiple drivers 2024-05-31 4:27 ` Sagi Grimberg @ 2024-06-04 0:06 ` Kamaljit Singh 0 siblings, 0 replies; 9+ messages in thread From: Kamaljit Singh @ 2024-06-04 0:06 UTC (permalink / raw) To: Sagi Grimberg, Chaitanya Kulkarni; +Cc: kbusch, linux-kernel, linux-nvme >>> I see, what is the mdts? >> MDTS=3, i.e. 32K >> >>> and what are the r2t lengths the controller is sending to the host? >> r2t lengths are 24 bytes or 28 bytes. >> >Umm, so your controller requests 24 or 28 bytes from the host at a time?? >That's unlikely. by "r2t lengths" I meant r2t_length field in the r2t, >which tells >the host how much to send in a h2cdata... You couldn't read my mind ;) ? JK The controller used requested data length = 0x2000 (8192) bytes Sorry for any confusion. ^ permalink raw reply [flat|nested] 9+ messages in thread
end of thread, other threads:[~2024-06-04 0:06 UTC | newest]
Thread overview: 9+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2024-03-18 22:33 WQ_UNBOUND workqueue warnings from multiple drivers Kamaljit Singh
2024-03-20 9:11 ` Sagi Grimberg
2024-03-21 17:36 ` Chaitanya Kulkarni
2024-04-02 23:50 ` Kamaljit Singh
2024-04-07 20:08 ` Sagi Grimberg
2024-05-08 23:16 ` Kamaljit Singh
2024-05-09 6:36 ` Sagi Grimberg
[not found] ` <BYAPR04MB41514F3EA640750D6C68AF56BCF12@BYAPR04MB4151.namprd04.prod.outlook.com>
2024-05-31 4:27 ` Sagi Grimberg
2024-06-04 0:06 ` Kamaljit Singh
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®