* [PATCH] xen-blkfront: Fix IO race during unplug
@ 2026-10-06 13:58 Ross Lagerwall
2026-10-06 14:14 ` Roger Pau Monné
0 siblings, 1 reply; 4+ messages in thread
From: Ross Lagerwall @ 2026-10-06 13:58 UTC (permalink / raw)
To: xen-devel, linux-block, linux-kernel
Cc: Ross Lagerwall, Juergen Gross, Stefano Stabellini,
Oleksandr Tyshchenko, Roger Pau Monné,
Jens Axboe, stable
During unplug, blkfront stops the hw queues and marks the disk as dead,
then later during removal calls del_gendisk(). However, IO issued after
the hw queues are stopped but before the call to del_gendisk() will be
queued but never handled. This causes del_gendisk() to hang forever
waiting for the queue refcount to drop to zero.
This can be reproduced by issuing IO during an artificial delay after
stopping the hw queues.
Fix this by simply not stopping the hw queues directly. Marking the disk
as dead also freezes the queue which prevents new requests being added
and it synchronously runs the hw queues to clear anything pending.
Fixes: 8e141f9eb803 ("block: drain file system I/O on del_gendisk")
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Ross Lagerwall <ross.lagerwall@citrix.com>
---
I'm not sure about the Fixes tag. It's the most likely looking candidate
to me but I didn't confirm whether it actually introduced the
regression.
drivers/block/xen-blkfront.c | 4 +---
1 file changed, 1 insertion(+), 3 deletions(-)
diff --git a/drivers/block/xen-blkfront.c b/drivers/block/xen-blkfront.c
index 8dad7bf5f664..69a2315a1b20 100644
--- a/drivers/block/xen-blkfront.c
+++ b/drivers/block/xen-blkfront.c
@@ -2138,10 +2138,8 @@ static void blkfront_closing(struct blkfront_info *info)
return;
/* No more blkif_request(). */
- if (info->rq && info->gd) {
- blk_mq_stop_hw_queues(info->rq);
+ if (info->gd)
blk_mark_disk_dead(info->gd);
- }
for_each_rinfo(info, rinfo, i) {
/* No more gnttab callback work. */
--
2.55.0
^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [PATCH] xen-blkfront: Fix IO race during unplug
2026-10-06 13:58 [PATCH] xen-blkfront: Fix IO race during unplug Ross Lagerwall
@ 2026-10-06 14:14 ` Roger Pau Monné
2026-10-06 14:43 ` Ross Lagerwall
0 siblings, 1 reply; 4+ messages in thread
From: Roger Pau Monné @ 2026-10-06 14:14 UTC (permalink / raw)
To: Ross Lagerwall
Cc: xen-devel, linux-block, linux-kernel, Juergen Gross,
Stefano Stabellini, Oleksandr Tyshchenko, Jens Axboe, stable
On Tue, Oct 06, 2026 at 02:58:06PM +0100, Ross Lagerwall wrote:
> During unplug, blkfront stops the hw queues and marks the disk as dead,
> then later during removal calls del_gendisk(). However, IO issued after
> the hw queues are stopped but before the call to del_gendisk() will be
> queued but never handled. This causes del_gendisk() to hang forever
> waiting for the queue refcount to drop to zero.
>
> This can be reproduced by issuing IO during an artificial delay after
> stopping the hw queues.
So the window is between the blk_mq_stop_hw_queues() and
blk_mark_disk_dead() calls where requests would be queued and never
processed?
> Fix this by simply not stopping the hw queues directly. Marking the disk
> as dead also freezes the queue which prevents new requests being added
> and it synchronously runs the hw queues to clear anything pending.
I think this likely needs expanding a bit: the disk is already marked
as dead with the current logic, and in the same place in the code.
> Fixes: 8e141f9eb803 ("block: drain file system I/O on del_gendisk")
> Cc: stable@vger.kernel.org
> Assisted-by: LLM
> Signed-off-by: Ross Lagerwall <ross.lagerwall@citrix.com>
> ---
>
> I'm not sure about the Fixes tag. It's the most likely looking candidate
> to me but I didn't confirm whether it actually introduced the
> regression.
I was wondering the same, and even then someone might argue this was a
latent bug in blkfront itself, and the reference commit just exposed
it. I don't have a strong opinion.
>
> drivers/block/xen-blkfront.c | 4 +---
> 1 file changed, 1 insertion(+), 3 deletions(-)
>
> diff --git a/drivers/block/xen-blkfront.c b/drivers/block/xen-blkfront.c
> index 8dad7bf5f664..69a2315a1b20 100644
> --- a/drivers/block/xen-blkfront.c
> +++ b/drivers/block/xen-blkfront.c
> @@ -2138,10 +2138,8 @@ static void blkfront_closing(struct blkfront_info *info)
> return;
>
> /* No more blkif_request(). */
> - if (info->rq && info->gd) {
> - blk_mq_stop_hw_queues(info->rq);
> + if (info->gd)
> blk_mark_disk_dead(info->gd);
So blk_mark_disk_dead() behaves differently when called with the
queues still active?
Thanks, Roger.
^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [PATCH] xen-blkfront: Fix IO race during unplug
2026-10-06 14:14 ` Roger Pau Monné
@ 2026-10-06 14:43 ` Ross Lagerwall
2026-10-06 15:27 ` Roger Pau Monné
0 siblings, 1 reply; 4+ messages in thread
From: Ross Lagerwall @ 2026-10-06 14:43 UTC (permalink / raw)
To: Roger Pau Monné
Cc: xen-devel, linux-block, linux-kernel, Juergen Gross,
Stefano Stabellini, Oleksandr Tyshchenko, Jens Axboe, stable
On 10/6/26 3:14 PM, Roger Pau Monné wrote:
> On Tue, Oct 06, 2026 at 02:58:06PM +0100, Ross Lagerwall wrote:
>> During unplug, blkfront stops the hw queues and marks the disk as dead,
>> then later during removal calls del_gendisk(). However, IO issued after
>> the hw queues are stopped but before the call to del_gendisk() will be
>> queued but never handled. This causes del_gendisk() to hang forever
>> waiting for the queue refcount to drop to zero.
>>
>> This can be reproduced by issuing IO during an artificial delay after
>> stopping the hw queues.
>
> So the window is between the blk_mq_stop_hw_queues() and
> blk_mark_disk_dead() calls where requests would be queued and never
> processed?
Yes.
>
>> Fix this by simply not stopping the hw queues directly. Marking the disk
>> as dead also freezes the queue which prevents new requests being added
>> and it synchronously runs the hw queues to clear anything pending.
>
> I think this likely needs expanding a bit: the disk is already marked
> as dead with the current logic, and in the same place in the code.
>
>> Fixes: 8e141f9eb803 ("block: drain file system I/O on del_gendisk")
>> Cc: stable@vger.kernel.org
>> Assisted-by: LLM
>> Signed-off-by: Ross Lagerwall <ross.lagerwall@citrix.com>
>> ---
>>
>> I'm not sure about the Fixes tag. It's the most likely looking candidate
>> to me but I didn't confirm whether it actually introduced the
>> regression.
>
> I was wondering the same, and even then someone might argue this was a
> latent bug in blkfront itself, and the reference commit just exposed
> it. I don't have a strong opinion.
Good point, I might be inclined to drop it then.
>
>>
>> drivers/block/xen-blkfront.c | 4 +---
>> 1 file changed, 1 insertion(+), 3 deletions(-)
>>
>> diff --git a/drivers/block/xen-blkfront.c b/drivers/block/xen-blkfront.c
>> index 8dad7bf5f664..69a2315a1b20 100644
>> --- a/drivers/block/xen-blkfront.c
>> +++ b/drivers/block/xen-blkfront.c
>> @@ -2138,10 +2138,8 @@ static void blkfront_closing(struct blkfront_info *info)
>> return;
>>
>> /* No more blkif_request(). */
>> - if (info->rq && info->gd) {
>> - blk_mq_stop_hw_queues(info->rq);
>> + if (info->gd)
>> blk_mark_disk_dead(info->gd);
>
> So blk_mark_disk_dead() behaves differently when called with the
> queues still active?
Yes. In both cases, the queues are frozen and the hw queues synchronously run,
but the latter is a no-op when the hw queues are in the stopped state,
therefore it could leave requests unprocessed.
How about this for the final paragraph of the commit message?
"""
Fix this by simply not stopping the hw queues directly. Marking the disk
as dead already freezes the queue which prevents new requests being added
and it synchronously runs the hw queues to clear anything pending. If the hw
queues are stopped when calling blk_mark_disk_dead(), running the hw queues
is a no-op and can leave queued requests unprocessed.
"""
Ross
^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [PATCH] xen-blkfront: Fix IO race during unplug
2026-10-06 14:43 ` Ross Lagerwall
@ 2026-10-06 15:27 ` Roger Pau Monné
0 siblings, 0 replies; 4+ messages in thread
From: Roger Pau Monné @ 2026-10-06 15:27 UTC (permalink / raw)
To: Ross Lagerwall
Cc: xen-devel, linux-block, linux-kernel, Juergen Gross,
Stefano Stabellini, Oleksandr Tyshchenko, Jens Axboe, stable
On Tue, Oct 06, 2026 at 03:43:44PM +0100, Ross Lagerwall wrote:
> On 10/6/26 3:14 PM, Roger Pau Monné wrote:
> > On Tue, Oct 06, 2026 at 02:58:06PM +0100, Ross Lagerwall wrote:
> > > During unplug, blkfront stops the hw queues and marks the disk as dead,
> > > then later during removal calls del_gendisk(). However, IO issued after
> > > the hw queues are stopped but before the call to del_gendisk() will be
> > > queued but never handled. This causes del_gendisk() to hang forever
> > > waiting for the queue refcount to drop to zero.
> > >
> > > This can be reproduced by issuing IO during an artificial delay after
> > > stopping the hw queues.
> >
> > So the window is between the blk_mq_stop_hw_queues() and
> > blk_mark_disk_dead() calls where requests would be queued and never
> > processed?
>
> Yes.
>
> >
> > > Fix this by simply not stopping the hw queues directly. Marking the disk
> > > as dead also freezes the queue which prevents new requests being added
> > > and it synchronously runs the hw queues to clear anything pending.
> >
> > I think this likely needs expanding a bit: the disk is already marked
> > as dead with the current logic, and in the same place in the code.
> >
> > > Fixes: 8e141f9eb803 ("block: drain file system I/O on del_gendisk")
> > > Cc: stable@vger.kernel.org
> > > Assisted-by: LLM
> > > Signed-off-by: Ross Lagerwall <ross.lagerwall@citrix.com>
> > > ---
> > >
> > > I'm not sure about the Fixes tag. It's the most likely looking candidate
> > > to me but I didn't confirm whether it actually introduced the
> > > regression.
> >
> > I was wondering the same, and even then someone might argue this was a
> > latent bug in blkfront itself, and the reference commit just exposed
> > it. I don't have a strong opinion.
>
> Good point, I might be inclined to drop it then.
>
> >
> > >
> > > drivers/block/xen-blkfront.c | 4 +---
> > > 1 file changed, 1 insertion(+), 3 deletions(-)
> > >
> > > diff --git a/drivers/block/xen-blkfront.c b/drivers/block/xen-blkfront.c
> > > index 8dad7bf5f664..69a2315a1b20 100644
> > > --- a/drivers/block/xen-blkfront.c
> > > +++ b/drivers/block/xen-blkfront.c
> > > @@ -2138,10 +2138,8 @@ static void blkfront_closing(struct blkfront_info *info)
> > > return;
> > > /* No more blkif_request(). */
> > > - if (info->rq && info->gd) {
> > > - blk_mq_stop_hw_queues(info->rq);
> > > + if (info->gd)
> > > blk_mark_disk_dead(info->gd);
> >
> > So blk_mark_disk_dead() behaves differently when called with the
> > queues still active?
>
> Yes. In both cases, the queues are frozen and the hw queues synchronously run,
> but the latter is a no-op when the hw queues are in the stopped state,
> therefore it could leave requests unprocessed.
>
> How about this for the final paragraph of the commit message?
>
> """
> Fix this by simply not stopping the hw queues directly. Marking the disk
> as dead already freezes the queue which prevents new requests being added
> and it synchronously runs the hw queues to clear anything pending. If the hw
> queues are stopped when calling blk_mark_disk_dead(), running the hw queues
> is a no-op and can leave queued requests unprocessed.
> """
LGTM.
I think it might be best if you send v2 with that adjustment and the
Fixes tag possibly dropped.
Thanks, Roger.
^ permalink raw reply [flat|nested] 4+ messages in thread
end of thread, other threads:[~2026-10-06 15:27 UTC | newest]
Thread overview: 4+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-06 13:58 [PATCH] xen-blkfront: Fix IO race during unplug Ross Lagerwall
2026-10-06 14:14 ` Roger Pau Monné
2026-10-06 14:43 ` Ross Lagerwall
2026-10-06 15:27 ` Roger Pau Monné
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®