From: Ross Lagerwall <ross.lagerwall@citrix.com>
To: "Roger Pau Monné" <roger@xenproject.org>
Cc: xen-devel@lists.xenproject.org, linux-block@vger.kernel.org,
linux-kernel@vger.kernel.org, Juergen Gross <jgross@suse.com>,
Stefano Stabellini <sstabellini@kernel.org>,
Oleksandr Tyshchenko <oleksandr_tyshchenko@epam.com>,
Jens Axboe <axboe@kernel.dk>,
stable@vger.kernel.org
Subject: Re: [PATCH] xen-blkfront: Fix IO race during unplug
Date: Tue, 6 Oct 2026 15:43:44 +0100 [thread overview]
Message-ID: <cf503bba-9f46-4cf6-9c42-710d0ffe4120@citrix.com> (raw)
In-Reply-To: <asUCYtIJkFchdbNH@Mac.lan>
On 10/6/26 3:14 PM, Roger Pau Monné wrote:
> On Tue, Oct 06, 2026 at 02:58:06PM +0100, Ross Lagerwall wrote:
>> During unplug, blkfront stops the hw queues and marks the disk as dead,
>> then later during removal calls del_gendisk(). However, IO issued after
>> the hw queues are stopped but before the call to del_gendisk() will be
>> queued but never handled. This causes del_gendisk() to hang forever
>> waiting for the queue refcount to drop to zero.
>>
>> This can be reproduced by issuing IO during an artificial delay after
>> stopping the hw queues.
>
> So the window is between the blk_mq_stop_hw_queues() and
> blk_mark_disk_dead() calls where requests would be queued and never
> processed?
Yes.
>
>> Fix this by simply not stopping the hw queues directly. Marking the disk
>> as dead also freezes the queue which prevents new requests being added
>> and it synchronously runs the hw queues to clear anything pending.
>
> I think this likely needs expanding a bit: the disk is already marked
> as dead with the current logic, and in the same place in the code.
>
>> Fixes: 8e141f9eb803 ("block: drain file system I/O on del_gendisk")
>> Cc: stable@vger.kernel.org
>> Assisted-by: LLM
>> Signed-off-by: Ross Lagerwall <ross.lagerwall@citrix.com>
>> ---
>>
>> I'm not sure about the Fixes tag. It's the most likely looking candidate
>> to me but I didn't confirm whether it actually introduced the
>> regression.
>
> I was wondering the same, and even then someone might argue this was a
> latent bug in blkfront itself, and the reference commit just exposed
> it. I don't have a strong opinion.
Good point, I might be inclined to drop it then.
>
>>
>> drivers/block/xen-blkfront.c | 4 +---
>> 1 file changed, 1 insertion(+), 3 deletions(-)
>>
>> diff --git a/drivers/block/xen-blkfront.c b/drivers/block/xen-blkfront.c
>> index 8dad7bf5f664..69a2315a1b20 100644
>> --- a/drivers/block/xen-blkfront.c
>> +++ b/drivers/block/xen-blkfront.c
>> @@ -2138,10 +2138,8 @@ static void blkfront_closing(struct blkfront_info *info)
>> return;
>>
>> /* No more blkif_request(). */
>> - if (info->rq && info->gd) {
>> - blk_mq_stop_hw_queues(info->rq);
>> + if (info->gd)
>> blk_mark_disk_dead(info->gd);
>
> So blk_mark_disk_dead() behaves differently when called with the
> queues still active?
Yes. In both cases, the queues are frozen and the hw queues synchronously run,
but the latter is a no-op when the hw queues are in the stopped state,
therefore it could leave requests unprocessed.
How about this for the final paragraph of the commit message?
"""
Fix this by simply not stopping the hw queues directly. Marking the disk
as dead already freezes the queue which prevents new requests being added
and it synchronously runs the hw queues to clear anything pending. If the hw
queues are stopped when calling blk_mark_disk_dead(), running the hw queues
is a no-op and can leave queued requests unprocessed.
"""
Ross
next prev parent reply other threads:[~2026-10-06 14:43 UTC|newest]
Thread overview: 4+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-06 13:58 Ross Lagerwall
2026-10-06 14:14 ` Roger Pau Monné
2026-10-06 14:43 ` Ross Lagerwall [this message]
2026-10-06 15:27 ` Roger Pau Monné
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=cf503bba-9f46-4cf6-9c42-710d0ffe4120@citrix.com \
--to=ross.lagerwall@citrix.com \
--cc=axboe@kernel.dk \
--cc=jgross@suse.com \
--cc=linux-block@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=oleksandr_tyshchenko@epam.com \
--cc=roger@xenproject.org \
--cc=sstabellini@kernel.org \
--cc=stable@vger.kernel.org \
--cc=xen-devel@lists.xenproject.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®