From: "Roger Pau Monné" <roger@xenproject.org>
To: Ross Lagerwall <ross.lagerwall@citrix.com>
Cc: xen-devel@lists.xenproject.org, linux-block@vger.kernel.org,
linux-kernel@vger.kernel.org, Juergen Gross <jgross@suse.com>,
Stefano Stabellini <sstabellini@kernel.org>,
Oleksandr Tyshchenko <oleksandr_tyshchenko@epam.com>,
Jens Axboe <axboe@kernel.dk>,
stable@vger.kernel.org
Subject: Re: [PATCH] xen-blkfront: Fix IO race during unplug
Date: Tue, 6 Oct 2026 17:27:39 +0200 [thread overview]
Message-ID: <asUTa3lIC0gkNAJB@Mac.lan> (raw)
In-Reply-To: <cf503bba-9f46-4cf6-9c42-710d0ffe4120@citrix.com>
On Tue, Oct 06, 2026 at 03:43:44PM +0100, Ross Lagerwall wrote:
> On 10/6/26 3:14 PM, Roger Pau Monné wrote:
> > On Tue, Oct 06, 2026 at 02:58:06PM +0100, Ross Lagerwall wrote:
> > > During unplug, blkfront stops the hw queues and marks the disk as dead,
> > > then later during removal calls del_gendisk(). However, IO issued after
> > > the hw queues are stopped but before the call to del_gendisk() will be
> > > queued but never handled. This causes del_gendisk() to hang forever
> > > waiting for the queue refcount to drop to zero.
> > >
> > > This can be reproduced by issuing IO during an artificial delay after
> > > stopping the hw queues.
> >
> > So the window is between the blk_mq_stop_hw_queues() and
> > blk_mark_disk_dead() calls where requests would be queued and never
> > processed?
>
> Yes.
>
> >
> > > Fix this by simply not stopping the hw queues directly. Marking the disk
> > > as dead also freezes the queue which prevents new requests being added
> > > and it synchronously runs the hw queues to clear anything pending.
> >
> > I think this likely needs expanding a bit: the disk is already marked
> > as dead with the current logic, and in the same place in the code.
> >
> > > Fixes: 8e141f9eb803 ("block: drain file system I/O on del_gendisk")
> > > Cc: stable@vger.kernel.org
> > > Assisted-by: LLM
> > > Signed-off-by: Ross Lagerwall <ross.lagerwall@citrix.com>
> > > ---
> > >
> > > I'm not sure about the Fixes tag. It's the most likely looking candidate
> > > to me but I didn't confirm whether it actually introduced the
> > > regression.
> >
> > I was wondering the same, and even then someone might argue this was a
> > latent bug in blkfront itself, and the reference commit just exposed
> > it. I don't have a strong opinion.
>
> Good point, I might be inclined to drop it then.
>
> >
> > >
> > > drivers/block/xen-blkfront.c | 4 +---
> > > 1 file changed, 1 insertion(+), 3 deletions(-)
> > >
> > > diff --git a/drivers/block/xen-blkfront.c b/drivers/block/xen-blkfront.c
> > > index 8dad7bf5f664..69a2315a1b20 100644
> > > --- a/drivers/block/xen-blkfront.c
> > > +++ b/drivers/block/xen-blkfront.c
> > > @@ -2138,10 +2138,8 @@ static void blkfront_closing(struct blkfront_info *info)
> > > return;
> > > /* No more blkif_request(). */
> > > - if (info->rq && info->gd) {
> > > - blk_mq_stop_hw_queues(info->rq);
> > > + if (info->gd)
> > > blk_mark_disk_dead(info->gd);
> >
> > So blk_mark_disk_dead() behaves differently when called with the
> > queues still active?
>
> Yes. In both cases, the queues are frozen and the hw queues synchronously run,
> but the latter is a no-op when the hw queues are in the stopped state,
> therefore it could leave requests unprocessed.
>
> How about this for the final paragraph of the commit message?
>
> """
> Fix this by simply not stopping the hw queues directly. Marking the disk
> as dead already freezes the queue which prevents new requests being added
> and it synchronously runs the hw queues to clear anything pending. If the hw
> queues are stopped when calling blk_mark_disk_dead(), running the hw queues
> is a no-op and can leave queued requests unprocessed.
> """
LGTM.
I think it might be best if you send v2 with that adjustment and the
Fixes tag possibly dropped.
Thanks, Roger.
prev parent reply other threads:[~2026-10-06 15:27 UTC|newest]
Thread overview: 4+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-06 13:58 Ross Lagerwall
2026-10-06 14:14 ` Roger Pau Monné
2026-10-06 14:43 ` Ross Lagerwall
2026-10-06 15:27 ` Roger Pau Monné [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=asUTa3lIC0gkNAJB@Mac.lan \
--to=roger@xenproject.org \
--cc=axboe@kernel.dk \
--cc=jgross@suse.com \
--cc=linux-block@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=oleksandr_tyshchenko@epam.com \
--cc=ross.lagerwall@citrix.com \
--cc=sstabellini@kernel.org \
--cc=stable@vger.kernel.org \
--cc=xen-devel@lists.xenproject.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®