|
[Date Prev][Date Next][Thread Prev][Thread Next][Date Index][Thread Index] Re: [PATCH] xen-blkfront: Fix IO race during unplug
On Tue, Oct 06, 2026 at 03:43:44PM +0100, Ross Lagerwall wrote:
> On 10/6/26 3:14 PM, Roger Pau Monné wrote:
> > On Tue, Oct 06, 2026 at 02:58:06PM +0100, Ross Lagerwall wrote:
> > > During unplug, blkfront stops the hw queues and marks the disk as dead,
> > > then later during removal calls del_gendisk(). However, IO issued after
> > > the hw queues are stopped but before the call to del_gendisk() will be
> > > queued but never handled. This causes del_gendisk() to hang forever
> > > waiting for the queue refcount to drop to zero.
> > >
> > > This can be reproduced by issuing IO during an artificial delay after
> > > stopping the hw queues.
> >
> > So the window is between the blk_mq_stop_hw_queues() and
> > blk_mark_disk_dead() calls where requests would be queued and never
> > processed?
>
> Yes.
>
> >
> > > Fix this by simply not stopping the hw queues directly. Marking the disk
> > > as dead also freezes the queue which prevents new requests being added
> > > and it synchronously runs the hw queues to clear anything pending.
> >
> > I think this likely needs expanding a bit: the disk is already marked
> > as dead with the current logic, and in the same place in the code.
> >
> > > Fixes: 8e141f9eb803 ("block: drain file system I/O on del_gendisk")
> > > Cc: stable@xxxxxxxxxxxxxxx
> > > Assisted-by: LLM
> > > Signed-off-by: Ross Lagerwall <ross.lagerwall@xxxxxxxxxx>
> > > ---
> > >
> > > I'm not sure about the Fixes tag. It's the most likely looking candidate
> > > to me but I didn't confirm whether it actually introduced the
> > > regression.
> >
> > I was wondering the same, and even then someone might argue this was a
> > latent bug in blkfront itself, and the reference commit just exposed
> > it. I don't have a strong opinion.
>
> Good point, I might be inclined to drop it then.
>
> >
> > >
> > > drivers/block/xen-blkfront.c | 4 +---
> > > 1 file changed, 1 insertion(+), 3 deletions(-)
> > >
> > > diff --git a/drivers/block/xen-blkfront.c b/drivers/block/xen-blkfront.c
> > > index 8dad7bf5f664..69a2315a1b20 100644
> > > --- a/drivers/block/xen-blkfront.c
> > > +++ b/drivers/block/xen-blkfront.c
> > > @@ -2138,10 +2138,8 @@ static void blkfront_closing(struct blkfront_info
> > > *info)
> > > return;
> > > /* No more blkif_request(). */
> > > - if (info->rq && info->gd) {
> > > - blk_mq_stop_hw_queues(info->rq);
> > > + if (info->gd)
> > > blk_mark_disk_dead(info->gd);
> >
> > So blk_mark_disk_dead() behaves differently when called with the
> > queues still active?
>
> Yes. In both cases, the queues are frozen and the hw queues synchronously run,
> but the latter is a no-op when the hw queues are in the stopped state,
> therefore it could leave requests unprocessed.
>
> How about this for the final paragraph of the commit message?
>
> """
> Fix this by simply not stopping the hw queues directly. Marking the disk
> as dead already freezes the queue which prevents new requests being added
> and it synchronously runs the hw queues to clear anything pending. If the hw
> queues are stopped when calling blk_mark_disk_dead(), running the hw queues
> is a no-op and can leave queued requests unprocessed.
> """
LGTM.
I think it might be best if you send v2 with that adjustment and the
Fixes tag possibly dropped.
Thanks, Roger.
|
![]() |
Lists.xenproject.org is hosted with RackSpace, monitoring our |