[Date Prev][Date Next][Thread Prev][Thread Next][Date Index][Thread Index]

Re: [PATCH] xen-blkfront: Fix IO race during unplug


  • To: Roger Pau Monné <roger@xxxxxxxxxxxxxx>
  • From: Ross Lagerwall <ross.lagerwall@xxxxxxxxxx>
  • Date: Tue, 6 Oct 2026 15:43:44 +0100
  • Arc-authentication-results: i=1; mx.microsoft.com 1; spf=pass smtp.mailfrom=citrix.com; dmarc=pass action=none header.from=citrix.com; dkim=pass header.d=citrix.com; arc=none
  • Arc-message-signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=lhQaN5T1hU1rRx42cQYMLNsDzKy6RS7+ImA5ST0O1uA=; b=aKuV98wh4TuBGAktz9vBSRkpPCDMUxtMVIKqefc8C13NL/LHS8fqHAYSr9B/i1nzoo1tzCNcwqip44w9U4aL/QOBEuEDOjU43r21AtOFs6oOcnuKYcdrXYC29zDoXXjJ2VNmxX1GWiboXHuVcUECMXJ0Lt7T+Eeoexo5/e6Ap7+dyx3GxOeEIirur6EaFgfsrvAvEhjzFrTFkNSsxWi4VdtBSyhnJuFEDYwvqw7DC2w/y/DrJwhcGTGFQxXwN+u+5ekKZZRALNgZBdrbQa4Qw6yQbrH5Qw+bMZNzMtEFT0pJCCzAWH0WMN45mxgfmCYa5ugXdQRwhDXTz1Xcv0i1Vg==
  • Arc-seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=RSZG6PIOrWZ6sqwU6NOv8mVzS8R/kN/9FPpUj374XZUh3g6e5qCABFNXaLSDAiY6Bcy1r5DVfKBy2H6sxCpapDv0kSyv8nI5lA82W0mi1AaQdo/Vvrzmq6H1aeTi5avk7AWmRpG25Ar7z6BOYXLHU0KHAb+SKWcYEezR6gUit/gVxG0GFyIKmVooDuRX56M3W9eG8DDyIzKPB23E68Agr+ygFj3cRmdqpktpvp2fjld65SadIPI0dyEqoHcKmE+djf2xjhmZIMt0tkXrJQzggduKYFA8fz5nw67u/zK0bEELX6pgOJbrgk5T8WjanSGr1NzMfdg//wY0j7Gj7B0gIw==
  • Authentication-results: eu.smtp.expurgate.cloud; dkim=pass header.s=selector1 header.d=citrix.com header.i="@citrix.com" header.h="From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck"
  • Authentication-results: mx.microsoft.com 1; dkim=none (message not signed) header.d=none;dmarc=none action=none header.from=citrix.com;
  • Cc: xen-devel@xxxxxxxxxxxxxxxxxxxx, linux-block@xxxxxxxxxxxxxxx, linux-kernel@xxxxxxxxxxxxxxx, Juergen Gross <jgross@xxxxxxxx>, Stefano Stabellini <sstabellini@xxxxxxxxxx>, Oleksandr Tyshchenko <oleksandr_tyshchenko@xxxxxxxx>, Jens Axboe <axboe@xxxxxxxxx>, stable@xxxxxxxxxxxxxxx
  • Delivery-date: Tue, 06 Oct 2026 14:44:14 +0000
  • List-id: Xen developer discussion <xen-devel.lists.xenproject.org>

On 10/6/26 3:14 PM, Roger Pau Monné wrote:
On Tue, Oct 06, 2026 at 02:58:06PM +0100, Ross Lagerwall wrote:
During unplug, blkfront stops the hw queues and marks the disk as dead,
then later during removal calls del_gendisk(). However, IO issued after
the hw queues are stopped but before the call to del_gendisk() will be
queued but never handled. This causes del_gendisk() to hang forever
waiting for the queue refcount to drop to zero.

This can be reproduced by issuing IO during an artificial delay after
stopping the hw queues.

So the window is between the blk_mq_stop_hw_queues() and
blk_mark_disk_dead() calls where requests would be queued and never
processed?

Yes.


Fix this by simply not stopping the hw queues directly. Marking the disk
as dead also freezes the queue which prevents new requests being added
and it synchronously runs the hw queues to clear anything pending.

I think this likely needs expanding a bit: the disk is already marked
as dead with the current logic, and in the same place in the code.

Fixes: 8e141f9eb803 ("block: drain file system I/O on del_gendisk")
Cc: stable@xxxxxxxxxxxxxxx
Assisted-by: LLM
Signed-off-by: Ross Lagerwall <ross.lagerwall@xxxxxxxxxx>
---

I'm not sure about the Fixes tag. It's the most likely looking candidate
to me but I didn't confirm whether it actually introduced the
regression.

I was wondering the same, and even then someone might argue this was a
latent bug in blkfront itself, and the reference commit just exposed
it.  I don't have a strong opinion.

Good point, I might be inclined to drop it then.



  drivers/block/xen-blkfront.c | 4 +---
  1 file changed, 1 insertion(+), 3 deletions(-)

diff --git a/drivers/block/xen-blkfront.c b/drivers/block/xen-blkfront.c
index 8dad7bf5f664..69a2315a1b20 100644
--- a/drivers/block/xen-blkfront.c
+++ b/drivers/block/xen-blkfront.c
@@ -2138,10 +2138,8 @@ static void blkfront_closing(struct blkfront_info *info)
                return;
/* No more blkif_request(). */
-       if (info->rq && info->gd) {
-               blk_mq_stop_hw_queues(info->rq);
+       if (info->gd)
                blk_mark_disk_dead(info->gd);

So blk_mark_disk_dead() behaves differently when called with the
queues still active?

Yes. In both cases, the queues are frozen and the hw queues synchronously run,
but the latter is a no-op when the hw queues are in the stopped state,
therefore it could leave requests unprocessed.

How about this for the final paragraph of the commit message?

"""
Fix this by simply not stopping the hw queues directly. Marking the disk
as dead already freezes the queue which prevents new requests being added
and it synchronously runs the hw queues to clear anything pending. If the hw
queues are stopped when calling blk_mark_disk_dead(), running the hw queues
is a no-op and can leave queued requests unprocessed.
"""

Ross



 


Rackspace

Lists.xenproject.org is hosted with RackSpace, monitoring our
servers 24x7x365 and backed by RackSpace's Fanatical Support®.