From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pl1-f169.google.com (mail-pl1-f169.google.com [209.85.214.169]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 96818276049 for ; Wed, 27 May 2026 07:53:45 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.169 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1779868426; cv=none; b=QsiemCANBbsMBWnEy2eUU6AEvsrywdlGtNFZ12lJMsJT6FeIIoYtBec+N6mfyHVeHCupw0+cY2a5qHa9NTjp/zobEbwklcCE9pRK+hNWhWm8/yyGLDadfrMB+3eurj12bPOYg3S36SvNpUsvOmCrL9aRmsKFJ/VjlM0cpkclW/w= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1779868426; c=relaxed/simple; bh=uwc1Fafbf0eO6JSlyC6f97afRDZ/jkIMxtp+JLbFS/0=; h=From:To:Cc:Subject:Date:Message-Id:MIME-Version; b=ScC7a0PWfCzkoDmpj5D3FQGLXBI5rHKZ8buAndDairrSI42O2U/Ktgw3WNA46nUO4I/VeIyObhv2nDSJ0lqPmRiRRWZ5ivxhDvMfknkxLsRq6PVF6WQFF1zIfkwYnItmQTn3CLLmazv0cMU572nzQxVPZMM7/n+acQxxGcFy43c= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=JPqmHFDf; arc=none smtp.client-ip=209.85.214.169 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="JPqmHFDf" Received: by mail-pl1-f169.google.com with SMTP id d9443c01a7336-2ba6485d219so88136725ad.3 for ; Wed, 27 May 2026 00:53:45 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1779868425; x=1780473225; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:from:to:cc:subject:date:message-id:reply-to; bh=b+YxUb9tnY9+516KxTbXxWzTyfA5pPGtJElvbz03+zw=; b=JPqmHFDfsWsDNo7WH3eLhANjtOBbzTUHLG2Tbykc6ZUdEjo5zDZOvj70pwV8TcJ0OQ CN5zyxRbEp/L+8zEWTMa6Kez6jBoewcSjGxXyHRdclABYWcY/ij5E4YTaEUN4l4Acis0 oT1y33hHkszfz2xdfeTa15vL1JBf3bnHdcYXB3+rjJysOM7arIn5TG26YO3t/RDbX0vl iP6TLQ/9c3u1ri9a6rC4d+RzzN7Sa5QT9HSXB/gNUJ+oUWGN38jlliUkuHvwJ+eHDhcA uIDi3t8zaKXCb/LYZKbnNCrszYQvLH0irjNOLeo9wrUscQbOkanOKD4Jhko2WREzg4SD GrcA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1779868425; x=1780473225; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:x-gm-gg:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to; bh=b+YxUb9tnY9+516KxTbXxWzTyfA5pPGtJElvbz03+zw=; b=fbr6JT0r5Gadd724LUYvV6/y7D/EknUR7xr5EfFr5WtDUg+Mfd3RxUgPyMbMasv1VN UOjZ4o6X3B2phPmt3qHkpbYmnMSyVNLxWiX6F0Lf6xxDNIqypxak9SjydF0NhAkHIUJd 2xdsJ+Bej6OcrEZ8WPyjMdRxLW6EkZnGqeDy9pbFj0vkNykmmRHvdzeQbeyIt6ac2Cpx 5Z8RUsfR/qQK6gDBbw5AYfqKYc7curE5JxoazYtAZwVgGxfzpn9xq3PaQx5CN0Nfd8ym GLqpKFYR1WvAayYjnOO9q3KShW8iWlVRmX58i5y/gUdMrVjjzo2eN69rSl/8PbsuRk2D hE4w== X-Forwarded-Encrypted: i=1; AFNElJ/M8NRsNKg8oZGrjxAfqW9mmVqStcxHoOU2eZ16J6nV14yCGUseWyWxVqYX1WkM+vIpDJdkLiuBIfSoJeg=@vger.kernel.org X-Gm-Message-State: AOJu0YyRq56pEaF4wzMeZqVupz6hmrHcUMDti5+MKcm+aEhwIh5Gpi8d cQWgzTf2lU2FDSC5lihGVfIKXCAB7LwMbSDiWL318opCaZRNnlmxdkPX X-Gm-Gg: Acq92OEt/IR/K7pgT1RILzfjdyrfjfz11xNyffMwnj9WyMxK4JgSYzNJI1b30IFbEk3 U1xkVgEI1mMxxQzPrr5LwWqW3yt7Oy+uZV2EnRGhLUmW368fL2SZdYbraGnq+3q1ujl4robhfJF Duzy95Z9jzHRLHpkQ6dFSYj/NiQHMzu/A5P1ydr/s1bJt1c7JiPW+VVfNikMq+rBpUGIi/JA4DQ Sydd1FmW72jsKwZ1G8I1swDPNJQYCwbrJBKKBW89+V3Gl6xPGH1S7MR+onJkknjTU99ii0RE/kG +ENp5gCgzOkh3e3khBBkloloatlJOPJvaCsaHJYLFvDfFNeL/x+wd8/qo1V6czzrXdTwDMsYLFT DFu8rL6BijgGuhrBosDTGyWwZEH5YesO+P/rkB/R6qClzfU5EMkz2o3zCl9NnHnJRGAr6QsNY0B xG8BRFblGNpji9noUPHQq3s+0/jzmGhQn9eXWdBGQ= X-Received: by 2002:a17:902:db07:b0:2ba:3e50:e3f5 with SMTP id d9443c01a7336-2beb0634c8amr248897615ad.30.1779868424801; Wed, 27 May 2026 00:53:44 -0700 (PDT) Received: from n151-105-216.byted.org ([240e:b1:e401:3::a4]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2beb56bc151sm147558455ad.24.2026.05.27.00.53.42 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 27 May 2026 00:53:44 -0700 (PDT) From: guzebing To: kbusch@kernel.org, axboe@kernel.dk, hch@lst.de, sagi@grimberg.me Cc: linux-nvme@lists.infradead.org, linux-kernel@vger.kernel.org, Guzebing Subject: [RFC PATCH 0/1] nvme-pci: detect I/O queue depth changes after reset Date: Wed, 27 May 2026 15:53:19 +0800 Message-Id: <20260527075320.3178600-1-guzebing1612@gmail.com> X-Mailer: git-send-email 2.20.1 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit From: Guzebing We have hit a case where an NVMe firmware activation made the controller report a different CAP.MQES value after the following controller reset. This was seen in production on a Memblaze PBlaze5 510 device: model: P5510DS0384T00 old firmware: 224005A0, CAP.MQES-derived queue depth 1024 new firmware: 224005F0, CAP.MQES-derived queue depth 128 One way to hit this path is to activate the new firmware and then reset the controller: nvme fw-download /dev/nvmeX -f fw_bin.tar nvme fw-activate /dev/nvmeX -s 2 -a 1 nvme reset /dev/nvmeX When the I/O queue depth derived from CAP.MQES became smaller after that reset, the driver failed to recover any usable I/O queues. In our production kernel this was logged as: nvme nvme0: IO queues lost The namespaces were then removed, so the corresponding block device disappeared. The opposite direction is less visible: if the CAP-derived depth becomes larger, reset can complete without an error and the block device can remain usable, but the live queue depth state is not updated consistently. The reason is that reset updates only part of the live queue depth state. The nvme-pci reset path disables the controller, re-enables it, re-reads CAP, and recalculates: dev->q_depth ctrl->sqsize from the new CAP.MQES value. Later, however, nvme_create_io_queues() reuses the existing struct nvme_queue entries. nvme_alloc_queue() returns immediately when the queue already exists, so the old values remain in: nvmeq->q_depth nvmeq->cq_dma_addr nvmeq->sq_dma_addr Create CQ/SQ then requests queues with nvmeq->q_depth entries (encoded in the command as nvmeq->q_depth - 1) and uses the old SQ/CQ DMA addresses, not the newly computed dev->q_depth. The blk-mq side also keeps the old depth: the reset path updates the number of hardware queues through blk_mq_update_nr_hw_queues(), but it does not resize the existing tag set or update its queue_depth. This explains the observed shrink failure and the expected grow case: * If the CAP-derived depth becomes smaller, the driver may try to create an I/O queue with the old larger nvmeq->q_depth. A controller that now enforces the smaller CAP.MQES-derived limit can reject the Create CQ/SQ command. If no I/O queues are recovered, nvme-pci removes the namespaces, so the block device disappears. * If the CAP-derived depth becomes larger, the old nvmeq->q_depth is still within the new controller limit. Queue creation can therefore succeed and the device can remain usable, but the live state is inconsistent: dev->q_depth and ctrl->sqsize reflect the new capability while nvmeq queue resources and the blk-mq tag set still reflect the old depth. The larger depth is not used until the controller is removed and probed again. There are two broad ways to address this. The direct fix would be to make reset recovery handle a changed live queue depth. That would require updating or rebuilding the nvmeq depth and SQ/CQ DMA allocations, and resizing the block-layer depth state consistently, including the blk-mq tag set, scheduler tags when present, and queue->nr_requests. That is broader than an nvme-pci-only change and needs block layer review. This RFC instead takes the smaller approach of detecting the reset-time CAP.MQES change and making it visible. If the live I/O queue depth shrinks, reset recovery is failed before recreating I/O queues. If it grows, the driver warns and continues with the existing queue resources. Feedback would be appreciated on whether this detection is useful on its own, or whether nvme-pci should instead support full live queue-depth resizing together with the required blk-mq changes. Guzebing (1): nvme-pci: detect I/O queue depth changes after reset drivers/nvme/host/pci.c | 30 ++++++++++++++++++++++++++++++ 1 file changed, 30 insertions(+) -- 2.20.1