From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-vs1-f99.google.com (mail-vs1-f99.google.com [209.85.217.99]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8E68B36F916 for ; Thu, 25 Jun 2026 21:27:45 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.217.99 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1782422867; cv=none; b=Xt5zijzmBeLYibLjt6FvZr2m7cp4uKtj5DcDNrwa3DhN2VRkMShLmHfaXsb2CP60f4Cf4TM6+lO4DLaoZ3MyeU4OMXzS6j7PjSKKNI9ZdYPllEUhcqlVK7sHthHXrT3b5ppEFpPakMGkip22KNAsCIe+QLaxhzM0F9WDtI7aTVk= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1782422867; c=relaxed/simple; bh=oEuJD60fGH02qZlcsbhdJa3ivIUS8QZcVLD616moe8I=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=sOPbT2xWUsxY/z5j/L/BUYyaNb6ijltOOUgnhdwW7IQ8qu+2b9EfRFqchyxPaCG0rPlfpCFXJX933F7GxfieWBslvWHwlBVj5mPoWhi1N4dh1vfWp76583p7i4kAAQadghBCGnpM7WPDrokeusp5AbFm+fCN2rjUO8IVBLeY2wI= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=purestorage.com; spf=pass smtp.mailfrom=purestorage.com; dkim=pass (2048-bit key) header.d=purestorage.com header.i=@purestorage.com header.b=S2reP7fn; arc=none smtp.client-ip=209.85.217.99 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=purestorage.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=purestorage.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=purestorage.com header.i=@purestorage.com header.b="S2reP7fn" Received: by mail-vs1-f99.google.com with SMTP id ada2fe7eead31-734653f9899so131336137.3 for ; Thu, 25 Jun 2026 14:27:45 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=purestorage.com; s=google2022; t=1782422864; x=1783027664; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to; bh=Y4b2dFsTucb4aU0SZUuf4AYL/tgS35O39eqo6KSrHbE=; b=S2reP7fnT/wjDv789Vnp5WILep+ypvoO5OCc6dZjpOB9dCNEzn9FX87XMLRghMs8NM NOZwjqGyJJUyDW/rg2L1Bj24KohR9ZxMWs73A6wuts/4VNDgtBfrke9mAsAlusCw8z4B Zjxz2MNi9suqjHVs4zrA7VOfh+WMppAYJyEtq9ThvtTMgl7pDkCXLgGqmLs25CCqahLu txEmdWGrX3xgyUgJeYYOrgNh7c14Z1fP1zsvkHK8NYiJ9cT4hX615/I0sJQkaCxdHrzV a3NCazaeiLamN1cVGhRwZVUBzIfFkqdlQQs8Ie1D0Uap0oWuI4wVj6UgpNp28TVx9xgj 3A/g== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1782422864; x=1783027664; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to; bh=Y4b2dFsTucb4aU0SZUuf4AYL/tgS35O39eqo6KSrHbE=; b=W59tZ9rzYxOMhQBQPcVRIBZIa38f5LfcSh1M5u/n9nQPEpBD0CTGZGYxFyZNYwWPvG T+BwxO5DhDa+GbbpF1/gBG0wl3/EZU1LpBb5gNpWX852CHQLMAlq3zS7Sd4cWobwZAIP koyR+HDDhWFeE40dFTyhkYnOQINdPsHuX04qBDX/5pgdV/OC2N3aNND73wzi57Cj6Dqr WVOdjjQTnoRwzwWMrwEO4yJCCksh53ssPQiJPG8ra8V25wSp/qYrqNndQoG8hS7CB4h7 jN8yTFTE1NxDBd/8ngC1d1fJxZuuX5X6o+p8h4GG4UrDSdCqimNf0FO40P7ZxuP3myfu uqNA== X-Forwarded-Encrypted: i=1; AHgh+RpM35S7RyqReODUKw1ghEq5Fua0PG50P51DB/chPCqTgavPxWWTsYHITksMll708l4UQHgQ2Txrkt1g9ws=@vger.kernel.org X-Gm-Message-State: AOJu0YxghEs77J9e6z9SXfIYNs919OL6KJqs+igJ5ao+Jiqc/C+6eJ5t hzkHx83KdARODDFJMEeFmGQ5Oryb5zXY+IKMDO0sO0zM6HDkDsR1oy9ZFklGt2AKLi1U/4WzplG Lv3F7aC3izxBqhZ5Kz09fhx0ztv+zYSAVWMSo X-Gm-Gg: AfdE7ckrFGqFX1TA1CYehWQXwlXjyblMJ4XgIZS5MQIKCsdT952iBnOF5A+yubzutUV 5vwhfFrZ37yQoAQKxTG/ubB5Tb4Nwd8k+HF5dBggcFIitfNoRSt0kKxDB4IzPGyHyBywxd7qCCB I5yfhPy4WDMF89tsTFVLgkaubFVNKh0aVysmHyPjclz+3YqbnnkgVpcFkP8dENAQxI0r2ELWzgD FIUvJ+C7MRnQOyrl7B0JqTgprqSAmp/PxIpnrm65fdpgjx66xVZLNbt3Z7yGTZnVsHcTMW44JMl q2rYQ5ULLtEVfGhg3aBomQjjHwlYE0GDctAYY1UriCy67PmJvXHa09tPRecdmig+MHy3i+Qy341 Hv83jyAVBsLJ1KCcblKzJYdQYoQqCg5pZ61cAxYxyyg== X-Received: by 2002:a05:6102:c8a:b0:6c6:74d:e090 with SMTP id ada2fe7eead31-734349573a6mr2562285137.9.1782422864391; Thu, 25 Jun 2026 14:27:44 -0700 (PDT) Received: from c7-smtp-2026.dev.purestorage.com ([208.88.159.129]) by smtp-relay.gmail.com with ESMTPS id ada2fe7eead31-7356689f0aasm18944137.5.2026.06.25.14.27.44 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 25 Jun 2026 14:27:44 -0700 (PDT) X-Relaying-Domain: purestorage.com Received: from dev-sgogte.dev.purestorage.com (bond0.slc5-n22m24-k8s.dev.purestorage.com [IPv6:2620:125:9025:20::a31:429]) by c7-smtp-2026.dev.purestorage.com (Postfix) with ESMTP id C320F401FD; Thu, 25 Jun 2026 15:27:43 -0600 (MDT) Received: by dev-sgogte.dev.purestorage.com (Postfix, from userid 1557734945) id C0B8751222; Thu, 25 Jun 2026 15:27:43 -0600 (MDT) From: Surabhi Gogte To: Keith Busch , Jens Axboe , Christoph Hellwig , Sagi Grimberg Cc: linux-nvme@lists.infradead.org, linux-kernel@vger.kernel.org, mkhalfella@purestorage.com, randyj@purestorage.com, adailey@purestorage.com, Surabhi Gogte Subject: [PATCH v3 2/2] nvme-rdma: parallelize I/O queue allocation and startup Date: Thu, 25 Jun 2026 15:27:22 -0600 Message-ID: <20260625212722.1302344-3-sgogte@purestorage.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260625212722.1302344-1-sgogte@purestorage.com> References: <20260625212722.1302344-1-sgogte@purestorage.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit Refactor nvme rdma I/O queue setup to use async API, combining allocation and startup into a single parallel operation per queue. This reduces connection and reconnection setup time when there are delays in establishing connections, which is especially important for high-core-count hosts. Key changes: - Use async API to facilitate parallel calls for io queue setup. - Add nvme_rdma_setup_ctx for propagating errors from async workers. - Remove nvme_rdma_alloc_io_queues() and nvme_rdma_start_io_queues(); their logic is folded into nvme_rdma_setup_io_queues() and nvme_rdma_configure_io_queues(). - Move queue count negotiation (nvme_set_queue_count, nvmf_set_io_queues) from the removed nvme_rdma_alloc_io_queues() into nvme_rdma_configure_io_queues(). Testing on a 64-core host with 64 IO-queues shows nvme-rdma connection time reduced from ~1.4s to 416ms. Signed-off-by: Surabhi Gogte --- drivers/nvme/host/rdma.c | 124 ++++++++++++++++++++++++--------------- 1 file changed, 77 insertions(+), 47 deletions(-) diff --git a/drivers/nvme/host/rdma.c b/drivers/nvme/host/rdma.c index 6b0b0a3dea62..45d0ef8c4dd3 100644 --- a/drivers/nvme/host/rdma.c +++ b/drivers/nvme/host/rdma.c @@ -16,6 +16,7 @@ #include #include #include +#include #include #include #include @@ -100,6 +101,11 @@ struct nvme_rdma_queue { struct mutex queue_lock; }; +struct nvme_rdma_setup_ctx { + struct nvme_rdma_queue *queue; + int *err; +}; + struct nvme_rdma_ctrl { /* read only in the hot path */ struct nvme_rdma_queue *queues; @@ -690,60 +696,68 @@ static int nvme_rdma_start_queue(struct nvme_rdma_ctrl *ctrl, int idx) return ret; } -static int nvme_rdma_start_io_queues(struct nvme_rdma_ctrl *ctrl, - int first, int last) +static void nvme_rdma_setup_queue_async(void *data, async_cookie_t cookie) { - int i, ret = 0; + struct nvme_rdma_setup_ctx *ctx = data; + struct nvme_rdma_queue *queue; + int ret; - for (i = first; i < last; i++) { - ret = nvme_rdma_start_queue(ctrl, i); - if (ret) - goto out_stop_queues; - } + queue = ctx->queue; + ret = nvme_rdma_alloc_queue(queue); + if (ret) + goto out_err; - return 0; + ret = nvme_rdma_start_queue(queue->ctrl, nvme_rdma_queue_idx(queue)); + if (ret) + goto out_err; -out_stop_queues: - for (i--; i >= first; i--) - nvme_rdma_stop_queue(&ctrl->queues[i]); - return ret; + return; +out_err: + WRITE_ONCE(*ctx->err, ret); } -static int nvme_rdma_alloc_io_queues(struct nvme_rdma_ctrl *ctrl) +static int nvme_rdma_setup_io_queues(struct nvme_rdma_ctrl *ctrl, unsigned int first, + unsigned int last, size_t queue_size) { - struct nvmf_ctrl_options *opts = ctrl->ctrl.opts; - unsigned int nr_io_queues; - int i, ret; - - nr_io_queues = nvmf_nr_io_queues(opts); - ret = nvme_set_queue_count(&ctrl->ctrl, &nr_io_queues); - if (ret) - return ret; + ASYNC_DOMAIN_EXCLUSIVE(queue_domain); + struct nvme_rdma_setup_ctx *ctxs; + int nr_queues = last - first; + int err = 0, i, ret; - if (nr_io_queues == 0) { - dev_err(ctrl->ctrl.device, - "unable to set any I/O queues\n"); + ctxs = kmalloc_array(nr_queues, sizeof(*ctxs), GFP_KERNEL); + if (!ctxs) return -ENOMEM; - } - ctrl->ctrl.queue_count = nr_io_queues + 1; - dev_info(ctrl->ctrl.device, - "creating %d I/O queues.\n", nr_io_queues); - - nvmf_set_io_queues(opts, nr_io_queues, ctrl->io_queues); - for (i = 1; i < ctrl->ctrl.queue_count; i++) { - ctrl->queues[i].ctrl = ctrl; - ctrl->queues[i].queue_size = ctrl->ctrl.sqsize + 1; - ret = nvme_rdma_alloc_queue(&ctrl->queues[i]); - if (ret) - goto out_free_queues; + for (i = 0; i < nr_queues; i++) { + struct nvme_rdma_queue *queue = &ctrl->queues[first + i]; + + queue->ctrl = ctrl; + queue->queue_size = queue_size; + + ctxs[i].queue = queue; + ctxs[i].err = &err; + async_schedule_domain(nvme_rdma_setup_queue_async, &ctxs[i], + &queue_domain); } - return 0; + async_synchronize_full_domain(&queue_domain); + kfree(ctxs); + ret = READ_ONCE(err); + if (ret) + goto out_free_queues; + + return 0; out_free_queues: - for (i--; i >= 1; i--) - nvme_rdma_free_queue(&ctrl->queues[i]); + for (i = 0; i < nr_queues; i++) { + struct nvme_rdma_queue *queue = + &ctrl->queues[first + i]; + + if (test_bit(NVME_RDMA_Q_LIVE, &queue->flags)) + nvme_rdma_stop_queue(queue); + if (test_bit(NVME_RDMA_Q_ALLOCATED, &queue->flags)) + nvme_rdma_free_queue(queue); + } return ret; } @@ -862,12 +876,23 @@ static int nvme_rdma_configure_admin_queue(struct nvme_rdma_ctrl *ctrl, static int nvme_rdma_configure_io_queues(struct nvme_rdma_ctrl *ctrl, bool new) { + unsigned int nr_io_queues; int ret, nr_queues; - ret = nvme_rdma_alloc_io_queues(ctrl); + nr_io_queues = nvmf_nr_io_queues(ctrl->ctrl.opts); + ret = nvme_set_queue_count(&ctrl->ctrl, &nr_io_queues); if (ret) return ret; + if (nr_io_queues == 0) { + dev_err(ctrl->ctrl.device, "unable to set any I/O queues\n"); + return -ENOMEM; + } + + ctrl->ctrl.queue_count = nr_io_queues + 1; + dev_info(ctrl->ctrl.device, "creating %d I/O queues.\n", nr_io_queues); + nvmf_set_io_queues(ctrl->ctrl.opts, nr_io_queues, ctrl->io_queues); + if (new) { ret = nvme_rdma_alloc_tag_set(&ctrl->ctrl); if (ret) @@ -880,7 +905,9 @@ static int nvme_rdma_configure_io_queues(struct nvme_rdma_ctrl *ctrl, bool new) * queue number might have changed. */ nr_queues = min(ctrl->tag_set.nr_hw_queues + 1, ctrl->ctrl.queue_count); - ret = nvme_rdma_start_io_queues(ctrl, 1, nr_queues); + ret = nvme_rdma_setup_io_queues(ctrl, 1, nr_queues, + ctrl->ctrl.sqsize + 1); + if (ret) goto out_cleanup_tagset; @@ -904,12 +931,15 @@ static int nvme_rdma_configure_io_queues(struct nvme_rdma_ctrl *ctrl, bool new) /* * If the number of queues has increased (reconnect case) - * start all new queues now. + * setup all new queues now. */ - ret = nvme_rdma_start_io_queues(ctrl, nr_queues, - ctrl->tag_set.nr_hw_queues + 1); - if (ret) - goto out_wait_freeze_timed_out; + if (ctrl->tag_set.nr_hw_queues + 1 > nr_queues) { + ret = nvme_rdma_setup_io_queues(ctrl, nr_queues, + ctrl->tag_set.nr_hw_queues + 1, + ctrl->ctrl.sqsize + 1); + if (ret) + goto out_wait_freeze_timed_out; + } return 0; -- 2.54.0