From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-dy1-f227.google.com (mail-dy1-f227.google.com [74.125.82.227]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6932D219E8 for ; Sat, 27 Jun 2026 04:16:58 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.82.227 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1782533820; cv=none; b=H3IluCK0DL2a8kRUTb5vL7zzhFTpR2vVf5eOEYbwX4tZA72pqdMOJf9k4Zu/f/sHibab0tSpvXyUf6z5DTvyvQahF/T/vBsl5ciLQejimCIv4p65gDMbNDdiwrMKW/tegAM+IjOy769Ed5i4A9ovm3zb+EHVcT73AftLHbyMesA= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1782533820; c=relaxed/simple; bh=CsWWDpOsJ5M4JVd3SeuvCLDHlGuW1ynlLhkD+lNZItU=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=sVlEV0HPhxN+OZgElMjBbwkR+H/kIR2dpSu2dKeTf3P+UMoyxMPnvw6tZYHkZgdtTmsOnCjj///wYkqPbfNoGnrn5gEIAcdsHlMJyX48b6Qh1mph8YcJriXpx0wkU/ck/95E3VK15XQVP9n5RA5Ko9uvS8Ps7OJhhh0GZb4sRAw= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=purestorage.com; spf=pass smtp.mailfrom=purestorage.com; dkim=pass (2048-bit key) header.d=purestorage.com header.i=@purestorage.com header.b=fV623e+u; arc=none smtp.client-ip=74.125.82.227 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=purestorage.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=purestorage.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=purestorage.com header.i=@purestorage.com header.b="fV623e+u" Received: by mail-dy1-f227.google.com with SMTP id 5a478bee46e88-30b6dad2382so3367424eec.0 for ; Fri, 26 Jun 2026 21:16:58 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=purestorage.com; s=google2022; t=1782533817; x=1783138617; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to; bh=PaVrMx15yJC3UQAct8EhgWFhwmFAMOShPun1TdcWdQM=; b=fV623e+uWz0stMeTVzYMea7o3rKb3slgujv8vHOTTd0d2otj0Ba9Xu/NsoC4Xay3UV oyue6nGJwHCTJqGL0oDEQuXw7NyOW9xzRGBEPsGOwNrXcuHPxezkrz1b6mSeTLvN3xQU zeoEJd5bzou24/+Hex6xCFCTwS7Q2c8w8tM61wKULAbI+ernJlB/rBAHpQoo7hAQFEST 4/BoX130bT6nZvUMe4Vm4a41gxqKW2BKdOtO7G2NDfCF5ZIuRriEKCMoqXGCqzNkF+FC gM7adlLZHHFMoOjmlB62a2+9HW8XfCAA9n0jXmsQ+xDoMFdppwYp98rbOB0SJdjtlJoQ BcXw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1782533817; x=1783138617; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to; bh=PaVrMx15yJC3UQAct8EhgWFhwmFAMOShPun1TdcWdQM=; b=P/kXq5qXPz2iDYlzNbfClRyMa2YyNl+6Nk5oLmQFlpI21Gvm1y5e9x0u+gSk+TsbNO t96vIqsUyaJ1q54dRbVOxhiEVpT2ZW/urcXwVSlHdpdhLHPlf7JVbYojZC9SFhnaCMgz z9HiO3mfFxkJXw/e5wUUT4dMRnHyM2sIuiRMSa3oT1K0pu7aH5YheNgyJetQxJgnX5yY exRjNA92nDuXRxkkCuwcuiCZzABxOjZK7pXtfya6AhD19lB6G0umXhUUCTe970uvDk3V rKS8jlJnTnpnsInwGwZqmGfKQH4LWKuQ+PRFV91oJLHmbH8Mrm15pZhtmuB2pYYLJeo3 HHQQ== X-Forwarded-Encrypted: i=1; AHgh+RoJrhxVIIFbyynT6BCdPeSAmT8b+jEvcLdyXyqfrfAbUjFItH4TkjuJlx9GVpfrCMawpW3x0IP+huDutoI=@vger.kernel.org X-Gm-Message-State: AOJu0YzCIZBhZYFLZS0eP1pfRqy5RYzJK0u1m6hmvO7mYMJJ/z6NsJV0 UJvIdpuy5/xFINsRpVotNQyzifICM6YgSIqiOBIIQbG9MN4ycllrzELTe2gf8hfFa8A4jyWcTud IG+kAyVQLl87YwDW2lDIS5yEv1jRScRz6hIzY X-Gm-Gg: AfdE7cmN8hmjC2wFp7RQ01zsVbg8pVUAl7trDbRjxAW1Xjc5iztIV/ai2C2K99iHExd FP0zvmMZeLeXVW7qQpqPZ4uQnofFozvN9+OZU1WSX+/jXyDhJdpVoImZqqn5im67lYCnlVEsINt 1a0uTILkwapyG7dbAkQnKmSzjOprriWJEokFU6fUKaJyanJ02SJjmBxWIiPH3OJRpKUpzwAXoEk YG+oSma0filJ9XK9b3HcoSTcVmceG3yrGGdF3xPOtCoiK7UVRBUnivLxp+ESSMRO3/EB7AnmJ3J +jfSBZ0reeDmpMgE2pzwJ0CyAK2LEILO88D+6UBGSSdAghJeykq0dfwKhl00J64vv3dqbR8uObl 0bV6fxL6WrOUpRU+zqpDOax4C/WlrcL5Jx5ZukIey6g== X-Received: by 2002:a05:7301:7bc2:b0:30c:ab4d:da43 with SMTP id 5a478bee46e88-30cab4ddbdemr2560368eec.39.1782533817219; Fri, 26 Jun 2026 21:16:57 -0700 (PDT) Received: from c7-smtp-2026.dev.purestorage.com ([208.88.159.129]) by smtp-relay.gmail.com with ESMTPS id 5a478bee46e88-30c7c443596sm617443eec.4.2026.06.26.21.16.57 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 26 Jun 2026 21:16:57 -0700 (PDT) X-Relaying-Domain: purestorage.com Received: from dev-sgogte.dev.purestorage.com (bond0.slc5-n22m24-k8s.dev.purestorage.com [IPv6:2620:125:9025:20::a31:429]) by c7-smtp-2026.dev.purestorage.com (Postfix) with ESMTP id 4DCD840146; Fri, 26 Jun 2026 22:16:56 -0600 (MDT) Received: by dev-sgogte.dev.purestorage.com (Postfix, from userid 1557734945) id 45C8C51219; Fri, 26 Jun 2026 22:16:56 -0600 (MDT) From: Surabhi Gogte To: Christoph Hellwig , Keith Busch , Jens Axboe , Sagi Grimberg Cc: linux-nvme@lists.infradead.org, linux-kernel@vger.kernel.org, mkhalfella@purestorage.com, randyj@purestorage.com, adailey@purestorage.com, Surabhi Gogte Subject: [PATCH v4 2/2] nvme-rdma: parallelize I/O queue allocation and startup Date: Fri, 26 Jun 2026 22:15:51 -0600 Message-ID: <20260627041551.1981256-3-sgogte@purestorage.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260627041551.1981256-1-sgogte@purestorage.com> References: <20260627041551.1981256-1-sgogte@purestorage.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit Refactor nvme rdma I/O queue setup to use async API, combining allocation and startup into a single parallel operation per queue. This reduces connection and reconnection setup time when there are delays in establishing connections, which is especially important for high-core-count hosts. Key changes: - Use async API to facilitate parallel calls for io queue setup. - Add nvme_rdma_setup_ctx for propagating errors from async workers. - Remove nvme_rdma_alloc_io_queues() and nvme_rdma_start_io_queues(); their logic is folded into nvme_rdma_setup_io_queues() and nvme_rdma_configure_io_queues(). - Move queue count negotiation (nvme_set_queue_count, nvmf_set_io_queues) from the removed nvme_rdma_alloc_io_queues() into nvme_rdma_configure_io_queues(). Testing on a 64-core host with 64 IO-queues shows nvme-rdma connection time reduced from ~1.4s to 416ms. Signed-off-by: Surabhi Gogte --- drivers/nvme/host/rdma.c | 124 ++++++++++++++++++++++++--------------- 1 file changed, 77 insertions(+), 47 deletions(-) diff --git a/drivers/nvme/host/rdma.c b/drivers/nvme/host/rdma.c index 6b0b0a3dea62..52933d11ea03 100644 --- a/drivers/nvme/host/rdma.c +++ b/drivers/nvme/host/rdma.c @@ -16,6 +16,7 @@ #include #include #include +#include #include #include #include @@ -100,6 +101,11 @@ struct nvme_rdma_queue { struct mutex queue_lock; }; +struct nvme_rdma_setup_ctx { + struct nvme_rdma_queue *queue; + int *err; +}; + struct nvme_rdma_ctrl { /* read only in the hot path */ struct nvme_rdma_queue *queues; @@ -690,60 +696,68 @@ static int nvme_rdma_start_queue(struct nvme_rdma_ctrl *ctrl, int idx) return ret; } -static int nvme_rdma_start_io_queues(struct nvme_rdma_ctrl *ctrl, - int first, int last) +static void nvme_rdma_setup_queue_async(void *data, async_cookie_t cookie) { - int i, ret = 0; + struct nvme_rdma_setup_ctx *ctx = data; + struct nvme_rdma_queue *queue; + int ret; - for (i = first; i < last; i++) { - ret = nvme_rdma_start_queue(ctrl, i); - if (ret) - goto out_stop_queues; - } + queue = ctx->queue; + ret = nvme_rdma_alloc_queue(queue); + if (ret) + goto out_err; - return 0; + ret = nvme_rdma_start_queue(queue->ctrl, nvme_rdma_queue_idx(queue)); + if (ret) + goto out_err; -out_stop_queues: - for (i--; i >= first; i--) - nvme_rdma_stop_queue(&ctrl->queues[i]); - return ret; + return; +out_err: + WRITE_ONCE(*ctx->err, ret); } -static int nvme_rdma_alloc_io_queues(struct nvme_rdma_ctrl *ctrl) +static int nvme_rdma_setup_io_queues(struct nvme_rdma_ctrl *ctrl, + unsigned int first, unsigned int last, size_t queue_size) { - struct nvmf_ctrl_options *opts = ctrl->ctrl.opts; - unsigned int nr_io_queues; - int i, ret; - - nr_io_queues = nvmf_nr_io_queues(opts); - ret = nvme_set_queue_count(&ctrl->ctrl, &nr_io_queues); - if (ret) - return ret; + ASYNC_DOMAIN_EXCLUSIVE(queue_domain); + struct nvme_rdma_setup_ctx *ctxs; + int nr_queues = last - first; + int err = 0, i, ret; - if (nr_io_queues == 0) { - dev_err(ctrl->ctrl.device, - "unable to set any I/O queues\n"); + ctxs = kmalloc_objs(*ctxs, nr_queues); + if (!ctxs) return -ENOMEM; - } - ctrl->ctrl.queue_count = nr_io_queues + 1; - dev_info(ctrl->ctrl.device, - "creating %d I/O queues.\n", nr_io_queues); - - nvmf_set_io_queues(opts, nr_io_queues, ctrl->io_queues); - for (i = 1; i < ctrl->ctrl.queue_count; i++) { - ctrl->queues[i].ctrl = ctrl; - ctrl->queues[i].queue_size = ctrl->ctrl.sqsize + 1; - ret = nvme_rdma_alloc_queue(&ctrl->queues[i]); - if (ret) - goto out_free_queues; + for (i = 0; i < nr_queues; i++) { + struct nvme_rdma_queue *queue = &ctrl->queues[first + i]; + + queue->ctrl = ctrl; + queue->queue_size = queue_size; + + ctxs[i].queue = queue; + ctxs[i].err = &err; + async_schedule_domain(nvme_rdma_setup_queue_async, &ctxs[i], + &queue_domain); } - return 0; + async_synchronize_full_domain(&queue_domain); + kfree(ctxs); + ret = READ_ONCE(err); + if (ret) + goto out_free_queues; + + return 0; out_free_queues: - for (i--; i >= 1; i--) - nvme_rdma_free_queue(&ctrl->queues[i]); + for (i = 0; i < nr_queues; i++) { + struct nvme_rdma_queue *queue = + &ctrl->queues[first + i]; + + if (test_bit(NVME_RDMA_Q_LIVE, &queue->flags)) + nvme_rdma_stop_queue(queue); + if (test_bit(NVME_RDMA_Q_ALLOCATED, &queue->flags)) + nvme_rdma_free_queue(queue); + } return ret; } @@ -862,12 +876,23 @@ static int nvme_rdma_configure_admin_queue(struct nvme_rdma_ctrl *ctrl, static int nvme_rdma_configure_io_queues(struct nvme_rdma_ctrl *ctrl, bool new) { + unsigned int nr_io_queues; int ret, nr_queues; - ret = nvme_rdma_alloc_io_queues(ctrl); + nr_io_queues = nvmf_nr_io_queues(ctrl->ctrl.opts); + ret = nvme_set_queue_count(&ctrl->ctrl, &nr_io_queues); if (ret) return ret; + if (nr_io_queues == 0) { + dev_err(ctrl->ctrl.device, "unable to set any I/O queues\n"); + return -ENOMEM; + } + + ctrl->ctrl.queue_count = nr_io_queues + 1; + dev_info(ctrl->ctrl.device, "creating %d I/O queues.\n", nr_io_queues); + nvmf_set_io_queues(ctrl->ctrl.opts, nr_io_queues, ctrl->io_queues); + if (new) { ret = nvme_rdma_alloc_tag_set(&ctrl->ctrl); if (ret) @@ -880,7 +905,9 @@ static int nvme_rdma_configure_io_queues(struct nvme_rdma_ctrl *ctrl, bool new) * queue number might have changed. */ nr_queues = min(ctrl->tag_set.nr_hw_queues + 1, ctrl->ctrl.queue_count); - ret = nvme_rdma_start_io_queues(ctrl, 1, nr_queues); + ret = nvme_rdma_setup_io_queues(ctrl, 1, nr_queues, + ctrl->ctrl.sqsize + 1); + if (ret) goto out_cleanup_tagset; @@ -904,12 +931,15 @@ static int nvme_rdma_configure_io_queues(struct nvme_rdma_ctrl *ctrl, bool new) /* * If the number of queues has increased (reconnect case) - * start all new queues now. + * setup all new queues now. */ - ret = nvme_rdma_start_io_queues(ctrl, nr_queues, - ctrl->tag_set.nr_hw_queues + 1); - if (ret) - goto out_wait_freeze_timed_out; + if (ctrl->tag_set.nr_hw_queues + 1 > nr_queues) { + ret = nvme_rdma_setup_io_queues(ctrl, nr_queues, + ctrl->tag_set.nr_hw_queues + 1, + ctrl->ctrl.sqsize + 1); + if (ret) + goto out_wait_freeze_timed_out; + } return 0; -- 2.54.0