From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-ed1-f49.google.com (mail-ed1-f49.google.com [209.85.208.49]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 71CD330FF30 for ; Sat, 25 Jul 2026 13:51:24 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.208.49 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784987486; cv=none; b=AhNe1mmeDfsTSfP6Kqx0rSfYa/YUOXY7qv33pnsJd67rSDcscY7hIsdlehVokxjb4Lvo60HoaDiPUI0R4w8b+ft6tvrLtsWgjr73Fi9WBgIRVbGCJXBXLW3lROz657CUX8MFyVtlAxCugUGOqrfAaAE/ZgwYAEpFaqNIprnskvw= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784987486; c=relaxed/simple; bh=oNS4a3LbPNWdItRtx0eJfsJx8XTdUjD/bcdzSBec0Ck=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=mTKMU0g+1rbI2q5MdGIbmgrKgINwQQBLUwPKw3D7LOodImPKb5JoST9U7xffRIg8+g8UcllBrqbbz5+L/3IVyrYs5fsbvP+Kuo7EMIX4y1p2Ijl+Ln/01MjT0r4qmLUdnSCSwBNqyi9nL8XI94n9ynFjibJ4uqZzIJjD5SDA5/g= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=hOFkLCBo; arc=none smtp.client-ip=209.85.208.49 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="hOFkLCBo" Received: by mail-ed1-f49.google.com with SMTP id 4fb4d7f45d1cf-69fc9f25118so185602a12.2 for ; Sat, 25 Jul 2026 06:51:24 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784987482; x=1785592282; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=BNlCqWN9rYlM7R3wxguszrDK0t4TEHNES9xZZYsPYhg=; b=hOFkLCBobby4y1cZZwaz79Gxmgn78DDBhia6S337+z0GSg3XSwd1UnO71o0JnCznie ngQvtPFX9wuPKesFiS31En1WFUmaPJor2YWXiQNMiYnVgTYmcC3a/vGdTMj5EhGl2okW mafD3XnwusG7aBjbR0oTr89bC0krsOtOGt9zd3A6qS9X86BPo4YkOq3dv+02sVxjI/yo kXFtBR5kxEgYJAtzrFld3sO5oNPFjldGG/+jygOUHJcDuQUwobqLIYAtoGhFaFLKTnmi SIr1vXG4fmHyf5dC82p0ykc+Uc3chsU1y7aXNKOX8lW12wzqAA1BgxeGwDULJNSv9kXy +b7g== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784987482; x=1785592282; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=BNlCqWN9rYlM7R3wxguszrDK0t4TEHNES9xZZYsPYhg=; b=rFiNiEfU+K+y/dJ4Bbo83n8/dQGNGtz8c0kUBXgePjfuj6LPIJy7jQfDIoI3olZhPf 1A0+5DW6BWeZ2z735YdNKAn4je70QBmU3srvwhtllqDEx7+5ScMVyjClV90mvbO0ApqQ ffmwZMH5UfEV7zz7h2F+m35CVs6b8dbbZ+kihAoo1ymkhHsy8DU1Hci2uCRyE0DqtfKG tiDC2CP4jbNIwNBivCADAdNu6YrYfTndDukw0+pmQ7NeZ+zs5DdNpGohQMoWDYPoKx5l 7bO4DMK6B4CzEbxekY2DetfN6iG4pJJ8T5sYBIC+3kDu4amKaHqMXGytUVGayFA+zSmd hsZw== X-Forwarded-Encrypted: i=1; AHgh+RqFX1/Lrfl07vM7MK2bRgY8+yUzTnBQCKnNQYdp+H2L8eJ91UjmV6JL0Nv7CaQwzJUwyYNYOlofDtra9Z4=@vger.kernel.org X-Gm-Message-State: AOJu0Yw9SrYG88uzbGtO2axdvsCwbitEmug4HkPQ51siYRUB+mRDlDEM SLSCbU0RYR3f9slRhVereEmHqsHDJkKzaEkXDe+VQ3Qdv0XvnXCIt3nF X-Gm-Gg: AR+sD10SWYEaq5hF6qIwwpEaXPMlfFM6z2tcGgqoTqZBdoC7m6k9Az+ncLPZuL8HS0+ WbbHtwoY5QjXJ2eT5qaHs+aP2vGiRDtQEnrI3p8VDggxQF4ycdJZTcR6iOXP8wzstADOvHYVMRe F1+RR1KUbQJgqY+CZt/M6FsYr7KKtW08IOUAu4D3ofALya9/9kmhM96tnegmUWPbdLdY0Eu1Ght QGsT0lB/McSdoeL+u+CWhdMrnGlqtcdQA1wct6V/Zbw57So9buxJBYDbh8wPJQAlQrwiHsuLctO tmUI1aZRfJo4tdMhY4tacawztzytKXxTXZ4gDLdClR6rHEerhamU+RBYq/O08ehY7WrrmL3co3/ QpPkY+EL+SvqaMPD5gYXz2TSpETudU1RidTX5g+WaRNywiJsD/3KMx3CSqxRLZYjDZc9pwh0gX1 Tlf3YWdbTF9/jHUJ2ie8ApttdbgEyfpK14k8iQN5bF2hsEq6qDzQeXapVgZqRXEKs= X-Received: by 2002:a05:6402:360f:b0:69f:aca0:3b9f with SMTP id 4fb4d7f45d1cf-69fc1009d18mr832970a12.15.1784987482213; Sat, 25 Jul 2026 06:51:22 -0700 (PDT) Received: from misharu.home (2a02-a463-a071-0-7475-eec2-a79e-2db.fixed6.kpn.net. [2a02:a463:a071:0:7475:eec2:a79e:2db]) by smtp.gmail.com with ESMTPSA id 4fb4d7f45d1cf-69fb5543aacsm836431a12.14.2026.07.25.06.51.20 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sat, 25 Jul 2026 06:51:20 -0700 (PDT) From: Hari Mishal To: Keith Busch , Jens Axboe , Christoph Hellwig , Sagi Grimberg Cc: Hannes Reinecke , Kanchan Joshi , Nitesh Shetty , Greg Kroah-Hartman , linux-nvme@lists.infradead.org, linux-kernel@vger.kernel.org, Hari Mishal Subject: [PATCH 1/2] nvme: fix racy access to FDP placement ID array Date: Sat, 25 Jul 2026 15:51:10 +0200 Message-ID: <20260725135111.14041-2-harimishal1@gmail.com> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20260725135111.14041-1-harimishal1@gmail.com> References: <20260725135111.14041-1-harimishal1@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit nvme_query_fdp_info() populates head->nr_plids and head->plids the first time a namespace's FDP configuration is registered, guarded only by a check-then-act "if (head->nr_plids) return 0" with no locking. Since a namespace's nvme_ns_head can be shared across multiple nvme_ns paths, two controller paths scanning the same namespace at the same time can race to populate this pair concurrently: - Two unsynchronized writers can each set nr_plids/plids independently, so the last writer of each field can differ, producing a count that doesn't match the actual size of the published array. - A concurrent reader in nvme_setup_rw() or nvme_update_ns_info_block() can observe a non-zero nr_plids while plids is still NULL, or sized for a different count, leading to a NULL dereference or an out-of-bounds read of ns->head->plids[]. Add a spinlock to nvme_ns_head and take it around every access to nr_plids/plids, both the writer in nvme_query_fdp_info() and the readers, so the pair is always observed and updated as a single consistent unit. Use scoped_guard() so the lock covers the entire read-and-use in both readers rather than being released before the values are actually used. Signed-off-by: Hari Mishal --- drivers/nvme/host/core.c | 63 ++++++++++++++++++++++++++-------------- drivers/nvme/host/nvme.h | 1 + 2 files changed, 43 insertions(+), 21 deletions(-) diff --git a/drivers/nvme/host/core.c b/drivers/nvme/host/core.c index 453c1f0b2dd0..bdc5f07f5bf0 100644 --- a/drivers/nvme/host/core.c +++ b/drivers/nvme/host/core.c @@ -1018,15 +1018,19 @@ static inline blk_status_t nvme_setup_rw(struct nvme_ns *ns, if (req->cmd_flags & REQ_RAHEAD) dsmgmt |= NVME_RW_DSM_FREQ_PREFETCH; - if (op == nvme_cmd_write && ns->head->nr_plids) { - u16 write_stream = req->bio->bi_write_stream; - - if (WARN_ON_ONCE(write_stream > ns->head->nr_plids)) - return BLK_STS_INVAL; - - if (write_stream) { - dsmgmt |= ns->head->plids[write_stream - 1] << 16; - control |= NVME_RW_DTYPE_DPLCMT; + if (op == nvme_cmd_write) { + scoped_guard(spinlock, &ns->head->fdp_lock) { + u16 write_stream = req->bio->bi_write_stream; + + if (ns->head->nr_plids) { + if (WARN_ON_ONCE(write_stream > ns->head->nr_plids)) + return BLK_STS_INVAL; + + if (write_stream) { + dsmgmt |= ns->head->plids[write_stream - 1] << 16; + control |= NVME_RW_DTYPE_DPLCMT; + } + } } } @@ -2317,6 +2321,8 @@ static int nvme_query_fdp_info(struct nvme_ns *ns, struct nvme_ns_info *info) struct nvme_fdp_ruh_status *ruhs; struct nvme_fdp_config fdp; struct nvme_command c = {}; + u16 nr_plids; + u16 *plids; size_t size; int i, ret; @@ -2325,8 +2331,10 @@ static int nvme_query_fdp_info(struct nvme_ns *ns, struct nvme_ns_info *info) * so return immediately if we've already registered this namespace's * streams. */ - if (head->nr_plids) - return 0; + scoped_guard(spinlock, &head->fdp_lock) { + if (head->nr_plids) + return 0; + } ret = nvme_get_features(ctrl, NVME_FEAT_FDP, info->endgid, NULL, 0, &fdp); @@ -2357,23 +2365,34 @@ static int nvme_query_fdp_info(struct nvme_ns *ns, struct nvme_ns_info *info) goto free; } - head->nr_plids = le16_to_cpu(ruhs->nruhsd); - if (!head->nr_plids) + nr_plids = le16_to_cpu(ruhs->nruhsd); + if (!nr_plids) goto free; - head->plids = kcalloc(head->nr_plids, sizeof(*head->plids), - GFP_KERNEL); - if (!head->plids) { + plids = kcalloc(nr_plids, sizeof(*plids), GFP_KERNEL); + if (!plids) { dev_warn(ctrl->device, "failed to allocate %u FDP placement IDs\n", - head->nr_plids); - head->nr_plids = 0; + nr_plids); ret = -ENOMEM; goto free; } - for (i = 0; i < head->nr_plids; i++) - head->plids[i] = le16_to_cpu(ruhs->ruhsd[i].pid); + for (i = 0; i < nr_plids; i++) + plids[i] = le16_to_cpu(ruhs->ruhsd[i].pid); + + /* + * Publish the fully-populated array; if another path already won + * the race, drop our redundant copy. + */ + scoped_guard(spinlock, &head->fdp_lock) { + if (head->nr_plids) { + kfree(plids); + goto free; + } + head->plids = plids; + head->nr_plids = nr_plids; + } free: kfree(ruhs); return ret; @@ -2468,7 +2487,8 @@ static int nvme_update_ns_info_block(struct nvme_ns *ns, if (!nvme_init_integrity(ns->head, &lim, info)) capacity = 0; - lim.max_write_streams = ns->head->nr_plids; + scoped_guard(spinlock, &ns->head->fdp_lock) + lim.max_write_streams = ns->head->nr_plids; if (lim.max_write_streams) lim.write_stream_granularity = min(info->runs, U32_MAX); else @@ -3991,6 +4011,7 @@ static struct nvme_ns_head *nvme_alloc_ns_head(struct nvme_ctrl *ctrl, head->ids = info->ids; head->shared = info->is_shared; head->rotational = info->is_rotational; + spin_lock_init(&head->fdp_lock); ratelimit_state_init(&head->rs_nuse, 5 * HZ, 1); ratelimit_set_flags(&head->rs_nuse, RATELIMIT_MSG_ON_RELEASE); kref_init(&head->ref); diff --git a/drivers/nvme/host/nvme.h b/drivers/nvme/host/nvme.h index 824651cc898d..22a68e09b065 100644 --- a/drivers/nvme/host/nvme.h +++ b/drivers/nvme/host/nvme.h @@ -560,6 +560,7 @@ struct nvme_ns_head { u16 nr_plids; u16 *plids; + spinlock_t fdp_lock; /* protects nr_plids and plids */ #ifdef CONFIG_NVME_MULTIPATH struct bio_list requeue_list; spinlock_t requeue_lock; -- 2.43.0