From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-dl1-f41.google.com (mail-dl1-f41.google.com [74.125.82.41]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8E61438A700 for ; Fri, 30 Jan 2026 22:36:24 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.82.41 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1769812586; cv=none; b=nT1xyswfbUz4rI/pJPVV78HOHU05l/JBMKrXs+x+w4JFJISh6WgBswaNtequn6Cofl/SbaH39DtFMNdCE4fi4/59HUZUpl8v7SibULjrwrNkCbNjJjSYQhiDFbkhsYadapjqrMWxL5vfKDk7IyDUSMXdR/ePWMz4nVFy1VgJD/0= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1769812586; c=relaxed/simple; bh=WtHtc6ApIDXsfgZIM12a9+PZYY7paUQZIBM0757h5bA=; h=From:To:Cc:Subject:Date:Message-ID:MIME-Version; b=ek+orUFSU1D0gixoNhNiYfRtsGz4uCxIHMjxNv5ZUXkQOnP2ZEmf22XxKZAD9FahXBh2PqYxHPUvsicpof0c0NMjzkMlQ/sI0hSpcAy9bbNTVL9hmoYP7NPokqiTkH3OEbTDUnCtaFv0KCLHC3SknTZVDHACX9T9k4Z8S5ALY2g= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=purestorage.com; spf=fail smtp.mailfrom=purestorage.com; dkim=pass (2048-bit key) header.d=purestorage.com header.i=@purestorage.com header.b=UILdALy+; arc=none smtp.client-ip=74.125.82.41 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=purestorage.com Authentication-Results: smtp.subspace.kernel.org; spf=fail smtp.mailfrom=purestorage.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=purestorage.com header.i=@purestorage.com header.b="UILdALy+" Received: by mail-dl1-f41.google.com with SMTP id a92af1059eb24-124899ee9d3so1955966c88.0 for ; Fri, 30 Jan 2026 14:36:24 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=purestorage.com; s=google2022; t=1769812584; x=1770417384; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:from:to:cc:subject:date:message-id:reply-to; bh=yRCnfJL1AhUwVhtaNTyqMZcDiqAPROP54eRaRo/hfy8=; b=UILdALy+4QhJ5pFD4FmPFTP6ABT+VwnIBN5GX5ylFLCz1++KOxnbpLsxPBg+fO0+F+ 9DoTiKCgu/abptVomZJtg3G+ZetOCYvaXjpfaROOQEUShTwhP6dI7G/KvDzptKDT1GpS 7vU7FWPo+M1TrGzKd7RqXC0oM6HmyJD7AN++I0p+x+Ie/yzddRca84rN7mWPHeHaSlyS D95GuRYT1AJsJEQJr5zfbFiLQGTHwXiXyNyKDULDp9PPV02eP5nPgI14Uj1pp6xOqm/m /KuaW9DGtR/omEDWmtP+/PHvwB9apGGeXaifZgZK2IV7Hbv2S7TwZ72HMyInWxALWihv eF0w== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1769812584; x=1770417384; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:x-gm-gg:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to; bh=yRCnfJL1AhUwVhtaNTyqMZcDiqAPROP54eRaRo/hfy8=; b=nGBxRUsk5YfYvjQHFSn1ppF/6Cb3lX8epY6pcBUS24z55rZ2n5mlkN499dXb5rQ8l6 qVCExn/PB6n5+GOUnch+3ggqQDfKqGoIw2qwwvhWNBU20LjHMYZP/9oLwZb8RLpfg2/j NCK3FTZWEAyFCz56n+RjfFlGk2L7MytRVeXQdpsX5Ny/v1QjX0g9i1HzG0VC98pwK3Cm 1f+Ro41VGyEuhQqMRBBray/O67mN99GohWUDN0xfaEuo+6VR3rEN4Vvcmg69pwAG1gAX zapEUSjGURETqjX4AsvQjh2WBiUP/eUgX2HIENR6PwZFnA78XqqfddoYd79PRpj/4JgC 6VwQ== X-Forwarded-Encrypted: i=1; AJvYcCWuIfcTuZn+lY81INjL23nvgTgHsDOZ5+9IqSdctv2SnDt+s2Gx15CRQwyvY7qwLsZerPIn8FLeUO57DhE=@vger.kernel.org X-Gm-Message-State: AOJu0YxpsPEC87/iP6ncZsQK2g2NerhKaWUy4tMfZBMAjhJLhcLezKvV KbS4+3+LazxTq9maV9Rtu5r6IbSVjgRMxgPiDyJaKTYYbJ9b6lbvSYmhj+jMl2vyhmY= X-Gm-Gg: AZuq6aKkfbYmo3GueO0KwAiX9PwuRl/QW9YFtUZiIGxq+24w3R5GCszR9+UsMSA3iW6 HMo08DetDoaJQYV459amTYwIoG0J4kJYNjs13HG4S4jbjR897H3JygrszIKADCWEch698QNUnDy RiuV+OuCzwi8S2J5bJ7JzOe9pzou3TDbRgJRiUe02Z8uuquqrZpavhBPD/Tv8dLX5AppYHajMUN j8fQGbvtmixDST/RJenELoPE2vpeKsA4+aUmaAZ8gT3kboY/WoBXMPCbemhRWdJlDhPBbAHH3lJ 1jI4DqLcEWD9f08TnIeX1/nknknKJhnGl8dzw1kV/MjL3nhvruLOWNGbJ2ownpvLd0VRdbl+P62 4+Lq8X8BSLMkmd/S78niJZO93fq+5Mce3NFdA/H9cTpdeBAUc1wDvovGlFWS3IzL7evbazMhaCN 6coEt7dBMMzif4df/cKUumLRW08luTnpCDqA== X-Received: by 2002:a05:7022:6886:b0:124:9f56:832e with SMTP id a92af1059eb24-124b1093e80mr4183627c88.18.1769812583554; Fri, 30 Jan 2026 14:36:23 -0800 (PST) Received: from apollo.purestorage.com ([208.88.152.253]) by smtp.googlemail.com with ESMTPSA id a92af1059eb24-124a9d6b906sm13161717c88.4.2026.01.30.14.36.22 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 30 Jan 2026 14:36:23 -0800 (PST) From: Mohamed Khalfella To: Justin Tee , Naresh Gottumukkala , Paul Ely , Chaitanya Kulkarni , Christoph Hellwig , Jens Axboe , Keith Busch , Sagi Grimberg Cc: Aaron Dailey , Randy Jennings , Dhaval Giani , Hannes Reinecke , linux-nvme@lists.infradead.org, linux-kernel@vger.kernel.org, Mohamed Khalfella Subject: [PATCH v2 00/14] TP8028 Rapid Path Failure Recovery Date: Fri, 30 Jan 2026 14:34:04 -0800 Message-ID: <20260130223531.2478849-1-mkhalfella@purestorage.com> X-Mailer: git-send-email 2.52.0 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit This patchset adds support for TP8028 Rapid Path Failure Recovery for both nvme target and initiator. Rapid Path Failure Recovery brings Cross-Controller Reset (CCR) functionality to nvme. This allows nvme host to send an nvme command to source nvme controller to reset impacted nvme controller. Provided that both source and impacted controllers are in the same nvme subsystem. The main use of CCR is when one path to nvme subsystem fails. Inflight IOs on impacted nvme controller need to be terminated first before they can be retried on another path. Otherwise data corruption may happen. CCR provides a quick way to terminate these IOs on the unreachable nvme controller allowing recovery to move quickly and avoiding unnecessary delays. In case of CCR is not possible, then inflight requests are held for duration defined by TP4129 KATO Corrections and Clarifications before they are allowed to be retried. On the target side: * New struct members have been added to support CCR. struct nvme_id_ctrl has been updated with CIU (Controller Instance Uniquifier), CIRN (Controller Instance Random Number), and CQT (Command Quiesce Time). The combination of CIU, CNTLID, and CIRN is used to identify impacted controller in CCR command. * CCR nvme command implemented on the target causes impacted controller to fail and drop connections to host. * CCR logpage contains the status of pending CCR requests. An entry is added to the logpage after CCR request is validated. Completed CCR requests are removed from the logpage when controller becomes ready or when requested in get logpage command. * An AEN is sent when CCR completes to let the host know that it is safe to retry inflight requests. On the host side: * CIU, CIRN, and CQT have been added to struct nvme_ctrl. CIU and CIRN have been added to sysfs to make the values visible to user. CIU and CIRN can be used to construct and manually send admin-passthru CCR commands. * New controller state NVME_CTRL_RECOVERING has been added to prevent cancelling timed out inflight requests while CCR is in progress. Controller flag NVME_CTRL_RECOVERED was also added to signal end of time-based recovery. * Controller recovery in nvme_recover_ctrl() is invoked when LIVE controller hits an error or when a request times out. CCR is attempted to reset impacted controller. * Updated nvme fabric transports nvme-tcp, nvme-rdma, and nvme-fc to use CCR recovery. Ideally all inflight requests should be held during controller recovery and only retried after recovery is done. However, there are known situations that is not the case in this implementation. These gaps will be addressed in future patches: * Manual controller reset from sysfs will result in controller going to RESETTING state and all inflight requests to be canceled immediately and maybe retried on another path. * Manual controller delete from sysfs will also result in all inflight requests to be canceled immediately and maybe retried on another path. * In nvme-fc nvme controller will be deleted if remote port disappears with no timeout specified. This results in immediate cancellation of requests that maybe retried on another path. * In nvme-rdma if HCA is removed all nvme controllers will be deleted. This results in canceling inflight IOs and maybe they will be retred on another path. Changes from v1: * nvmet: Rapid Path Failure Recovery set controller identify fields - Added subsys->cqt defaults to 0 to maintain current behavior. - subsys->cqt is configurable via configfs - Added ctrl->cqt initialized from subsys->cqt. - Renamed ctrl->uniquifier to ctrl->ciu, ctrl->random to ctrl->cirn. * nvmet: Implement CCR nvme command - Refactored nvmet_execute_cross_ctrl_reset() for simpler error handling - Renamed CCR list from ctrl->ccrs to ctrl->ccr_list. * nvmet: Implement CCR logpage - Added CCR status and flags enums * nvme: Rapid Path Failure Recovery read controller identify fields - Renamed ctrl sysfs attributes uniquifier -> ciu, random -> cirn * nvme: Introduce FENCING and FENCED controller states - Added two states (FENCING and FENCED) instead of (RECOVERING and controller flag RECOVERED) - Updated __nvme_check_ready() such that fabric controller in FENCING state is not ready to send requests. Also a request sent while controller in FENCING state is completed with host path error instead of returning BLK_STS_RESOURCE. * nvme: Implement cross-controller reset recovery - Renamed nvme_find_ccr_ctrl() to *nvme_find_ctrl_ccr() to pair with newly added nvme_put_ctrl_ccr(). The later handles releasing source controller used to issue CCR command. - Renamed nvme_recover_ctrl() to nvme_fence_ctrl(). - Deleted nvme_end_ctrl_recovery() because the state change has been moved to nvme_change_ctrl_state(). - Renamed CCR list from ctrl->ccrs to ctrl->ccr_list. * nvme-tcp: Use CCR to recover controller that hits an error - Added ctrl->fencing_work and ctrl->fenced_work instead of changing ctrl->err_work and using it for fencing purpose. * nvme-rdma: Use CCR to recover controller that hits an error - Similar change to nvme-tcp. * nvme-fc: Use CCR to recover controller that hits an error - Similar to nvme-rdma and nvme-tcp. * nvme-fc: Hold inflight requests while in RECOVERING state - Updated nvme_fc_fcpio_done() to hold the first request that starts error recovery. That was one of the limitations mentioned in the cover letter of v1. v1: https://lore.kernel.org/all/20251126021250.2583630-1-mkhalfella@purestorage.com/ Mohamed Khalfella (14): nvmet: Rapid Path Failure Recovery set controller identify fields nvmet/debugfs: Add ctrl uniquifier and random values nvmet: Implement CCR nvme command nvmet: Implement CCR logpage nvmet: Send an AEN on CCR completion nvme: Rapid Path Failure Recovery read controller identify fields nvme: Introduce FENCING and FENCED controller states nvme: Implement cross-controller reset recovery nvme: Implement cross-controller reset completion nvme-tcp: Use CCR to recover controller that hits an error nvme-rdma: Use CCR to recover controller that hits an error nvme-fc: Decouple error recovery from controller reset nvme-fc: Use CCR to recover controller that hits an error nvme-fc: Hold inflight requests while in FENCING state drivers/nvme/host/constants.c | 1 + drivers/nvme/host/core.c | 208 +++++++++++++++++++++++++++++- drivers/nvme/host/fc.c | 215 ++++++++++++++++++++++---------- drivers/nvme/host/nvme.h | 25 ++++ drivers/nvme/host/rdma.c | 62 ++++++++- drivers/nvme/host/sysfs.c | 25 ++++ drivers/nvme/host/tcp.c | 62 ++++++++- drivers/nvme/target/admin-cmd.c | 124 ++++++++++++++++++ drivers/nvme/target/configfs.c | 31 +++++ drivers/nvme/target/core.c | 108 +++++++++++++++- drivers/nvme/target/debugfs.c | 21 ++++ drivers/nvme/target/nvmet.h | 20 ++- include/linux/nvme.h | 70 ++++++++++- 13 files changed, 897 insertions(+), 75 deletions(-) base-commit: 8dfce8991b95d8625d0a1d2896e42f93b9d7f68d -- 2.52.0