* [PATCH v4 0/2] nvme-multipath: fail I/O when no usable path exists
@ 2026-10-01 9:48 Krishna Iyer
2026-10-01 9:48 ` [PATCH v4 1/2] nvme-multipath: fix path state evaluation for failfast and ANA Krishna Iyer
2026-10-01 9:48 ` [PATCH v4 2/2] nvme-multipath: add fail_if_no_path sysfs attribute Krishna Iyer
0 siblings, 2 replies; 9+ messages in thread
From: Krishna Iyer @ 2026-10-01 9:48 UTC (permalink / raw)
To: kbusch, axboe, hch, sagi
Cc: linux-nvme, linux-kernel, nilay, hare, saravanand, sjpark, Krishna Iyer
On virtualization hosts we run NVMe/TCP targets with ctrl_loss_tmo=-1.
During a long fabric outage, I/O on a multipath namespace is queued
until a path returns, so a SIGKILLed VM process cannot exit while it
drains I/O to an unreachable target and is stuck in D state.
Patch 1 fixes nvme_available_path() so it stops queueing I/O that should
be failed, by default and without any flag:
- respect fast_io_fail_tmo: when a path exists but its failfast timer
has expired, fail instead of falling through to queue_if_no_path
- a LIVE path whose ANA state is inaccessible or persistent-loss no
longer counts as available (ANA change stays transient and queues)
- kick the requeue work from nvme_update_ns_ana_state() so parked I/O
is re-evaluated when an ANA transition leaves a namespace inaccessible
Patch 2 adds the fail_if_no_path attribute, an opt-in per-namespace
policy that fails parked and newly arriving I/O instead of queueing it
when no usable path exists. It is the namespace-scoped counterpart to
the controller-scoped fast_io_fail_tmo, needed because sibling
namespaces behind the same controllers must keep queueing while one
namespace whose consumer is gone releases its parked I/O. It sits
alongside delayed_removal_secs as a per-namespace policy in
nvme-multipath sysfs. Controller state is untouched and reconnects
continue.
This answers the question from the v3 review. The intent is exactly per
namespace differentiation, some namespaces stop waiting while others
keep waiting for the same controllers to reconnect. Patch 1 fails the
conditions the kernel can detect, an expired failfast timer or an
inaccessible ANA state. It cannot detect that the consumer of a
namespace has exited while its paths are still connecting, so that case
still queues by default and only an explicit per namespace opt in can
release it.
Changes since v3 [3]:
- Split the single v3 patch into a two-patch series: the
default-behavior fixes first, the opt-in attribute on top (Hannes,
Keith).
- The ANA inaccessible and persistent-loss handling is now the default
and no longer gated behind the flag (Keith).
- New: respect fast_io_fail_tmo. When a path is present but its failfast
timer has expired, nvme_available_path() now fails instead of falling
through to the queue_if_no_path window (Keith).
- New: kick the requeue work from nvme_update_ns_ana_state() so parked
I/O is re-evaluated when an ANA transition leaves a namespace
inaccessible (Keith).
- Patch 2 now carries only the fail_if_no_path attribute, rebased on
patch 1's corrected defaults: it gates the CONNECTING case and
overrides delayed_removal_secs, taking precedence over
queue_if_no_path when both are set.
- Document the fail_if_no_path vs delayed_removal_secs (queue-if-no-path)
interaction in the ABI entry (Hannes).
Changes since v2 [2]:
- Use the nvme_state_is_live() helper for the ANA state check, keeping
the explicit NVME_ANA_CHANGE carve-out so transient ANA transitions
still queue (Nilay).
- Return early when the stored value matches the current setting,
skipping synchronize_srcu() and the requeue kick (Nilay).
Changes since v1 [1]:
- Rename the attribute from fail_io_now to fail_if_no_path (Nilay).
- Make the policy persistent instead of self-clearing. Drop the clear
in nvme_mpath_set_live() so it stays set until userspace clears it.
- Enforce it inside nvme_available_path() by controller state. A
CONNECTING controller and a LIVE controller whose path ANA state is
inaccessible or persistent-loss stop counting as available. RESETTING
and ANA change keep queueing. With no controllers left it overrides
the delayed_removal_secs window.
Validated on hardware with a 6.17 backport of this series: parked I/O on
a SIGKILLed VM process failed within a second of enabling the policy and
the process was reaped, new I/O failed fast during the outage, the
policy persisted across path recovery, and disabling it restored
queueing.
[1] v1: https://lore.kernel.org/linux-nvme/20260904032605.65758-1-kiyer@crusoe.ai/
[2] v2: https://lore.kernel.org/linux-nvme/20260917231647.79956-1-kiyer@crusoe.ai/
[3] v3: https://lore.kernel.org/linux-nvme/20260923004959.88440-1-kiyer@crusoe.ai/
Krishna Iyer (2):
nvme-multipath: fix path state evaluation for failfast and ANA
nvme-multipath: add fail_if_no_path sysfs attribute
Documentation/ABI/stable/sysfs-nvme | 16 +++++++
drivers/nvme/host/multipath.c | 74 +++++++++++++++++++++++++----
drivers/nvme/host/nvme.h | 2 +
drivers/nvme/host/sysfs.c | 4 +-
4 files changed, 87 insertions(+), 9 deletions(-)
base-commit: 2ee54f01f07c0307deaf90ca8691a4643ae0357b
--
2.54.0
^ permalink raw reply [flat|nested] 9+ messages in thread
* [PATCH v4 1/2] nvme-multipath: fix path state evaluation for failfast and ANA
2026-10-01 9:48 [PATCH v4 0/2] nvme-multipath: fail I/O when no usable path exists Krishna Iyer
@ 2026-10-01 9:48 ` Krishna Iyer
2026-10-02 9:36 ` Hannes Reinecke
2026-10-02 11:53 ` Nilay Shroff
2026-10-01 9:48 ` [PATCH v4 2/2] nvme-multipath: add fail_if_no_path sysfs attribute Krishna Iyer
1 sibling, 2 replies; 9+ messages in thread
From: Krishna Iyer @ 2026-10-01 9:48 UTC (permalink / raw)
To: kbusch, axboe, hch, sagi
Cc: linux-nvme, linux-kernel, nilay, hare, saravanand, sjpark, Krishna Iyer
nvme_available_path() has two problems that keep I/O queued when it
should be failed:
1. When fast_io_fail_tmo expires, NVME_CTRL_FAILFAST_EXPIRED is set and
the path is skipped, but the function still falls through to
nvme_mpath_queue_if_no_path(). If delayed_removal_secs is configured
the I/O is requeued indefinitely, defeating fast_io_fail_tmo. Only
fall through to queue_if_no_path when there really are no paths;
if a path exists but its failfast timer has expired, fail instead.
2. A LIVE controller counts as a usable path regardless of the namespace
ANA state. A path whose ANA state is inaccessible or persistent-loss
cannot serve I/O, so it should not count as available. ANA change is
transient and bounded by ANATT, so keep queueing while it resolves.
Also kick the requeue work from nvme_update_ns_ana_state() when a path
does not transition to live, so parked I/O is re-evaluated when an ANA
transition leaves the namespace inaccessible.
Move nvme_state_is_live() above nvme_available_path() so it can be used
there.
Assisted-by: Claude:claude-fable-5
Signed-off-by: Krishna Iyer <kiyer@crusoe.ai>
---
drivers/nvme/host/multipath.c | 26 +++++++++++++++++++-------
1 file changed, 19 insertions(+), 7 deletions(-)
diff --git a/drivers/nvme/host/multipath.c b/drivers/nvme/host/multipath.c
index 11871f5f18c2..0d6c05f0808b 100644
--- a/drivers/nvme/host/multipath.c
+++ b/drivers/nvme/host/multipath.c
@@ -509,9 +509,15 @@ inline struct nvme_ns *nvme_find_path(struct nvme_ns_head *head)
}
}
+static inline bool nvme_state_is_live(enum nvme_ana_state state)
+{
+ return state == NVME_ANA_OPTIMIZED || state == NVME_ANA_NONOPTIMIZED;
+}
+
static bool nvme_available_path(struct nvme_ns_head *head)
__must_hold_shared(&head->srcu)
{
+ bool failfast = false;
struct nvme_ns *ns;
if (!test_bit(NVME_NSHEAD_DISK_LIVE, &head->flags))
@@ -519,10 +525,16 @@ static bool nvme_available_path(struct nvme_ns_head *head)
list_for_each_entry_srcu(ns, &head->list, siblings,
srcu_read_lock_held(&head->srcu)) {
- if (test_bit(NVME_CTRL_FAILFAST_EXPIRED, &ns->ctrl->flags))
+ if (test_bit(NVME_CTRL_FAILFAST_EXPIRED, &ns->ctrl->flags)) {
+ failfast = true;
continue;
+ }
switch (nvme_ctrl_state(ns->ctrl)) {
case NVME_CTRL_LIVE:
+ if (!nvme_state_is_live(ns->ana_state) &&
+ ns->ana_state != NVME_ANA_CHANGE)
+ continue;
+ return true;
case NVME_CTRL_RESETTING:
case NVME_CTRL_CONNECTING:
return true;
@@ -531,6 +543,9 @@ static bool nvme_available_path(struct nvme_ns_head *head)
}
}
+ if (failfast)
+ return false;
+
/*
* If "head->delayed_removal_secs" is configured (i.e., non-zero), do
* not immediately fail I/O. Instead, requeue the I/O for the configured
@@ -887,11 +902,6 @@ static int nvme_parse_ana_log(struct nvme_ctrl *ctrl, void *data,
return 0;
}
-static inline bool nvme_state_is_live(enum nvme_ana_state state)
-{
- return state == NVME_ANA_OPTIMIZED || state == NVME_ANA_NONOPTIMIZED;
-}
-
static void nvme_update_ns_ana_state(struct nvme_ana_group_desc *desc,
struct nvme_ns *ns)
{
@@ -926,8 +936,10 @@ static void nvme_update_ns_ana_state(struct nvme_ana_group_desc *desc,
* is not live but still create the sysfs link to this path from
* head node if head node of the path has already come alive.
*/
- if (test_bit(NVME_NSHEAD_DISK_LIVE, &ns->head->flags))
+ if (test_bit(NVME_NSHEAD_DISK_LIVE, &ns->head->flags)) {
nvme_mpath_add_sysfs_link(ns->head);
+ kblockd_schedule_work(&ns->head->requeue_work);
+ }
}
}
--
2.54.0
^ permalink raw reply [flat|nested] 9+ messages in thread
* [PATCH v4 2/2] nvme-multipath: add fail_if_no_path sysfs attribute
2026-10-01 9:48 [PATCH v4 0/2] nvme-multipath: fail I/O when no usable path exists Krishna Iyer
2026-10-01 9:48 ` [PATCH v4 1/2] nvme-multipath: fix path state evaluation for failfast and ANA Krishna Iyer
@ 2026-10-01 9:48 ` Krishna Iyer
2026-10-02 9:39 ` Hannes Reinecke
1 sibling, 1 reply; 9+ messages in thread
From: Krishna Iyer @ 2026-10-01 9:48 UTC (permalink / raw)
To: kbusch, axboe, hch, sagi
Cc: linux-nvme, linux-kernel, nilay, hare, saravanand, sjpark, Krishna Iyer
When no usable path exists, I/O on a multipath namespace is queued
until a path returns. With ctrl_loss_tmo=-1 that can be forever:
during a long fabric outage any process waiting on the I/O is stuck in
D state. We hit this on virtualization hosts, where a SIGKILLed VM
process cannot exit while draining I/O to an unreachable NVMe/TCP
target.
Nothing can fail this I/O without tearing something down: controller
deletion takes every namespace on the controller with it.
Add a fail_if_no_path attribute on the ns-head disk: a persistent
per-namespace policy to fail parked and newly arriving I/O instead of
queueing it when no usable path exists. It is the namespace-scoped
counterpart to the controller-scoped fast_io_fail_tmo: the trigger is
an event userspace observes (a consumer known to be gone, for us a
SIGKILLed VM the host must reap), not a duration picked up front, and
sibling namespaces behind the same controllers keep queueing and ride
out the outage.
A CONNECTING controller stops counting as an available path, and with
no controllers left the policy overrides the delayed_removal_secs
queue-if-no-path window; the two are opposites, so fail_if_no_path takes
precedence when both are set. ANA change and controller resetting still
queue, since both are transient. Controller state is untouched and
reconnects continue. Like dm's fail_if_no_path, the policy is transport
agnostic.
Assisted-by: Claude:claude-fable-5
Signed-off-by: Krishna Iyer <kiyer@crusoe.ai>
---
Documentation/ABI/stable/sysfs-nvme | 16 +++++++++
drivers/nvme/host/multipath.c | 50 +++++++++++++++++++++++++++--
drivers/nvme/host/nvme.h | 2 ++
drivers/nvme/host/sysfs.c | 4 ++-
4 files changed, 69 insertions(+), 3 deletions(-)
diff --git a/Documentation/ABI/stable/sysfs-nvme b/Documentation/ABI/stable/sysfs-nvme
index a0bb88ca1694..448692dcb117 100644
--- a/Documentation/ABI/stable/sysfs-nvme
+++ b/Documentation/ABI/stable/sysfs-nvme
@@ -358,6 +358,22 @@ Description:
is deferred. Only visible on multipath head devices.
Requires CONFIG_NVME_MULTIPATH.
+What: /sys/block/nvmeXnY/fail_if_no_path
+Date: October 2026
+KernelVersion: 7.4
+Contact: Krishna Iyer <kiyer@crusoe.ai>
+Description:
+ Shows or sets the fail-if-no-path policy of the multipath
+ head device ("on" or "off", default "off"). When on, I/O
+ queued or arriving while no usable path exists is failed
+ immediately instead of being queued. This is the opposite of
+ the queue-if-no-path behavior enabled by a nonzero
+ delayed_removal_secs: when both are set fail_if_no_path takes
+ precedence and the parked I/O is failed rather than held for
+ the delayed_removal_secs window. Reconnect attempts are not
+ affected. Only visible on multipath head devices.
+ Requires CONFIG_NVME_MULTIPATH.
+
What: /sys/block/nvmeXnY/csi
What: /sys/block/nvmeXnY/metadata_bytes
What: /sys/block/nvmeXnY/nuse
diff --git a/drivers/nvme/host/multipath.c b/drivers/nvme/host/multipath.c
index 0d6c05f0808b..e43b0f43e73e 100644
--- a/drivers/nvme/host/multipath.c
+++ b/drivers/nvme/host/multipath.c
@@ -517,6 +517,8 @@ static inline bool nvme_state_is_live(enum nvme_ana_state state)
static bool nvme_available_path(struct nvme_ns_head *head)
__must_hold_shared(&head->srcu)
{
+ bool fail_if_no_path = test_bit(NVME_NSHEAD_FAIL_IF_NO_PATH,
+ &head->flags);
bool failfast = false;
struct nvme_ns *ns;
@@ -536,14 +538,17 @@ static bool nvme_available_path(struct nvme_ns_head *head)
continue;
return true;
case NVME_CTRL_RESETTING:
- case NVME_CTRL_CONNECTING:
return true;
+ case NVME_CTRL_CONNECTING:
+ if (!fail_if_no_path)
+ return true;
+ continue;
default:
break;
}
}
- if (failfast)
+ if (failfast || fail_if_no_path)
return false;
/*
@@ -1242,6 +1247,47 @@ static ssize_t delayed_removal_secs_store(struct device *dev,
DEVICE_ATTR_RW(delayed_removal_secs);
+static ssize_t fail_if_no_path_show(struct device *dev,
+ struct device_attribute *attr, char *buf)
+{
+ struct gendisk *disk = dev_to_disk(dev);
+ struct nvme_ns_head *head = disk->private_data;
+
+ return sysfs_emit(buf, test_bit(NVME_NSHEAD_FAIL_IF_NO_PATH,
+ &head->flags) ? "on\n" : "off\n");
+}
+
+static ssize_t fail_if_no_path_store(struct device *dev,
+ struct device_attribute *attr, const char *buf, size_t count)
+{
+ struct gendisk *disk = dev_to_disk(dev);
+ struct nvme_ns_head *head = disk->private_data;
+ bool enable;
+ int ret;
+
+ ret = kstrtobool(buf, &enable);
+ if (ret < 0)
+ return ret;
+
+ if (enable) {
+ if (test_and_set_bit(NVME_NSHEAD_FAIL_IF_NO_PATH, &head->flags))
+ return count;
+ } else {
+ if (!test_and_clear_bit(NVME_NSHEAD_FAIL_IF_NO_PATH,
+ &head->flags))
+ return count;
+ }
+
+ synchronize_srcu(&head->srcu);
+
+ if (enable)
+ kblockd_schedule_work(&head->requeue_work);
+
+ return count;
+}
+
+DEVICE_ATTR_RW(fail_if_no_path);
+
static ssize_t multipath_failover_count_show(struct device *dev,
struct device_attribute *attr, char *buf)
{
diff --git a/drivers/nvme/host/nvme.h b/drivers/nvme/host/nvme.h
index bac25b287d25..71ec72306c9c 100644
--- a/drivers/nvme/host/nvme.h
+++ b/drivers/nvme/host/nvme.h
@@ -589,6 +589,7 @@ struct nvme_ns_head {
#define NVME_NSHEAD_DISK_LIVE 0
#define NVME_NSHEAD_QUEUE_IF_NO_PATH 1
#define NVME_NSHEAD_CDEV_LIVE 2
+#define NVME_NSHEAD_FAIL_IF_NO_PATH 3
struct nvme_ns __rcu_guarded *current_path[];
#endif
};
@@ -1097,6 +1098,7 @@ extern struct device_attribute dev_attr_queue_depth;
extern struct device_attribute dev_attr_numa_nodes;
extern struct device_attribute dev_attr_path_state;
extern struct device_attribute dev_attr_delayed_removal_secs;
+extern struct device_attribute dev_attr_fail_if_no_path;
extern struct device_attribute dev_attr_multipath_failover_count;
extern struct device_attribute dev_attr_io_requeue_no_usable_path_count;
extern struct device_attribute dev_attr_io_fail_no_available_path_count;
diff --git a/drivers/nvme/host/sysfs.c b/drivers/nvme/host/sysfs.c
index 4ac3ea3a86fd..fe1e5305880e 100644
--- a/drivers/nvme/host/sysfs.c
+++ b/drivers/nvme/host/sysfs.c
@@ -265,6 +265,7 @@ static struct attribute *nvme_ns_attrs[] = {
&dev_attr_numa_nodes.attr,
&dev_attr_path_state.attr,
&dev_attr_delayed_removal_secs.attr,
+ &dev_attr_fail_if_no_path.attr,
#endif
&dev_attr_io_passthru_err_log_enabled.attr,
NULL,
@@ -302,7 +303,8 @@ static umode_t nvme_ns_attrs_are_visible(struct kobject *kobj,
if (nvme_disk_is_ns_head(dev_to_disk(dev)))
return 0;
}
- if (a == &dev_attr_delayed_removal_secs.attr) {
+ if (a == &dev_attr_delayed_removal_secs.attr ||
+ a == &dev_attr_fail_if_no_path.attr) {
struct gendisk *disk = dev_to_disk(dev);
if (!nvme_disk_is_ns_head(disk))
--
2.54.0
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [PATCH v4 1/2] nvme-multipath: fix path state evaluation for failfast and ANA
2026-10-01 9:48 ` [PATCH v4 1/2] nvme-multipath: fix path state evaluation for failfast and ANA Krishna Iyer
@ 2026-10-02 9:36 ` Hannes Reinecke
2026-10-02 11:53 ` Nilay Shroff
1 sibling, 0 replies; 9+ messages in thread
From: Hannes Reinecke @ 2026-10-02 9:36 UTC (permalink / raw)
To: Krishna Iyer, kbusch, axboe, hch, sagi
Cc: linux-nvme, linux-kernel, nilay, saravanand, sjpark
On 10/1/26 11:48 AM, Krishna Iyer wrote:
> nvme_available_path() has two problems that keep I/O queued when it
> should be failed:
>
> 1. When fast_io_fail_tmo expires, NVME_CTRL_FAILFAST_EXPIRED is set and
> the path is skipped, but the function still falls through to
> nvme_mpath_queue_if_no_path(). If delayed_removal_secs is configured
> the I/O is requeued indefinitely, defeating fast_io_fail_tmo. Only
> fall through to queue_if_no_path when there really are no paths;
> if a path exists but its failfast timer has expired, fail instead.
>
> 2. A LIVE controller counts as a usable path regardless of the namespace
> ANA state. A path whose ANA state is inaccessible or persistent-loss
> cannot serve I/O, so it should not count as available. ANA change is
> transient and bounded by ANATT, so keep queueing while it resolves.
>
> Also kick the requeue work from nvme_update_ns_ana_state() when a path
> does not transition to live, so parked I/O is re-evaluated when an ANA
> transition leaves the namespace inaccessible.
>
> Move nvme_state_is_live() above nvme_available_path() so it can be used
> there.
>
> Assisted-by: Claude:claude-fable-5
> Signed-off-by: Krishna Iyer <kiyer@crusoe.ai>
> ---
> drivers/nvme/host/multipath.c | 26 +++++++++++++++++++-------
> 1 file changed, 19 insertions(+), 7 deletions(-)
>
> diff --git a/drivers/nvme/host/multipath.c b/drivers/nvme/host/multipath.c
> index 11871f5f18c2..0d6c05f0808b 100644
> --- a/drivers/nvme/host/multipath.c
> +++ b/drivers/nvme/host/multipath.c
> @@ -509,9 +509,15 @@ inline struct nvme_ns *nvme_find_path(struct nvme_ns_head *head)
> }
> }
>
> +static inline bool nvme_state_is_live(enum nvme_ana_state state)
> +{
> + return state == NVME_ANA_OPTIMIZED || state == NVME_ANA_NONOPTIMIZED;
> +}
> +
> static bool nvme_available_path(struct nvme_ns_head *head)
> __must_hold_shared(&head->srcu)
> {
> + bool failfast = false;
> struct nvme_ns *ns;
>
> if (!test_bit(NVME_NSHEAD_DISK_LIVE, &head->flags))
> @@ -519,10 +525,16 @@ static bool nvme_available_path(struct nvme_ns_head *head)
>
> list_for_each_entry_srcu(ns, &head->list, siblings,
> srcu_read_lock_held(&head->srcu)) {
> - if (test_bit(NVME_CTRL_FAILFAST_EXPIRED, &ns->ctrl->flags))
> + if (test_bit(NVME_CTRL_FAILFAST_EXPIRED, &ns->ctrl->flags)) {
> + failfast = true;
> continue;
> + }
> switch (nvme_ctrl_state(ns->ctrl)) {
> case NVME_CTRL_LIVE:
> + if (!nvme_state_is_live(ns->ana_state) &&
> + ns->ana_state != NVME_ANA_CHANGE)
> + continue;
> + return true;
> case NVME_CTRL_RESETTING:
> case NVME_CTRL_CONNECTING:
> return true;
> @@ -531,6 +543,9 @@ static bool nvme_available_path(struct nvme_ns_head *head)
> }
> }
>
> + if (failfast)
> + return false;
> +
> /*
> * If "head->delayed_removal_secs" is configured (i.e., non-zero), do
> * not immediately fail I/O. Instead, requeue the I/O for the configured
> @@ -887,11 +902,6 @@ static int nvme_parse_ana_log(struct nvme_ctrl *ctrl, void *data,
> return 0;
> }
>
> -static inline bool nvme_state_is_live(enum nvme_ana_state state)
> -{
> - return state == NVME_ANA_OPTIMIZED || state == NVME_ANA_NONOPTIMIZED;
> -}
> -
> static void nvme_update_ns_ana_state(struct nvme_ana_group_desc *desc,
> struct nvme_ns *ns)
> {
> @@ -926,8 +936,10 @@ static void nvme_update_ns_ana_state(struct nvme_ana_group_desc *desc,
> * is not live but still create the sysfs link to this path from
> * head node if head node of the path has already come alive.
> */
> - if (test_bit(NVME_NSHEAD_DISK_LIVE, &ns->head->flags))
> + if (test_bit(NVME_NSHEAD_DISK_LIVE, &ns->head->flags)) {
> nvme_mpath_add_sysfs_link(ns->head);
> + kblockd_schedule_work(&ns->head->requeue_work);
> + }
> }
> }
>
Reviewed-by: Hannes Reinecke <hare@kernel.org>
Cheers,
Hannes
--
Dr. Hannes Reinecke Kernel Storage Architect
hare@suse.de +49 911 74053 688
SUSE Software Solutions GmbH, Frankenstr. 146, 90461 Nürnberg
HRB 36809 (AG Nürnberg), GF: I. Totev, A. McDonald, W. Knoblich
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [PATCH v4 2/2] nvme-multipath: add fail_if_no_path sysfs attribute
2026-10-01 9:48 ` [PATCH v4 2/2] nvme-multipath: add fail_if_no_path sysfs attribute Krishna Iyer
@ 2026-10-02 9:39 ` Hannes Reinecke
2026-10-02 10:56 ` Krishna Iyer
0 siblings, 1 reply; 9+ messages in thread
From: Hannes Reinecke @ 2026-10-02 9:39 UTC (permalink / raw)
To: Krishna Iyer, kbusch, axboe, hch, sagi
Cc: linux-nvme, linux-kernel, nilay, saravanand, sjpark
On 10/1/26 11:48 AM, Krishna Iyer wrote:
> When no usable path exists, I/O on a multipath namespace is queued
> until a path returns. With ctrl_loss_tmo=-1 that can be forever:
> during a long fabric outage any process waiting on the I/O is stuck in
> D state. We hit this on virtualization hosts, where a SIGKILLed VM
> process cannot exit while draining I/O to an unreachable NVMe/TCP
> target.
>
> Nothing can fail this I/O without tearing something down: controller
> deletion takes every namespace on the controller with it.
>
> Add a fail_if_no_path attribute on the ns-head disk: a persistent
> per-namespace policy to fail parked and newly arriving I/O instead of
> queueing it when no usable path exists. It is the namespace-scoped
> counterpart to the controller-scoped fast_io_fail_tmo: the trigger is
> an event userspace observes (a consumer known to be gone, for us a
> SIGKILLed VM the host must reap), not a duration picked up front, and
> sibling namespaces behind the same controllers keep queueing and ride
> out the outage.
>
> A CONNECTING controller stops counting as an available path, and with
> no controllers left the policy overrides the delayed_removal_secs
> queue-if-no-path window; the two are opposites, so fail_if_no_path takes
> precedence when both are set. ANA change and controller resetting still
> queue, since both are transient. Controller state is untouched and
> reconnects continue. Like dm's fail_if_no_path, the policy is transport
> agnostic.
>
> Assisted-by: Claude:claude-fable-5
> Signed-off-by: Krishna Iyer <kiyer@crusoe.ai>
> ---
> Documentation/ABI/stable/sysfs-nvme | 16 +++++++++
> drivers/nvme/host/multipath.c | 50 +++++++++++++++++++++++++++--
> drivers/nvme/host/nvme.h | 2 ++
> drivers/nvme/host/sysfs.c | 4 ++-
> 4 files changed, 69 insertions(+), 3 deletions(-)
>
I do get the problem, but I think that 'fail_if_no_path' is a misnomer.
Problem is that we already have a 'QUEUE_IF_NO_PATH' setting, so
'FAIL_IF_NO_PATH' really sounds like the inverstion of that.
Only that it isn't.
Maybe rename to 'FAIL_ON_CTRL_LOSS' to make it clear that the two
settings really describe different use-cases?
Cheers,
Hannes
--
Dr. Hannes Reinecke Kernel Storage Architect
hare@suse.de +49 911 74053 688
SUSE Software Solutions GmbH, Frankenstr. 146, 90461 Nürnberg
HRB 36809 (AG Nürnberg), GF: I. Totev, A. McDonald, W. Knoblich
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [PATCH v4 2/2] nvme-multipath: add fail_if_no_path sysfs attribute
2026-10-02 9:39 ` Hannes Reinecke
@ 2026-10-02 10:56 ` Krishna Iyer
2026-10-02 11:52 ` Nilay Shroff
0 siblings, 1 reply; 9+ messages in thread
From: Krishna Iyer @ 2026-10-02 10:56 UTC (permalink / raw)
To: Hannes Reinecke, Nilay Shroff
Cc: Krishna Iyer, Keith Busch, Jens Axboe, Christoph Hellwig,
Sagi Grimberg, linux-nvme, linux-kernel, saravanand, sjpark
On 10/2/26 11:39 AM, Hannes Reinecke wrote:
> I do get the problem, but I think that 'fail_if_no_path' is a misnomer.
> Problem is that we already have a 'QUEUE_IF_NO_PATH' setting, so
> 'FAIL_IF_NO_PATH' really sounds like the inverstion of that.
> Only that it isn't.
> Maybe rename to 'FAIL_ON_CTRL_LOSS' to make it clear that the two
> settings really describe different use-cases?
Agreed, and thanks. You are right that fail_if_no_path reads as the
inverse of queue_if_no_path when it is not, and the v4 split makes that
clearer. Patch 1 now fails the cases where a path exists but cannot
serve I/O, so the only thing this attribute still does is stop waiting
for a controller that is gone or reconnecting. I am happy to rename it,
and fail_on_ctrl_loss describes that well.
One piece of history worth surfacing first. The fail_if_no_path name was
Nilay's suggestion in v2, so I would request we settle on one name you
both agree on rather than change it twice. Nilay, does fail_on_ctrl_loss
work for you?
Cheers,
Krishna
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [PATCH v4 2/2] nvme-multipath: add fail_if_no_path sysfs attribute
2026-10-02 10:56 ` Krishna Iyer
@ 2026-10-02 11:52 ` Nilay Shroff
2026-10-02 12:38 ` Krishna Iyer
0 siblings, 1 reply; 9+ messages in thread
From: Nilay Shroff @ 2026-10-02 11:52 UTC (permalink / raw)
To: Krishna Iyer, Hannes Reinecke
Cc: Keith Busch, Jens Axboe, Christoph Hellwig, Sagi Grimberg,
linux-nvme, linux-kernel, saravanand, sjpark
On 10/2/26 4:26 PM, Krishna Iyer wrote:
> On 10/2/26 11:39 AM, Hannes Reinecke wrote:
>> I do get the problem, but I think that 'fail_if_no_path' is a misnomer.
>> Problem is that we already have a 'QUEUE_IF_NO_PATH' setting, so
>> 'FAIL_IF_NO_PATH' really sounds like the inverstion of that.
>> Only that it isn't.
>> Maybe rename to 'FAIL_ON_CTRL_LOSS' to make it clear that the two
>> settings really describe different use-cases?
>
> Agreed, and thanks. You are right that fail_if_no_path reads as the
> inverse of queue_if_no_path when it is not, and the v4 split makes that
> clearer. Patch 1 now fails the cases where a path exists but cannot
> serve I/O, so the only thing this attribute still does is stop waiting
> for a controller that is gone or reconnecting. I am happy to rename it,
> and fail_on_ctrl_loss describes that well.
>
> One piece of history worth surfacing first. The fail_if_no_path name was
> Nilay's suggestion in v2, so I would request we settle on one name you
> both agree on rather than change it twice. Nilay, does fail_on_ctrl_loss
> work for you?
> Yes, I also like fail_on_ctrl_loss as Hannes suggested.
The earlier fail_if_no_path name came from the dm-multipath policy, which
uses the same name for a similar use case. However, I agree that fail_if_no_path
could be misleading here since it sounds like the inverse of QUEUE_IF_NO_PATH.
So I'd vote for fail_on_ctrl_loss for this change.
Thanks,
--Nilay
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [PATCH v4 1/2] nvme-multipath: fix path state evaluation for failfast and ANA
2026-10-01 9:48 ` [PATCH v4 1/2] nvme-multipath: fix path state evaluation for failfast and ANA Krishna Iyer
2026-10-02 9:36 ` Hannes Reinecke
@ 2026-10-02 11:53 ` Nilay Shroff
1 sibling, 0 replies; 9+ messages in thread
From: Nilay Shroff @ 2026-10-02 11:53 UTC (permalink / raw)
To: Krishna Iyer, kbusch, axboe, hch, sagi
Cc: linux-nvme, linux-kernel, hare, saravanand, sjpark
On 10/1/26 3:18 PM, Krishna Iyer wrote:
> nvme_available_path() has two problems that keep I/O queued when it
> should be failed:
>
> 1. When fast_io_fail_tmo expires, NVME_CTRL_FAILFAST_EXPIRED is set and
> the path is skipped, but the function still falls through to
> nvme_mpath_queue_if_no_path(). If delayed_removal_secs is configured
> the I/O is requeued indefinitely, defeating fast_io_fail_tmo. Only
> fall through to queue_if_no_path when there really are no paths;
> if a path exists but its failfast timer has expired, fail instead.
>
> 2. A LIVE controller counts as a usable path regardless of the namespace
> ANA state. A path whose ANA state is inaccessible or persistent-loss
> cannot serve I/O, so it should not count as available. ANA change is
> transient and bounded by ANATT, so keep queueing while it resolves.
>
> Also kick the requeue work from nvme_update_ns_ana_state() when a path
> does not transition to live, so parked I/O is re-evaluated when an ANA
> transition leaves the namespace inaccessible.
>
> Move nvme_state_is_live() above nvme_available_path() so it can be used
> there.
>
> Assisted-by: Claude:claude-fable-5
> Signed-off-by: Krishna Iyer<kiyer@crusoe.ai>
Looks good to me.
Reviewed-by: Nilay Shroff <nilay@linux.ibm.com>
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [PATCH v4 2/2] nvme-multipath: add fail_if_no_path sysfs attribute
2026-10-02 11:52 ` Nilay Shroff
@ 2026-10-02 12:38 ` Krishna Iyer
0 siblings, 0 replies; 9+ messages in thread
From: Krishna Iyer @ 2026-10-02 12:38 UTC (permalink / raw)
To: nilay, hare
Cc: kbusch, axboe, hch, sagi, linux-nvme, linux-kernel, saravanand,
sjpark, Krishna Iyer
On 10/2/26 5:22 PM, Nilay Shroff wrote:
> The earlier fail_if_no_path name came from the dm-multipath policy, which
> uses the same name for a similar use case. However, I agree that
> fail_if_no_path could be misleading here since it sounds like the inverse
> of QUEUE_IF_NO_PATH.
>
> So I'd vote for fail_on_ctrl_loss for this change.
Thanks both. We have consensus, so I'll send a v5 with the rename:
the attribute, the NVME_NSHEAD flag, the ABI entry, and all references
become fail_on_ctrl_loss / NVME_NSHEAD_FAIL_ON_CTRL_LOSS. No logic
changes otherwise, and I'll carry the Reviewed-by tags on patch 1.
Cheers,
Krishna
^ permalink raw reply [flat|nested] 9+ messages in thread
end of thread, other threads:[~2026-10-02 12:38 UTC | newest]
Thread overview: 9+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-01 9:48 [PATCH v4 0/2] nvme-multipath: fail I/O when no usable path exists Krishna Iyer
2026-10-01 9:48 ` [PATCH v4 1/2] nvme-multipath: fix path state evaluation for failfast and ANA Krishna Iyer
2026-10-02 9:36 ` Hannes Reinecke
2026-10-02 11:53 ` Nilay Shroff
2026-10-01 9:48 ` [PATCH v4 2/2] nvme-multipath: add fail_if_no_path sysfs attribute Krishna Iyer
2026-10-02 9:39 ` Hannes Reinecke
2026-10-02 10:56 ` Krishna Iyer
2026-10-02 11:52 ` Nilay Shroff
2026-10-02 12:38 ` Krishna Iyer
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®