From: Hannes Reinecke <hare@suse.de>
To: Krishna Iyer <kiyer@crusoe.ai>,
kbusch@kernel.org, axboe@kernel.dk, hch@lst.de, sagi@grimberg.me
Cc: linux-nvme@lists.infradead.org, linux-kernel@vger.kernel.org,
nilay@linux.ibm.com, saravanand@crusoe.ai, sjpark@crusoe.ai
Subject: Re: [PATCH v4 1/2] nvme-multipath: fix path state evaluation for failfast and ANA
Date: Fri, 2 Oct 2026 11:36:36 +0200 [thread overview]
Message-ID: <09c64ac0-8f05-4815-bb17-481614d2db91@suse.de> (raw)
In-Reply-To: <20261001094857.74567-2-kiyer@crusoe.ai>
On 10/1/26 11:48 AM, Krishna Iyer wrote:
> nvme_available_path() has two problems that keep I/O queued when it
> should be failed:
>
> 1. When fast_io_fail_tmo expires, NVME_CTRL_FAILFAST_EXPIRED is set and
> the path is skipped, but the function still falls through to
> nvme_mpath_queue_if_no_path(). If delayed_removal_secs is configured
> the I/O is requeued indefinitely, defeating fast_io_fail_tmo. Only
> fall through to queue_if_no_path when there really are no paths;
> if a path exists but its failfast timer has expired, fail instead.
>
> 2. A LIVE controller counts as a usable path regardless of the namespace
> ANA state. A path whose ANA state is inaccessible or persistent-loss
> cannot serve I/O, so it should not count as available. ANA change is
> transient and bounded by ANATT, so keep queueing while it resolves.
>
> Also kick the requeue work from nvme_update_ns_ana_state() when a path
> does not transition to live, so parked I/O is re-evaluated when an ANA
> transition leaves the namespace inaccessible.
>
> Move nvme_state_is_live() above nvme_available_path() so it can be used
> there.
>
> Assisted-by: Claude:claude-fable-5
> Signed-off-by: Krishna Iyer <kiyer@crusoe.ai>
> ---
> drivers/nvme/host/multipath.c | 26 +++++++++++++++++++-------
> 1 file changed, 19 insertions(+), 7 deletions(-)
>
> diff --git a/drivers/nvme/host/multipath.c b/drivers/nvme/host/multipath.c
> index 11871f5f18c2..0d6c05f0808b 100644
> --- a/drivers/nvme/host/multipath.c
> +++ b/drivers/nvme/host/multipath.c
> @@ -509,9 +509,15 @@ inline struct nvme_ns *nvme_find_path(struct nvme_ns_head *head)
> }
> }
>
> +static inline bool nvme_state_is_live(enum nvme_ana_state state)
> +{
> + return state == NVME_ANA_OPTIMIZED || state == NVME_ANA_NONOPTIMIZED;
> +}
> +
> static bool nvme_available_path(struct nvme_ns_head *head)
> __must_hold_shared(&head->srcu)
> {
> + bool failfast = false;
> struct nvme_ns *ns;
>
> if (!test_bit(NVME_NSHEAD_DISK_LIVE, &head->flags))
> @@ -519,10 +525,16 @@ static bool nvme_available_path(struct nvme_ns_head *head)
>
> list_for_each_entry_srcu(ns, &head->list, siblings,
> srcu_read_lock_held(&head->srcu)) {
> - if (test_bit(NVME_CTRL_FAILFAST_EXPIRED, &ns->ctrl->flags))
> + if (test_bit(NVME_CTRL_FAILFAST_EXPIRED, &ns->ctrl->flags)) {
> + failfast = true;
> continue;
> + }
> switch (nvme_ctrl_state(ns->ctrl)) {
> case NVME_CTRL_LIVE:
> + if (!nvme_state_is_live(ns->ana_state) &&
> + ns->ana_state != NVME_ANA_CHANGE)
> + continue;
> + return true;
> case NVME_CTRL_RESETTING:
> case NVME_CTRL_CONNECTING:
> return true;
> @@ -531,6 +543,9 @@ static bool nvme_available_path(struct nvme_ns_head *head)
> }
> }
>
> + if (failfast)
> + return false;
> +
> /*
> * If "head->delayed_removal_secs" is configured (i.e., non-zero), do
> * not immediately fail I/O. Instead, requeue the I/O for the configured
> @@ -887,11 +902,6 @@ static int nvme_parse_ana_log(struct nvme_ctrl *ctrl, void *data,
> return 0;
> }
>
> -static inline bool nvme_state_is_live(enum nvme_ana_state state)
> -{
> - return state == NVME_ANA_OPTIMIZED || state == NVME_ANA_NONOPTIMIZED;
> -}
> -
> static void nvme_update_ns_ana_state(struct nvme_ana_group_desc *desc,
> struct nvme_ns *ns)
> {
> @@ -926,8 +936,10 @@ static void nvme_update_ns_ana_state(struct nvme_ana_group_desc *desc,
> * is not live but still create the sysfs link to this path from
> * head node if head node of the path has already come alive.
> */
> - if (test_bit(NVME_NSHEAD_DISK_LIVE, &ns->head->flags))
> + if (test_bit(NVME_NSHEAD_DISK_LIVE, &ns->head->flags)) {
> nvme_mpath_add_sysfs_link(ns->head);
> + kblockd_schedule_work(&ns->head->requeue_work);
> + }
> }
> }
>
Reviewed-by: Hannes Reinecke <hare@kernel.org>
Cheers,
Hannes
--
Dr. Hannes Reinecke Kernel Storage Architect
hare@suse.de +49 911 74053 688
SUSE Software Solutions GmbH, Frankenstr. 146, 90461 Nürnberg
HRB 36809 (AG Nürnberg), GF: I. Totev, A. McDonald, W. Knoblich
next prev parent reply other threads:[~2026-10-02 9:36 UTC|newest]
Thread overview: 9+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-01 9:48 [PATCH v4 0/2] nvme-multipath: fail I/O when no usable path exists Krishna Iyer
2026-10-01 9:48 ` [PATCH v4 1/2] nvme-multipath: fix path state evaluation for failfast and ANA Krishna Iyer
2026-10-02 9:36 ` Hannes Reinecke [this message]
2026-10-02 11:53 ` Nilay Shroff
2026-10-01 9:48 ` [PATCH v4 2/2] nvme-multipath: add fail_if_no_path sysfs attribute Krishna Iyer
2026-10-02 9:39 ` Hannes Reinecke
2026-10-02 10:56 ` Krishna Iyer
2026-10-02 11:52 ` Nilay Shroff
2026-10-02 12:38 ` Krishna Iyer
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=09c64ac0-8f05-4815-bb17-481614d2db91@suse.de \
--to=hare@suse.de \
--cc=axboe@kernel.dk \
--cc=hch@lst.de \
--cc=kbusch@kernel.org \
--cc=kiyer@crusoe.ai \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-nvme@lists.infradead.org \
--cc=nilay@linux.ibm.com \
--cc=sagi@grimberg.me \
--cc=saravanand@crusoe.ai \
--cc=sjpark@crusoe.ai \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®