From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-wm1-f50.google.com (mail-wm1-f50.google.com [209.85.128.50]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 61FCD19D062 for ; Tue, 24 Dec 2024 10:31:39 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.50 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1735036301; cv=none; b=q84bD0fIlLFo/fj3KAFOzM9loUPPq1H2ylVG+iexhXxd3cAVQwqlZYP/UqHF0uo6ryeOM+XbTizm0WqrZ8ZPqijIV0121Z2fJn8WPxDGR7TR72TCMpUICImyx/QpsjVoJ56yGYb+hDfruHg7zI1o0LjVfycDZjuzdQe/GezEs0g= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1735036301; c=relaxed/simple; bh=cltn/7Y2q2fUhbiT38jQSccCb+XHNKPJiY7Se4R7SiA=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=CBu5tyqhszhSKsP//h9+ksjSwV7VFhBX58JYz5oBQm6a1HX7BnwUndR2t0AMdoEpDJdzOo1mChUH09z5aHGbbypJZa3wVLxhJitb8ugHDODxrmiNFk5xiQE57aLBmDk5syISfH1uVkDdddoMDVW7L1vJjYw6M+qVAySxFpGB8sM= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=grimberg.me; spf=pass smtp.mailfrom=gmail.com; arc=none smtp.client-ip=209.85.128.50 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=grimberg.me Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Received: by mail-wm1-f50.google.com with SMTP id 5b1f17b1804b1-4364a37a1d7so52295345e9.3 for ; Tue, 24 Dec 2024 02:31:39 -0800 (PST) X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1735036297; x=1735641097; h=content-transfer-encoding:in-reply-to:from:content-language :references:cc:to:subject:user-agent:mime-version:date:message-id :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to; bh=Un1QRK3Mtuw3SeHg4TzziEYwXQ3lUmtGd5ycEZTLM3c=; b=bedFR0nmBdtnDFzSywNj983OYxx69/L39iU4cVgWf19gnVftZoTg8DP96z5M7HypcP 2RAaneh0dKY0wX3bgi/EcPNdgSm3iRs/zTJ82YknCOhCjpFOmrjlokEO6bC++GI8Zw5P +BOun/Z2x/sDeX38CTWEhWl6+qHjc5ZIeLmp0ahDlXnluunN7ueoHr3tbdhkGcFCsoS4 wacchQ21NYFb37keD8Xg8kNX1pAHSpZDHsRbaZqYtouvFrUJnpnoK/4DPuf7phOSw4Zi eSkJcYDeuc6HCYWxJbHlLYzj0NlhOgBIi/lwHqwrCRwZ0xPJ+yDhSuF5bSxkem4rYbgU k3OQ== X-Forwarded-Encrypted: i=1; AJvYcCWfXTnj040jOJmOPNgGtH3ygYxr9YukZbcfDjFllCgGZsjUugVXAYcP5U1gMCHGEBNnqrlWIEaTRkU/3mQ=@vger.kernel.org X-Gm-Message-State: AOJu0YwoVunEiouqhYK4pwVfXsIc4CdJhCMJ3MBQADE8yK1pXF33KG5T u6HhwDBmI/6F74z6ao/rUy4fr7povHmkiPj3ohmF6CuXKaKSh/ov X-Gm-Gg: ASbGncvS8FGIIVkBY9lpmp9BZ2eaVGy4CjaRQ2XQ2skGgprww5GLm8j6k9vG8+SpKu+ PjZDt238j1l15w2ms5jrxEGvPBJ78SU07cs2KVEyvn72uLMo7YS6On+Uc+LzjwLaNbSZ5r1kBoW At/bLSqD4cVkQkb/fdED3uV0lB6hLNCqWR81l9XYQzmIm0+EbhmlyfF19tWsCc0LF8eSWuxjlH7 1AMRTZtsd58YCCRABzhXMgixoHdTZUtqmAqLtfrfPEQCBA1r8Kg00FzhIwGNZeuwanA79JmjIf4 HWyWli69O7WCgPck2ITgeHI= X-Google-Smtp-Source: AGHT+IELhZBLifRj9VyixbpOXsOxovDC9lByi6ix7Gv6MXH6pgWyVTduZCE+m2i7IKkE/947S7R57w== X-Received: by 2002:a5d:6d84:0:b0:385:f1d9:4b90 with SMTP id ffacd0b85a97d-38a221ea720mr14479446f8f.13.1735036297394; Tue, 24 Dec 2024 02:31:37 -0800 (PST) Received: from [10.50.4.206] (bzq-84-110-32-226.static-ip.bezeqint.net. [84.110.32.226]) by smtp.gmail.com with ESMTPSA id ffacd0b85a97d-38a1c89e3d9sm13544042f8f.67.2024.12.24.02.31.36 (version=TLS1_3 cipher=TLS_AES_128_GCM_SHA256 bits=128/128); Tue, 24 Dec 2024 02:31:37 -0800 (PST) Message-ID: Date: Tue, 24 Dec 2024 12:31:35 +0200 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH v3 2/3] nvme: trigger reset when keep alive fails To: Daniel Wagner , James Smart , Keith Busch , Christoph Hellwig , Hannes Reinecke , Paul Ely Cc: linux-nvme@lists.infradead.org, linux-kernel@vger.kernel.org References: <20241129-nvme-fc-handle-com-lost-v3-0-d8967b3cae54@kernel.org> <20241129-nvme-fc-handle-com-lost-v3-2-d8967b3cae54@kernel.org> Content-Language: en-US From: Sagi Grimberg In-Reply-To: <20241129-nvme-fc-handle-com-lost-v3-2-d8967b3cae54@kernel.org> Content-Type: text/plain; charset=UTF-8; format=flowed Content-Transfer-Encoding: 7bit On 29/11/2024 11:28, Daniel Wagner wrote: > nvme_keep_alive_work setups a keep alive command and uses > blk_execute_rq_nowait to send out the command in an asynchronously > manner. Eventually, nvme_keep_alive_end_io is called. If the status > argument is 0, a new keep alive is send out. When the status argument is > not 0, only an error is logged. The keep alive machinery does not > trigger the error recovery. > > The FC driver is relying on the keep alive machinery to trigger recovery > when an error is detected. Whenever an error happens during the creation > of the association the idea is that the operation is aborted and retried > later. Though there is a window where an error happens and > nvme_fc_create_assocation can't detect the error. > > 1) nvme nvme10: NVME-FC{10}: create association : ... > 2) nvme nvme10: NVME-FC{10}: controller connectivity lost. Awaiting Reconnect > nvme nvme10: queue_size 128 > ctrl maxcmd 32, reducing to maxcmd > 3) nvme nvme10: Could not set queue count (880) > nvme nvme10: Failed to configure AEN (cfg 900) > 4) nvme nvme10: NVME-FC{10}: controller connect complete > 5) nvme nvme10: failed nvme_keep_alive_end_io error=4 > > A connection attempt starts 1) and the ctrl is in state CONNECTING. > Shortly after the LLDD driver detects a connection lost event and calls > nvme_fc_ctrl_connectivity_loss 2). Because we are still in CONNECTING > state, this event is ignored. > > nvme_fc_create_association continues to run in parallel and tries to > communicate with the controller and those commands fail. Though these > errors are filtered out, e.g in 3) setting the I/O queues numbers fails > which leads to an early exit in nvme_fc_create_io_queues. Because the > number of IO queues is 0 at this point, there is nothing left in > nvme_fc_create_association which could detected the connection drop. > Thus the ctrl enters LIVE state 4). > > The keep alive timer fires and a keep alive command is send off but > gets rejected by nvme_fc_queue_rq and the rq status is set to > NVME_SC_HOST_PATH_ERROR. The nvme status is then mapped to a block layer > status BLK_STS_TRANSPORT/4 in nvme_end_req. Eventually, > nvme_keep_alive_end_io sees the status != 0 and just logs an error 5). > > We should obviously detect the problem in 3) and abort there (will > address this later), but that still leaves a race window open. There is > a race window open in nvme_fc_create_association after starting the IO > queues and setting the ctrl state to LIVE. > > Thus trigger a reset from the keep alive handler when an error is > reported. > > Signed-off-by: Daniel Wagner > --- > drivers/nvme/host/core.c | 6 ++++++ > 1 file changed, 6 insertions(+) > > diff --git a/drivers/nvme/host/core.c b/drivers/nvme/host/core.c > index bfd71511c85f8b1a9508c6ea062475ff51bf27fe..2a07c2c540b26c8cbe886711abaf6f0afbe6c4df 100644 > --- a/drivers/nvme/host/core.c > +++ b/drivers/nvme/host/core.c > @@ -1320,6 +1320,12 @@ static enum rq_end_io_ret nvme_keep_alive_end_io(struct request *rq, > dev_err(ctrl->device, > "failed nvme_keep_alive_end_io error=%d\n", > status); > + /* > + * The driver reports that we lost the connection, > + * trigger a recovery. > + */ > + if (status == BLK_STS_TRANSPORT) > + nvme_reset_ctrl(ctrl); > return RQ_END_IO_NONE; > } > > A lengthy explanation that results in nvme core behavior that assumes a very specific driver behavior. Isn't the root of the problem that FC is willing to live peacefully with a controller without any queues/connectivity to it without periodically reconnecting?