From: Viacheslav Dubeyko <Slava.Dubeyko@ibm.com>
To: "ceph-devel@vger.kernel.org" <ceph-devel@vger.kernel.org>,
"ionut.nechita@windriver.com" <ionut.nechita@windriver.com>
Cc: "idryomov@gmail.com" <idryomov@gmail.com>,
Xiubo Li <xiubli@redhat.com>,
"linux-kernel@vger.kernel.org" <linux-kernel@vger.kernel.org>,
"ionut_n2001@yahoo.com" <ionut_n2001@yahoo.com>
Subject: Re: [PATCH v1 05/13] ceph: add timeout protection to ceph_lock_wait_for_completion()
Date: Thu, 12 Mar 2026 19:38:34 +0000 [thread overview]
Message-ID: <81a8db5a7de48ec04435b9c31ec63514636c47e8.camel@ibm.com> (raw)
In-Reply-To: <20260312081619.40854-6-ionut.nechita@windriver.com>
On Thu, 2026-03-12 at 10:16 +0200, Ionut Nechita (Wind River) wrote:
> From: Ionut Nechita <ionut.nechita@windriver.com>
>
> When a file lock operation is interrupted and an unlock request is
> sent to cancel it, ceph_lock_wait_for_completion() waits indefinitely
> for r_safe_completion using wait_for_completion_killable().
> If the MDS becomes unreachable after the unlock request is sent,
> this wait will block indefinitely, causing hung task warnings:
> INFO: task flock:12345 blocked for more than 122 seconds.
> Call Trace:
> wait_for_completion_killable+0x...
> ceph_lock_wait_for_completion+0x...
> ceph_flock+0x...
> This is similar to the issue fixed in ceph_mdsc_sync() where
> indefinite waits on r_safe_completion can hang when MDS is
> unavailable.
> Fix this by using wait_for_completion_killable_timeout() with
> mount_timeout instead of the indefinite wait. On timeout, return
> -ETIMEDOUT to the caller. The lock state remains consistent because:
> 1. If the unlock succeeded on MDS, the lock is released
> 2. If the unlock didn't reach MDS, the original lock request
> was already aborted (CEPH_MDS_R_ABORTED set), so MDS will
> clean it up on reconnect
> This follows the same timeout pattern used throughout the ceph
> client for MDS operations.
> Signed-off-by: Ionut Nechita <ionut.nechita@windriver.com>
> ---
> fs/ceph/locks.c | 14 +++++++++++++-
> 1 file changed, 13 insertions(+), 1 deletion(-)
>
> diff --git a/fs/ceph/locks.c b/fs/ceph/locks.c
> index ebf4ac0055ddc..55dd99460b81a 100644
> --- a/fs/ceph/locks.c
> +++ b/fs/ceph/locks.c
> @@ -160,6 +160,8 @@ static int ceph_lock_wait_for_completion(struct ceph_mds_client *mdsc,
> struct ceph_mds_request *req)
> {
> struct ceph_client *cl = mdsc->fsc->client;
> + struct ceph_options *opts = mdsc->fsc->client->options;
> + unsigned long timeout = ceph_timeout_jiffies(opts->mount_timeout);
The opts->mount_timeout could be configured unreasonably. Should we do something
about it?
> struct ceph_mds_request *intr_req;
> struct inode *inode = req->r_inode;
> int err, lock_type;
> @@ -221,7 +223,17 @@ static int ceph_lock_wait_for_completion(struct ceph_mds_client *mdsc,
> if (err && err != -ERESTARTSYS)
> return err;
>
> - wait_for_completion_killable(&req->r_safe_completion);
> + err = wait_for_completion_killable_timeout(&req->r_safe_completion,
> + timeout);
> + if (err == -ERESTARTSYS) {
Interesting... You didn't check this in other patches. Why? :)
> + /* Interrupted again, just return the error */
> + return err;
> + }
> + if (err == 0) {
> + pr_warn_client(cl, "lock request tid %llu safe completion timed out\n",
> + req->r_tid);
The same concern about sending warning into system log here.
> + return -ETIMEDOUT;
> + }
Maybe, some style cleanup here:
if (err == -ERESTARTSYS) {
<logic_1>
} else if (err == 0) {
<logic_2>
}
Thanks,
Slava.
> return 0;
> }
>
next prev parent reply other threads:[~2026-03-12 19:38 UTC|newest]
Thread overview: 31+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-03-12 8:16 [PATCH v1 00/13] ceph/libceph: fix hung tasks and connection recovery during network disruptions Ionut Nechita (Wind River)
2026-03-12 8:16 ` [PATCH v1 01/13] libceph: handle EADDRNOTAVAIL more gracefully Ionut Nechita (Wind River)
2026-03-12 18:51 ` Viacheslav Dubeyko
2026-03-12 8:16 ` [PATCH v1 02/13] ceph: add timeout protection to ceph_mdsc_sync() path Ionut Nechita (Wind River)
2026-03-12 19:19 ` Viacheslav Dubeyko
2026-03-12 8:16 ` [PATCH v1 03/13] ceph: add timeout protection to ceph_osdc_sync() path Ionut Nechita (Wind River)
2026-03-12 19:26 ` Viacheslav Dubeyko
2026-03-12 8:16 ` [PATCH v1 04/13] ceph: fix race condition in cleanup_session_requests() Ionut Nechita (Wind River)
2026-03-12 19:32 ` Viacheslav Dubeyko
2026-03-12 8:16 ` [PATCH v1 05/13] ceph: add timeout protection to ceph_lock_wait_for_completion() Ionut Nechita (Wind River)
2026-03-12 19:38 ` Viacheslav Dubeyko [this message]
2026-03-12 8:16 ` [PATCH v1 06/13] ceph: set default timeout for MDS requests Ionut Nechita (Wind River)
2026-03-12 19:41 ` Viacheslav Dubeyko
2026-03-12 8:16 ` [PATCH v1 07/13] ceph: add timeout to caps wait in __ceph_get_caps() Ionut Nechita (Wind River)
2026-03-12 19:52 ` Viacheslav Dubeyko
2026-03-12 8:16 ` [PATCH v1 08/13] ceph: make ceph_start_io_write() killable Ionut Nechita (Wind River)
2026-03-12 20:02 ` Viacheslav Dubeyko
2026-03-12 20:45 ` Ionut Nechita (Wind River)
2026-03-13 18:28 ` Viacheslav Dubeyko
2026-03-12 8:16 ` [PATCH v1 09/13] ceph: make remaining I/O lock functions killable Ionut Nechita (Wind River)
2026-03-12 20:05 ` Viacheslav Dubeyko
2026-03-12 8:16 ` [PATCH v1 10/13] ceph: force mdsmap refresh on persistent MDS connection failures Ionut Nechita (Wind River)
2026-03-12 21:23 ` Viacheslav Dubeyko
2026-03-12 8:16 ` [PATCH v1 11/13] libceph: reset source address on persistent EADDRNOTAVAIL Ionut Nechita (Wind River)
2026-03-12 21:39 ` Viacheslav Dubeyko
2026-03-12 8:16 ` [PATCH v1 12/13] libceph: force monitor reconnect " Ionut Nechita (Wind River)
2026-03-12 8:16 ` [PATCH v1 13/13] libceph: force host network namespace for kernel CephFS mounts Ionut Nechita (Wind River)
2026-03-16 15:28 ` Ilya Dryomov
2026-03-16 21:20 ` Ionut Nechita (Wind River)
2026-04-02 17:06 ` Ionut Nechita (Wind River)
2026-04-03 15:05 ` Ionut Nechita (Wind River)
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=81a8db5a7de48ec04435b9c31ec63514636c47e8.camel@ibm.com \
--to=slava.dubeyko@ibm.com \
--cc=ceph-devel@vger.kernel.org \
--cc=idryomov@gmail.com \
--cc=ionut.nechita@windriver.com \
--cc=ionut_n2001@yahoo.com \
--cc=linux-kernel@vger.kernel.org \
--cc=xiubli@redhat.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®