mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Imran Khan <imran.f.khan@oracle.com>
To: Al Viro <viro@zeniv.linux.org.uk>
Cc: tj@kernel.org, gregkh@linuxfoundation.org,
	akpm@linux-foundation.org, linux-kernel@vger.kernel.org
Subject: Re: [RESEND PATCH v7 7/8] kernfs: Replace per-fs rwsem with hashed rwsems.
Date: Mon, 21 Mar 2022 12:57:07 +1100	[thread overview]
Message-ID: <536f2392-45d2-2f43-5e9d-01ef50e33126@oracle.com> (raw)
In-Reply-To: <YjPNOQJf/Wxa4YeV@zeniv-ca.linux.org.uk>

Hello Al,
Thanks again for reviewing this.

On 18/3/22 11:07 am, Al Viro wrote:
> On Thu, Mar 17, 2022 at 06:26:11PM +1100, Imran Khan wrote:
> 
>> diff --git a/fs/kernfs/symlink.c b/fs/kernfs/symlink.c
>> index 9d4103602554..cbdd1be5f0a8 100644
>> --- a/fs/kernfs/symlink.c
>> +++ b/fs/kernfs/symlink.c
>> @@ -113,12 +113,19 @@ static int kernfs_getlink(struct inode *inode, char *path)
>>  	struct kernfs_node *kn = inode->i_private;
>>  	struct kernfs_node *parent = kn->parent;
>>  	struct kernfs_node *target = kn->symlink.target_kn;
>> -	struct rw_semaphore *rwsem;
>> +	struct kernfs_rwsem_token token;
>>  	int error;
>>  
>> -	rwsem = kernfs_down_read(parent);
>> +	/**
>> +	 * Lock both parent and target, to avoid their movement
>> +	 * or removal in the middle of path construction.
>> +	 * If a competing remove or rename for parent or target
>> +	 * wins, it will be reflected in result returned from
>> +	 * kernfs_get_target_path.
>> +	 */
>> +	kernfs_down_read_double_nodes(target, parent, &token);
>>  	error = kernfs_get_target_path(parent, target, path);
>> -	kernfs_up_read(rwsem);
>> +	kernfs_up_read_double_nodes(target, parent, &token);
>>  
>>  	return error;
>>  }
> 
> No.  Read through the kernfs_get_target_path().  Why would locking these
> two specific nodes be sufficient for anything useful?  That code relies
> upon ->parent of *many* nodes being stable.  Which is not going to be
> guaranteed by anything of that sort.
> 
> And it's not just "we might get garbage if we race" - it's "we might
> walk into kfree'd object and proceed to walk the pointer chain".
> 
> Or have this loop
> 	kn = target;
> 	while (kn->parent && kn != base) {
> 		len += strlen(kn->name) + 1;
> 		kn = kn->parent;
> 	}
> see the names that are not identical to what we see in
> 	kn = target;
> 	while (kn->parent && kn != base) {
> 		int slen = strlen(kn->name);
> 
> 		len -= slen;
> 		memcpy(s + len, kn->name, slen);
> 		if (len)
> 			s[--len] = '/';
> 
> 		kn = kn->parent;
> 	}
> done later in the same function.  With obvious unpleasant effects.
> Or a different set of nodes, for that matter.
> 
> This code really depends upon the tree being stable.  No renames of
> any sort allowed during that thing.

Yes. My earlier approach is wrong.

This patch set has also introduced a per-fs mutex (kernfs_rm_mutex)
which should fix the problem of inconsistent tree view as far as
kernfs_get_path is concerned.
Acquiring kernfs_rm_mutex before invoking kernfs_get_path in
kernfs_getlink will ensure that kernfs_get_path will get a consistent
view of ->parent of nodes from root to target. This is because acquiring
kernfs_rm_mutex will ensure that __kernfs_remove does not remove any
kernfs_node(or parent of kernfs_node). Further it ensures that
kernfs_rename_ns does not move any kernfs_node. So far I have not used
per-fs mutex in kernfs_rename_ns but I can make this change in next
version. So following change on top of current patch set should fix
this issue of ->parent change in the middle of kernfs_get_path.


diff --git a/fs/kernfs/dir.c b/fs/kernfs/dir.c
index 1b28d99ff1c3..8095dcdd437c 100644
--- a/fs/kernfs/dir.c
+++ b/fs/kernfs/dir.c
@@ -1672,11 +1672,13 @@ int kernfs_rename_ns(struct kernfs_node *kn,
struct kernfs_node *new_parent,
        const char *old_name = NULL;
        struct kernfs_rwsem_token token;
        int error;
+       struct kernfs_root *root = kernfs_root(kn);

        /* can't move or rename root */
        if (!kn->parent)
                return -EINVAL;

+       mutex_lock(&root->kernfs_rm_mutex);
        old_parent = kn->parent;
        kernfs_get(old_parent);
        kernfs_down_write_triple_nodes(kn, old_parent, new_parent, &token);
@@ -1741,6 +1743,7 @@ int kernfs_rename_ns(struct kernfs_node *kn,
struct kernfs_node *new_parent,
        error = 0;
  out:
        kernfs_up_write_triple_nodes(kn, new_parent, old_parent, &token);
+       mutex_unlock(&root->kernfs_rm_mutex);
        return error;
 }

diff --git a/fs/kernfs/symlink.c b/fs/kernfs/symlink.c
index cbdd1be5f0a8..805543d7a1f2 100644
--- a/fs/kernfs/symlink.c
+++ b/fs/kernfs/symlink.c
@@ -113,19 +113,22 @@ static int kernfs_getlink(struct inode *inode,
char *path)
        struct kernfs_node *kn = inode->i_private;
        struct kernfs_node *parent = kn->parent;
        struct kernfs_node *target = kn->symlink.target_kn;
-       struct kernfs_rwsem_token token;
+       struct kernfs_root *root;
        int error;

+       root = kernfs_root(kn);
+
        /**
-        * Lock both parent and target, to avoid their movement
-        * or removal in the middle of path construction.
-        * If a competing remove or rename for parent or target
-        * wins, it will be reflected in result returned from
-        * kernfs_get_target_path.
+        * Acquire kernfs_rm_mutex to ensure that kernfs_get_path
+        * sees correct ->parent for all nodes.
+        * We need to make sure that during kernfs_get_path parent
+        * of any node from target to root does not change. Acquiring
+        * kernfs_rm_mutex ensure that there are no concurrent remove
+        * or rename operations.
         */
-       kernfs_down_read_double_nodes(target, parent, &token);
+       mutex_lock(&root->kernfs_rm_mutex);
        error = kernfs_get_target_path(parent, target, path);
-       kernfs_up_read_double_nodes(target, parent, &token);
+       mutex_unlock(&root->kernfs_rm_mutex);

        return error;
 }

Could you please let me know if you see some issues with this approach ?

Thanks
-- Imran



  reply	other threads:[~2022-03-21  1:57 UTC|newest]

Thread overview: 31+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2022-03-17  7:26 [RESEND PATCH v7 0/8] kernfs: Introduce interface to access global kernfs_open_file_mutex Imran Khan
2022-03-17  7:26 ` [RESEND PATCH v7 1/8] " Imran Khan
2022-03-17 21:34   ` Al Viro
2022-04-05  5:36     ` Imran Khan
2022-04-05 14:24       ` Al Viro
2022-04-06  4:54         ` Imran Khan
2022-04-06 14:54           ` Al Viro
2022-04-06 15:18             ` Tejun Heo
2022-04-14  0:01           ` Imran Khan
2022-03-18 17:10   ` Eric W. Biederman
2022-03-21  0:10     ` Imran Khan
2022-03-17  7:26 ` [RESEND PATCH v7 2/8] kernfs: Replace global kernfs_open_file_mutex with hashed mutexes Imran Khan
2022-03-17  7:26 ` [RESEND PATCH v7 3/8] kernfs: Introduce interface to access kernfs_open_node_lock Imran Khan
2022-03-17  7:26 ` [RESEND PATCH v7 4/8] kernfs: Replace global kernfs_open_node_lock with hashed spinlocks Imran Khan
2022-03-17  7:26 ` [RESEND PATCH v7 5/8] kernfs: Use a per-fs rwsem to protect per-fs list of kernfs_super_info Imran Khan
2022-03-17  7:26 ` [RESEND PATCH v7 6/8] kernfs: Introduce interface to access per-fs rwsem Imran Khan
2022-03-17  7:26 ` [RESEND PATCH v7 7/8] kernfs: Replace per-fs rwsem with hashed rwsems Imran Khan
2022-03-18  0:07   ` Al Viro
2022-03-21  1:57     ` Imran Khan [this message]
2022-03-21  7:29       ` Al Viro
2022-03-21 16:46         ` Tejun Heo
2022-03-21 17:55           ` Al Viro
2022-03-21 19:20             ` Tejun Heo
2022-03-22  2:40               ` Al Viro
2022-03-22 17:08                 ` Tejun Heo
2022-03-22 20:26                   ` Al Viro
2022-03-22 21:20                     ` Tejun Heo
2022-03-28  0:15                 ` Imran Khan
2022-03-28 17:30                   ` Tejun Heo
2022-03-30  2:23                 ` Imran Khan
2022-03-17  7:26 ` [RESEND PATCH v7 8/8] kernfs: Add a document to describe hashed locks used in kernfs Imran Khan

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=536f2392-45d2-2f43-5e9d-01ef50e33126@oracle.com \
    --to=imran.f.khan@oracle.com \
    --cc=akpm@linux-foundation.org \
    --cc=gregkh@linuxfoundation.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=tj@kernel.org \
    --cc=viro@zeniv.linux.org.uk \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®