mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v4 0/7] VFS: prepare for changes to directory locking
@ 2026-09-04 21:48 NeilBrown
  2026-09-04 21:48 ` [PATCH v4 1/7] VFS: fix various typos in documentation for start_creating start_removing etc NeilBrown
                   ` (8 more replies)
  0 siblings, 9 replies; 13+ messages in thread
From: NeilBrown @ 2026-09-04 21:48 UTC (permalink / raw)
  To: Alexander Viro, Christian Brauner
  Cc: Jan Kara, linux-fsdevel, Jeff Layton, Amir Goldstein,
	Miklos Szeredi, linux-kernel

The only change in this v4 over v3 are to replace lock_sync with an
acquire/release pair which builds correctly when lockdep isn't enabled,
to add the 7th patch which I previously mentioned, and to include
linux-kernel so that sashiko gets to see the series.

Previous intro:

Hi,
 As you know I am working on changes to locking for directory operations
such as lookup/create/remove/rename.  The ultimate goal is for the VFS
to lock the dentry, not the parent directory, and to push the i_rwsem
parent locking down into the filesystems where it can be kept, removed,
or adjusted as best fits each filesystem.

The next step is to change the order of locking for d_alloc_parallel() -
which allocates a locked (in-lookup) dentry or waits for an existing
dentry to be unlocked.  Currently d_alloc_parallel() is ordered below
i_rwsem on parent, so you cannot wait on i_rwsem while holding a locked
(in-lookup) dentry.

To push i_rwsem into filesystems I need to push i_rwsem below d_alloc_parallel(),
so code can wait for i_rwsem while holding an in-lookup dentry, but won't be
able to wait in d_alloc_parallel() while holding i_rwsem.

There are two broad changes that are needed before that order can be
swapped.  This series sets the direction.  Subsequent patches which I
hope can also land this cycle spread those changes throughout
filesystems.

The two changes are:
1 - don't call d_alloc_parallel() while holding i_rwsem.
   This involves introducing d_alloc_trylock() and changing
   some places to drop i_rwsem before taking d_alloc_parallel()
   for this to work they will need to know if the lock is shared
   or exclusive so LOOKUP_SHARED is added.

2 - don't d_drop() a dentry while it being worked on.  This might
   not be entirely needed yet, but it will be needed soon and doing it
   now is a convenient time, and it helps make the end result clearer.
   Calling d_drop() effectively unlocks an in-lookup denty.
   It is OK for a filesystem to do this *after* an operation has
   completed, whether in success or failure.  Doing it before completion
   will allow another dentry to be allocated maybe too early.

   Enhancing d_splice_alias() to handle hashed dentries is key to
   removing the need to d_drop() in filesystems.  Though not strictly
   necessary, this allows d_add() to be removed as there will be nothing
   that d_add() does which cannot be done with d_splice_alias()

This set of 6 patches makes core-VFS changes. After this I have
 - 7 nfs patches
 - 6 afs patches
 - 4 smb/client patches
 - 3 cephfs patches
 - 2 fuse patches
 - one each for shmem, code, configfs, hostfs, procfs, ovl
and a patch to d_add_ci() which affects xfs and ntfs.

Once all of those have landed I have 4 patches to remove deprecated interfaces.
All of these can be found at
   https://github.com/neilbrown/linux/commits/pdirops


If all goes to plan, then in the next cycle a few patches will invert the
lock order and remove LOOKUP_SHARED.  Then we can get on to the really
fun stuff! (branch pdirops-next)

Thanks,
NeilBrown


 [PATCH v4 1/7] VFS: fix various typos in documentation for
 [PATCH v4 2/7] VFS: enhance d_splice_alias() to handle hashed
 [PATCH v4 3/7] VFS: introduce d_alloc_trylock()
 [PATCH v4 4/7] VFS: add d_duplicate()
 [PATCH v4 5/7] VFS: Add LOOKUP_SHARED flag.
 [PATCH v4 6/7] VFS: add lockdep monitoring of DCACHE_PAR_LOOKUP lock.
 [PATCH v4 7/7] VFS: reserve a d_flags bit for fs-specific usage

^ permalink raw reply	[flat|nested] 13+ messages in thread

* [PATCH v4 1/7] VFS: fix various typos in documentation for start_creating start_removing etc
  2026-09-04 21:48 [PATCH v4 0/7] VFS: prepare for changes to directory locking NeilBrown
@ 2026-09-04 21:48 ` NeilBrown
  2026-09-04 21:48 ` [PATCH v4 2/7] VFS: enhance d_splice_alias() to handle hashed dentries NeilBrown
                   ` (7 subsequent siblings)
  8 siblings, 0 replies; 13+ messages in thread
From: NeilBrown @ 2026-09-04 21:48 UTC (permalink / raw)
  To: Alexander Viro, Christian Brauner
  Cc: Jan Kara, linux-fsdevel, Jeff Layton, Amir Goldstein,
	Miklos Szeredi, linux-kernel

From: NeilBrown <neil@brown.name>

Various typos fixes.
start_creating_dentry() now documented as *creating*, not *removing* the
entry.
Unwanted spaces in Documentation/filesystems/porting.rst removed.

Signed-off-by: NeilBrown <neil@brown.name>
---
 Documentation/filesystems/porting.rst | 10 ++++----
 fs/namei.c                            | 34 +++++++++++++--------------
 2 files changed, 22 insertions(+), 22 deletions(-)

diff --git a/Documentation/filesystems/porting.rst b/Documentation/filesystems/porting.rst
index 60880eb0c49d..4e015f1bf1f8 100644
--- a/Documentation/filesystems/porting.rst
+++ b/Documentation/filesystems/porting.rst
@@ -1203,16 +1203,16 @@ will fail-safe.
 
 ---
 
-** mandatory**
+**mandatory**
 
 lookup_one(), lookup_one_unlocked(), lookup_one_positive_unlocked() now
 take a qstr instead of a name and len.  These, not the "one_len"
 versions, should be used whenever accessing a filesystem from outside
-that filesysmtem, through a mount point - which will have a mnt_idmap.
+that filesystem, through a mount point - which will have a mnt_idmap.
 
 ---
 
-** mandatory**
+**mandatory**
 
 Functions try_lookup_one_len(), lookup_one_len(),
 lookup_one_len_unlocked() and lookup_positive_unlocked() have been
@@ -1229,7 +1229,7 @@ already been performed such as after vfs_path_parent_lookup()
 
 ---
 
-** mandatory**
+**mandatory**
 
 d_hash_and_lookup() is no longer exported or available outside the VFS.
 Use try_lookup_noperm() instead.  This adds name validation and takes
@@ -1370,7 +1370,7 @@ similar.
 
 ---
 
-** mandatory**
+**mandatory**
 
 lock_rename(), lock_rename_child(), unlock_rename() are no
 longer available.  Use start_renaming() or similar.
diff --git a/fs/namei.c b/fs/namei.c
index 20a6534ea3ef..8d8d2af185f1 100644
--- a/fs/namei.c
+++ b/fs/namei.c
@@ -2946,8 +2946,8 @@ struct dentry *start_dirop(struct dentry *parent, struct qstr *name,
  * end_dirop - signal completion of a dirop
  * @de: the dentry which was returned by start_dirop or similar.
  *
- * If the de is an error, nothing happens. Otherwise any lock taken to
- * protect the dentry is dropped and the dentry itself is release (dput()).
+ * If the @de is an error, nothing happens. Otherwise any lock taken to
+ * protect the dentry is dropped and the dentry itself is released (dput()).
  */
 void end_dirop(struct dentry *de)
 {
@@ -3210,7 +3210,7 @@ EXPORT_SYMBOL(lookup_one);
 /**
  * lookup_one_unlocked - lookup single pathname component
  * @idmap:	idmap of the mount the lookup is performed from
- * @name:	qstr olding pathname component to lookup
+ * @name:	qstr holding pathname component to lookup
  * @base:	base directory to lookup from
  *
  * This can be used for in-kernel filesystem clients such as file servers.
@@ -3243,7 +3243,7 @@ EXPORT_SYMBOL(lookup_one_unlocked);
 /**
  * lookup_one_positive_killable - lookup single pathname component
  * @idmap:	idmap of the mount the lookup is performed from
- * @name:	qstr olding pathname component to lookup
+ * @name:	qstr holding pathname component to lookup
  * @base:	base directory to lookup from
  *
  * This helper will yield ERR_PTR(-ENOENT) on negatives. The helper returns
@@ -3259,7 +3259,7 @@ EXPORT_SYMBOL(lookup_one_unlocked);
  * the i_rwsem itself if necessary.  If a fatal signal is pending or
  * delivered, it will return %-EINTR if the lock is needed.
  *
- * Returns: A dentry, possibly negative, or
+ * Returns: A positive dentry, or
  *	   - same errors as lookup_one_unlocked() or
  *	   - ERR_PTR(-EINTR) if a fatal signal is pending.
  */
@@ -3381,7 +3381,7 @@ struct dentry *lookup_noperm_positive_unlocked(struct qstr *name,
 EXPORT_SYMBOL(lookup_noperm_positive_unlocked);
 
 /**
- * start_creating - prepare to create a given name with permission checking
+ * start_creating - prepare to access or create a given name with permission checking
  * @idmap:  idmap of the mount
  * @parent: directory in which to prepare to create the name
  * @name:   the name to be created
@@ -3413,8 +3413,8 @@ EXPORT_SYMBOL(start_creating);
  * @parent: directory in which to find the name
  * @name:   the name to be removed
  *
- * Locks are taken and a lookup in performed prior to removing
- * an object from a directory.  Permission checking (MAY_EXEC) is performed
+ * Locks are taken and a lookup is performed prior to removing an object
+ * from a directory.  Permission checking (MAY_EXEC) is performed
  * against @idmap.
  *
  * If the name doesn't exist, an error is returned.
@@ -3440,7 +3440,7 @@ EXPORT_SYMBOL(start_removing);
  * @parent: directory in which to prepare to create the name
  * @name:   the name to be created
  *
- * Locks are taken and a lookup in performed prior to creating
+ * Locks are taken and a lookup is performed prior to creating
  * an object in a directory.  Permission checking (MAY_EXEC) is performed
  * against @idmap.
  *
@@ -3469,7 +3469,7 @@ EXPORT_SYMBOL(start_creating_killable);
  * @parent: directory in which to find the name
  * @name:   the name to be removed
  *
- * Locks are taken and a lookup in performed prior to removing
+ * Locks are taken and a lookup is performed prior to removing
  * an object from a directory.  Permission checking (MAY_EXEC) is performed
  * against @idmap.
  *
@@ -3499,7 +3499,7 @@ EXPORT_SYMBOL(start_removing_killable);
  * @parent: directory in which to prepare to create the name
  * @name:   the name to be created
  *
- * Locks are taken and a lookup in performed prior to creating
+ * Locks are taken and a lookup is performed prior to creating
  * an object in a directory.
  *
  * If the name already exists, a positive dentry is returned.
@@ -3522,7 +3522,7 @@ EXPORT_SYMBOL(start_creating_noperm);
  * @parent: directory in which to find the name
  * @name:   the name to be removed
  *
- * Locks are taken and a lookup in performed prior to removing
+ * Locks are taken and a lookup is performed prior to removing
  * an object from a directory.
  *
  * If the name doesn't exist, an error is returned.
@@ -3543,11 +3543,11 @@ struct dentry *start_removing_noperm(struct dentry *parent,
 EXPORT_SYMBOL(start_removing_noperm);
 
 /**
- * start_creating_dentry - prepare to create a given dentry
- * @parent: directory from which dentry should be removed
- * @child:  the dentry to be removed
+ * start_creating_dentry - prepare to access or create a given dentry
+ * @parent: directory of dentry
+ * @child:  the dentry to be prepared
  *
- * A lock is taken to protect the dentry again other dirops and
+ * A lock is taken to protect the dentry against other dirops and
  * the validity of the dentry is checked: correct parent and still hashed.
  *
  * If the dentry is valid and negative a reference is taken and
@@ -3580,7 +3580,7 @@ EXPORT_SYMBOL(start_creating_dentry);
  * @parent: directory from which dentry should be removed
  * @child:  the dentry to be removed
  *
- * A lock is taken to protect the dentry again other dirops and
+ * A lock is taken to protect the dentry against other dirops and
  * the validity of the dentry is checked: correct parent and still hashed.
  *
  * If the dentry is valid and positive, a reference is taken and

base-commit: 89a312991dc6e638a36adc43ccb91dbc25504c04
-- 
2.50.0.107.gf914562f5916.dirty


^ permalink raw reply	[flat|nested] 13+ messages in thread

* [PATCH v4 2/7] VFS: enhance d_splice_alias() to handle hashed dentries
  2026-09-04 21:48 [PATCH v4 0/7] VFS: prepare for changes to directory locking NeilBrown
  2026-09-04 21:48 ` [PATCH v4 1/7] VFS: fix various typos in documentation for start_creating start_removing etc NeilBrown
@ 2026-09-04 21:48 ` NeilBrown
  2026-09-04 21:48 ` [PATCH v4 3/7] VFS: introduce d_alloc_trylock() NeilBrown
                   ` (6 subsequent siblings)
  8 siblings, 0 replies; 13+ messages in thread
From: NeilBrown @ 2026-09-04 21:48 UTC (permalink / raw)
  To: Alexander Viro, Christian Brauner
  Cc: Jan Kara, linux-fsdevel, Jeff Layton, Amir Goldstein,
	Miklos Szeredi, linux-kernel

From: NeilBrown <neil@brown.name>

We currently have three interfaces for attaching existing inodes to
normal filesystems(*).
- d_add() requires an unhashed or in-lookup dentry and doesn't handle
  splicing in case a directory already has dentry
- d_instantiate() requires a hashed dentry, and also doesn't handle
  splicing.
- d_splice_alias() requires unhashed or in-lookup and does handle
  splicing, and can return an alternate dentry.

So there is no interface that supports both hashed and in-lookup, which
is what ->atomic_open needs to deal with.

Some filesystems check for in-lookup in their atomic_open and if found,
perform a ->lookup and can subsequently use d_instantiate() if the
dentry is still negative.  Others d_drop() the dentry so they can use
d_splice_alias().

This last will cause a problem for proposed changes to locking which
require the dentry to remain hashed while an operation proceeds on it.

There is also no interface which splices a directory (which might
already have a dentry) to a hashed dentry.  Filesystems which need to do
this d_drop() first.

Some filesystems (NFS) skip ->lookup processing for
  LOOKUP_CREATE|LOOKUP_EXCL
which includes mknod, link, symlink etc.  So these inode operations
might get an unhashed or a hashed-negative dentry.  There is no
interface for instantiating these so again they need to unhash
first (nfs_link)

So with this patch d_splice_alias() can handle hashed, unhashed, or
in-lookup dentries.  This makes it suitable for ->lookup, ->atomic_open,
and ->mkdir as well as others.

As a side effect d_add() will also now handle hashed dentries, but
I have plans to remove d_add() as there is no benefit having it as
well as the others.

Once updated to handle nr_dentry_negative as is required for hashed
dentrties, __d_add() contains code that is identical to
__d_instantiate(), so the former is changed to call the later so now:

- d_add() calls __d_add() which hashes and might call __d_instantiate.
- d_instantiate() calls __d_instantiate() with appropriate locks.

It was suggested by Al Viro
   https://lore.kernel.org/all/20250813050717.GD222315@ZenIV/
that rather than allow d_splice_alias() to handle both hashed and
unhashed, we should have a new d_splice_alias_hashed().
I chose not to follow this path because, as noted above, there
are several cases where the filesystem has no a priori knowledge
of the state of the dentry, and so would need
  if (d_unhashed(dentry))
      alias = d_splice_alias(inode, dentry);
  else
      alias = d_splice_alias_hashed(inode, dentry);
which is clumsy.
Also I hope to minimise the distinction between hashed and in-lookup
(they will both be hashed, just with different DCACHE_ENTRY_TYPE)
and reduce the use of unhashed dentries.
Unhashed dentries would only be created by d_alloc_name() and all
of those are passed to d_make_persistent() (though configfs passes
some to d_add(dentry, NULL) first!).  So the use-case for
d_splice_alias() on unhashed dentries would disappear.
Note that d_make_persistent() already handles both hashed and
unhashed dentries - just not in-lookup.

* There is also d_make_persistent() for filesystems which are
  dcache-based and don't support mkdir, create etc, and
  d_instantiate_new() for newly created inodes that are still locked.

Signed-off-by: NeilBrown <neil@brown.name>
---
 Documentation/filesystems/vfs.rst |  4 ++--
 fs/dcache.c                       | 30 ++++++++++++------------------
 2 files changed, 14 insertions(+), 20 deletions(-)

diff --git a/Documentation/filesystems/vfs.rst b/Documentation/filesystems/vfs.rst
index d3a93eec3945..de8f8502056e 100644
--- a/Documentation/filesystems/vfs.rst
+++ b/Documentation/filesystems/vfs.rst
@@ -507,8 +507,8 @@ otherwise noted.
 	dentry before the first mkdir returns.
 
 	If there is any chance this could happen, then the new inode
-	should be d_drop()ed and attached with d_splice_alias().  The
-	returned dentry (if any) should be returned by ->mkdir().
+	should be attached with d_splice_alias().  The returned
+	dentry (if any) should be returned by ->mkdir().
 
 ``rmdir``
 	called by the rmdir(2) system call.  Only required if you want
diff --git a/fs/dcache.c b/fs/dcache.c
index 1b1a81f10da6..fd74753e7715 100644
--- a/fs/dcache.c
+++ b/fs/dcache.c
@@ -2168,7 +2168,6 @@ static void __d_instantiate(struct dentry *dentry, struct inode *inode)
  * (or otherwise set) by the caller to indicate that it is now
  * in use by the dcache.
  */
- 
 void d_instantiate(struct dentry *entry, struct inode * inode)
 {
 	BUG_ON(d_really_is_positive(entry));
@@ -2931,15 +2930,10 @@ static inline void __d_add(struct dentry *dentry, struct inode *inode,
 	}
 	if (unlikely(ops))
 		d_set_d_op(dentry, ops);
-	if (inode) {
-		unsigned add_flags = d_flags_for_inode(inode);
-		hlist_add_head(&dentry->d_alias, &inode->i_dentry);
-		raw_write_seqcount_begin(&dentry->d_seq);
-		__d_set_inode_and_type(dentry, inode, add_flags);
-		raw_write_seqcount_end(&dentry->d_seq);
-		fsnotify_update_flags(dentry);
-	}
-	__d_rehash(dentry);
+	if (inode)
+		__d_instantiate(dentry, inode);
+	if (d_unhashed(dentry))
+		__d_rehash(dentry);
 	if (dir) {
 		end_dir_add(dir, n);
 		__d_wake_in_lookup_waiters(dentry);
@@ -3241,7 +3235,7 @@ struct dentry *d_splice_alias_ops(struct inode *inode, struct dentry *dentry,
 	if (IS_ERR(inode))
 		return ERR_CAST(inode);
 
-	BUG_ON(!d_unhashed(dentry));
+	BUG_ON(d_really_is_positive(dentry));
 
 	if (!inode)
 		goto out;
@@ -3297,6 +3291,8 @@ struct dentry *d_splice_alias_ops(struct inode *inode, struct dentry *dentry,
  * @inode:  the inode which may have a disconnected dentry
  * @dentry: a negative dentry which we want to point to the inode.
  *
+ * @dentry must be negative and may be in-lookup or unhashed or hashed.
+ *
  * If inode is a directory and has an IS_ROOT alias, then d_move that in
  * place of the given dentry and return it, else simply d_add the inode
  * to the dentry and return NULL.
@@ -3304,16 +3300,14 @@ struct dentry *d_splice_alias_ops(struct inode *inode, struct dentry *dentry,
  * If a non-IS_ROOT directory is found, the filesystem is corrupt, and
  * we should error out: directories can't have multiple aliases.
  *
- * This is needed in the lookup routine of any filesystem that is exportable
- * (via knfsd) so that we can build dcache paths to directories effectively.
+ * This should be used to return the result of ->lookup() and to
+ * instantiate the result of ->mkdir(), is often useful for
+ * ->atomic_open, and may be used to instantiate other objects.
  *
  * If a dentry was found and moved, then it is returned.  Otherwise NULL
- * is returned.  This matches the expected return value of ->lookup.
+ * is returned.  This matches the expected return value of ->lookup and
+ * ->mkdir.
  *
- * Cluster filesystems may call this function with a negative, hashed dentry.
- * In that case, we know that the inode will be a regular file, and also this
- * will only occur during atomic_open. So we need to check for the dentry
- * being already hashed only in the final case.
  */
 struct dentry *d_splice_alias(struct inode *inode, struct dentry *dentry)
 {
-- 
2.50.0.107.gf914562f5916.dirty


^ permalink raw reply	[flat|nested] 13+ messages in thread

* [PATCH v4 3/7] VFS: introduce d_alloc_trylock()
  2026-09-04 21:48 [PATCH v4 0/7] VFS: prepare for changes to directory locking NeilBrown
  2026-09-04 21:48 ` [PATCH v4 1/7] VFS: fix various typos in documentation for start_creating start_removing etc NeilBrown
  2026-09-04 21:48 ` [PATCH v4 2/7] VFS: enhance d_splice_alias() to handle hashed dentries NeilBrown
@ 2026-09-04 21:48 ` NeilBrown
  2026-09-04 21:48 ` [PATCH v4 4/7] VFS: add d_duplicate() NeilBrown
                   ` (5 subsequent siblings)
  8 siblings, 0 replies; 13+ messages in thread
From: NeilBrown @ 2026-09-04 21:48 UTC (permalink / raw)
  To: Alexander Viro, Christian Brauner
  Cc: Jan Kara, linux-fsdevel, Jeff Layton, Amir Goldstein,
	Miklos Szeredi, linux-kernel

From: NeilBrown <neil@brown.name>

Several filesystems use the results of readdir to prime the dcache.
These filesystems use d_alloc_parallel() which can block if there is a
concurrent lookup.  Blocking in that case is pointless as the lookup
will add info to the dcache and there is no value in the readdir waiting
to see if it should add the info too.

Also these calls to d_alloc_parallel() are made while the parent
directory is locked.  A proposed change to locking will lock the parent
later, after d_alloc_parallel().  This means it won't be safe to wait in
d_alloc_parallel() while holding the directory lock.

So this patch introduces d_alloc_trylock() which doesn't block but
instead returns ERR_PTR(-EWOULDBLOCK).  Filesystems that prime the
dcache (smb/client, nfs, fuse, cephfs) can now use that and ignore
-EWOULDBLOCK errors as harmless.

Unlike d_alloc_parallel(), d_alloc_trylock() calculates the hash and
performs a lookup before an allocation, as that is what all callers
want.  This is done using try_lookup_noperm(), necessitating the
inclusion of namei.h in dcache.c.

Signed-off-by: NeilBrown <neil@brown.name>
---
 fs/dcache.c            | 81 ++++++++++++++++++++++++++++++++++++++++--
 include/linux/dcache.h |  1 +
 2 files changed, 80 insertions(+), 2 deletions(-)

diff --git a/fs/dcache.c b/fs/dcache.c
index fd74753e7715..43149c7849d9 100644
--- a/fs/dcache.c
+++ b/fs/dcache.c
@@ -32,6 +32,7 @@
 #include <linux/bit_spinlock.h>
 #include <linux/rculist_bl.h>
 #include <linux/list_lru.h>
+#include <linux/namei.h>
 #include "internal.h"
 #include "mount.h"
 
@@ -2756,8 +2757,16 @@ static void d_wait_lookup(struct dentry *dentry)
 	}
 }
 
-struct dentry *d_alloc_parallel(struct dentry *parent,
-				const struct qstr *name)
+/* What to do when __d_alloc_parallel finds a d_in_lookup dentry */
+enum alloc_para {
+	ALLOC_PARA_WAIT,
+	ALLOC_PARA_FAIL,
+};
+
+static inline
+struct dentry *__d_alloc_parallel(struct dentry *parent,
+				  const struct qstr *name,
+				  enum alloc_para how)
 {
 	unsigned int hash = name->hash;
 	struct hlist_bl_head *b = in_lookup_hash(parent, hash);
@@ -2830,6 +2839,12 @@ struct dentry *d_alloc_parallel(struct dentry *parent,
 			spin_unlock(&dentry->d_lock);
 			goto retry;
 		}
+		if (unlikely(how == ALLOC_PARA_FAIL)) {
+			/* mustn't wait for concurrent lookup to complete */
+			spin_unlock(&dentry->d_lock);
+			dput(new);
+			return ERR_PTR(-EWOULDBLOCK);
+		}
 		/*
 		 * somebody is likely to be still doing lookup for it;
 		 * pin it and wait for them to finish
@@ -2863,8 +2878,70 @@ struct dentry *d_alloc_parallel(struct dentry *parent,
 	dput(dentry);
 	goto retry;
 }
+
+/**
+ * d_alloc_parallel() - allocate a new dentry and ensure uniqueness
+ * @parent: dentry of the parent
+ * @name:   name of the dentry within that parent.
+ *
+ * A new dentry is allocated and, providing it is unique, added to the
+ * relevant index.
+ * If an existing dentry is found with the same parent/name that is
+ * not d_in_lookup(), then that is returned instead.
+ * If the existing dentry is d_in_lookup(), d_alloc_parallel() waits for
+ * that lookup to complete before returning the dentry and then ensures the
+ * match is still valid.
+ * Thus if the returned dentry is d_in_lookup() then the caller has
+ * exclusive access until it completes the lookup.
+ * If the returned dentry is not d_in_lookup() then a lookup has
+ * already completed.
+ *
+ * The @name must already have ->hash set, as can be achieved
+ * by e.g. try_lookup_noperm().
+ *
+ * Returns: the dentry, whether found or allocated, or an error %-ENOMEM.
+ */
+struct dentry *d_alloc_parallel(struct dentry *parent,
+				const struct qstr *name)
+{
+	return __d_alloc_parallel(parent, name, ALLOC_PARA_WAIT);
+}
 EXPORT_SYMBOL(d_alloc_parallel);
 
+/**
+ * d_alloc_trylock() - find or allocate a new dentry
+ * @parent: dentry of the parent
+ * @name:   name of the dentry within that parent.
+ *
+ * A new dentry is allocated and, providing it is unique, added to the
+ * relevant index.
+ * If an existing dentry is found with the same parent/name that is
+ * not d_in_lookup() then that is returned instead.
+ * If the existing dentry is d_in_lookup(), d_alloc_trylock()
+ * returns with error %-EWOULDBLOCK.
+ * Thus if the returned dentry is d_in_lookup() then the caller has
+ * exclusive access until it completes the lookup.
+ * If the returned dentry is not d_in_lookup() then a lookup has
+ * already completed.
+ *
+ * The @name need not already have ->hash set.
+ *
+ * Returns: the dentry, whether found or allocated, or an error
+ *    %-ENOMEM, %-EWOULDBLOCK, %-EACCES (for a bad name) or
+ *    anything returned by ->d_hash().
+ */
+struct dentry *d_alloc_trylock(struct dentry *parent,
+			       struct qstr *name)
+{
+	struct dentry *de;
+
+	de = try_lookup_noperm(name, parent);
+	if (!de)
+		de = __d_alloc_parallel(parent, name, ALLOC_PARA_FAIL);
+	return de;
+}
+EXPORT_SYMBOL(d_alloc_trylock);
+
 /*
  * Move dentry from in-lookup state to busy-negative one.
  *
diff --git a/include/linux/dcache.h b/include/linux/dcache.h
index 4b1ff99608e0..7afe16d4664d 100644
--- a/include/linux/dcache.h
+++ b/include/linux/dcache.h
@@ -257,6 +257,7 @@ extern void d_delete(struct dentry *);
 extern struct dentry * d_alloc(struct dentry *, const struct qstr *);
 extern struct dentry * d_alloc_anon(struct super_block *);
 extern struct dentry * d_alloc_parallel(struct dentry *, const struct qstr *);
+extern struct dentry * d_alloc_trylock(struct dentry *, struct qstr *);
 extern struct dentry * d_splice_alias(struct inode *, struct dentry *);
 /* weird procfs mess; *NOT* exported */
 extern struct dentry * d_splice_alias_ops(struct inode *, struct dentry *,
-- 
2.50.0.107.gf914562f5916.dirty


^ permalink raw reply	[flat|nested] 13+ messages in thread

* [PATCH v4 4/7] VFS: add d_duplicate()
  2026-09-04 21:48 [PATCH v4 0/7] VFS: prepare for changes to directory locking NeilBrown
                   ` (2 preceding siblings ...)
  2026-09-04 21:48 ` [PATCH v4 3/7] VFS: introduce d_alloc_trylock() NeilBrown
@ 2026-09-04 21:48 ` NeilBrown
  2026-09-04 21:48 ` [PATCH v4 5/7] VFS: Add LOOKUP_SHARED flag NeilBrown
                   ` (4 subsequent siblings)
  8 siblings, 0 replies; 13+ messages in thread
From: NeilBrown @ 2026-09-04 21:48 UTC (permalink / raw)
  To: Alexander Viro, Christian Brauner
  Cc: Jan Kara, linux-fsdevel, Jeff Layton, Amir Goldstein,
	Miklos Szeredi, linux-kernel

From: NeilBrown <neil@brown.name>

Occasionally a single operation can require two sub-operations on the
same name, and it is important that a d_alloc_parallel() (once that can
be run unlocked) does not create another dentry with the same name
between the operations.

Two examples:
1/ rename where the target name (a positive dentry) needs to be
  "silly-renamed" to a temporary name so it will remain available on the
  server (NFS and AFS).  Here the same name needs to be the subject
  of one rename, and the target of another.
2/ rename where the subject needs to be replaced with a white-out
  (shmemfs).  Here the same name need to be the subject of a rename
  and the target of a mknod()

In both cases the original dentry is renamed to something else, and a
replacement is instantiated, possibly as the target of d_move(), possibly
by d_instantiate().

Currently d_alloc() is used to create the dentry and the exclusive lock
on the parent ensures no other dentry is created.  When
d_alloc_parallel() is moved out of the parent lock, this will no longer
be sufficient.  In particular if the original is renamed away before the
new is instantiated, there is a window where d_alloc_parallel() could
create another name.  "silly-rename" does work in this order.  shmemfs
whiteout doesn't open this hole but is essentially the same pattern and
should use the same approach.

The new d_duplicate() creates an in-lookup dentry with the same name as
the original dentry, which must be hashed.  There is no need to check if
an in-lookup dentry exists with the same name as d_alloc_parallel() will
never try add one while the hashed dentry exists.  Once the new
in-lookup is created, d_alloc_parallel() will find it and wait for it to
complete, then use it.

Signed-off-by: NeilBrown <neil@brown.name>
---
 fs/dcache.c            | 51 ++++++++++++++++++++++++++++++++++++++++++
 include/linux/dcache.h |  1 +
 2 files changed, 52 insertions(+)

diff --git a/fs/dcache.c b/fs/dcache.c
index 43149c7849d9..cbd5738de168 100644
--- a/fs/dcache.c
+++ b/fs/dcache.c
@@ -2000,6 +2000,57 @@ struct dentry *d_alloc(struct dentry * parent, const struct qstr *name)
 }
 EXPORT_SYMBOL(d_alloc);
 
+/**
+ * d_duplicate - duplicate a dentry for combined atomic operation
+ * @dentry: the dentry to duplicate
+ *
+ * Some rename operations need to be combined with another operation
+ * inside the filesystem.
+ * 1/ A cluster filesystem when renaming to an in-use file might need to
+ *   first "silly-rename" that target out of the way before the main rename
+ * 2/ A filesystem that supports white-out might want to create a whiteout
+ *   in place of the file being moved.
+ *
+ * For this they need two dentries which temporarily have the same name,
+ * before one is renamed.  d_duplicate() provides for this.  Given a
+ * positive hashed dentry, it creates a second in-lookup dentry.
+ * Because the original dentry exists, no other thread will try to
+ * create an in-lookup dentry, so there can be no race in this create.
+ *
+ * The caller should d_move() the original to a new name, often via a
+ * rename request, and should call d_lookup_done() on the newly created
+ * dentry.  If the new is instantiated then the old MUST either be moved
+ * or dropped.
+ *
+ * Parent must be locked.
+ *
+ * Returns: an in-lookup dentry, or -ENOMEM.
+ */
+struct dentry *d_duplicate(struct dentry *dentry)
+{
+	unsigned int hash = dentry->d_name.hash;
+	struct dentry *parent = dentry->d_parent;
+	struct hlist_bl_head *b = in_lookup_hash(parent, hash);
+	struct dentry *new = __d_alloc(parent->d_sb, &dentry->d_name);
+
+	if (unlikely(!new))
+		return ERR_PTR(-ENOMEM);
+
+	new->d_flags |= DCACHE_PAR_LOOKUP;
+	spin_lock(&parent->d_lock);
+	new->d_parent = dget_dlock(parent);
+	hlist_add_head(&new->d_sib, &parent->d_children);
+	if (parent->d_flags & DCACHE_DISCONNECTED)
+		new->d_flags |= DCACHE_DISCONNECTED;
+	spin_unlock(&parent->d_lock);
+
+	hlist_bl_lock(b);
+	hlist_bl_add_head(&new->d_in_lookup_hash, b);
+	hlist_bl_unlock(b);
+	return new;
+}
+EXPORT_SYMBOL(d_duplicate);
+
 struct dentry *d_alloc_anon(struct super_block *sb)
 {
 	return __d_alloc(sb, NULL);
diff --git a/include/linux/dcache.h b/include/linux/dcache.h
index 7afe16d4664d..2b7d99ec9306 100644
--- a/include/linux/dcache.h
+++ b/include/linux/dcache.h
@@ -259,6 +259,7 @@ extern struct dentry * d_alloc_anon(struct super_block *);
 extern struct dentry * d_alloc_parallel(struct dentry *, const struct qstr *);
 extern struct dentry * d_alloc_trylock(struct dentry *, struct qstr *);
 extern struct dentry * d_splice_alias(struct inode *, struct dentry *);
+struct dentry *d_duplicate(struct dentry *dentry);
 /* weird procfs mess; *NOT* exported */
 extern struct dentry * d_splice_alias_ops(struct inode *, struct dentry *,
 					  const struct dentry_operations *);
-- 
2.50.0.107.gf914562f5916.dirty


^ permalink raw reply	[flat|nested] 13+ messages in thread

* [PATCH v4 5/7] VFS: Add LOOKUP_SHARED flag.
  2026-09-04 21:48 [PATCH v4 0/7] VFS: prepare for changes to directory locking NeilBrown
                   ` (3 preceding siblings ...)
  2026-09-04 21:48 ` [PATCH v4 4/7] VFS: add d_duplicate() NeilBrown
@ 2026-09-04 21:48 ` NeilBrown
  2026-09-04 21:48 ` [PATCH v4 6/7] VFS: add lockdep monitoring of DCACHE_PAR_LOOKUP lock NeilBrown
                   ` (3 subsequent siblings)
  8 siblings, 0 replies; 13+ messages in thread
From: NeilBrown @ 2026-09-04 21:48 UTC (permalink / raw)
  To: Alexander Viro, Christian Brauner
  Cc: Jan Kara, linux-fsdevel, Jeff Layton, Amir Goldstein,
	Miklos Szeredi, linux-kernel

From: NeilBrown <neil@brown.name>

Some ->lookup handlers will need to drop and retake the parent lock, so
they can safely use d_alloc_parallel().

->lookup can be called with the parent lock either exclusive or shared.

A new flag, LOOKUP_SHARED, tells ->lookup how the parent is locked.

This is rather ugly, but will be gone soon after we move
d_alloc_parallel() out of the directory lock as ->lookup() will *always*
called with a shared lock on the parent.

Signed-off-by: NeilBrown <neil@brown.name>
---
 fs/namei.c            | 20 +++++++++++---------
 include/linux/namei.h |  3 ++-
 2 files changed, 13 insertions(+), 10 deletions(-)

diff --git a/fs/namei.c b/fs/namei.c
index 8d8d2af185f1..e978a75eb8a1 100644
--- a/fs/namei.c
+++ b/fs/namei.c
@@ -1933,7 +1933,7 @@ static noinline struct dentry *lookup_slow(const struct qstr *name,
 	struct inode *inode = dir->d_inode;
 	struct dentry *res;
 	inode_lock_shared(inode);
-	res = __lookup_slow(name, dir, flags);
+	res = __lookup_slow(name, dir, flags | LOOKUP_SHARED);
 	inode_unlock_shared(inode);
 	return res;
 }
@@ -1947,7 +1947,7 @@ static struct dentry *lookup_slow_killable(const struct qstr *name,
 
 	if (inode_lock_shared_killable(inode))
 		return ERR_PTR(-EINTR);
-	res = __lookup_slow(name, dir, flags);
+	res = __lookup_slow(name, dir, flags | LOOKUP_SHARED);
 	inode_unlock_shared(inode);
 	return res;
 }
@@ -4440,12 +4440,14 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file,
 	int error, create_error;
 	umode_t mode;
 	bool got_write;
+	unsigned int shared_flag;
 
 retry:
 	open_flag = op->open_flag;
 	got_write = false;
 	mode = op->mode;
 	create_error = 0;
+	shared_flag = (open_flag & O_CREAT) ? 0 : LOOKUP_SHARED;
 
 	if (open_flag & (O_CREAT | O_TRUNC | O_WRONLY | O_RDWR)) {
 		got_write = !mnt_want_write(nd->path.mnt);
@@ -4454,10 +4456,10 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file,
 		 * a different error; we'll be dropping this one anyway.
 		 */
 	}
-	if (open_flag & O_CREAT)
-		inode_lock(dir_inode);
-	else
+	if (shared_flag)
 		inode_lock_shared(dir_inode);
+	else
+		inode_lock(dir_inode);
 
 	if (unlikely(IS_DEADDIR(dir_inode))) {
 		dentry = ERR_PTR(-ENOENT);
@@ -4526,7 +4528,7 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file,
 
 	if (d_in_lookup(dentry)) {
 		struct dentry *res = dir_inode->i_op->lookup(dir_inode, dentry,
-							     nd->flags);
+							     nd->flags | shared_flag);
 		d_lookup_done(dentry);
 		if (unlikely(res)) {
 			if (IS_ERR(res)) {
@@ -4574,10 +4576,10 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file,
 		if (file->f_mode & FMODE_OPENED)
 			fsnotify_open(file);
 	}
-	if ((open_flag & O_CREAT) || create_error)
-		inode_unlock(dir_inode);
-	else
+	if (shared_flag)
 		inode_unlock_shared(dir_inode);
+	else
+		inode_unlock(dir_inode);
 
 	if (got_write)
 		mnt_drop_write(nd->path.mnt);
diff --git a/include/linux/namei.h b/include/linux/namei.h
index 86d657b24fc6..65f0f5712885 100644
--- a/include/linux/namei.h
+++ b/include/linux/namei.h
@@ -32,8 +32,9 @@ enum { MAX_NESTED_LINKS = 8 };
 #define LOOKUP_CREATE		BIT(17)	/* ... in object creation */
 #define LOOKUP_EXCL		BIT(18)	/* ... in target must not exist */
 #define LOOKUP_RENAME_TARGET	BIT(19)	/* ... in destination of rename() */
+#define LOOKUP_SHARED		BIT(20) /* Parent lock is held shared */
 
-/* 4 spare bits for intent */
+/* 3 spare bits for intent */
 
 /* Scoping flags for lookup. */
 #define LOOKUP_NO_SYMLINKS	BIT(24) /* No symlink crossing. */
-- 
2.50.0.107.gf914562f5916.dirty


^ permalink raw reply	[flat|nested] 13+ messages in thread

* [PATCH v4 6/7] VFS: add lockdep monitoring of DCACHE_PAR_LOOKUP lock.
  2026-09-04 21:48 [PATCH v4 0/7] VFS: prepare for changes to directory locking NeilBrown
                   ` (4 preceding siblings ...)
  2026-09-04 21:48 ` [PATCH v4 5/7] VFS: Add LOOKUP_SHARED flag NeilBrown
@ 2026-09-04 21:48 ` NeilBrown
  2026-09-30 16:31   ` Borah, Chaitanya Kumar
  2026-09-04 21:48 ` [PATCH v4 7/7] VFS: reserve a d_flags bit for fs-specific usage NeilBrown
                   ` (2 subsequent siblings)
  8 siblings, 1 reply; 13+ messages in thread
From: NeilBrown @ 2026-09-04 21:48 UTC (permalink / raw)
  To: Alexander Viro, Christian Brauner
  Cc: Jan Kara, linux-fsdevel, Jeff Layton, Amir Goldstein,
	Miklos Szeredi, linux-kernel

From: NeilBrown <neil@brown.name>

DCACHE_PAR_LOOKUP acts like a lock in that threads can block waiting for
it to clear.  As we plan to make changes to lock order for this lock,
teach lockdep to monitor it so as to help detect bugs early.

As NFS allocates an in-lookup dentry to unlink a silly-renamed file, and
completes the lookup in a different thread, we need interfaces to
release and the acquire ownership of the lock.  This avoids lockdep
complaining that a lock is still held on return to user-space.

Signed-off-by: NeilBrown <neil@brown.name>
---
 fs/dcache.c            | 15 +++++++++++++++
 fs/nfs/unlink.c        |  3 +++
 include/linux/dcache.h | 32 ++++++++++++++++++++++++++++++++
 3 files changed, 50 insertions(+)

diff --git a/fs/dcache.c b/fs/dcache.c
index cbd5738de168..83790c7a4dee 100644
--- a/fs/dcache.c
+++ b/fs/dcache.c
@@ -1901,6 +1901,7 @@ EXPORT_SYMBOL(d_invalidate);
  
 static struct dentry *__d_alloc(struct super_block *sb, const struct qstr *name)
 {
+	static struct lock_class_key __lookup_key;
 	struct dentry *dentry;
 	char *dname;
 	int err;
@@ -1958,6 +1959,8 @@ static struct dentry *__d_alloc(struct super_block *sb, const struct qstr *name)
 	dentry->waiters = NULL;
 	INIT_HLIST_NODE(&dentry->d_sib);
 
+	lockdep_init_map(&dentry->lookup_map, "DCACHE_PAR_LOOKUP", &__lookup_key, 0);
+
 	if (dentry->d_op && dentry->d_op->d_init) {
 		err = dentry->d_op->d_init(dentry);
 		if (err) {
@@ -2037,6 +2040,7 @@ struct dentry *d_duplicate(struct dentry *dentry)
 		return ERR_PTR(-ENOMEM);
 
 	new->d_flags |= DCACHE_PAR_LOOKUP;
+	lock_map_acquire_try(&new->lookup_map);
 	spin_lock(&parent->d_lock);
 	new->d_parent = dget_dlock(parent);
 	hlist_add_head(&new->d_sib, &parent->d_children);
@@ -2801,6 +2805,15 @@ static inline void end_dir_add(struct inode *dir, unsigned int n)
 static void d_wait_lookup(struct dentry *dentry)
 {
 	if (likely(d_in_lookup(dentry))) {
+		/*
+		 * Tell lockdep we will wait for the lookup lock, after
+		 * dropping ->d_lock, but won't actually take it.
+		 */
+		spin_release(&dentry->d_lock.dep_map, _THIS_IP_);
+		lock_map_acquire(&dentry->lookup_map);
+		lock_map_release(&dentry->lookup_map);
+		spin_acquire(&dentry->d_lock.dep_map, 0, 1, _THIS_IP_);
+
 		dentry->d_flags |= DCACHE_LOOKUP_WAITERS;
 		wait_var_event_spinlock(&dentry->d_flags,
 					!d_in_lookup(dentry),
@@ -2923,6 +2936,7 @@ struct dentry *__d_alloc_parallel(struct dentry *parent,
 	}
 	hlist_bl_add_head(&new->d_in_lookup_hash, b);
 	hlist_bl_unlock(b);
+	lock_map_acquire_try(&new->lookup_map);
 	return new;
 mismatch:
 	spin_unlock(&dentry->d_lock);
@@ -3021,6 +3035,7 @@ static void __d_lookup_unhash(struct dentry *dentry)
 	b = in_lookup_hash(dentry->d_parent, dentry->d_name.hash);
 	hlist_bl_lock(b);
 	dentry->d_flags &= ~DCACHE_PAR_LOOKUP;
+	lock_map_release(&dentry->lookup_map);
 	__hlist_bl_del(&dentry->d_in_lookup_hash);
 	hlist_bl_unlock(b);
 	dentry->waiters = NULL;
diff --git a/fs/nfs/unlink.c b/fs/nfs/unlink.c
index b57cfaa4d516..c8d712204e64 100644
--- a/fs/nfs/unlink.c
+++ b/fs/nfs/unlink.c
@@ -67,6 +67,7 @@ static void nfs_async_unlink_release(void *calldata)
 	struct super_block *sb = dentry->d_sb;
 
 	up_read_non_owner(&NFS_I(d_inode(dentry->d_parent))->rmdir_sem);
+	d_lookup_acquire(dentry);
 	d_lookup_done(dentry);
 	nfs_free_unlinkdata(data);
 	dput(dentry);
@@ -159,6 +160,8 @@ static int nfs_call_unlink(struct dentry *dentry, struct inode *inode, struct nf
 		return ret;
 	}
 	data->dentry = alias;
+	d_lookup_release(alias);
+
 	nfs_do_call_unlink(inode, data);
 	return 1;
 }
diff --git a/include/linux/dcache.h b/include/linux/dcache.h
index 2b7d99ec9306..e7e3ef05313b 100644
--- a/include/linux/dcache.h
+++ b/include/linux/dcache.h
@@ -116,6 +116,8 @@ struct dentry {
 					 * possible!
 					 */
 
+	/* lockdep tracking of DCACHE_PAR_LOOKUP locks */
+	struct lockdep_map		lookup_map;
 	struct list_head d_lru;		/* LRU list */
 	struct hlist_node d_sib;	/* child of parent list */
 	struct hlist_head d_children;	/* our children */
@@ -554,6 +556,36 @@ static inline int simple_positive(const struct dentry *dentry)
 
 unsigned long vfs_pressure_ratio(unsigned long val);
 
+/**
+ * d_lookup_release - release ownership of DCACHE_PAR_LOOKUP lock
+ * @dentry: dentry that is locked
+ *
+ * If an in-lookup dentry is to be passed to another thread which
+ * will drop the in-lookup lock, then d_lookup_release() must be called
+ * to tell lockdep that this thread no lock holds the lock.  The
+ * thread that receives the lock must call d_lookup_acquire() to
+ * acquire the lock.
+ */
+static inline void d_lookup_release(struct dentry *dentry)
+{
+	if (d_in_lookup(dentry))
+		lock_map_release(&dentry->lookup_map);
+}
+
+/**
+ * d_lookup_acquire - acquire ownership of DCACHE_PAR_LOOKUP lock
+ * @dentry: dentry that is locked
+ *
+ * If an in-lookup dentry was passed to this thread, the
+ * d_lookup_acquire() must be called to tell lockdep that this
+ * thread now owns the DCACHE_PAR_LOOKUP lock.
+ */
+static inline void d_lookup_acquire(struct dentry *dentry)
+{
+	if (d_in_lookup(dentry))
+		lock_map_acquire_try(&dentry->lookup_map);
+}
+
 /**
  * d_inode - Get the actual inode of this dentry
  * @dentry: The dentry to query
-- 
2.50.0.107.gf914562f5916.dirty


^ permalink raw reply	[flat|nested] 13+ messages in thread

* [PATCH v4 7/7] VFS: reserve a d_flags bit for fs-specific usage
  2026-09-04 21:48 [PATCH v4 0/7] VFS: prepare for changes to directory locking NeilBrown
                   ` (5 preceding siblings ...)
  2026-09-04 21:48 ` [PATCH v4 6/7] VFS: add lockdep monitoring of DCACHE_PAR_LOOKUP lock NeilBrown
@ 2026-09-04 21:48 ` NeilBrown
  2026-09-15 21:11 ` [PATCH v4 0/7] VFS: prepare for changes to directory locking NeilBrown
  2026-09-25 14:28 ` Christian Brauner
  8 siblings, 0 replies; 13+ messages in thread
From: NeilBrown @ 2026-09-04 21:48 UTC (permalink / raw)
  To: Alexander Viro, Christian Brauner
  Cc: Jan Kara, linux-fsdevel, Jeff Layton, Amir Goldstein,
	Miklos Szeredi, linux-kernel

From: NeilBrown <neil@brown.name>

DCACHE_PRIVATE may be used by any filesystem for its own purposes, much
like d_fsdata and d_time.
I plan to use this in a similar manner the way nfs stores
NFS_FSDATA_BLOCKED in d_fsdata.

Signed-off-by: NeilBrown <neil@brown.name>
---
 include/linux/dcache.h | 4 +++-
 1 file changed, 3 insertions(+), 1 deletion(-)

diff --git a/include/linux/dcache.h b/include/linux/dcache.h
index e7e3ef05313b..adf239f8205f 100644
--- a/include/linux/dcache.h
+++ b/include/linux/dcache.h
@@ -238,7 +238,9 @@ enum dentry_flags {
 	DCACHE_PAR_LOOKUP		= BIT(24),	/* being looked up (with parent locked shared) */
 	DCACHE_DENTRY_CURSOR		= BIT(25),
 	DCACHE_NORCU			= BIT(26),	/* No RCU delay for freeing */
-	DCACHE_PERSISTENT		= BIT(27)
+	DCACHE_PERSISTENT		= BIT(27),
+/* 28, 29, 30 free */
+	DCACHE_PRIVATE			= BIT(31)	/* fs-specific flag */
 };
 
 #define DCACHE_MANAGED_DENTRY \
-- 
2.50.0.107.gf914562f5916.dirty


^ permalink raw reply	[flat|nested] 13+ messages in thread

* Re: [PATCH v4 0/7] VFS: prepare for changes to directory locking
  2026-09-04 21:48 [PATCH v4 0/7] VFS: prepare for changes to directory locking NeilBrown
                   ` (6 preceding siblings ...)
  2026-09-04 21:48 ` [PATCH v4 7/7] VFS: reserve a d_flags bit for fs-specific usage NeilBrown
@ 2026-09-15 21:11 ` NeilBrown
  2026-09-25 14:28 ` Christian Brauner
  8 siblings, 0 replies; 13+ messages in thread
From: NeilBrown @ 2026-09-15 21:11 UTC (permalink / raw)
  To: Alexander Viro, Christian Brauner
  Cc: Jan Kara, linux-fsdevel, Jeff Layton, Amir Goldstein,
	Miklos Szeredi, linux-kernel


Hi,
 is there any chance these might land this cycle?  I was hoping to add a
 bunch of follow-on per-fs patches, but I cannot without this prep.

Thanks,
NeilBrown


On Sat, 05 Sep 2026, NeilBrown wrote:
> The only change in this v4 over v3 are to replace lock_sync with an
> acquire/release pair which builds correctly when lockdep isn't enabled,
> to add the 7th patch which I previously mentioned, and to include
> linux-kernel so that sashiko gets to see the series.
> 
> Previous intro:
> 
> Hi,
>  As you know I am working on changes to locking for directory operations
> such as lookup/create/remove/rename.  The ultimate goal is for the VFS
> to lock the dentry, not the parent directory, and to push the i_rwsem
> parent locking down into the filesystems where it can be kept, removed,
> or adjusted as best fits each filesystem.
> 
> The next step is to change the order of locking for d_alloc_parallel() -
> which allocates a locked (in-lookup) dentry or waits for an existing
> dentry to be unlocked.  Currently d_alloc_parallel() is ordered below
> i_rwsem on parent, so you cannot wait on i_rwsem while holding a locked
> (in-lookup) dentry.
> 
> To push i_rwsem into filesystems I need to push i_rwsem below d_alloc_parallel(),
> so code can wait for i_rwsem while holding an in-lookup dentry, but won't be
> able to wait in d_alloc_parallel() while holding i_rwsem.
> 
> There are two broad changes that are needed before that order can be
> swapped.  This series sets the direction.  Subsequent patches which I
> hope can also land this cycle spread those changes throughout
> filesystems.
> 
> The two changes are:
> 1 - don't call d_alloc_parallel() while holding i_rwsem.
>    This involves introducing d_alloc_trylock() and changing
>    some places to drop i_rwsem before taking d_alloc_parallel()
>    for this to work they will need to know if the lock is shared
>    or exclusive so LOOKUP_SHARED is added.
> 
> 2 - don't d_drop() a dentry while it being worked on.  This might
>    not be entirely needed yet, but it will be needed soon and doing it
>    now is a convenient time, and it helps make the end result clearer.
>    Calling d_drop() effectively unlocks an in-lookup denty.
>    It is OK for a filesystem to do this *after* an operation has
>    completed, whether in success or failure.  Doing it before completion
>    will allow another dentry to be allocated maybe too early.
> 
>    Enhancing d_splice_alias() to handle hashed dentries is key to
>    removing the need to d_drop() in filesystems.  Though not strictly
>    necessary, this allows d_add() to be removed as there will be nothing
>    that d_add() does which cannot be done with d_splice_alias()
> 
> This set of 6 patches makes core-VFS changes. After this I have
>  - 7 nfs patches
>  - 6 afs patches
>  - 4 smb/client patches
>  - 3 cephfs patches
>  - 2 fuse patches
>  - one each for shmem, code, configfs, hostfs, procfs, ovl
> and a patch to d_add_ci() which affects xfs and ntfs.
> 
> Once all of those have landed I have 4 patches to remove deprecated interfaces.
> All of these can be found at
>    https://github.com/neilbrown/linux/commits/pdirops
> 
> 
> If all goes to plan, then in the next cycle a few patches will invert the
> lock order and remove LOOKUP_SHARED.  Then we can get on to the really
> fun stuff! (branch pdirops-next)
> 
> Thanks,
> NeilBrown
> 
> 
>  [PATCH v4 1/7] VFS: fix various typos in documentation for
>  [PATCH v4 2/7] VFS: enhance d_splice_alias() to handle hashed
>  [PATCH v4 3/7] VFS: introduce d_alloc_trylock()
>  [PATCH v4 4/7] VFS: add d_duplicate()
>  [PATCH v4 5/7] VFS: Add LOOKUP_SHARED flag.
>  [PATCH v4 6/7] VFS: add lockdep monitoring of DCACHE_PAR_LOOKUP lock.
>  [PATCH v4 7/7] VFS: reserve a d_flags bit for fs-specific usage
> 
> 


^ permalink raw reply	[flat|nested] 13+ messages in thread

* Re: [PATCH v4 0/7] VFS: prepare for changes to directory locking
  2026-09-04 21:48 [PATCH v4 0/7] VFS: prepare for changes to directory locking NeilBrown
                   ` (7 preceding siblings ...)
  2026-09-15 21:11 ` [PATCH v4 0/7] VFS: prepare for changes to directory locking NeilBrown
@ 2026-09-25 14:28 ` Christian Brauner
  8 siblings, 0 replies; 13+ messages in thread
From: Christian Brauner @ 2026-09-25 14:28 UTC (permalink / raw)
  To: NeilBrown
  Cc: Christian Brauner, Alexander Viro, Jan Kara, linux-fsdevel,
	Jeff Layton, Amir Goldstein, Miklos Szeredi, linux-kernel

On Sat, 05 Sep 2026 07:48:09 +1000, NeilBrown wrote:
> The only change in this v4 over v3 are to replace lock_sync with an
> acquire/release pair which builds correctly when lockdep isn't enabled,
> to add the 7th patch which I previously mentioned, and to include
> linux-kernel so that sashiko gets to see the series.
> 
> Previous intro:
> 
> [...]

Applied to the vfs-7.4.lookup branch of the vfs/vfs.git tree.
Patches in the vfs-7.4.lookup branch should appear in linux-next soon.

Please report any outstanding bugs that were missed during review in a
new review to the original patch series allowing us to drop it.

It's encouraged to provide Acked-bys and Reviewed-bys even though the
patch has now been applied. If possible patch trailers will be updated.

Note that commit hashes shown below are subject to change due to rebase,
trailer updates or similar. If in doubt, please check the listed branch.

tree:   https://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs.git
branch: vfs-7.4.lookup

[1/7] VFS: fix various typos in documentation for start_creating start_removing etc
      https://git.kernel.org/vfs/vfs/c/f8550827bf3d
[2/7] VFS: enhance d_splice_alias() to handle hashed dentries
      https://git.kernel.org/vfs/vfs/c/f1ac607f786f
[3/7] VFS: introduce d_alloc_trylock()
      https://git.kernel.org/vfs/vfs/c/17c7d109d8c9
[4/7] VFS: add d_duplicate()
      https://git.kernel.org/vfs/vfs/c/33d76519f0cc
[5/7] VFS: Add LOOKUP_SHARED flag.
      https://git.kernel.org/vfs/vfs/c/cc47a1ba1983
[6/7] VFS: add lockdep monitoring of DCACHE_PAR_LOOKUP lock.
      https://git.kernel.org/vfs/vfs/c/59492f9991dc
[7/7] VFS: reserve a d_flags bit for fs-specific usage
      https://git.kernel.org/vfs/vfs/c/fd96e30426ff

^ permalink raw reply	[flat|nested] 13+ messages in thread

* Re: [PATCH v4 6/7] VFS: add lockdep monitoring of DCACHE_PAR_LOOKUP lock.
  2026-09-04 21:48 ` [PATCH v4 6/7] VFS: add lockdep monitoring of DCACHE_PAR_LOOKUP lock NeilBrown
@ 2026-09-30 16:31   ` Borah, Chaitanya Kumar
  2026-09-30 21:08     ` NeilBrown
  0 siblings, 1 reply; 13+ messages in thread
From: Borah, Chaitanya Kumar @ 2026-09-30 16:31 UTC (permalink / raw)
  To: NeilBrown, Alexander Viro, Christian Brauner
  Cc: Jan Kara, linux-fsdevel, Jeff Layton, Amir Goldstein,
	Miklos Szeredi, linux-kernel, ravitejax.veesam, intel-gfx,
	intel-xe

Hello Neil,

On 9/5/2026 3:18 AM, NeilBrown wrote:
> From: NeilBrown <neil@brown.name>
> 
> DCACHE_PAR_LOOKUP acts like a lock in that threads can block waiting for
> it to clear.  As we plan to make changes to lock order for this lock,
> teach lockdep to monitor it so as to help detect bugs early.
> 
> As NFS allocates an in-lookup dentry to unlink a silly-renamed file, and
> completes the lookup in a different thread, we need interfaces to
> release and the acquire ownership of the lock.  This avoids lockdep
> complaining that a lock is still held on return to user-space.
> 

This seems to be causing regression in our linux-next CI [1] since
next-20260928.

<4>[   10.773931] ======================================================
<4>[   10.780181] WARNING: possible circular locking dependency detected
<4>[   10.786407] 7.3.0-rc5-next-20260928-next-20260928-g6375e61c01e9+ 
#1 Not tainted
<4>[   10.793773] ------------------------------------------------------
<4>[   10.800017] podman/794 is trying to acquire lock:
<4>[   10.804801] ffff888133cb8388 
(&type->i_mutex_dir_key#3){++++}-{4:4}, at: lookup_slow+0x31/0x60
<4>[   10.814849]
                   but task is already holding lock:
<4>[   10.820743] ffff88812dbd3048 (DCACHE_PAR_LOOKUP){+.+.}-{0:0}, at: 
__d_alloc_parallel+0x53a/0x920
<4>[   10.830976]
                   which lock already depends on the new lock.

Detailed log can be seen found in [2].

We confirmed that reverting the patch solves the issue.

Could you please check why the patch causes this regression and provide
a fix if necessary?

Regards
Chaitanya

[1] https://intel-gfx-ci.01.org/tree/linux-next/combined-alt.html?
[2] 
https://intel-gfx-ci.01.org/tree/linux-next/next-20260928/bat-arls-6/boot0.txt

--Bisect Logs--

git bisect start
# status: waiting for both good and bad commits
# bad: [6375e61c01e93e35ee7acd336a689ac1fae4b509] Add linux-next 
specific files for 20260928
git bisect bad 6375e61c01e93e35ee7acd336a689ac1fae4b509
# status: waiting for good commit(s), bad commit known
# good: [5dad87615c9861cfa366ca984b52f581e861df20] tty: add break_wait 
kernel-doc
git bisect good 5dad87615c9861cfa366ca984b52f581e861df20
# bad: [3bd2d86b7cc084d0a610760cfa40349497ea8920] Merge branch 
'for-next' of https://git.kernel.org/pub/scm/linux/kernel/git/rdma/rdma.git
git bisect bad 3bd2d86b7cc084d0a610760cfa40349497ea8920
# good: [3cc80c8dbcd8475e98ae86759750d0a140b063fd] Merge branch 
'for-next' of 
https://git.kernel.org/pub/scm/linux/kernel/git/gclement/mvebu.git
git bisect good 3cc80c8dbcd8475e98ae86759750d0a140b063fd
# bad: [3ac0b1643b3b216ce37d5ec4352fee8a01ecd37d] Merge branch 'fs-next' 
of linux-next
git bisect bad 3ac0b1643b3b216ce37d5ec4352fee8a01ecd37d
# good: [48a0400362e2648ba3ab56df8eceb782abfb7d17] Merge branch 
'for-next' of 
https://git.kernel.org/pub/scm/linux/kernel/git/geert/linux-m68k.git
git bisect good 48a0400362e2648ba3ab56df8eceb782abfb7d17
# good: [1ba62e451b02da61d4786a11da46501cd253cd2e] Merge branch 
'for-next' of 
https://git.kernel.org/pub/scm/linux/kernel/git/hubcap/linux.git
git bisect good 1ba62e451b02da61d4786a11da46501cd253cd2e
# bad: [b298f749884547ec4b746bdffc293fdd3b5a64ec] Merge branch 
'vfs-7.4.misc' into vfs.all
git bisect bad b298f749884547ec4b746bdffc293fdd3b5a64ec
# good: [cb70d8c361808b08d13dc4bdddb59cd92c48f84c] Merge branch 
'vfs-7.4.file' into vfs.all
git bisect good cb70d8c361808b08d13dc4bdddb59cd92c48f84c
# good: [d2b6b01e5969a762e92832cb6c9e2a470a842fcb] Merge branch 
'vfs-7.4.iomap' into vfs.all
git bisect good d2b6b01e5969a762e92832cb6c9e2a470a842fcb
# good: [a647a53bc89a046bee4a12a0a5627f2dfaaef807] dcache: report a 
Tasks-RCU quiescent state in dentry_kill()
git bisect good a647a53bc89a046bee4a12a0a5627f2dfaaef807
# good: [a2fb05f5133f83b2359b4b15d091bcf21e8b02bc] Merge patch series 
"kernfs: don't hold kernfs_rwsem across dir_emit()"
git bisect good a2fb05f5133f83b2359b4b15d091bcf21e8b02bc
# bad: [3879f51857325da9bf3cfb073280257cd16ae067] Merge patch series 
"VFS: prepare for changes to directory locking"
git bisect bad 3879f51857325da9bf3cfb073280257cd16ae067
# good: [17c7d109d8c9a6744b40ded06cc596cd6a80ec49] VFS: introduce 
d_alloc_trylock()
git bisect good 17c7d109d8c9a6744b40ded06cc596cd6a80ec49
# good: [cc47a1ba1983116ee7a720a79197eb5299ae5d31] VFS: Add 
LOOKUP_SHARED flag.
git bisect good cc47a1ba1983116ee7a720a79197eb5299ae5d31
# bad: [fd96e30426ff3e1352309ccf6843dc7fd9d38fde] VFS: reserve a d_flags 
bit for fs-specific usage
git bisect bad fd96e30426ff3e1352309ccf6843dc7fd9d38fde
# bad: [59492f9991dc19d5cc2c34295f1403d0f211e734] VFS: add lockdep 
monitoring of DCACHE_PAR_LOOKUP lock.
git bisect bad 59492f9991dc19d5cc2c34295f1403d0f211e734
# first bad commit: [59492f9991dc19d5cc2c34295f1403d0f211e734] VFS: add 
lockdep monitoring of DCACHE_PAR_LOOKUP lock.


> Signed-off-by: NeilBrown <neil@brown.name>
> ---
>   fs/dcache.c            | 15 +++++++++++++++
>   fs/nfs/unlink.c        |  3 +++
>   include/linux/dcache.h | 32 ++++++++++++++++++++++++++++++++
>   3 files changed, 50 insertions(+)
> 
> diff --git a/fs/dcache.c b/fs/dcache.c
> index cbd5738de168..83790c7a4dee 100644
> --- a/fs/dcache.c
> +++ b/fs/dcache.c
> @@ -1901,6 +1901,7 @@ EXPORT_SYMBOL(d_invalidate);
>    
>   static struct dentry *__d_alloc(struct super_block *sb, const struct qstr *name)
>   {
> +	static struct lock_class_key __lookup_key;
>   	struct dentry *dentry;
>   	char *dname;
>   	int err;
> @@ -1958,6 +1959,8 @@ static struct dentry *__d_alloc(struct super_block *sb, const struct qstr *name)
>   	dentry->waiters = NULL;
>   	INIT_HLIST_NODE(&dentry->d_sib);
>   
> +	lockdep_init_map(&dentry->lookup_map, "DCACHE_PAR_LOOKUP", &__lookup_key, 0);
> +
>   	if (dentry->d_op && dentry->d_op->d_init) {
>   		err = dentry->d_op->d_init(dentry);
>   		if (err) {
> @@ -2037,6 +2040,7 @@ struct dentry *d_duplicate(struct dentry *dentry)
>   		return ERR_PTR(-ENOMEM);
>   
>   	new->d_flags |= DCACHE_PAR_LOOKUP;
> +	lock_map_acquire_try(&new->lookup_map);
>   	spin_lock(&parent->d_lock);
>   	new->d_parent = dget_dlock(parent);
>   	hlist_add_head(&new->d_sib, &parent->d_children);
> @@ -2801,6 +2805,15 @@ static inline void end_dir_add(struct inode *dir, unsigned int n)
>   static void d_wait_lookup(struct dentry *dentry)
>   {
>   	if (likely(d_in_lookup(dentry))) {
> +		/*
> +		 * Tell lockdep we will wait for the lookup lock, after
> +		 * dropping ->d_lock, but won't actually take it.
> +		 */
> +		spin_release(&dentry->d_lock.dep_map, _THIS_IP_);
> +		lock_map_acquire(&dentry->lookup_map);
> +		lock_map_release(&dentry->lookup_map);
> +		spin_acquire(&dentry->d_lock.dep_map, 0, 1, _THIS_IP_);
> +
>   		dentry->d_flags |= DCACHE_LOOKUP_WAITERS;
>   		wait_var_event_spinlock(&dentry->d_flags,
>   					!d_in_lookup(dentry),
> @@ -2923,6 +2936,7 @@ struct dentry *__d_alloc_parallel(struct dentry *parent,
>   	}
>   	hlist_bl_add_head(&new->d_in_lookup_hash, b);
>   	hlist_bl_unlock(b);
> +	lock_map_acquire_try(&new->lookup_map);
>   	return new;
>   mismatch:
>   	spin_unlock(&dentry->d_lock);
> @@ -3021,6 +3035,7 @@ static void __d_lookup_unhash(struct dentry *dentry)
>   	b = in_lookup_hash(dentry->d_parent, dentry->d_name.hash);
>   	hlist_bl_lock(b);
>   	dentry->d_flags &= ~DCACHE_PAR_LOOKUP;
> +	lock_map_release(&dentry->lookup_map);
>   	__hlist_bl_del(&dentry->d_in_lookup_hash);
>   	hlist_bl_unlock(b);
>   	dentry->waiters = NULL;
> diff --git a/fs/nfs/unlink.c b/fs/nfs/unlink.c
> index b57cfaa4d516..c8d712204e64 100644
> --- a/fs/nfs/unlink.c
> +++ b/fs/nfs/unlink.c
> @@ -67,6 +67,7 @@ static void nfs_async_unlink_release(void *calldata)
>   	struct super_block *sb = dentry->d_sb;
>   
>   	up_read_non_owner(&NFS_I(d_inode(dentry->d_parent))->rmdir_sem);
> +	d_lookup_acquire(dentry);
>   	d_lookup_done(dentry);
>   	nfs_free_unlinkdata(data);
>   	dput(dentry);
> @@ -159,6 +160,8 @@ static int nfs_call_unlink(struct dentry *dentry, struct inode *inode, struct nf
>   		return ret;
>   	}
>   	data->dentry = alias;
> +	d_lookup_release(alias);
> +
>   	nfs_do_call_unlink(inode, data);
>   	return 1;
>   }
> diff --git a/include/linux/dcache.h b/include/linux/dcache.h
> index 2b7d99ec9306..e7e3ef05313b 100644
> --- a/include/linux/dcache.h
> +++ b/include/linux/dcache.h
> @@ -116,6 +116,8 @@ struct dentry {
>   					 * possible!
>   					 */
>   
> +	/* lockdep tracking of DCACHE_PAR_LOOKUP locks */
> +	struct lockdep_map		lookup_map;
>   	struct list_head d_lru;		/* LRU list */
>   	struct hlist_node d_sib;	/* child of parent list */
>   	struct hlist_head d_children;	/* our children */
> @@ -554,6 +556,36 @@ static inline int simple_positive(const struct dentry *dentry)
>   
>   unsigned long vfs_pressure_ratio(unsigned long val);
>   
> +/**
> + * d_lookup_release - release ownership of DCACHE_PAR_LOOKUP lock
> + * @dentry: dentry that is locked
> + *
> + * If an in-lookup dentry is to be passed to another thread which
> + * will drop the in-lookup lock, then d_lookup_release() must be called
> + * to tell lockdep that this thread no lock holds the lock.  The
> + * thread that receives the lock must call d_lookup_acquire() to
> + * acquire the lock.
> + */
> +static inline void d_lookup_release(struct dentry *dentry)
> +{
> +	if (d_in_lookup(dentry))
> +		lock_map_release(&dentry->lookup_map);
> +}
> +
> +/**
> + * d_lookup_acquire - acquire ownership of DCACHE_PAR_LOOKUP lock
> + * @dentry: dentry that is locked
> + *
> + * If an in-lookup dentry was passed to this thread, the
> + * d_lookup_acquire() must be called to tell lockdep that this
> + * thread now owns the DCACHE_PAR_LOOKUP lock.
> + */
> +static inline void d_lookup_acquire(struct dentry *dentry)
> +{
> +	if (d_in_lookup(dentry))
> +		lock_map_acquire_try(&dentry->lookup_map);
> +}
> +
>   /**
>    * d_inode - Get the actual inode of this dentry
>    * @dentry: The dentry to query


^ permalink raw reply	[flat|nested] 13+ messages in thread

* Re: [PATCH v4 6/7] VFS: add lockdep monitoring of DCACHE_PAR_LOOKUP lock.
  2026-09-30 16:31   ` Borah, Chaitanya Kumar
@ 2026-09-30 21:08     ` NeilBrown
  2026-10-01 10:39       ` Borah, Chaitanya Kumar
  0 siblings, 1 reply; 13+ messages in thread
From: NeilBrown @ 2026-09-30 21:08 UTC (permalink / raw)
  To: Borah, Chaitanya Kumar
  Cc: Alexander Viro, Christian Brauner, Jan Kara, linux-fsdevel,
	Jeff Layton, Amir Goldstein, Miklos Szeredi, linux-kernel,
	ravitejax.veesam, intel-gfx, intel-xe

On Thu, 01 Oct 2026, Borah, Chaitanya Kumar wrote:
> Hello Neil,
> 
> On 9/5/2026 3:18 AM, NeilBrown wrote:
> > From: NeilBrown <neil@brown.name>
> > 
> > DCACHE_PAR_LOOKUP acts like a lock in that threads can block waiting for
> > it to clear.  As we plan to make changes to lock order for this lock,
> > teach lockdep to monitor it so as to help detect bugs early.
> > 
> > As NFS allocates an in-lookup dentry to unlink a silly-renamed file, and
> > completes the lookup in a different thread, we need interfaces to
> > release and the acquire ownership of the lock.  This avoids lockdep
> > complaining that a lock is still held on return to user-space.
> > 
> 
> This seems to be causing regression in our linux-next CI [1] since
> next-20260928.
> 
> <4>[   10.773931] ======================================================
> <4>[   10.780181] WARNING: possible circular locking dependency detected
> <4>[   10.786407] 7.3.0-rc5-next-20260928-next-20260928-g6375e61c01e9+ 
> #1 Not tainted
> <4>[   10.793773] ------------------------------------------------------
> <4>[   10.800017] podman/794 is trying to acquire lock:
> <4>[   10.804801] ffff888133cb8388 
> (&type->i_mutex_dir_key#3){++++}-{4:4}, at: lookup_slow+0x31/0x60
> <4>[   10.814849]
>                    but task is already holding lock:
> <4>[   10.820743] ffff88812dbd3048 (DCACHE_PAR_LOOKUP){+.+.}-{0:0}, at: 
> __d_alloc_parallel+0x53a/0x920
> <4>[   10.830976]
>                    which lock already depends on the new lock.
> 
> Detailed log can be seen found in [2].
> 
> We confirmed that reverting the patch solves the issue.
> 
> Could you please check why the patch causes this regression and provide
> a fix if necessary?
> 
> Regards
> Chaitanya
> 
> [1] https://intel-gfx-ci.01.org/tree/linux-next/combined-alt.html?
> [2] 
> https://intel-gfx-ci.01.org/tree/linux-next/next-20260928/bat-arls-6/boot0.txt

Thanks for the report!
The log shows that overlayfs is involved.  overlayfs has special needs
with respect to lock nesting which I hadn't allowed for.

I think this patch should fix it.  Please let me know.

Thanks,
NeilBrown


diff --git a/fs/overlayfs/super.c b/fs/overlayfs/super.c
index e487597337e8..e1448dd776b4 100644
--- a/fs/overlayfs/super.c
+++ b/fs/overlayfs/super.c
@@ -163,10 +163,34 @@ static int ovl_dentry_weak_revalidate(struct dentry *dentry, unsigned int flags)
 	return ovl_dentry_revalidate_common(dentry, flags, true);
 }
 
+#ifdef CONFIG_LOCKDEP
+#define OVL_MAX_NESTING FILESYSTEM_MAX_STACK_DEPTH
+static int ovl_dentry_init(struct dentry *dentry)
+{
+	static struct lock_class_key ovl_d_lock_key[OVL_MAX_NESTING];
+	int depth = dentry->d_sb->s_stack_depth - 1;
+
+	if (WARN_ON_ONCE(depth < 0 || depth >= OVL_MAX_NESTING))
+		depth = 0;
+
+	/* Based on lockdep_set_class() */
+	lockdep_init_map_type(&dentry->lookup_map, "DCACHE_PAR_LOOKUP_OVL",
+			      &ovl_d_lock_key[depth], 0,
+			      dentry->lookup_map.wait_type_inner,
+			      dentry->lookup_map.wait_type_outer,
+			      dentry->lookup_map.lock_type);
+
+	return 0;
+}
+#else
+#define ovl_dentry_init NULL
+#endif
+
 static const struct dentry_operations ovl_dentry_operations = {
 	.d_real = ovl_d_real,
 	.d_revalidate = ovl_dentry_revalidate,
 	.d_weak_revalidate = ovl_dentry_weak_revalidate,
+	.d_init = ovl_dentry_init,
 };
 
 #if IS_ENABLED(CONFIG_UNICODE)
@@ -176,6 +200,7 @@ static const struct dentry_operations ovl_dentry_ci_operations = {
 	.d_weak_revalidate = ovl_dentry_weak_revalidate,
 	.d_hash = generic_ci_d_hash,
 	.d_compare = generic_ci_d_compare,
+	.d_init = ovl_dentry_init,
 };
 #endif
 

^ permalink raw reply	[flat|nested] 13+ messages in thread

* Re: [PATCH v4 6/7] VFS: add lockdep monitoring of DCACHE_PAR_LOOKUP lock.
  2026-09-30 21:08     ` NeilBrown
@ 2026-10-01 10:39       ` Borah, Chaitanya Kumar
  0 siblings, 0 replies; 13+ messages in thread
From: Borah, Chaitanya Kumar @ 2026-10-01 10:39 UTC (permalink / raw)
  To: NeilBrown
  Cc: Alexander Viro, Christian Brauner, Jan Kara, linux-fsdevel,
	Jeff Layton, Amir Goldstein, Miklos Szeredi, linux-kernel,
	ravitejax.veesam, intel-gfx, intel-xe



On 10/1/2026 2:38 AM, NeilBrown wrote:
> On Thu, 01 Oct 2026, Borah, Chaitanya Kumar wrote:
>> Hello Neil,
>>
>> On 9/5/2026 3:18 AM, NeilBrown wrote:
>>> From: NeilBrown <neil@brown.name>
>>>
>>> DCACHE_PAR_LOOKUP acts like a lock in that threads can block waiting for
>>> it to clear.  As we plan to make changes to lock order for this lock,
>>> teach lockdep to monitor it so as to help detect bugs early.
>>>
>>> As NFS allocates an in-lookup dentry to unlink a silly-renamed file, and
>>> completes the lookup in a different thread, we need interfaces to
>>> release and the acquire ownership of the lock.  This avoids lockdep
>>> complaining that a lock is still held on return to user-space.
>>>
>>
>> This seems to be causing regression in our linux-next CI [1] since
>> next-20260928.
>>
>> <4>[   10.773931] ======================================================
>> <4>[   10.780181] WARNING: possible circular locking dependency detected
>> <4>[   10.786407] 7.3.0-rc5-next-20260928-next-20260928-g6375e61c01e9+
>> #1 Not tainted
>> <4>[   10.793773] ------------------------------------------------------
>> <4>[   10.800017] podman/794 is trying to acquire lock:
>> <4>[   10.804801] ffff888133cb8388
>> (&type->i_mutex_dir_key#3){++++}-{4:4}, at: lookup_slow+0x31/0x60
>> <4>[   10.814849]
>>                     but task is already holding lock:
>> <4>[   10.820743] ffff88812dbd3048 (DCACHE_PAR_LOOKUP){+.+.}-{0:0}, at:
>> __d_alloc_parallel+0x53a/0x920
>> <4>[   10.830976]
>>                     which lock already depends on the new lock.
>>
>> Detailed log can be seen found in [2].
>>
>> We confirmed that reverting the patch solves the issue.
>>
>> Could you please check why the patch causes this regression and provide
>> a fix if necessary?
>>
>> Regards
>> Chaitanya
>>
>> [1] https://intel-gfx-ci.01.org/tree/linux-next/combined-alt.html?
>> [2]
>> https://intel-gfx-ci.01.org/tree/linux-next/next-20260928/bat-arls-6/boot0.txt
> 
> Thanks for the report!
> The log shows that overlayfs is involved.  overlayfs has special needs
> with respect to lock nesting which I hadn't allowed for.
> 
> I think this patch should fix it.  Please let me know.
> 

This works! Thank you. Hopefully it gets into linux-next soon.

Feel free to use

Tested-by: Chaitanya Kumar Borah <chaitanya.kumar.borah@intel.com>

> Thanks,
> NeilBrown
> 
> 
> diff --git a/fs/overlayfs/super.c b/fs/overlayfs/super.c
> index e487597337e8..e1448dd776b4 100644
> --- a/fs/overlayfs/super.c
> +++ b/fs/overlayfs/super.c
> @@ -163,10 +163,34 @@ static int ovl_dentry_weak_revalidate(struct dentry *dentry, unsigned int flags)
>   	return ovl_dentry_revalidate_common(dentry, flags, true);
>   }
>   
> +#ifdef CONFIG_LOCKDEP
> +#define OVL_MAX_NESTING FILESYSTEM_MAX_STACK_DEPTH
> +static int ovl_dentry_init(struct dentry *dentry)
> +{
> +	static struct lock_class_key ovl_d_lock_key[OVL_MAX_NESTING];
> +	int depth = dentry->d_sb->s_stack_depth - 1;
> +
> +	if (WARN_ON_ONCE(depth < 0 || depth >= OVL_MAX_NESTING))
> +		depth = 0;
> +
> +	/* Based on lockdep_set_class() */
> +	lockdep_init_map_type(&dentry->lookup_map, "DCACHE_PAR_LOOKUP_OVL",
> +			      &ovl_d_lock_key[depth], 0,
> +			      dentry->lookup_map.wait_type_inner,
> +			      dentry->lookup_map.wait_type_outer,
> +			      dentry->lookup_map.lock_type);
> +
> +	return 0;
> +}
> +#else
> +#define ovl_dentry_init NULL
> +#endif
> +
>   static const struct dentry_operations ovl_dentry_operations = {
>   	.d_real = ovl_d_real,
>   	.d_revalidate = ovl_dentry_revalidate,
>   	.d_weak_revalidate = ovl_dentry_weak_revalidate,
> +	.d_init = ovl_dentry_init,
>   };
>   
>   #if IS_ENABLED(CONFIG_UNICODE)
> @@ -176,6 +200,7 @@ static const struct dentry_operations ovl_dentry_ci_operations = {
>   	.d_weak_revalidate = ovl_dentry_weak_revalidate,
>   	.d_hash = generic_ci_d_hash,
>   	.d_compare = generic_ci_d_compare,
> +	.d_init = ovl_dentry_init,
>   };
>   #endif
>   


^ permalink raw reply	[flat|nested] 13+ messages in thread

end of thread, other threads:[~2026-10-01 10:39 UTC | newest]

Thread overview: 13+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-04 21:48 [PATCH v4 0/7] VFS: prepare for changes to directory locking NeilBrown
2026-09-04 21:48 ` [PATCH v4 1/7] VFS: fix various typos in documentation for start_creating start_removing etc NeilBrown
2026-09-04 21:48 ` [PATCH v4 2/7] VFS: enhance d_splice_alias() to handle hashed dentries NeilBrown
2026-09-04 21:48 ` [PATCH v4 3/7] VFS: introduce d_alloc_trylock() NeilBrown
2026-09-04 21:48 ` [PATCH v4 4/7] VFS: add d_duplicate() NeilBrown
2026-09-04 21:48 ` [PATCH v4 5/7] VFS: Add LOOKUP_SHARED flag NeilBrown
2026-09-04 21:48 ` [PATCH v4 6/7] VFS: add lockdep monitoring of DCACHE_PAR_LOOKUP lock NeilBrown
2026-09-30 16:31   ` Borah, Chaitanya Kumar
2026-09-30 21:08     ` NeilBrown
2026-10-01 10:39       ` Borah, Chaitanya Kumar
2026-09-04 21:48 ` [PATCH v4 7/7] VFS: reserve a d_flags bit for fs-specific usage NeilBrown
2026-09-15 21:11 ` [PATCH v4 0/7] VFS: prepare for changes to directory locking NeilBrown
2026-09-25 14:28 ` Christian Brauner

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®