From: Christian Brauner <brauner@kernel.org>
To: linux-fsdevel@vger.kernel.org
Cc: Alexander Viro <viro@zeniv.linux.org.uk>, Jan Kara <jack@suse.cz>,
linux-kernel@vger.kernel.org, Jeff Layton <jlayton@kernel.org>,
Jann Horn <jannh@google.com>, Neil Brown <neil@brown.name>,
Amir Goldstein <amir73il@gmail.com>,
"Christian Brauner (Amutable)" <brauner@kernel.org>,
stable@vger.kernel.org
Subject: [PATCH 19/21] readdir: take no inode lock on an immutable directory
Date: Fri, 02 Oct 2026 15:52:50 +0200 [thread overview]
Message-ID: <20261002-work-mount-fixes-4-v1-19-dd44b89d44ce@kernel.org> (raw)
In-Reply-To: <20261002-work-mount-fixes-4-v1-0-dd44b89d44ce@kernel.org>
iterate_dir() takes the directory's i_rwsem shared and holds it across
->iterate_shared(). For a directory that never has an entry and is
never removed the lock keeps nothing still, it only orders every reader
and every writer of that inode behind each other.
For the directory of a nullfs instance that matters. The instance of
the initial mount namespace is the root of every empty mount namespace
and the private instance is the root of every kernel thread, so one
inode is shared across users who have nothing else in common. And a
reader can hold the lock for as long as it likes: back the getdents()
buffer with a mapping of a file on a FUSE mount of your own, let the
copy of "." and ".." fault and let the server wait. Queue an exclusive
taker behind it, a mkdir() in that directory goes through start_dirop()
before the read-only mount is reported, and from then on every lookup
that misses the dcache in that directory, every create and every mount
on it waits until the server answers. One user of an empty mount
namespace stalls all the others.
Add FOP_IMMUTABLE for the file operations of a directory that never
changes and is never removed and let iterate_dir() skip the lock for
it. The flag never changes for a file, ->f_pos is protected by
f_pos_lock since directories are FMODE_ATOMIC_POS, IS_DEADDIR can't be
set on such a directory and neither touch_atime() nor fsnotify take
i_rwsem. Set it on the nullfs directory. The placeholder directories of
libfs never have an entry either but their owners remove them, so they
keep the lock.
Fixes: 9d4e752a24f7 ("namespace: allow creating empty mount namespaces")
Cc: stable@vger.kernel.org # v7.1+
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
fs/nullfs.c | 1 +
fs/readdir.c | 13 +++++++++----
include/linux/fs.h | 2 ++
3 files changed, 12 insertions(+), 4 deletions(-)
diff --git a/fs/nullfs.c b/fs/nullfs.c
index f76b87cf1841..bfc04bca3940 100644
--- a/fs/nullfs.c
+++ b/fs/nullfs.c
@@ -44,6 +44,7 @@ static const struct file_operations nullfs_dir_operations = {
.lock = nullfs_nolock,
.flock = nullfs_nolock,
.setlease = nullfs_nolease,
+ .fop_flags = FOP_IMMUTABLE,
};
static int nullfs_fs_fill_super(struct super_block *s, struct fs_context *fc)
diff --git a/fs/readdir.c b/fs/readdir.c
index 76bb1ae3a450..f2288841337a 100644
--- a/fs/readdir.c
+++ b/fs/readdir.c
@@ -87,6 +87,8 @@ EXPORT_SYMBOL(wrap_directory_iterator);
int iterate_dir(struct file *file, struct dir_context *ctx)
{
struct inode *inode = file_inode(file);
+ /* never an entry, never removed: nothing for the lock to keep still */
+ bool locked = !(file->f_op->fop_flags & FOP_IMMUTABLE);
int res = -ENOTDIR;
if (!file->f_op->iterate_shared)
@@ -100,9 +102,11 @@ int iterate_dir(struct file *file, struct dir_context *ctx)
if (res)
goto out;
- res = down_read_killable(&inode->i_rwsem);
- if (res)
- goto out;
+ if (locked) {
+ res = down_read_killable(&inode->i_rwsem);
+ if (res)
+ goto out;
+ }
res = -ENOENT;
if (!IS_DEADDIR(inode)) {
@@ -112,7 +116,8 @@ int iterate_dir(struct file *file, struct dir_context *ctx)
fsnotify_access(file);
file_accessed(file);
}
- inode_unlock_shared(inode);
+ if (locked)
+ inode_unlock_shared(inode);
out:
return res;
}
diff --git a/include/linux/fs.h b/include/linux/fs.h
index 784fa20217c4..deb411e86661 100644
--- a/include/linux/fs.h
+++ b/include/linux/fs.h
@@ -1978,6 +1978,8 @@ struct file_operations {
#define FOP_ASYNC_LOCK ((__force fop_flags_t)(1 << 6))
/* File system supports uncached read/write buffered IO */
#define FOP_DONTCACHE ((__force fop_flags_t)(1 << 7))
+/* Never changes and is never removed, readdir of a directory takes no lock */
+#define FOP_IMMUTABLE ((__force fop_flags_t)(1 << 8))
/* Wrap a directory iterator that needs exclusive inode access */
int wrap_directory_iterator(struct file *, struct dir_context *,
--
2.53.0
next prev parent reply other threads:[~2026-10-02 13:54 UTC|newest]
Thread overview: 24+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-02 13:52 [PATCH 00/21] mount: more bugfixes, trapped in the Black Lodge edition Christian Brauner
2026-10-02 13:52 ` [PATCH 01/21] namespace: unhash a dentry before detaching the mounts on it Christian Brauner
2026-10-02 13:52 ` [PATCH 02/21] namei: don't reveal overmounted entries in refwalk Christian Brauner
2026-10-02 13:52 ` [PATCH 03/21] fcntl: refuse F_SET_RW_HINT on an immutable inode Christian Brauner
2026-10-02 13:52 ` [PATCH 04/21] selftests/filesystems: check that an immutable inode takes no write hint Christian Brauner
2026-10-02 13:52 ` [PATCH 05/21] namespace: refuse an automount below a mount that is in no namespace Christian Brauner
2026-10-02 13:52 ` [PATCH 06/21] namespace: handle mount locking for automounts correctly Christian Brauner
2026-10-02 13:52 ` [PATCH 07/21] nullfs: don't update the access time Christian Brauner
2026-10-02 13:52 ` [PATCH 08/21] namespace: never expire a locked mount Christian Brauner
2026-10-02 13:52 ` [PATCH 09/21] namespace: keep the lock on a mount that a propagated copy is moved beneath Christian Brauner
2026-10-02 13:52 ` [PATCH 10/21] selftests/filesystems: check that a lock lands on the right mount and stays Christian Brauner
2026-10-02 13:52 ` [PATCH 11/21] selftests/filesystems: check the atime of the empty mount namespace root Christian Brauner
2026-10-02 13:52 ` [PATCH 12/21] selftests/filesystems: check that an automount below an overlay layer is refused Christian Brauner
2026-10-02 13:52 ` [PATCH 13/21] fhandle: decide the subtree check under mount_lock Christian Brauner
2026-10-02 13:52 ` [PATCH 14/21] namespace: keep the private nullfs instance in knullfs Christian Brauner
2026-10-02 13:52 ` [PATCH 15/21] namespace: nothing is mounted on or written through knullfs Christian Brauner
2026-10-02 13:52 ` [PATCH 16/21] fsnotify: let a filesystem refuse marks on its objects Christian Brauner
2026-10-02 14:26 ` Amir Goldstein
2026-10-02 13:52 ` [PATCH 17/21] nullfs: refuse file locks Christian Brauner
2026-10-02 13:52 ` [PATCH 18/21] nullfs: refuse leases and delegations Christian Brauner
2026-10-03 8:20 ` Jeff Layton
2026-10-02 13:52 ` Christian Brauner [this message]
2026-10-02 13:52 ` [PATCH 20/21] selftests/filesystems: add a helper that holds a readdir in a page fault Christian Brauner
2026-10-02 13:52 ` [PATCH 21/21] selftests/filesystems: check that reading the root of an empty mount namespace stalls nobody Christian Brauner
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261002-work-mount-fixes-4-v1-19-dd44b89d44ce@kernel.org \
--to=brauner@kernel.org \
--cc=amir73il@gmail.com \
--cc=jack@suse.cz \
--cc=jannh@google.com \
--cc=jlayton@kernel.org \
--cc=linux-fsdevel@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=neil@brown.name \
--cc=stable@vger.kernel.org \
--cc=viro@zeniv.linux.org.uk \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®