mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH 00/21] mount: more bugfixes, trapped in the Black Lodge edition
@ 2026-10-02 13:52 Christian Brauner
  2026-10-02 13:52 ` [PATCH 01/21] namespace: unhash a dentry before detaching the mounts on it Christian Brauner
                   ` (20 more replies)
  0 siblings, 21 replies; 23+ messages in thread
From: Christian Brauner @ 2026-10-02 13:52 UTC (permalink / raw)
  To: linux-fsdevel
  Cc: Alexander Viro, Jan Kara, linux-kernel, Jeff Layton, Jann Horn,
	Neil Brown, Amir Goldstein, Christian Brauner (Amutable),
	stable

More bugfixes from work in this area:

- unhash a dentry before detaching the mounts on it to avoid leaking
  sensitive data
- don't reveal overmounted entries in refwalk
- refuse F_SET_RW_HINT on an immutable inode
- refuse an automount below a mount that is in no namespace
- handle mount locking for automounts correctly
- don't update the access time on nullfs
- never expire a locked mount
- keep the lock on a mount that a propagated copy is moved beneath
- decide the subtree check under mount_lock
- keep the private nullfs instance in knullfs
- nothing is mounted on or written through knullfs
- let a filesystem refuse fsnotify marks on its objects
- refuse file locks on nullfs
- refuse leases and delegations on nullfs
- take no inode lock on an immutable directory

Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
Christian Brauner (21):
      namespace: unhash a dentry before detaching the mounts on it
      namei: don't reveal overmounted entries in refwalk
      fcntl: refuse F_SET_RW_HINT on an immutable inode
      selftests/filesystems: check that an immutable inode takes no write hint
      namespace: refuse an automount below a mount that is in no namespace
      namespace: handle mount locking for automounts correctly
      nullfs: don't update the access time
      namespace: never expire a locked mount
      namespace: keep the lock on a mount that a propagated copy is moved beneath
      selftests/filesystems: check that a lock lands on the right mount and stays
      selftests/filesystems: check the atime of the empty mount namespace root
      selftests/filesystems: check that an automount below an overlay layer is refused
      fhandle: decide the subtree check under mount_lock
      namespace: keep the private nullfs instance in knullfs
      namespace: nothing is mounted on or written through knullfs
      fsnotify: let a filesystem refuse marks on its objects
      nullfs: refuse file locks
      nullfs: refuse leases and delegations
      readdir: take no inode lock on an immutable directory
      selftests/filesystems: add a helper that holds a readdir in a page fault
      selftests/filesystems: check that reading the root of an empty mount namespace stalls nobody

 fs/fcntl.c                                         |   3 +
 fs/fhandle.c                                       |  20 +-
 fs/mount.h                                         |   1 +
 fs/namei.c                                         |  29 ++
 fs/namespace.c                                     |  78 ++--
 fs/notify/mark.c                                   |   4 +
 fs/nullfs.c                                        |  43 +-
 fs/readdir.c                                       |  13 +-
 include/linux/fs.h                                 |   3 +
 tools/testing/selftests/filesystems/.gitignore     |   1 +
 tools/testing/selftests/filesystems/Makefile       |   2 +-
 .../selftests/filesystems/empty_mntns/.gitignore   |   2 +
 .../selftests/filesystems/empty_mntns/Makefile     |   5 +-
 .../filesystems/empty_mntns/nullfs_atime_test.c    | 129 ++++++
 .../filesystems/empty_mntns/root_readdir_test.c    |  45 +++
 .../selftests/filesystems/overlayfs/.gitignore     |   1 +
 .../selftests/filesystems/overlayfs/Makefile       |   1 +
 .../filesystems/overlayfs/automount_in_layer.c     | 174 +++++++++
 tools/testing/selftests/filesystems/readdir_hold.h | 224 +++++++++++
 tools/testing/selftests/filesystems/rw_hint_test.c | 129 ++++++
 .../filesystems/umount_propagation/Makefile        |   2 +-
 .../umount_propagation/locked_mount_test.c         | 432 +++++++++++++++++++++
 22 files changed, 1305 insertions(+), 36 deletions(-)
---
base-commit: cbade031e2ebb3ca0a423b1804ce87908b27c229
change-id: 20261002-work-mount-fixes-4-b28d7d43102e


^ permalink raw reply	[flat|nested] 23+ messages in thread

* [PATCH 01/21] namespace: unhash a dentry before detaching the mounts on it
  2026-10-02 13:52 [PATCH 00/21] mount: more bugfixes, trapped in the Black Lodge edition Christian Brauner
@ 2026-10-02 13:52 ` Christian Brauner
  2026-10-02 13:52 ` [PATCH 02/21] namei: don't reveal overmounted entries in refwalk Christian Brauner
                   ` (19 subsequent siblings)
  20 siblings, 0 replies; 23+ messages in thread
From: Christian Brauner @ 2026-10-02 13:52 UTC (permalink / raw)
  To: linux-fsdevel
  Cc: Alexander Viro, Jan Kara, linux-kernel, Jeff Layton, Jann Horn,
	Neil Brown, Amir Goldstein, Christian Brauner (Amutable),
	stable

unlink(), rmdir() and rename() remove the entry from the filesystem,
call detach_mounts() on the entry with the inode locked and then call
d_delete() once the inode is unlocked.

The thing is that between detach_mounts() and d_delete() the dentry is
still hashed and positive. But after detach_mounts() nothing covers the
dentry anymore. A lookup that finds this dentry in the dcache can
uncover what the mounts hid.

That's a problem when unlinking files or directories that are
mountpoints in other mount namespaces. Everybody who had the underlying
entry covered can race the detach_mount() call until d_delete() has run.

The race window isn't all that small. It encompasses namespace_unlock()
with a full synchronize_rcu_expedited() grace period and the
inode_unlock() of the entry.

In my experiments three walkers that kept trying read an overmounted
file in 11938 of the 16140 unlinks that removed its mountpoint from a
bind mount of the filesystem within a minute.

Fun fact, d_invalidate() has the same ordering problem but gets it
right. It unhashes the dentry first and detaches the mounts afterwards.
Let's do the same in __detach_mounts():

- A lookup that hasn't found the dentry yet misses it in the dcache and
  waits for the directory lock that the caller holds until the name is
  gone for good.

- A lockless lookup that found it already fails the mount_lock check
  while the dentry still counts as a mountpoint and the d_seq check once
  it doesn't, and retries.

This fixes the lockless path. We still need to fix the reference count
lookup in a follow-up patch.

So d_drop() the dentry. The dentry stays positive and held but can't be
found anymore. Then proceed with the detach and unlink.

Reproducer:

The reads were counted with the program below, three walkers for 60 s
in a VM with 4 CPUs. The race window is stretched with the debug patch
pasted here. fs.detach_race_walk_us sleeps in lookup_fast() once it
found a mountpoint dentry and before its mounts are crossed.
fs.detach_race_unlink_us sleeps in vfs_unlink() between detach_mounts()
and d_delete().

  // SPDX-License-Identifier: GPL-2.0
  /*
   * unlink_covered: unlink a file that is a mountpoint in a detached copy of
   * its mount, against lookups that open it through that copy.
   *
   * A is a tmpfs with the file f1 (content SECRET_F) and the plain file
   * MARK_A. A' is an open_tree(OPEN_TREE_CLONE) copy of A with MARK_A bound
   * on A'/f1, held through an O_PATH fd on its root once the tree fd is
   * closed. Walkers open f1 through that fd while the driver unlinks A/f1,
   * where nothing is mounted on it. A walker may read MARK_A or get ENOENT.
   * A read of SECRET_F is a hit: the name was found after its mount was
   * gone. The lockless walks use openat2(RESOLVE_CACHED).
   *
   * usage: unlink_covered [-t seconds] [-w walkers]
   */
  #ifndef _GNU_SOURCE
  #define _GNU_SOURCE
  #endif
  #include <errno.h>
  #include <fcntl.h>
  #include <pthread.h>
  #include <stdatomic.h>
  #include <stdio.h>
  #include <stdlib.h>
  #include <string.h>
  #include <unistd.h>
  #include <sys/mount.h>
  #include <sys/stat.h>
  #include <sys/syscall.h>
  #include <linux/openat2.h>

  #ifndef __NR_open_tree
  #define __NR_open_tree 428
  #endif
  #ifndef __NR_move_mount
  #define __NR_move_mount 429
  #endif
  #ifndef __NR_openat2
  #define __NR_openat2 437
  #endif
  #ifndef OPEN_TREE_CLONE
  #define OPEN_TREE_CLONE 1
  #endif
  #ifndef OPEN_TREE_CLOEXEC
  #define OPEN_TREE_CLOEXEC O_CLOEXEC
  #endif
  #ifndef MOVE_MOUNT_F_EMPTY_PATH
  #define MOVE_MOUNT_F_EMPTY_PATH 0x00000004
  #endif

  #define WORK "/tmp/uc"
  #define ADIR WORK "/A"

  static int duration = 60, nwalkers = 3;
  static atomic_int stop, writer_waiting;
  static pthread_rwlock_t cur_lock = PTHREAD_RWLOCK_INITIALIZER;
  static int cur_fd = -1;		/* the root of A', -1 while there is none */
  static atomic_long n_unlink, n_walk, n_mark, n_enoent, n_other, n_secret,
  		   n_secret_cached;

  static void die(const char *what)
  {
  	fprintf(stderr, "FATAL %s: %s\n", what, strerror(errno));
  	exit(2);
  }

  static void put_file(int dfd, const char *name, const char *content)
  {
  	int fd = openat(dfd, name, O_CREAT | O_WRONLY | O_TRUNC | O_CLOEXEC,
  			0644);

  	if (fd < 0 || write(fd, content, strlen(content)) < 0)
  		die(name);
  	close(fd);
  }

  /* "plain/../" @n times, then f1: a longer walk that checks nothing on its way */
  static char *longpath(int n)
  {
  	char *p = malloc(n * 9 + 3), *q = p;

  	for (int i = 0; i < n; i++, q += 9)
  		memcpy(q, "plain/../", 9);
  	strcpy(q, "f1");
  	return p;
  }

  static void try_read(int dfd, const char *path, int cached)
  {
  	struct open_how how = { .flags = O_RDONLY | O_CLOEXEC,
  				.resolve = RESOLVE_CACHED };
  	char buf[32] = "";
  	long n;
  	int fd;

  	if (cached)
  		fd = syscall(__NR_openat2, dfd, path, &how, sizeof(how));
  	else
  		fd = openat(dfd, path, O_RDONLY | O_CLOEXEC);
  	atomic_fetch_add(&n_walk, 1);
  	if (fd < 0) {
  		if (errno == ENOENT)
  			atomic_fetch_add(&n_enoent, 1);
  		else if (!cached || errno != EAGAIN)
  			atomic_fetch_add(&n_other, 1);
  		return;
  	}
  	n = read(fd, buf, sizeof(buf) - 1);
  	close(fd);
  	if (n >= 6 && !strncmp(buf, "SECRET", 6)) {
  		if (cached)
  			atomic_fetch_add(&n_secret_cached, 1);
  		if (!atomic_fetch_add(&n_secret, 1))
  			printf("HIT: read \"%s\" through %s (%s walk)\n", buf,
  			       path, cached ? "lockless" : "any");
  	} else {
  		atomic_fetch_add(&n_mark, 1);
  	}
  }

  static void *walker(void *arg)
  {
  	unsigned int r = (long)arg * 2654435761u;

  	while (!atomic_load(&stop)) {
  		char *p_long, *p_short;
  		int dfd;

  		while (atomic_load(&writer_waiting) && !atomic_load(&stop))
  			usleep(20);
  		pthread_rwlock_rdlock(&cur_lock);
  		dfd = cur_fd;
  		if (dfd < 0) {
  			pthread_rwlock_unlock(&cur_lock);
  			usleep(100);
  			continue;
  		}
  		r = r * 1103515245u + 12345u;
  		p_long = longpath(1 + (r >> 8) % 400);
  		p_short = longpath(0);
  		try_read(dfd, p_long, 0);
  		try_read(dfd, p_short, 0);
  		try_read(dfd, p_long, 1);
  		pthread_rwlock_unlock(&cur_lock);
  		free(p_long);
  		free(p_short);
  	}
  	return NULL;
  }

  /* hand the walkers a new A' (or none), close the old one */
  static void publish(int fd)
  {
  	int old;

  	atomic_fetch_add(&writer_waiting, 1);
  	pthread_rwlock_wrlock(&cur_lock);
  	old = cur_fd;
  	cur_fd = fd;
  	pthread_rwlock_unlock(&cur_lock);
  	atomic_fetch_sub(&writer_waiting, 1);
  	if (old >= 0)
  		close(old);
  }

  static void *driver(void *arg __attribute__((unused)))
  {
  	int a;

  	if (mkdir(ADIR, 0755) && errno != EEXIST)
  		die("mkdir A");
  	if (mount("A", ADIR, "tmpfs", 0, "size=4M"))
  		die("mount A");
  	a = open(ADIR, O_PATH | O_DIRECTORY | O_CLOEXEC);
  	if (a < 0 || mkdirat(a, "plain", 0755))
  		die("A/plain");
  	put_file(a, "MARK_A", "MARK_A");
  	while (!atomic_load(&stop)) {
  		int t, m, fd_a;

  		put_file(a, "f1", "SECRET_F");
  		t = syscall(__NR_open_tree, a, "",
  			    OPEN_TREE_CLONE | OPEN_TREE_CLOEXEC | AT_EMPTY_PATH);
  		if (t < 0)
  			die("open_tree A");
  		m = syscall(__NR_open_tree, a, "MARK_A",
  			    OPEN_TREE_CLONE | OPEN_TREE_CLOEXEC);
  		if (m < 0)
  			die("open_tree MARK_A");
  		if (syscall(__NR_move_mount, m, "", t, "f1",
  			    MOVE_MOUNT_F_EMPTY_PATH))
  			die("move_mount");
  		close(m);
  		fd_a = openat(t, ".", O_PATH | O_DIRECTORY | O_CLOEXEC);
  		if (fd_a < 0)
  			die("open A'");
  		publish(fd_a);
  		usleep(200 + rand() % 1000);
  		close(t);			/* A' is unmounted, fd_a holds it */
  		usleep(200 + rand() % 1000);
  		if (unlinkat(a, "f1", 0))	/* through A, a plain file there */
  			die("unlink f1");
  		atomic_fetch_add(&n_unlink, 1);
  		usleep(rand() % 300);
  		publish(-1);
  	}
  	close(a);
  	umount2(ADIR, MNT_DETACH);
  	return NULL;
  }

  int main(int argc, char **argv)
  {
  	pthread_t d, *w;
  	int c, i;

  	setvbuf(stdout, NULL, _IOLBF, 0);
  	while ((c = getopt(argc, argv, "t:w:")) != -1) {
  		switch (c) {
  		case 't':
  			duration = atoi(optarg);
  			break;
  		case 'w':
  			nwalkers = atoi(optarg);
  			break;
  		default:
  			fprintf(stderr, "usage: unlink_covered [-t seconds] [-w walkers]\n");
  			return 2;
  		}
  	}
  	if (unshare(CLONE_NEWNS) ||
  	    mount(NULL, "/", NULL, MS_REC | MS_PRIVATE, NULL))
  		die("unshare");
  	if (mkdir(WORK, 0755) && errno != EEXIST)
  		die("mkdir");
  	w = calloc(nwalkers, sizeof(*w));
  	for (i = 0; i < nwalkers; i++)
  		if (pthread_create(&w[i], NULL, walker, (void *)(long)i))
  			die("pthread_create");
  	if (pthread_create(&d, NULL, driver, NULL))
  		die("pthread_create");
  	sleep(duration);
  	atomic_store(&stop, 1);
  	pthread_join(d, NULL);
  	for (i = 0; i < nwalkers; i++)
  		pthread_join(w[i], NULL);
  	printf("unlink_covered: %d s, %d walkers: unlinks %ld walks %ld mark %ld enoent %ld other %ld SECRET %ld (lockless %ld)\n",
  	       duration, nwalkers, atomic_load(&n_unlink), atomic_load(&n_walk),
  	       atomic_load(&n_mark), atomic_load(&n_enoent),
  	       atomic_load(&n_other), atomic_load(&n_secret),
  	       atomic_load(&n_secret_cached));
  	return atomic_load(&n_secret) ? 1 : 0;
  }

  diff --git a/fs/namei.c b/fs/namei.c
  --- a/fs/namei.c
  +++ b/fs/namei.c
  @@ -35,6 +35,7 @@
   #include <linux/fcntl.h>
   #include <linux/device_cgroup.h>
   #include <linux/fs_struct.h>
  +#include <linux/delay.h>
   #include <linux/posix_acl.h>
   #include <linux/hash.h>
   #include <linux/bitops.h>
  @@ -1205,9 +1206,26 @@ static int sysctl_protected_symlinks __read_mostly;
   static int sysctl_protected_hardlinks __read_mostly;
   static int sysctl_protected_fifos __read_mostly;
   static int sysctl_protected_regular __read_mostly;
  +/* debug: widen the two windows of the detach_mounts() race */
  +static int sysctl_detach_race_walk_us __read_mostly;
  +static int sysctl_detach_race_unlink_us __read_mostly;

   #ifdef CONFIG_SYSCTL
   static const struct ctl_table namei_sysctls[] = {
  +	{
  +		.procname	= "detach_race_walk_us",
  +		.data		= &sysctl_detach_race_walk_us,
  +		.maxlen		= sizeof(int),
  +		.mode		= 0644,
  +		.proc_handler	= proc_dointvec,
  +	},
  +	{
  +		.procname	= "detach_race_unlink_us",
  +		.data		= &sysctl_detach_race_unlink_us,
  +		.maxlen		= sizeof(int),
  +		.mode		= 0644,
  +		.proc_handler	= proc_dointvec,
  +	},
   	{
   		.procname	= "protected_symlinks",
   		.data		= &sysctl_protected_symlinks,
  @@ -1878,6 +1896,10 @@ static struct dentry *lookup_fast(struct nameidata *nd)
   		dentry = __d_lookup(parent, &nd->last);
   		if (unlikely(!dentry))
   			return NULL;
  +		/* debug: between finding a mountpoint and crossing its mounts */
  +		if (unlikely(sysctl_detach_race_walk_us) && d_mountpoint(dentry))
  +			usleep_range(sysctl_detach_race_walk_us,
  +				     sysctl_detach_race_walk_us + 10);
   		status = d_revalidate(nd->inode, &nd->last, dentry, nd->flags);
   	}
   	if (unlikely(status <= 0)) {
  @@ -5693,6 +5715,10 @@ int vfs_unlink(struct mnt_idmap *idmap, struct inode *dir,
   			if (!error) {
   				dont_mount(dentry);
   				detach_mounts(dentry);
  +				/* debug: the mounts are gone, d_delete() is still to come */
  +				if (unlikely(sysctl_detach_race_unlink_us))
  +					usleep_range(sysctl_detach_race_unlink_us,
  +						     sysctl_detach_race_unlink_us + 10);
   			}
   		}
   	}

Fixes: 8ed936b5671b ("vfs: Lazily remove mounts on unlinked files and directories.")
Cc: stable@vger.kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 fs/namespace.c | 5 +++++
 1 file changed, 5 insertions(+)

diff --git a/fs/namespace.c b/fs/namespace.c
index fcf42f192aae..e576a5d6eff0 100644
--- a/fs/namespace.c
+++ b/fs/namespace.c
@@ -1996,6 +1996,9 @@ static int do_umount(struct mount *mnt, int flags)
  * detach_mounts allows lazily unmounting those mounts instead of
  * leaking them.
  *
+ * The dentry is unhashed before the mounts go so that no lookup finds
+ * what they covered. The caller removes it for good afterwards.
+ *
  * The caller may hold dentry->d_inode->i_rwsem.
  */
 void __detach_mounts(struct dentry *dentry)
@@ -2009,6 +2012,8 @@ void __detach_mounts(struct dentry *dentry)
 	if (!lookup_mountpoint(dentry, &mp))
 		return;
 
+	/* the name goes first, what covered it goes second */
+	d_drop(dentry);
 	event++;
 	while (mp.node.next) {
 		mnt = hlist_entry(mp.node.next, struct mount, mnt_mp_list);

-- 
2.53.0


^ permalink raw reply	[flat|nested] 23+ messages in thread

* [PATCH 02/21] namei: don't reveal overmounted entries in refwalk
  2026-10-02 13:52 [PATCH 00/21] mount: more bugfixes, trapped in the Black Lodge edition Christian Brauner
  2026-10-02 13:52 ` [PATCH 01/21] namespace: unhash a dentry before detaching the mounts on it Christian Brauner
@ 2026-10-02 13:52 ` Christian Brauner
  2026-10-02 13:52 ` [PATCH 03/21] fcntl: refuse F_SET_RW_HINT on an immutable inode Christian Brauner
                   ` (18 subsequent siblings)
  20 siblings, 0 replies; 23+ messages in thread
From: Christian Brauner @ 2026-10-02 13:52 UTC (permalink / raw)
  To: linux-fsdevel
  Cc: Alexander Viro, Jan Kara, linux-kernel, Jeff Layton, Jann Horn,
	Neil Brown, Amir Goldstein, Christian Brauner (Amutable),
	stable

In rcuwalk the dentry is validated before it is accepted. step_into()
step_into() rechecks d_seq and __follow_mount_rcu() rechecks mount_lock.
An entry that gets unlinked in between causes the lookup to retry and
miss.

A refwalk doesn't do this. lookup_fast() takes a reference on the hashed
dentry and simply accepts it. So an unlink that happens after the
reference was taken isn't seen by refwalk. That's fine for a simple
file. We just happened to open it before it was unlinked, no problem.

For a mountpoint and specifically a locked mountpoint it very much
isn't. The unlink detaches all mounts and then removes the name. Any
refwalk that hasn't traversed the mounts yet simply reveals the
underlying entry. It's a very narrow window but it can be hit:

  reads of a covered file in 60 s, 3 walkers, ~15000 unlinks
    no widening            27
    that step + 200 us   3741

See the appended patch for a more reliable reproducer.

So check the entry after step_into(). unlink(), rmdir() and rename()
mark the dentry with dont_mount() before they detach the mounts and
remove the dentry.

So a dentry that is marked with DCACHE_CANT_MOUNT and is unhashed by the
time its mounts were looked at is a name that was unlinked under the
refwalk. DCACHE_CANT_MOUNT is read after the mounts were looked up and
it is set before they are detached by detach_mounts(). A refwalk that
missed the mounts will see DCACHE_CANT_MOUNT.

Plain d_unlinked() is fine. A rename takes the dentry off its hash chain
but ___d_drop() leaves d_hash.pprev set. So only __d_drop() and a rename
over the dentry unhash it. The only move that flips IS_ROOT splices in a
disconnected alias. That doesn't have DCACHE_CANT_MOUNT set.

Add the to step_into_slowpath(). ".." and LOOKUP_DOWN may legitimately
land on an unhashed directory. A refwalk that crossed onto a mount has
path.mnt different from nd->path.mnt. A dentry that a filesystem dropped
on its own (d_invalidate(), d_drop()) doesn't have the flag set and is
treated as before.

  // SPDX-License-Identifier: GPL-2.0
  /*
   * unlink_covered: unlink a file that is a mountpoint in a detached copy of
   * its mount, against walkers that open it through that copy.
   *
   * A is a tmpfs with the file f1 (content SECRET_F) and the plain file
   * MARK_A. A' is an open_tree(OPEN_TREE_CLONE) copy of A with MARK_A bound
   * on A'/f1, held through an O_PATH fd on its root once the tree fd is
   * closed. Walkers open f1 through that fd while the driver unlinks A/f1,
   * where nothing is mounted on it. A walker may read MARK_A or get ENOENT.
   * A read of SECRET_F is a hit: the name was found after its mount was
   * gone. The lockless walks use openat2(RESOLVE_CACHED).
   *
   * usage: unlink_covered [-t seconds] [-w walkers]
   */
  #ifndef _GNU_SOURCE
  #define _GNU_SOURCE
  #endif
  #include <errno.h>
  #include <fcntl.h>
  #include <pthread.h>
  #include <stdatomic.h>
  #include <stdio.h>
  #include <stdlib.h>
  #include <string.h>
  #include <unistd.h>
  #include <sys/mount.h>
  #include <sys/stat.h>
  #include <sys/syscall.h>
  #include <linux/openat2.h>

  #ifndef __NR_open_tree
  #define __NR_open_tree 428
  #endif
  #ifndef __NR_move_mount
  #define __NR_move_mount 429
  #endif
  #ifndef __NR_openat2
  #define __NR_openat2 437
  #endif
  #ifndef OPEN_TREE_CLONE
  #define OPEN_TREE_CLONE 1
  #endif
  #ifndef OPEN_TREE_CLOEXEC
  #define OPEN_TREE_CLOEXEC O_CLOEXEC
  #endif
  #ifndef MOVE_MOUNT_F_EMPTY_PATH
  #define MOVE_MOUNT_F_EMPTY_PATH 0x00000004
  #endif

  #define WORK "/tmp/uc"
  #define ADIR WORK "/A"

  static int duration = 60, nwalkers = 3;
  static atomic_int stop, writer_waiting;
  static pthread_rwlock_t cur_lock = PTHREAD_RWLOCK_INITIALIZER;
  static int cur_fd = -1;		/* the root of A', -1 while there is none */
  static atomic_long n_unlink, n_walk, n_mark, n_enoent, n_other, n_secret,
  		   n_secret_cached;

  static void die(const char *what)
  {
  	fprintf(stderr, "FATAL %s: %s\n", what, strerror(errno));
  	exit(2);
  }

  static void put_file(int dfd, const char *name, const char *content)
  {
  	int fd = openat(dfd, name, O_CREAT | O_WRONLY | O_TRUNC | O_CLOEXEC,
  			0644);

  	if (fd < 0 || write(fd, content, strlen(content)) < 0)
  		die(name);
  	close(fd);
  }

  /* "plain/../" @n times, then f1: a longer walk that checks nothing on its way */
  static char *longpath(int n)
  {
  	char *p = malloc(n * 9 + 3), *q = p;

  	for (int i = 0; i < n; i++, q += 9)
  		memcpy(q, "plain/../", 9);
  	strcpy(q, "f1");
  	return p;
  }

  static void try_read(int dfd, const char *path, int cached)
  {
  	struct open_how how = { .flags = O_RDONLY | O_CLOEXEC,
  				.resolve = RESOLVE_CACHED };
  	char buf[32] = "";
  	long n;
  	int fd;

  	if (cached)
  		fd = syscall(__NR_openat2, dfd, path, &how, sizeof(how));
  	else
  		fd = openat(dfd, path, O_RDONLY | O_CLOEXEC);
  	atomic_fetch_add(&n_walk, 1);
  	if (fd < 0) {
  		if (errno == ENOENT)
  			atomic_fetch_add(&n_enoent, 1);
  		else if (!cached || errno != EAGAIN)
  			atomic_fetch_add(&n_other, 1);
  		return;
  	}
  	n = read(fd, buf, sizeof(buf) - 1);
  	close(fd);
  	if (n >= 6 && !strncmp(buf, "SECRET", 6)) {
  		if (cached)
  			atomic_fetch_add(&n_secret_cached, 1);
  		if (!atomic_fetch_add(&n_secret, 1))
  			printf("HIT: read \"%s\" through %s (%s walk)\n", buf,
  			       path, cached ? "lockless" : "any");
  	} else {
  		atomic_fetch_add(&n_mark, 1);
  	}
  }

  static void *walker(void *arg)
  {
  	unsigned int r = (long)arg * 2654435761u;

  	while (!atomic_load(&stop)) {
  		char *p_long, *p_short;
  		int dfd;

  		while (atomic_load(&writer_waiting) && !atomic_load(&stop))
  			usleep(20);
  		pthread_rwlock_rdlock(&cur_lock);
  		dfd = cur_fd;
  		if (dfd < 0) {
  			pthread_rwlock_unlock(&cur_lock);
  			usleep(100);
  			continue;
  		}
  		r = r * 1103515245u + 12345u;
  		p_long = longpath(1 + (r >> 8) % 400);
  		p_short = longpath(0);
  		try_read(dfd, p_long, 0);
  		try_read(dfd, p_short, 0);
  		try_read(dfd, p_long, 1);
  		pthread_rwlock_unlock(&cur_lock);
  		free(p_long);
  		free(p_short);
  	}
  	return NULL;
  }

  /* hand the walkers a new A' (or none), close the old one */
  static void publish(int fd)
  {
  	int old;

  	atomic_fetch_add(&writer_waiting, 1);
  	pthread_rwlock_wrlock(&cur_lock);
  	old = cur_fd;
  	cur_fd = fd;
  	pthread_rwlock_unlock(&cur_lock);
  	atomic_fetch_sub(&writer_waiting, 1);
  	if (old >= 0)
  		close(old);
  }

  static void *driver(void *arg __attribute__((unused)))
  {
  	int a;

  	if (mkdir(ADIR, 0755) && errno != EEXIST)
  		die("mkdir A");
  	if (mount("A", ADIR, "tmpfs", 0, "size=4M"))
  		die("mount A");
  	a = open(ADIR, O_PATH | O_DIRECTORY | O_CLOEXEC);
  	if (a < 0 || mkdirat(a, "plain", 0755))
  		die("A/plain");
  	put_file(a, "MARK_A", "MARK_A");
  	while (!atomic_load(&stop)) {
  		int t, m, fd_a;

  		put_file(a, "f1", "SECRET_F");
  		t = syscall(__NR_open_tree, a, "",
  			    OPEN_TREE_CLONE | OPEN_TREE_CLOEXEC | AT_EMPTY_PATH);
  		if (t < 0)
  			die("open_tree A");
  		m = syscall(__NR_open_tree, a, "MARK_A",
  			    OPEN_TREE_CLONE | OPEN_TREE_CLOEXEC);
  		if (m < 0)
  			die("open_tree MARK_A");
  		if (syscall(__NR_move_mount, m, "", t, "f1",
  			    MOVE_MOUNT_F_EMPTY_PATH))
  			die("move_mount");
  		close(m);
  		fd_a = openat(t, ".", O_PATH | O_DIRECTORY | O_CLOEXEC);
  		if (fd_a < 0)
  			die("open A'");
  		publish(fd_a);
  		usleep(200 + rand() % 1000);
  		close(t);			/* A' is unmounted, fd_a holds it */
  		usleep(200 + rand() % 1000);
  		if (unlinkat(a, "f1", 0))	/* through A, a plain file there */
  			die("unlink f1");
  		atomic_fetch_add(&n_unlink, 1);
  		usleep(rand() % 300);
  		publish(-1);
  	}
  	close(a);
  	umount2(ADIR, MNT_DETACH);
  	return NULL;
  }

  int main(int argc, char **argv)
  {
  	pthread_t d, *w;
  	int c, i;

  	setvbuf(stdout, NULL, _IOLBF, 0);
  	while ((c = getopt(argc, argv, "t:w:")) != -1) {
  		switch (c) {
  		case 't':
  			duration = atoi(optarg);
  			break;
  		case 'w':
  			nwalkers = atoi(optarg);
  			break;
  		default:
  			fprintf(stderr, "usage: unlink_covered [-t seconds] [-w walkers]\n");
  			return 2;
  		}
  	}
  	if (unshare(CLONE_NEWNS) ||
  	    mount(NULL, "/", NULL, MS_REC | MS_PRIVATE, NULL))
  		die("unshare");
  	if (mkdir(WORK, 0755) && errno != EEXIST)
  		die("mkdir");
  	w = calloc(nwalkers, sizeof(*w));
  	for (i = 0; i < nwalkers; i++)
  		if (pthread_create(&w[i], NULL, walker, (void *)(long)i))
  			die("pthread_create");
  	if (pthread_create(&d, NULL, driver, NULL))
  		die("pthread_create");
  	sleep(duration);
  	atomic_store(&stop, 1);
  	pthread_join(d, NULL);
  	for (i = 0; i < nwalkers; i++)
  		pthread_join(w[i], NULL);
  	printf("unlink_covered: %d s, %d walkers: unlinks %ld walks %ld mark %ld enoent %ld other %ld SECRET %ld (lockless %ld)\n",
  	       duration, nwalkers, atomic_load(&n_unlink), atomic_load(&n_walk),
  	       atomic_load(&n_mark), atomic_load(&n_enoent),
  	       atomic_load(&n_other), atomic_load(&n_secret),
  	       atomic_load(&n_secret_cached));
  	return atomic_load(&n_secret) ? 1 : 0;
  }

  diff --git a/fs/namei.c b/fs/namei.c
  --- a/fs/namei.c
  +++ b/fs/namei.c
  @@ -35,6 +35,7 @@
   #include <linux/fcntl.h>
   #include <linux/device_cgroup.h>
   #include <linux/fs_struct.h>
  +#include <linux/delay.h>
   #include <linux/posix_acl.h>
   #include <linux/hash.h>
   #include <linux/bitops.h>
  @@ -1205,9 +1206,26 @@ static int sysctl_protected_symlinks __read_mostly;
   static int sysctl_protected_hardlinks __read_mostly;
   static int sysctl_protected_fifos __read_mostly;
   static int sysctl_protected_regular __read_mostly;
  +/* debug: widen the two windows of the detach_mounts() race */
  +static int sysctl_detach_race_walk_us __read_mostly;
  +static int sysctl_detach_race_unlink_us __read_mostly;

   #ifdef CONFIG_SYSCTL
   static const struct ctl_table namei_sysctls[] = {
  +	{
  +		.procname	= "detach_race_walk_us",
  +		.data		= &sysctl_detach_race_walk_us,
  +		.maxlen		= sizeof(int),
  +		.mode		= 0644,
  +		.proc_handler	= proc_dointvec,
  +	},
  +	{
  +		.procname	= "detach_race_unlink_us",
  +		.data		= &sysctl_detach_race_unlink_us,
  +		.maxlen		= sizeof(int),
  +		.mode		= 0644,
  +		.proc_handler	= proc_dointvec,
  +	},
   	{
   		.procname	= "protected_symlinks",
   		.data		= &sysctl_protected_symlinks,
  @@ -1878,6 +1896,10 @@ static struct dentry *lookup_fast(struct nameidata *nd)
   		dentry = __d_lookup(parent, &nd->last);
   		if (unlikely(!dentry))
   			return NULL;
  +		/* debug: between finding a mountpoint and crossing its mounts */
  +		if (unlikely(sysctl_detach_race_walk_us) && d_mountpoint(dentry))
  +			usleep_range(sysctl_detach_race_walk_us,
  +				     sysctl_detach_race_walk_us + 10);
   		status = d_revalidate(nd->inode, &nd->last, dentry, nd->flags);
   	}
   	if (unlikely(status <= 0)) {
  @@ -5693,6 +5715,10 @@ int vfs_unlink(struct mnt_idmap *idmap, struct inode *dir,
   			if (!error) {
   				dont_mount(dentry);
   				detach_mounts(dentry);
  +				/* debug: the mounts are gone, d_delete() is still to come */
  +				if (unlikely(sysctl_detach_race_unlink_us))
  +					usleep_range(sysctl_detach_race_unlink_us,
  +						     sysctl_detach_race_unlink_us + 10);
   			}
   		}
   	}

Fixes: 8ed936b5671b ("vfs: Lazily remove mounts on unlinked files and directories.")
Cc: stable@vger.kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 fs/namei.c | 29 +++++++++++++++++++++++++++++
 1 file changed, 29 insertions(+)

diff --git a/fs/namei.c b/fs/namei.c
index 20a6534ea3ef..c87c14ee1a25 100644
--- a/fs/namei.c
+++ b/fs/namei.c
@@ -2086,6 +2086,33 @@ static noinline const char *pick_link(struct nameidata *nd, struct path *link,
 	return NULL;
 }
 
+/*
+ * Be careful in case the dentry is unlinked or renamed. Any mounts
+ * stacked on top of it are going away. We need to make sure that we
+ * don't reveal the underyling dentry during refwalk. In rcuwalk we
+ * catch this via d_seq and another lookup for the name. Give the same
+ * guarantee in refwalk.
+ */
+static bool unlink_may_reveal(struct nameidata *nd, int flags,
+			      struct path *path, struct dentry *dentry)
+{
+	/* ".." and LOOKUP_DOWN may land on an unhashed directory */
+	if (flags & WALK_NOFOLLOW)
+		return false;
+	if (nd->flags & LOOKUP_REVAL)
+		return false;
+	/* We crossed onto a mount and the name led us here while it still existed */
+	if (path->mnt != nd->path.mnt)
+		return false;
+	/* only a name on its way out is flagged */
+	if (likely(!cant_mount(dentry)))
+		return false;
+	if (!d_unlinked(dentry))
+		return false;
+	dput(no_free_ptr(path->dentry));
+	return true;
+}
+
 /*
  * Do we need to follow links? We _really_ want to be able
  * to do this check without having to look at inode->i_op,
@@ -2115,6 +2142,8 @@ static noinline const char *step_into_slowpath(struct nameidata *nd, int flags,
 			if (unlikely(!inode))
 				return ERR_PTR(-ENOENT);
 		} else {
+			if (unlikely(unlink_may_reveal(nd, flags, &path, dentry)))
+				return ERR_PTR(-ESTALE);
 			dput(nd->path.dentry);
 			if (nd->path.mnt != path.mnt)
 				mntput(nd->path.mnt);

-- 
2.53.0


^ permalink raw reply	[flat|nested] 23+ messages in thread

* [PATCH 03/21] fcntl: refuse F_SET_RW_HINT on an immutable inode
  2026-10-02 13:52 [PATCH 00/21] mount: more bugfixes, trapped in the Black Lodge edition Christian Brauner
  2026-10-02 13:52 ` [PATCH 01/21] namespace: unhash a dentry before detaching the mounts on it Christian Brauner
  2026-10-02 13:52 ` [PATCH 02/21] namei: don't reveal overmounted entries in refwalk Christian Brauner
@ 2026-10-02 13:52 ` Christian Brauner
  2026-10-02 13:52 ` [PATCH 04/21] selftests/filesystems: check that an immutable inode takes no write hint Christian Brauner
                   ` (17 subsequent siblings)
  20 siblings, 0 replies; 23+ messages in thread
From: Christian Brauner @ 2026-10-02 13:52 UTC (permalink / raw)
  To: linux-fsdevel
  Cc: Alexander Viro, Jan Kara, linux-kernel, Jeff Layton, Jann Horn,
	Neil Brown, Amir Goldstein, Christian Brauner (Amutable),
	stable

An immutable inode is never written to so write hints are pointless.
Refuse the hint on an immutable inode the way setattr(), fallocate() and
etxattr() refuse their changes.

Fixes: c75b1d9421f8 ("fs: add fcntl() interface for setting/getting write life time hints")
Cc: stable@vger.kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 fs/fcntl.c | 3 +++
 1 file changed, 3 insertions(+)

diff --git a/fs/fcntl.c b/fs/fcntl.c
index c158f082f1da..9b71bc9d71d1 100644
--- a/fs/fcntl.c
+++ b/fs/fcntl.c
@@ -372,6 +372,9 @@ static long fcntl_set_rw_hint(struct file *file, unsigned long arg)
 	u64 __user *argp = (u64 __user *)arg;
 	u64 hint;
 
+	/* nothing is ever written to it */
+	if (IS_IMMUTABLE(inode))
+		return -EPERM;
 	if (!inode_owner_or_capable(file_mnt_idmap(file), inode))
 		return -EPERM;
 

-- 
2.53.0


^ permalink raw reply	[flat|nested] 23+ messages in thread

* [PATCH 04/21] selftests/filesystems: check that an immutable inode takes no write hint
  2026-10-02 13:52 [PATCH 00/21] mount: more bugfixes, trapped in the Black Lodge edition Christian Brauner
                   ` (2 preceding siblings ...)
  2026-10-02 13:52 ` [PATCH 03/21] fcntl: refuse F_SET_RW_HINT on an immutable inode Christian Brauner
@ 2026-10-02 13:52 ` Christian Brauner
  2026-10-02 13:52 ` [PATCH 05/21] namespace: refuse an automount below a mount that is in no namespace Christian Brauner
                   ` (16 subsequent siblings)
  20 siblings, 0 replies; 23+ messages in thread
From: Christian Brauner @ 2026-10-02 13:52 UTC (permalink / raw)
  To: linux-fsdevel
  Cc: Alexander Viro, Jan Kara, linux-kernel, Jeff Layton, Jann Horn,
	Neil Brown, Amir Goldstein, Christian Brauner (Amutable)

Check that an immutable inode takes no write hint.

Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 tools/testing/selftests/filesystems/.gitignore     |   1 +
 tools/testing/selftests/filesystems/Makefile       |   2 +-
 tools/testing/selftests/filesystems/rw_hint_test.c | 129 +++++++++++++++++++++
 3 files changed, 131 insertions(+), 1 deletion(-)

diff --git a/tools/testing/selftests/filesystems/.gitignore b/tools/testing/selftests/filesystems/.gitignore
index 9eb185fb2f9d..62f7b1c46649 100644
--- a/tools/testing/selftests/filesystems/.gitignore
+++ b/tools/testing/selftests/filesystems/.gitignore
@@ -7,3 +7,4 @@ anon_inode_test
 kernfs_test
 idmapped_tmpfile
 ustat_test
+rw_hint_test
diff --git a/tools/testing/selftests/filesystems/Makefile b/tools/testing/selftests/filesystems/Makefile
index 03be337c1f35..93e2cc9123b0 100644
--- a/tools/testing/selftests/filesystems/Makefile
+++ b/tools/testing/selftests/filesystems/Makefile
@@ -1,7 +1,7 @@
 # SPDX-License-Identifier: GPL-2.0
 
 CFLAGS += $(KHDR_INCLUDES)
-TEST_GEN_PROGS := devpts_pts file_stressor anon_inode_test kernfs_test fclog ustat_test
+TEST_GEN_PROGS := devpts_pts file_stressor anon_inode_test kernfs_test fclog ustat_test rw_hint_test
 TEST_GEN_PROGS += idmapped_tmpfile
 TEST_GEN_PROGS_EXTENDED := dnotify_test
 
diff --git a/tools/testing/selftests/filesystems/rw_hint_test.c b/tools/testing/selftests/filesystems/rw_hint_test.c
new file mode 100644
index 000000000000..d1930f82f63b
--- /dev/null
+++ b/tools/testing/selftests/filesystems/rw_hint_test.c
@@ -0,0 +1,129 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * F_SET_RW_HINT is refused on an immutable inode. Nothing is ever written
+ * to it and it may be shared with everybody, like a namespace file or the
+ * root of an empty mount namespace.
+ */
+#define _GNU_SOURCE
+#include <errno.h>
+#include <fcntl.h>
+#include <sched.h>
+#include <stdint.h>
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <unistd.h>
+#include <sys/ioctl.h>
+#include <sys/stat.h>
+#include <sys/wait.h>
+
+#include "../kselftest_harness.h"
+
+/* <linux/fs.h> and <linux/fcntl.h> don't mix with the libc headers */
+#ifndef FS_IOC_GETFLAGS
+#define FS_IOC_GETFLAGS		_IOR('f', 1, long)
+#define FS_IOC_SETFLAGS		_IOW('f', 2, long)
+#endif
+#ifndef FS_IMMUTABLE_FL
+#define FS_IMMUTABLE_FL		0x00000010
+#endif
+#ifndef F_LINUX_SPECIFIC_BASE
+#define F_LINUX_SPECIFIC_BASE	1024
+#endif
+#ifndef F_GET_RW_HINT
+#define F_GET_RW_HINT		(F_LINUX_SPECIFIC_BASE + 11)
+#define F_SET_RW_HINT		(F_LINUX_SPECIFIC_BASE + 12)
+#endif
+#ifndef RWH_WRITE_LIFE_SHORT
+#define RWH_WRITE_LIFE_SHORT	2
+#endif
+#ifndef UNSHARE_EMPTY_MNTNS
+#define UNSHARE_EMPTY_MNTNS	0x00100000
+#endif
+
+static int set_hint(int fd, uint64_t hint)
+{
+	return fcntl(fd, F_SET_RW_HINT, &hint);
+}
+
+static long get_hint(int fd)
+{
+	uint64_t hint;
+
+	if (fcntl(fd, F_GET_RW_HINT, &hint))
+		return -1;
+	return hint;
+}
+
+TEST(immutable_file)
+{
+	char path[] = "/tmp/rw_hint.XXXXXX";
+	int fd, flags;
+
+	if (geteuid())
+		SKIP(return, "test requires root");
+
+	fd = mkstemp(path);
+	ASSERT_GE(fd, 0);
+	unlink(path);
+	ASSERT_EQ(set_hint(fd, RWH_WRITE_LIFE_SHORT), 0);
+	EXPECT_EQ(get_hint(fd), RWH_WRITE_LIFE_SHORT);
+
+	if (ioctl(fd, FS_IOC_GETFLAGS, &flags)) {
+		close(fd);
+		SKIP(return, "no file attributes on this filesystem");
+	}
+	flags |= FS_IMMUTABLE_FL;
+	ASSERT_EQ(ioctl(fd, FS_IOC_SETFLAGS, &flags), 0);
+	EXPECT_EQ(set_hint(fd, RWH_WRITE_LIFE_SHORT), -1);
+	EXPECT_EQ(errno, EPERM);
+	flags &= ~FS_IMMUTABLE_FL;
+	ASSERT_EQ(ioctl(fd, FS_IOC_SETFLAGS, &flags), 0);
+	EXPECT_EQ(set_hint(fd, RWH_WRITE_LIFE_SHORT), 0);
+	close(fd);
+}
+
+TEST(namespace_file)
+{
+	int fd;
+
+	if (geteuid())
+		SKIP(return, "test requires root");
+
+	fd = open("/proc/self/ns/mnt", O_RDONLY | O_CLOEXEC);
+	ASSERT_GE(fd, 0);
+	EXPECT_EQ(set_hint(fd, RWH_WRITE_LIFE_SHORT), -1);
+	EXPECT_EQ(errno, EPERM);
+	close(fd);
+}
+
+TEST(empty_mntns_root)
+{
+	int status;
+	pid_t pid;
+
+	if (geteuid())
+		SKIP(return, "test requires root");
+
+	pid = fork();
+	ASSERT_GE(pid, 0);
+	if (pid == 0) {
+		int fd;
+
+		if (unshare(UNSHARE_EMPTY_MNTNS))
+			_exit(errno == EINVAL ? 100 : 1);
+		fd = open("/", O_RDONLY | O_DIRECTORY | O_CLOEXEC);
+		if (fd < 0)
+			_exit(2);
+		if (set_hint(fd, RWH_WRITE_LIFE_SHORT) == 0)
+			_exit(3);
+		_exit(errno == EPERM ? 0 : 4);
+	}
+	ASSERT_EQ(waitpid(pid, &status, 0), pid);
+	ASSERT_TRUE(WIFEXITED(status));
+	if (WEXITSTATUS(status) == 100)
+		SKIP(return, "UNSHARE_EMPTY_MNTNS not supported");
+	EXPECT_EQ(WEXITSTATUS(status), 0);
+}
+
+TEST_HARNESS_MAIN

-- 
2.53.0


^ permalink raw reply	[flat|nested] 23+ messages in thread

* [PATCH 05/21] namespace: refuse an automount below a mount that is in no namespace
  2026-10-02 13:52 [PATCH 00/21] mount: more bugfixes, trapped in the Black Lodge edition Christian Brauner
                   ` (3 preceding siblings ...)
  2026-10-02 13:52 ` [PATCH 04/21] selftests/filesystems: check that an immutable inode takes no write hint Christian Brauner
@ 2026-10-02 13:52 ` Christian Brauner
  2026-10-02 13:52 ` [PATCH 06/21] namespace: handle mount locking for automounts correctly Christian Brauner
                   ` (15 subsequent siblings)
  20 siblings, 0 replies; 23+ messages in thread
From: Christian Brauner @ 2026-10-02 13:52 UTC (permalink / raw)
  To: linux-fsdevel
  Cc: Alexander Viro, Jan Kara, linux-kernel, Jeff Layton, Jann Horn,
	Neil Brown, Amir Goldstein, Christian Brauner (Amutable),
	stable

It's possible to add autmounts even when the parent mount isn't in the
mount namespace of the caller. The only requirement we have is that the
parent mount namespace must not be NULL, i.e., unmounted.

Problem is that clone_private_mount() has MNT_NS_INTERNAL which makes
that trivially true. So that passes the test and attach_recursive_mnt()
accepts that as a mount point and funny enough, count_mounts()
dereferences MNT_NS_INTERNAL. The problem is it is an error pointer...

So we can reach this in userspace via fanotify. A filesystem mark on the
lower filesystem of an overlay reports paths on the layer clone and
reading the event hands out a descriptor on it. For example with debugfs
as the lower layer it goes kaboom:

  openat(evfd, "tracing", O_DIRECTORY)

  Oops: general protection fault
  KASAN: null-ptr-deref in range [0x1d0-0x1d7]
  RIP: 0010:count_mounts+0x35/0x200
   attach_recursive_mnt
   finish_automount

And since that sleeping beauty happens under namespace_sem held for
writing every mount operation on the system blocks from then on.
Congrats.

Use is_mounted() instead which rejects unmounted and internal mounts
alike. The open fails with EINVAL just as it did before
clone_private_mount() used MNT_NS_INTERNAL.

Fixes: df820f8de4e4 ("ovl: make private mounts longterm")
Cc: stable@vger.kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 fs/namespace.c | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/fs/namespace.c b/fs/namespace.c
index e576a5d6eff0..60b57572fc64 100644
--- a/fs/namespace.c
+++ b/fs/namespace.c
@@ -3806,7 +3806,7 @@ static int do_add_mount(struct mount *newmnt, const struct pinned_mountpoint *mp
 		if (!(mnt_flags & MNT_SHRINKABLE))
 			return -EINVAL;
 		/* ... and for those we'd better have mountpoint still alive */
-		if (!parent->mnt_ns)
+		if (!is_mounted(&parent->mnt))
 			return -EINVAL;
 	}
 

-- 
2.53.0


^ permalink raw reply	[flat|nested] 23+ messages in thread

* [PATCH 06/21] namespace: handle mount locking for automounts correctly
  2026-10-02 13:52 [PATCH 00/21] mount: more bugfixes, trapped in the Black Lodge edition Christian Brauner
                   ` (4 preceding siblings ...)
  2026-10-02 13:52 ` [PATCH 05/21] namespace: refuse an automount below a mount that is in no namespace Christian Brauner
@ 2026-10-02 13:52 ` Christian Brauner
  2026-10-02 13:52 ` [PATCH 07/21] nullfs: don't update the access time Christian Brauner
                   ` (14 subsequent siblings)
  20 siblings, 0 replies; 23+ messages in thread
From: Christian Brauner @ 2026-10-02 13:52 UTC (permalink / raw)
  To: linux-fsdevel
  Cc: Alexander Viro, Jan Kara, linux-kernel, Jeff Layton, Jann Horn,
	Neil Brown, Amir Goldstein, Christian Brauner (Amutable),
	stable

If mounts are propagated across user namespaces, attach_recursive_mnt()
locks every copy of the source mount to protect overmounts from
vanishing and revealing the underlying files or directories.

The user namespace is taken from the caller's mount namespaces since
this is where the mounts end up. Except, that's not always true.
Automounts may legitimately get popped in by tasks located in a
different mount and user namespace during path lookup.

Then check doesn't make sense at that point. The copy in the namespace
of the parent - which may be the host's - is now locked and the host
cannot change the flags of its own mounts anymore.

Here's the reproducer:

- host has debugfs mounted nosuid,nodev,noexec and shared
- tracefs automount below it is not yet active
- hand a child process a directory descriptor on that mount
- child process enters auser and mount namespace
- child process' namespace now holds a locked copy of the host's debugfs
  mount that receives propagation from it
- child process stats "tracing/." through the descriptor
- lookup runs on the host's mount so the automount lands below the host's mount
- propagation puts a copy below the child's copy
- both try to clear the flags on the mount they got, with a bind remount:

    host, on its own automount:            MS_REMOUNT|MS_BIND = EPERM
    child, on the copy in its namespace:   MS_REMOUNT|MS_BIND = 0

So the lock landed on the host's mount instead of the child's copy.
Congrats. So we need to compare with the owner of the namespace the
mount actually gets mounted on. For all regular cases that is the
caller's mount namespace and so nothing changes.

Detached trees in anonymous mount namespaces by be handed over via
SCM_RIGHTS or inherited in other ways on purpose so the attaching task's
mount namespace is authoritative, not the creator of the detached tree.

Fixes: 132c94e31b8b ("vfs: Carefully propogate mounts across user namespaces")
Cc: stable@vger.kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 fs/namespace.c | 18 +++++++++++++++++-
 1 file changed, 17 insertions(+), 1 deletion(-)

diff --git a/fs/namespace.c b/fs/namespace.c
index 60b57572fc64..27bf8665ed58 100644
--- a/fs/namespace.c
+++ b/fs/namespace.c
@@ -2604,11 +2604,11 @@ enum mnt_tree_flags_t {
 static int attach_recursive_mnt(struct mount *source_mnt,
 				const struct pinned_mountpoint *dest)
 {
-	struct user_namespace *user_ns = current->nsproxy->mnt_ns->user_ns;
 	struct mount *dest_mnt = dest->parent;
 	struct mountpoint *dest_mp = dest->mp;
 	HLIST_HEAD(tree_list);
 	struct mnt_namespace *ns = dest_mnt->mnt_ns;
+	struct user_namespace *user_ns = ns->user_ns;
 	struct pinned_mountpoint root = {};
 	struct mountpoint *shorter = NULL;
 	struct mount *child, *p;
@@ -2617,6 +2617,22 @@ static int attach_recursive_mnt(struct mount *source_mnt,
 	int err = 0;
 	bool moving = mnt_has_parent(source_mnt);
 
+	/*
+	 * A caller in an unprivileged mount namespaces may trigger an
+	 * automount and propagate locked mounts into privileged mount
+	 * namespaces. Take ownership from the target mount namespace.
+	 * It's equivalent for everything but the automount case.
+	 *
+	 * Detached trees in anonymous mount namespaces by be handed
+	 * over via SCM_RIGHTS or inherited in other ways on purpose
+	 * the attaching task's mount namespace is authoritative, not
+	 * the creator of the detached tree.
+	 */
+	if (is_anon_ns(ns))
+		user_ns = current->nsproxy->mnt_ns->user_ns;
+	else
+		user_ns = ns->user_ns;
+
 	/*
 	 * Preallocate a mountpoint in case the new mounts need to be
 	 * mounted beneath mounts on the same mountpoint.

-- 
2.53.0


^ permalink raw reply	[flat|nested] 23+ messages in thread

* [PATCH 07/21] nullfs: don't update the access time
  2026-10-02 13:52 [PATCH 00/21] mount: more bugfixes, trapped in the Black Lodge edition Christian Brauner
                   ` (5 preceding siblings ...)
  2026-10-02 13:52 ` [PATCH 06/21] namespace: handle mount locking for automounts correctly Christian Brauner
@ 2026-10-02 13:52 ` Christian Brauner
  2026-10-02 13:52 ` [PATCH 08/21] namespace: never expire a locked mount Christian Brauner
                   ` (13 subsequent siblings)
  20 siblings, 0 replies; 23+ messages in thread
From: Christian Brauner @ 2026-10-02 13:52 UTC (permalink / raw)
  To: linux-fsdevel
  Cc: Alexander Viro, Jan Kara, linux-kernel, Jeff Layton, Jann Horn,
	Neil Brown, Amir Goldstein, Christian Brauner (Amutable),
	stable

There's no point in updating access times of the nullfs instance. Raise
S_IMMUTABLE.

Fixes: 9d4e752a24f7 ("namespace: allow creating empty mount namespaces")
Cc: stable@vger.kernel.org # v7.1+
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 fs/nullfs.c | 4 ++--
 1 file changed, 2 insertions(+), 2 deletions(-)

diff --git a/fs/nullfs.c b/fs/nullfs.c
index e06352c7b2cc..40aa228bd81a 100644
--- a/fs/nullfs.c
+++ b/fs/nullfs.c
@@ -32,8 +32,8 @@ static int nullfs_fs_fill_super(struct super_block *s, struct fs_context *fc)
 	make_empty_dir_inode(inode);
 	simple_inode_init_ts(inode);
 	inode->i_ino	= 1;
-	/* ... and immutable. */
-	inode->i_flags |= S_IMMUTABLE;
+	/* ... and immutable, reading it leaves no trace either. */
+	inode->i_flags |= S_IMMUTABLE | S_NOATIME;
 
 	s->s_root = d_make_root(inode);
 	if (!s->s_root)

-- 
2.53.0


^ permalink raw reply	[flat|nested] 23+ messages in thread

* [PATCH 08/21] namespace: never expire a locked mount
  2026-10-02 13:52 [PATCH 00/21] mount: more bugfixes, trapped in the Black Lodge edition Christian Brauner
                   ` (6 preceding siblings ...)
  2026-10-02 13:52 ` [PATCH 07/21] nullfs: don't update the access time Christian Brauner
@ 2026-10-02 13:52 ` Christian Brauner
  2026-10-02 13:52 ` [PATCH 09/21] namespace: keep the lock on a mount that a propagated copy is moved beneath Christian Brauner
                   ` (12 subsequent siblings)
  20 siblings, 0 replies; 23+ messages in thread
From: Christian Brauner @ 2026-10-02 13:52 UTC (permalink / raw)
  To: linux-fsdevel
  Cc: Alexander Viro, Jan Kara, linux-kernel, Jeff Layton, Jann Horn,
	Neil Brown, Amir Goldstein, Christian Brauner (Amutable),
	stable

Locked mounts are special. They protect the underlying files and
directories from being revealed. do_umount() refuses to unmount locked
mounts but shrink_submounts() doesn't.

A shrinkable mount can become locked once the owner of a user namespace
puts it beneath a locked mount with MOVE_MOUNT_BENEATH. The locked
property now moves to the mount at the bottom making it possible to
unmount the top mount.

So umount() of an unlocked ancestor now expires the bottom mount first
and the covered directory is revealed:

  move_mount(c -> x/hidden, MOVE_MOUNT_BENEATH) = 0
  the cover after the lock moved:  umount2(x/hidden) = 0
  the holder of the lock:          umount2(x/hidden) = EINVAL
  the root of the copy, busy:      umount2(x) = EBUSY
  reads x/hidden/secret: "covered-by-root"

So leave a locked mount alone as it dies together with its parent. A
plain umount() of an unlocked mount with a locked one below it is EBUSY
from now on. It's the same for any other locked child. A lazy umount
still takes the whole tree.

mark_mounts_for_expiry() never sees a locked mount. lock_mnt_tree()
leaves a mount on an expiry list alone and nothing else puts a locked
one on such a list. Add an assert for this.

Fixes: 5ff9d8a65ce8 ("vfs: Lock in place mounts from more privileged users")
Fixes: c62a4766937e ("move_mount: transfer MNT_LOCKED")
Cc: stable@vger.kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 fs/namespace.c | 5 ++++-
 1 file changed, 4 insertions(+), 1 deletion(-)

diff --git a/fs/namespace.c b/fs/namespace.c
index 27bf8665ed58..e74e63466c24 100644
--- a/fs/namespace.c
+++ b/fs/namespace.c
@@ -4032,6 +4032,8 @@ void mark_mounts_for_expiry(struct list_head *mounts)
 	list_for_each_entry_safe(mnt, next, mounts, mnt_expire) {
 		if (!is_mounted(&mnt->mnt))
 			continue;
+		/* lock_mnt_tree() leaves expirable mounts alone */
+		VFS_WARN_ON_ONCE(IS_MNT_LOCKED(mnt));
 		if (!xchg(&mnt->mnt_expiry_mark, 1) ||
 			propagate_mount_busy(mnt, 1))
 			continue;
@@ -4058,7 +4060,8 @@ EXPORT_SYMBOL_GPL(mark_mounts_for_expiry);
  */
 static bool shrink_submount(struct mount *mnt)
 {
-	if (propagate_mount_busy(mnt, 1))
+	/* not the kernel's to remove either */
+	if (IS_MNT_LOCKED(mnt) || propagate_mount_busy(mnt, 1))
 		return false;
 	touch_mnt_namespace(mnt->mnt_ns);
 	umount_tree(mnt, UMOUNT_PROPAGATE|UMOUNT_SYNC);

-- 
2.53.0


^ permalink raw reply	[flat|nested] 23+ messages in thread

* [PATCH 09/21] namespace: keep the lock on a mount that a propagated copy is moved beneath
  2026-10-02 13:52 [PATCH 00/21] mount: more bugfixes, trapped in the Black Lodge edition Christian Brauner
                   ` (7 preceding siblings ...)
  2026-10-02 13:52 ` [PATCH 08/21] namespace: never expire a locked mount Christian Brauner
@ 2026-10-02 13:52 ` Christian Brauner
  2026-10-02 13:52 ` [PATCH 10/21] selftests/filesystems: check that a lock lands on the right mount and stays Christian Brauner
                   ` (11 subsequent siblings)
  20 siblings, 0 replies; 23+ messages in thread
From: Christian Brauner @ 2026-10-02 13:52 UTC (permalink / raw)
  To: linux-fsdevel
  Cc: Alexander Viro, Jan Kara, linux-kernel, Jeff Layton, Jann Horn,
	Neil Brown, Amir Goldstein, Christian Brauner (Amutable),
	stable

attach_recursive_mnt() transfers MNT_LOCKED from the top mount to the
mount that is moved beneath it with MOVE_MOUNT_BENEATH. This allows the
owner of a user namespace to replace its locked /proc or its root. The
mount beneath takes on the job of covering the underlying mountpoint
allowing the top mount to be unmounted.

Consider two mount namespaces:

(H) The host H has a shared mount P with a secret in P/d
(Z) Z is a user namespace made from M. Its copy of P receives
    propagation from the host's P and its copy of M's cover on P/d is
    locked Z's owner must not get to see P/d.

Now (H) mounts X on P/d. This propagates into (Z). The copy of X lands
beneath (Z)'s locked cover. The locked property is transfered from (Z)'s
cover to the copy of X propagated beneath it. The cover is now unlocked.

Now (H) unmounts X again. The copy of X in (Z) gets unmounted and the
covering mount is left unlocked on top of P/d. (Z) can now unmount it:

    Z: umount2(P/d)             = EINVAL            /* the cover is locked */
    H: mount X on P/d, umount X                     /* both propagate into Z /*
    Z: umount2(P/d)             = 0                 /* the cover is now unlocked */
    Z: read P/d/secret          = "covered-by-root" /* secret revealed */

So only transfer the locked property to the mount beneath for mounts the
caller has placed. A propagated copy that lands beneath a locked mount
is locked as well so that the mount at the bottom of the stack carries a
lock the way every check expects. The mount on top of it remains locked
to ensure that it keeps covering even if the propagated mount is
unmounted again.

Fixes: c62a4766937e ("move_mount: transfer MNT_LOCKED")
Cc: stable@vger.kernel.org # v7.1+
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 fs/namespace.c | 14 ++++++++------
 1 file changed, 8 insertions(+), 6 deletions(-)

diff --git a/fs/namespace.c b/fs/namespace.c
index e74e63466c24..bb0183ec2aaf 100644
--- a/fs/namespace.c
+++ b/fs/namespace.c
@@ -2712,15 +2712,17 @@ static int attach_recursive_mnt(struct mount *source_mnt,
 			/*
 			 * If @q was locked it was meant to hide
 			 * whatever was under it. Let @child take over
-			 * that job and lock it, then we can unlock @q.
-			 * That'll allow another namespace to shed @q
-			 * and reveal @child. Clearly, that mounter
-			 * consented to this by not severing the mount
-			 * relationship. Otherwise, what's the point.
+			 * that job and lock it. If @child is the mount
+			 * the caller placed we can then unlock @q:
+			 * nothing another namespace does removes it
+			 * again. A propagated copy goes away when the
+			 * mounter of the original unmounts it, so @q
+			 * keeps its lock.
 			 */
 			if (IS_MNT_LOCKED(q)) {
 				child->mnt.mnt_flags |= MNT_LOCKED;
-				q->mnt.mnt_flags &= ~MNT_LOCKED;
+				if (child == source_mnt)
+					q->mnt.mnt_flags &= ~MNT_LOCKED;
 			}
 			mnt_change_mountpoint(r, mp, q);
 		}

-- 
2.53.0


^ permalink raw reply	[flat|nested] 23+ messages in thread

* [PATCH 10/21] selftests/filesystems: check that a lock lands on the right mount and stays
  2026-10-02 13:52 [PATCH 00/21] mount: more bugfixes, trapped in the Black Lodge edition Christian Brauner
                   ` (8 preceding siblings ...)
  2026-10-02 13:52 ` [PATCH 09/21] namespace: keep the lock on a mount that a propagated copy is moved beneath Christian Brauner
@ 2026-10-02 13:52 ` Christian Brauner
  2026-10-02 13:52 ` [PATCH 11/21] selftests/filesystems: check the atime of the empty mount namespace root Christian Brauner
                   ` (10 subsequent siblings)
  20 siblings, 0 replies; 23+ messages in thread
From: Christian Brauner @ 2026-10-02 13:52 UTC (permalink / raw)
  To: linux-fsdevel
  Cc: Alexander Viro, Jan Kara, linux-kernel, Jeff Layton, Jann Horn,
	Neil Brown, Amir Goldstein, Christian Brauner (Amutable)

Check MNT_LOCKED:

- A child in a user namespace of its own triggers the tracefs automount
  through a directory descriptor on the host's debugfs mount. The copy
  in the child's namespace has to be locked and the host's mount unlocked.

- A child moves a bind of that automount beneath a locked covering mount,
  unmounts the covering mount and asks for the umount of an unlocked
  ancestor. This must not expire the mount which holds the lock now.

- The host mounts and unmounts on a directory that a locked mount covers
  in a user namespace further down the propagation chain. That cover has
  to be locked afterwards as it was before.

Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 .../filesystems/umount_propagation/Makefile        |   2 +-
 .../umount_propagation/locked_mount_test.c         | 432 +++++++++++++++++++++
 2 files changed, 433 insertions(+), 1 deletion(-)

diff --git a/tools/testing/selftests/filesystems/umount_propagation/Makefile b/tools/testing/selftests/filesystems/umount_propagation/Makefile
index eb85612abf8d..31f783a92b23 100644
--- a/tools/testing/selftests/filesystems/umount_propagation/Makefile
+++ b/tools/testing/selftests/filesystems/umount_propagation/Makefile
@@ -1,5 +1,5 @@
 # SPDX-License-Identifier: GPL-2.0
-TEST_GEN_PROGS := umount_propagation_test shrink_submounts_test
+TEST_GEN_PROGS := umount_propagation_test shrink_submounts_test locked_mount_test
 
 CFLAGS += -Wall -O2 -g $(KHDR_INCLUDES)
 
diff --git a/tools/testing/selftests/filesystems/umount_propagation/locked_mount_test.c b/tools/testing/selftests/filesystems/umount_propagation/locked_mount_test.c
new file mode 100644
index 000000000000..0799b2ba545d
--- /dev/null
+++ b/tools/testing/selftests/filesystems/umount_propagation/locked_mount_test.c
@@ -0,0 +1,432 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * MNT_LOCKED keeps the owner of a user namespace from revealing what a
+ * mount covers. The lock has to be set on the copy in that namespace and
+ * only there, it has to survive the expiry of a mount placed beneath the
+ * locked one, and it has to survive the propagated umount of such a copy.
+ */
+#define _GNU_SOURCE
+#include <errno.h>
+#include <fcntl.h>
+#include <sched.h>
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <unistd.h>
+#include <sys/mount.h>
+#include <sys/prctl.h>
+#include <sys/stat.h>
+#include <sys/syscall.h>
+#include <sys/wait.h>
+
+#include "../../kselftest_harness.h"
+
+#ifndef MOVE_MOUNT_F_EMPTY_PATH
+#define MOVE_MOUNT_F_EMPTY_PATH	0x00000004
+#endif
+#ifndef MOVE_MOUNT_BENEATH
+#define MOVE_MOUNT_BENEATH	0x00000200
+#endif
+
+#define DIR_LEN		64
+#define PATH_LEN	192
+#define FLAGS		(MS_NOSUID | MS_NODEV | MS_NOEXEC)
+
+/* exit codes of the children, each test says what they mean */
+enum {
+	CHILD_OK,
+	CHILD_SETUP,
+	CHILD_STEP1,
+	CHILD_STEP2,
+	CHILD_STEP3,
+	CHILD_STEP4,
+};
+
+static int write_file(const char *path, const char *s)
+{
+	ssize_t n = -1;
+	int fd;
+
+	fd = open(path, O_WRONLY | O_CLOEXEC);
+	if (fd >= 0) {
+		n = write(fd, s, strlen(s));
+		close(fd);
+	}
+	return n == (ssize_t)strlen(s) ? 0 : -1;
+}
+
+static int create_file(const char *path, const char *s)
+{
+	ssize_t n = -1;
+	int fd;
+
+	fd = open(path, O_WRONLY | O_CREAT | O_EXCL | O_CLOEXEC, 0644);
+	if (fd >= 0) {
+		n = write(fd, s, strlen(s));
+		close(fd);
+	}
+	return n == (ssize_t)strlen(s) ? 0 : -1;
+}
+
+/* Root in a new user namespace with a copy of the mount namespace. */
+static int enter_userns(void)
+{
+	uid_t uid = getuid();
+	gid_t gid = getgid();
+	char map[32];
+
+	prctl(PR_SET_DUMPABLE, 1);
+	if (unshare(CLONE_NEWUSER | CLONE_NEWNS))
+		return -1;
+	if (write_file("/proc/self/setgroups", "deny") && errno != ENOENT)
+		return -1;
+	snprintf(map, sizeof(map), "0 %d 1", uid);
+	if (write_file("/proc/self/uid_map", map))
+		return -1;
+	snprintf(map, sizeof(map), "0 %d 1", gid);
+	if (write_file("/proc/self/gid_map", map))
+		return -1;
+	return setgid(0) || setuid(0) ? -1 : 0;
+}
+
+static int wait_child(pid_t pid)
+{
+	int status;
+
+	if (waitpid(pid, &status, 0) != pid || !WIFEXITED(status))
+		return -1;
+	return WEXITSTATUS(status);
+}
+
+static int wait_byte(int fd)
+{
+	char c;
+
+	return read(fd, &c, 1) == 1 ? 0 : -1;
+}
+
+static int send_byte(int fd)
+{
+	return write(fd, "x", 1) == 1 ? 0 : -1;
+}
+
+/* A read of @path fails: the file is covered. */
+static bool covered(const char *path)
+{
+	int fd = open(path, O_RDONLY | O_CLOEXEC);
+
+	if (fd >= 0)
+		close(fd);
+	return fd < 0;
+}
+
+FIXTURE(locked_mount) {
+	char base[DIR_LEN];
+	bool tracing;		/* debugfs with the tracefs automount is there */
+};
+
+FIXTURE_SETUP(locked_mount)
+{
+	char p[PATH_LEN];
+	struct stat st;
+
+	if (geteuid())
+		SKIP(return, "test requires root");
+
+	snprintf(self->base, sizeof(self->base), "/tmp/locked_mount.XXXXXX");
+	ASSERT_NE(mkdtemp(self->base), NULL);
+	ASSERT_EQ(unshare(CLONE_NEWNS), 0);
+	ASSERT_EQ(mount("", "/", NULL, MS_REC | MS_PRIVATE, NULL), 0);
+	ASSERT_EQ(mount("tmpfs", self->base, "tmpfs", 0, "mode=0755"), 0);
+
+	/* an automount that a user can name: the tracefs below debugfs */
+	snprintf(p, sizeof(p), "%s/dbg", self->base);
+	ASSERT_EQ(mkdir(p, 0755), 0);
+	self->tracing = !mount("debugfs", p, "debugfs", FLAGS, NULL);
+	if (self->tracing) {
+		/* the cases trigger it themselves */
+		snprintf(p, sizeof(p), "%s/dbg/tracing", self->base);
+		self->tracing = !fstatat(AT_FDCWD, p, &st, AT_NO_AUTOMOUNT) &&
+				S_ISDIR(st.st_mode);
+	}
+}
+
+FIXTURE_TEARDOWN(locked_mount)
+{
+	umount2(self->base, MNT_DETACH);
+	rmdir(self->base);
+}
+
+/*
+ * The child keeps a directory descriptor on the host's debugfs mount, moves
+ * to a user namespace of its own and triggers the automount through the
+ * descriptor. The mount goes below the host's mount and propagates into the
+ * child's copy. The child's copy has to be locked, the host's not.
+ */
+static int automount_child(const char *base, int dfd, int to_host, int from_host)
+{
+	char p[PATH_LEN];
+	struct stat st;
+
+	if (enter_userns())
+		return CHILD_SETUP;
+	if (fstatat(dfd, "tracing/.", &st, 0))
+		return CHILD_STEP1;
+	if (send_byte(to_host) || wait_byte(from_host))
+		return CHILD_SETUP;
+	/* the flags of the copy are locked: EPERM */
+	snprintf(p, sizeof(p), "%s/dbg/tracing", base);
+	if (!mount(NULL, p, NULL, MS_REMOUNT | MS_BIND, NULL) || errno != EPERM)
+		return CHILD_STEP2;
+	return CHILD_OK;
+}
+
+TEST_F(locked_mount, automount_locked_in_the_triggering_namespace)
+{
+	int to_host[2], from_host[2], dfd, ret;
+	char p[PATH_LEN];
+	pid_t pid;
+
+	if (!self->tracing)
+		SKIP(return, "test requires debugfs with the tracefs automount");
+
+	snprintf(p, sizeof(p), "%s/dbg", self->base);
+	ASSERT_EQ(mount(NULL, p, NULL, MS_SHARED, NULL), 0);
+	dfd = open(p, O_RDONLY | O_DIRECTORY | O_CLOEXEC);
+	ASSERT_GE(dfd, 0);
+	ASSERT_EQ(pipe(to_host), 0);
+	ASSERT_EQ(pipe(from_host), 0);
+
+	pid = fork();
+	ASSERT_GE(pid, 0);
+	if (pid == 0)
+		_exit(automount_child(self->base, dfd, to_host[1], from_host[0]));
+	close(dfd);
+	ASSERT_EQ(wait_byte(to_host[0]), 0);
+
+	/* the host's own mount isn't locked: the flags can go */
+	snprintf(p, sizeof(p), "%s/dbg/tracing", self->base);
+	EXPECT_EQ(mount(NULL, p, NULL, MS_REMOUNT | MS_BIND, NULL), 0);
+
+	ASSERT_EQ(send_byte(from_host[1]), 0);
+	ret = wait_child(pid);
+	TH_LOG("child exit code %d", ret);
+	EXPECT_EQ(ret, CHILD_OK);
+}
+
+/*
+ * A copy of a tree with a locked cover. The child puts a shrinkable mount
+ * beneath the cover, which hands the lock down, unmounts the cover and then
+ * asks for the umount of an unlocked ancestor. That expires the shrinkable
+ * mount on a kernel that doesn't look at the lock, and the covered
+ * directory is bare.
+ */
+static int expiry_child(const char *base)
+{
+	char srv[PATH_LEN], shr[PATH_LEN], x[PATH_LEN], c[PATH_LEN];
+	char hidden[PATH_LEN], secret[PATH_LEN];
+
+	snprintf(srv, sizeof(srv), "%s/srv", base);
+	snprintf(shr, sizeof(shr), "%s/shr", base);
+	snprintf(x, sizeof(x), "%s/x", base);
+	snprintf(c, sizeof(c), "%s/c", base);
+	snprintf(hidden, sizeof(hidden), "%s/x/hidden", base);
+	snprintf(secret, sizeof(secret), "%s/x/hidden/secret", base);
+
+	if (enter_userns())
+		return CHILD_SETUP;
+	if (mount(srv, x, NULL, MS_BIND | MS_REC, NULL) ||
+	    mount(shr, c, NULL, MS_BIND, NULL))
+		return CHILD_SETUP;
+	/* the cover is locked */
+	if (!umount2(hidden, 0) || errno != EINVAL)
+		return CHILD_STEP1;
+	if (syscall(__NR_move_mount, AT_FDCWD, c, AT_FDCWD, hidden, MOVE_MOUNT_BENEATH))
+		return CHILD_STEP2;
+	/* the lock moved down, the cover may go */
+	if (umount2(hidden, 0))
+		return CHILD_STEP3;
+	/* the holder of the lock may not, in any way */
+	if (!umount2(hidden, 0) || errno != EINVAL)
+		return CHILD_STEP4;
+	if (chdir(x))
+		return CHILD_SETUP;
+	umount2(x, 0);
+	return covered(secret) ? CHILD_OK : CHILD_STEP4;
+}
+
+TEST_F(locked_mount, expiry_leaves_a_locked_mount_alone)
+{
+	char p[PATH_LEN], q[PATH_LEN];
+	pid_t pid;
+	int ret;
+
+	if (!self->tracing)
+		SKIP(return, "test requires debugfs with the tracefs automount");
+
+	snprintf(p, sizeof(p), "%s/dbg/tracing", self->base);
+	snprintf(q, sizeof(q), "%s/shr", self->base);
+	ASSERT_EQ(mkdir(q, 0755), 0);
+	/* a bind of an automount is shrinkable as well */
+	ASSERT_EQ(mount(p, q, NULL, MS_BIND, NULL), 0);
+	snprintf(p, sizeof(p), "%s/srv", self->base);
+	ASSERT_EQ(mkdir(p, 0755), 0);
+	ASSERT_EQ(mount("tmpfs", p, "tmpfs", 0, "mode=0755"), 0);
+	snprintf(p, sizeof(p), "%s/srv/hidden", self->base);
+	ASSERT_EQ(mkdir(p, 0755), 0);
+	snprintf(p, sizeof(p), "%s/srv/hidden/secret", self->base);
+	ASSERT_EQ(create_file(p, "covered-by-root\n"), 0);
+	snprintf(p, sizeof(p), "%s/srv/hidden", self->base);
+	ASSERT_EQ(mount("tmpfs", p, "tmpfs", 0, "mode=0755"), 0);
+	snprintf(p, sizeof(p), "%s/x", self->base);
+	ASSERT_EQ(mkdir(p, 0755), 0);
+	snprintf(p, sizeof(p), "%s/c", self->base);
+	ASSERT_EQ(mkdir(p, 0755), 0);
+
+	pid = fork();
+	ASSERT_GE(pid, 0);
+	if (pid == 0)
+		_exit(expiry_child(self->base));
+	ret = wait_child(pid);
+	TH_LOG("child exit code %d", ret);
+	EXPECT_EQ(ret, CHILD_OK);
+}
+
+/*
+ * The copy of the mount tree that a user namespace gets at its creation has
+ * the automounts in it locked unless they are on an expiry list. A bind of
+ * such a tree inside the namespace copies the lock to the automount but not
+ * to the root of the bind, so the child may ask for the umount of that root.
+ * That must not expire the locked automount below it: the root is busy, the
+ * umount fails and the automount has to be there afterwards.
+ */
+static int copied_tree_child(const char *base)
+{
+	char dbg[PATH_LEN], x[PATH_LEN], tracing[PATH_LEN];
+	struct stat root, st;
+
+	snprintf(dbg, sizeof(dbg), "%s/dbg", base);
+	snprintf(x, sizeof(x), "%s/x", base);
+	snprintf(tracing, sizeof(tracing), "%s/x/tracing", base);
+
+	if (enter_userns())
+		return CHILD_SETUP;
+	if (mount(dbg, x, NULL, MS_BIND | MS_REC, NULL))
+		return CHILD_SETUP;
+	/* the copy of the automount carries the lock */
+	if (!umount2(tracing, 0) || errno != EINVAL)
+		return CHILD_STEP1;
+	/* the root of the bind doesn't; keep it busy */
+	if (chdir(x))
+		return CHILD_SETUP;
+	if (!umount2(x, 0) || errno != EBUSY)
+		return CHILD_STEP2;
+	/* the automount below it must not have gone */
+	if (stat(x, &root) || fstatat(AT_FDCWD, tracing, &st, AT_NO_AUTOMOUNT))
+		return CHILD_SETUP;
+	return st.st_dev != root.st_dev ? CHILD_OK : CHILD_STEP3;
+}
+
+TEST_F(locked_mount, umount_of_a_bind_leaves_a_locked_automount_alone)
+{
+	char p[PATH_LEN];
+	struct stat st;
+	pid_t pid;
+	int ret;
+
+	if (!self->tracing)
+		SKIP(return, "test requires debugfs with the tracefs automount");
+
+	/* the copy the child gets has to contain the automount */
+	snprintf(p, sizeof(p), "%s/dbg/tracing/.", self->base);
+	ASSERT_EQ(stat(p, &st), 0);
+	snprintf(p, sizeof(p), "%s/x", self->base);
+	ASSERT_EQ(mkdir(p, 0755), 0);
+
+	pid = fork();
+	ASSERT_GE(pid, 0);
+	if (pid == 0)
+		_exit(copied_tree_child(self->base));
+	ret = wait_child(pid);
+	TH_LOG("child exit code %d", ret);
+	EXPECT_EQ(ret, CHILD_OK);
+}
+
+/*
+ * Z is a user namespace with a copy of a tree in which P/d is covered by a
+ * locked mount. The host mounts X on P/d, which propagates beneath Z's
+ * cover, and unmounts it again. Z's cover has to be locked afterwards as
+ * it was before.
+ */
+static int propagation_child(const char *base, int to_host, int from_host)
+{
+	char d[PATH_LEN], secret[PATH_LEN];
+
+	snprintf(d, sizeof(d), "%s/P/d", base);
+	snprintf(secret, sizeof(secret), "%s/P/d/secret", base);
+
+	if (enter_userns())
+		return CHILD_SETUP;
+	/* the cover is locked */
+	if (!umount2(d, 0) || errno != EINVAL)
+		return CHILD_STEP1;
+	if (send_byte(to_host) || wait_byte(from_host))
+		return CHILD_SETUP;
+	/* X came and went beneath it: still locked */
+	if (!umount2(d, 0) || errno != EINVAL)
+		return CHILD_STEP2;
+	return covered(secret) ? CHILD_OK : CHILD_STEP3;
+}
+
+TEST_F(locked_mount, propagated_copy_keeps_the_cover_locked)
+{
+	int to_host[2], from_host[2], ret;
+	char p[PATH_LEN];
+	pid_t pid;
+
+	snprintf(p, sizeof(p), "%s/P", self->base);
+	ASSERT_EQ(mkdir(p, 0755), 0);
+	ASSERT_EQ(mount("tmpfs", p, "tmpfs", 0, "mode=0755"), 0);
+	ASSERT_EQ(mount(NULL, p, NULL, MS_SHARED, NULL), 0);
+	snprintf(p, sizeof(p), "%s/P/d", self->base);
+	ASSERT_EQ(mkdir(p, 0755), 0);
+	snprintf(p, sizeof(p), "%s/P/d/secret", self->base);
+	ASSERT_EQ(create_file(p, "covered-by-root\n"), 0);
+	ASSERT_EQ(pipe(to_host), 0);
+	ASSERT_EQ(pipe(from_host), 0);
+
+	pid = fork();
+	ASSERT_GE(pid, 0);
+	if (pid == 0) {
+		/* M: a manager's namespace that covers P/d for Z */
+		pid_t z;
+
+		if (unshare(CLONE_NEWNS))
+			_exit(CHILD_SETUP);
+		snprintf(p, sizeof(p), "%s/P", self->base);
+		if (mount(NULL, p, NULL, MS_SLAVE, NULL))
+			_exit(CHILD_SETUP);
+		snprintf(p, sizeof(p), "%s/P/d", self->base);
+		if (mount("tmpfs", p, "tmpfs", 0, "mode=0755"))
+			_exit(CHILD_SETUP);
+		z = fork();
+		if (z < 0)
+			_exit(CHILD_SETUP);
+		if (z == 0)
+			_exit(propagation_child(self->base, to_host[1], from_host[0]));
+		_exit(wait_child(z));
+	}
+	ASSERT_EQ(wait_byte(to_host[0]), 0);
+
+	/* the host mounts on P/d and unmounts again; both propagate */
+	snprintf(p, sizeof(p), "%s/P/d", self->base);
+	ASSERT_EQ(mount("tmpfs", p, "tmpfs", 0, "mode=0755"), 0);
+	ASSERT_EQ(umount2(p, 0), 0);
+
+	ASSERT_EQ(send_byte(from_host[1]), 0);
+	ret = wait_child(pid);
+	TH_LOG("child exit code %d", ret);
+	EXPECT_EQ(ret, CHILD_OK);
+}
+
+TEST_HARNESS_MAIN

-- 
2.53.0


^ permalink raw reply	[flat|nested] 23+ messages in thread

* [PATCH 11/21] selftests/filesystems: check the atime of the empty mount namespace root
  2026-10-02 13:52 [PATCH 00/21] mount: more bugfixes, trapped in the Black Lodge edition Christian Brauner
                   ` (9 preceding siblings ...)
  2026-10-02 13:52 ` [PATCH 10/21] selftests/filesystems: check that a lock lands on the right mount and stays Christian Brauner
@ 2026-10-02 13:52 ` Christian Brauner
  2026-10-02 13:52 ` [PATCH 12/21] selftests/filesystems: check that an automount below an overlay layer is refused Christian Brauner
                   ` (9 subsequent siblings)
  20 siblings, 0 replies; 23+ messages in thread
From: Christian Brauner @ 2026-10-02 13:52 UTC (permalink / raw)
  To: linux-fsdevel
  Cc: Alexander Viro, Jan Kara, linux-kernel, Jeff Layton, Jann Horn,
	Neil Brown, Amir Goldstein, Christian Brauner (Amutable)

No access time updates for immutable nullfs.

Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 .../selftests/filesystems/empty_mntns/.gitignore   |   1 +
 .../selftests/filesystems/empty_mntns/Makefile     |   2 +-
 .../filesystems/empty_mntns/nullfs_atime_test.c    | 129 +++++++++++++++++++++
 3 files changed, 131 insertions(+), 1 deletion(-)

diff --git a/tools/testing/selftests/filesystems/empty_mntns/.gitignore b/tools/testing/selftests/filesystems/empty_mntns/.gitignore
index 32125b3eaa80..8ea166068d6b 100644
--- a/tools/testing/selftests/filesystems/empty_mntns/.gitignore
+++ b/tools/testing/selftests/filesystems/empty_mntns/.gitignore
@@ -3,3 +3,4 @@ clone3_empty_mntns_test
 empty_mntns_test
 overmount_chroot_test
 internal_sb_reconfigure_test
+nullfs_atime_test
diff --git a/tools/testing/selftests/filesystems/empty_mntns/Makefile b/tools/testing/selftests/filesystems/empty_mntns/Makefile
index b64818b962ca..4b2ea6acec1c 100644
--- a/tools/testing/selftests/filesystems/empty_mntns/Makefile
+++ b/tools/testing/selftests/filesystems/empty_mntns/Makefile
@@ -4,7 +4,7 @@ CFLAGS += -Wall -O2 -g $(KHDR_INCLUDES) $(TOOLS_INCLUDES)
 LDLIBS += -lcap
 
 TEST_GEN_PROGS := empty_mntns_test overmount_chroot_test clone3_empty_mntns_test
-TEST_GEN_PROGS += internal_sb_reconfigure_test
+TEST_GEN_PROGS += internal_sb_reconfigure_test nullfs_atime_test
 
 include ../../lib.mk
 
diff --git a/tools/testing/selftests/filesystems/empty_mntns/nullfs_atime_test.c b/tools/testing/selftests/filesystems/empty_mntns/nullfs_atime_test.c
new file mode 100644
index 000000000000..51e34f3c4f54
--- /dev/null
+++ b/tools/testing/selftests/filesystems/empty_mntns/nullfs_atime_test.c
@@ -0,0 +1,129 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * The root of every empty mount namespace is the same nullfs inode. A read
+ * of it by one user must not change the access time another user sees.
+ */
+#define _GNU_SOURCE
+#include <dirent.h>
+#include <errno.h>
+#include <fcntl.h>
+#include <sched.h>
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <unistd.h>
+#include <sys/stat.h>
+#include <sys/wait.h>
+
+#include "../../kselftest_harness.h"
+
+#ifndef UNSHARE_EMPTY_MNTNS
+#define UNSHARE_EMPTY_MNTNS	0x00100000
+#endif
+
+enum {
+	CHILD_OK,
+	CHILD_UNSUPPORTED,
+	CHILD_SETUP,
+	CHILD_CHANGED,
+};
+
+static int wait_byte(int fd)
+{
+	char c;
+
+	return read(fd, &c, 1) == 1 ? 0 : -1;
+}
+
+static int send_byte(int fd)
+{
+	return write(fd, "x", 1) == 1 ? 0 : -1;
+}
+
+static int empty_mntns(void)
+{
+	if (!unshare(UNSHARE_EMPTY_MNTNS))
+		return 0;
+	return errno == EINVAL ? CHILD_UNSUPPORTED : CHILD_SETUP;
+}
+
+/* the watcher: stats its root before and after the reader read its own */
+static int watcher(int to_reader, int from_reader)
+{
+	struct stat before, after;
+	int ret;
+
+	ret = empty_mntns();
+	if (ret)
+		return ret;
+	if (stat("/", &before))
+		return CHILD_SETUP;
+	if (send_byte(to_reader) || wait_byte(from_reader))
+		return CHILD_SETUP;
+	if (stat("/", &after))
+		return CHILD_SETUP;
+	if (before.st_atim.tv_sec != after.st_atim.tv_sec ||
+	    before.st_atim.tv_nsec != after.st_atim.tv_nsec)
+		return CHILD_CHANGED;
+	return CHILD_OK;
+}
+
+/* the reader: lists its own root, which is the same inode */
+static int reader(int to_watcher, int from_watcher)
+{
+	struct dirent *de;
+	DIR *d;
+	int ret;
+
+	ret = empty_mntns();
+	if (ret)
+		return ret;
+	if (wait_byte(from_watcher))
+		return CHILD_SETUP;
+	d = opendir("/");
+	if (!d)
+		return CHILD_SETUP;
+	while ((de = readdir(d)))
+		;
+	closedir(d);
+	return send_byte(to_watcher) ? CHILD_SETUP : CHILD_OK;
+}
+
+static int wait_child(pid_t pid)
+{
+	int status;
+
+	if (waitpid(pid, &status, 0) != pid || !WIFEXITED(status))
+		return -1;
+	return WEXITSTATUS(status);
+}
+
+TEST(empty_mntns_root_atime)
+{
+	int to_reader[2], to_watcher[2], w, r;
+	pid_t watcher_pid, reader_pid;
+
+	if (geteuid())
+		SKIP(return, "test requires root");
+	ASSERT_EQ(pipe(to_reader), 0);
+	ASSERT_EQ(pipe(to_watcher), 0);
+
+	watcher_pid = fork();
+	ASSERT_GE(watcher_pid, 0);
+	if (watcher_pid == 0)
+		_exit(watcher(to_reader[1], to_watcher[0]));
+	reader_pid = fork();
+	ASSERT_GE(reader_pid, 0);
+	if (reader_pid == 0)
+		_exit(reader(to_watcher[1], to_reader[0]));
+
+	r = wait_child(reader_pid);
+	w = wait_child(watcher_pid);
+	if (r == CHILD_UNSUPPORTED || w == CHILD_UNSUPPORTED)
+		SKIP(return, "UNSHARE_EMPTY_MNTNS not supported");
+	EXPECT_EQ(r, CHILD_OK);
+	EXPECT_EQ(w, CHILD_OK)
+		TH_LOG("the access time of the root changed while this namespace did nothing");
+}
+
+TEST_HARNESS_MAIN

-- 
2.53.0


^ permalink raw reply	[flat|nested] 23+ messages in thread

* [PATCH 12/21] selftests/filesystems: check that an automount below an overlay layer is refused
  2026-10-02 13:52 [PATCH 00/21] mount: more bugfixes, trapped in the Black Lodge edition Christian Brauner
                   ` (10 preceding siblings ...)
  2026-10-02 13:52 ` [PATCH 11/21] selftests/filesystems: check the atime of the empty mount namespace root Christian Brauner
@ 2026-10-02 13:52 ` Christian Brauner
  2026-10-02 13:52 ` [PATCH 13/21] fhandle: decide the subtree check under mount_lock Christian Brauner
                   ` (8 subsequent siblings)
  20 siblings, 0 replies; 23+ messages in thread
From: Christian Brauner @ 2026-10-02 13:52 UTC (permalink / raw)
  To: linux-fsdevel
  Cc: Alexander Viro, Jan Kara, linux-kernel, Jeff Layton, Jann Horn,
	Neil Brown, Amir Goldstein, Christian Brauner (Amutable)

Make sure that we don't automount on top of internal things.

Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 .../selftests/filesystems/overlayfs/.gitignore     |   1 +
 .../selftests/filesystems/overlayfs/Makefile       |   1 +
 .../filesystems/overlayfs/automount_in_layer.c     | 174 +++++++++++++++++++++
 3 files changed, 176 insertions(+)

diff --git a/tools/testing/selftests/filesystems/overlayfs/.gitignore b/tools/testing/selftests/filesystems/overlayfs/.gitignore
index 077f7a128168..b343cc430051 100644
--- a/tools/testing/selftests/filesystems/overlayfs/.gitignore
+++ b/tools/testing/selftests/filesystems/overlayfs/.gitignore
@@ -2,3 +2,4 @@
 dev_in_maps
 set_layers_via_fds
 idmapped_mounts
+automount_in_layer
diff --git a/tools/testing/selftests/filesystems/overlayfs/Makefile b/tools/testing/selftests/filesystems/overlayfs/Makefile
index b3185f684add..382a59bda5e7 100644
--- a/tools/testing/selftests/filesystems/overlayfs/Makefile
+++ b/tools/testing/selftests/filesystems/overlayfs/Makefile
@@ -9,6 +9,7 @@ LOCAL_HDRS += ../wrappers.h log.h
 TEST_GEN_PROGS := dev_in_maps
 TEST_GEN_PROGS += set_layers_via_fds
 TEST_GEN_PROGS += idmapped_mounts
+TEST_GEN_PROGS += automount_in_layer
 
 include ../../lib.mk
 
diff --git a/tools/testing/selftests/filesystems/overlayfs/automount_in_layer.c b/tools/testing/selftests/filesystems/overlayfs/automount_in_layer.c
new file mode 100644
index 000000000000..0476c4582bdf
--- /dev/null
+++ b/tools/testing/selftests/filesystems/overlayfs/automount_in_layer.c
@@ -0,0 +1,174 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * A layer of an overlay is a private clone of the mount it was given and
+ * belongs to no mount namespace. fanotify hands out descriptors on it. An
+ * automount triggered through one has no namespace to go into: the open has
+ * to fail, not oops with namespace_sem held.
+ */
+#define _GNU_SOURCE
+#include <dirent.h>
+#include <errno.h>
+#include <fcntl.h>
+#include <poll.h>
+#include <sched.h>
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <unistd.h>
+#include <linux/magic.h>
+#include <sys/fanotify.h>
+#include <sys/mount.h>
+#include <sys/stat.h>
+#include <sys/vfs.h>
+
+#include "../../kselftest_harness.h"
+
+#define DIR_LEN		64
+#define PATH_LEN	192
+
+static bool have_fs(const char *name)
+{
+	char line[128];
+	bool found = false;
+	FILE *f;
+
+	f = fopen("/proc/filesystems", "re");
+	if (!f)
+		return false;
+	while (fgets(line, sizeof(line), f)) {
+		char *nl = strchr(line, '\n');
+		char *tab = strchr(line, '\t');
+
+		if (nl)
+			*nl = 0;
+		if (tab && !strcmp(tab + 1, name))
+			found = true;
+	}
+	fclose(f);
+	return found;
+}
+
+static int mnt_id_of(int fd)
+{
+	char path[64], buf[4096], *p;
+	ssize_t n;
+	int info;
+
+	snprintf(path, sizeof(path), "/proc/self/fdinfo/%d", fd);
+	info = open(path, O_RDONLY | O_CLOEXEC);
+	if (info < 0)
+		return -1;
+	n = read(info, buf, sizeof(buf) - 1);
+	close(info);
+	if (n <= 0)
+		return -1;
+	buf[n] = 0;
+	p = strstr(buf, "mnt_id:");
+	return p ? atoi(p + strlen("mnt_id:")) : -1;
+}
+
+static bool on_debugfs(int fd)
+{
+	struct statfs sf;
+
+	return !fstatfs(fd, &sf) && sf.f_type == DEBUGFS_MAGIC;
+}
+
+FIXTURE(layer) {
+	char base[DIR_LEN];
+	int fan;
+	int evfd;	/* on the layer clone of the lower debugfs */
+};
+
+FIXTURE_SETUP(layer)
+{
+	char lower[PATH_LEN], other[PATH_LEN], ovl[PATH_LEN], opts[2 * PATH_LEN + 16];
+	struct fanotify_event_metadata *ev;
+	struct pollfd pfd;
+	char buf[4096];
+	int lfd, lower_id;
+	ssize_t n;
+	DIR *d;
+
+	self->fan = -1;
+	self->evfd = -1;
+	if (geteuid())
+		SKIP(return, "test requires root");
+	if (!have_fs("debugfs") || !have_fs("tracefs"))
+		SKIP(return, "test requires debugfs with the tracefs automount");
+	if (!have_fs("overlay"))
+		SKIP(return, "test requires overlayfs");
+
+	snprintf(self->base, sizeof(self->base), "/tmp/layer.XXXXXX");
+	ASSERT_NE(mkdtemp(self->base), NULL);
+	ASSERT_EQ(unshare(CLONE_NEWNS), 0);
+	ASSERT_EQ(mount("", "/", NULL, MS_REC | MS_PRIVATE, NULL), 0);
+	ASSERT_EQ(mount("tmpfs", self->base, "tmpfs", 0, "mode=0755"), 0);
+	snprintf(lower, sizeof(lower), "%s/lower", self->base);
+	snprintf(other, sizeof(other), "%s/other", self->base);
+	snprintf(ovl, sizeof(ovl), "%s/ovl", self->base);
+	ASSERT_EQ(mkdir(lower, 0755), 0);
+	ASSERT_EQ(mkdir(other, 0755), 0);
+	ASSERT_EQ(mkdir(ovl, 0755), 0);
+	ASSERT_EQ(mount("debugfs", lower, "debugfs", 0, NULL), 0);
+	snprintf(opts, sizeof(opts), "lowerdir=%s:%s", lower, other);
+	ASSERT_EQ(mount("overlay", ovl, "overlay", MS_RDONLY, opts), 0);
+
+	self->fan = fanotify_init(FAN_CLASS_NOTIF | FAN_NONBLOCK | FAN_CLOEXEC,
+				  O_RDONLY | O_CLOEXEC);
+	ASSERT_GE(self->fan, 0);
+	ASSERT_EQ(fanotify_mark(self->fan, FAN_MARK_ADD | FAN_MARK_FILESYSTEM,
+				FAN_OPEN | FAN_ONDIR, AT_FDCWD, lower), 0);
+	lfd = open(lower, O_RDONLY | O_DIRECTORY | O_CLOEXEC);
+	ASSERT_GE(lfd, 0);
+	lower_id = mnt_id_of(lfd);
+	close(lfd);
+
+	/* the overlay opens its lower directory through the layer clone */
+	d = opendir(ovl);
+	ASSERT_NE(d, NULL);
+	closedir(d);
+	pfd.fd = self->fan;
+	pfd.events = POLLIN;
+	ASSERT_EQ(poll(&pfd, 1, 5000), 1);
+	n = read(self->fan, buf, sizeof(buf));
+	ASSERT_GT(n, 0);
+	for (ev = (void *)buf; FAN_EVENT_OK(ev, n); ev = FAN_EVENT_NEXT(ev, n)) {
+		if (ev->fd < 0)
+			continue;
+		if (self->evfd < 0 && on_debugfs(ev->fd) &&
+		    mnt_id_of(ev->fd) != lower_id)
+			self->evfd = ev->fd;
+		else
+			close(ev->fd);
+	}
+	ASSERT_GE(self->evfd, 0)
+		TH_LOG("no event on the layer clone");
+}
+
+FIXTURE_TEARDOWN(layer)
+{
+	if (self->evfd >= 0)
+		close(self->evfd);
+	if (self->fan >= 0)
+		close(self->fan);
+	umount2(self->base, MNT_DETACH);
+	rmdir(self->base);
+}
+
+TEST_F(layer, automount_below_the_clone_is_refused)
+{
+	struct stat st;
+	int fd;
+
+	/* the automount point is there */
+	ASSERT_EQ(fstatat(self->evfd, "tracing", &st, AT_NO_AUTOMOUNT), 0);
+	/* the clone is in no namespace, so the automount has nowhere to go */
+	fd = openat(self->evfd, "tracing", O_RDONLY | O_DIRECTORY | O_CLOEXEC);
+	EXPECT_LT(fd, 0);
+	EXPECT_EQ(errno, EINVAL);
+	if (fd >= 0)
+		close(fd);
+}
+
+TEST_HARNESS_MAIN

-- 
2.53.0


^ permalink raw reply	[flat|nested] 23+ messages in thread

* [PATCH 13/21] fhandle: decide the subtree check under mount_lock
  2026-10-02 13:52 [PATCH 00/21] mount: more bugfixes, trapped in the Black Lodge edition Christian Brauner
                   ` (11 preceding siblings ...)
  2026-10-02 13:52 ` [PATCH 12/21] selftests/filesystems: check that an automount below an overlay layer is refused Christian Brauner
@ 2026-10-02 13:52 ` Christian Brauner
  2026-10-02 13:52 ` [PATCH 14/21] namespace: keep the private nullfs instance in knullfs Christian Brauner
                   ` (7 subsequent siblings)
  20 siblings, 0 replies; 23+ messages in thread
From: Christian Brauner @ 2026-10-02 13:52 UTC (permalink / raw)
  To: linux-fsdevel
  Cc: Alexander Viro, Jan Kara, linux-kernel, Jeff Layton, Jann Horn,
	Neil Brown, Amir Goldstein, Christian Brauner (Amutable)

may_decode_fh() lets a caller who is privileged over the mount
namespace of @root decode handles below @root->dentry as long as the
mount is mounted and no locked child covers something below it. The
three parts are read one after the other without a lock: is_mounted()
and capable_wrt_mount() read ->mnt_ns and has_locked_children() walks
->mnt_mounts under a mount_lock of its own.

Today the gaps are harmless. A lazy umount in between clears ->mnt_ns,
but the locked children stay attached to their unmounted parent, so the
walk still finds them and the decode is refused either way. The
following patches detach every unmounted mount from its parent. Then
an umount between is_mounted() and the walk makes the walk come back
empty and the caller decodes into what a locked child covered.

So take mount_lock once and answer all three questions under it.
is_mounted() is stable there, umount_tree() clears ->mnt_ns on the
write side, and a mounted mount still has its locked children on its
list. ns_capable() under the spinlock is fine, generic_permission()
calls it in RCU walk already. has_locked_children() loses its locking
wrapper, its other callers hold namespace_sem or mount_lock anyway.

Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 fs/fhandle.c   | 20 +++++++++++++++++---
 fs/namespace.c | 14 ++++----------
 2 files changed, 21 insertions(+), 13 deletions(-)

diff --git a/fs/fhandle.c b/fs/fhandle.c
index f8829231e3d7..d22f2e065677 100644
--- a/fs/fhandle.c
+++ b/fs/fhandle.c
@@ -298,6 +298,22 @@ static bool capable_wrt_mount(struct mount *mount)
 	return mnt_ns && ns_capable(mnt_ns->user_ns, CAP_SYS_ADMIN);
 }
 
+/*
+ * Does the caller have an unobstructed way to everything below @root? Only
+ * if the mount is mounted, the caller is privileged over its mount namespace
+ * and no locked child covers something below @root->dentry. One answer from
+ * under mount_lock: an umount in between clears ->mnt_ns and takes the
+ * children off the mount, the locked ones too.
+ */
+static bool subtree_unobstructed(const struct path *root)
+{
+	struct mount *mnt = real_mount(root->mnt);
+
+	guard(mount_locked_reader)();
+	return is_mounted(root->mnt) && capable_wrt_mount(mnt) &&
+	       !has_locked_children(mnt, root->dentry);
+}
+
 static inline int may_decode_fh(struct handle_to_path_ctx *ctx,
 				unsigned int o_flags)
 {
@@ -332,9 +348,7 @@ static inline int may_decode_fh(struct handle_to_path_ctx *ctx,
 
 	if (ns_capable(root->mnt->mnt_sb->s_user_ns, CAP_SYS_ADMIN))
 		ctx->flags = HANDLE_CHECK_PERMS;
-	else if (is_mounted(root->mnt) &&
-		 capable_wrt_mount(real_mount(root->mnt)) &&
-		 !has_locked_children(real_mount(root->mnt), root->dentry))
+	else if (subtree_unobstructed(root))
 		ctx->flags = HANDLE_CHECK_PERMS | HANDLE_CHECK_SUBTREE;
 	else
 		return -EPERM;
diff --git a/fs/namespace.c b/fs/namespace.c
index bb0183ec2aaf..e116894c8667 100644
--- a/fs/namespace.c
+++ b/fs/namespace.c
@@ -2353,7 +2353,7 @@ void dissolve_on_fput(struct vfsmount *mnt)
 }
 
 /* locks: namespace_shared && pinned(mnt) || mount_locked_reader */
-static bool __has_locked_children(struct mount *mnt, struct dentry *dentry)
+bool has_locked_children(struct mount *mnt, struct dentry *dentry)
 {
 	struct mount *child;
 
@@ -2367,12 +2367,6 @@ static bool __has_locked_children(struct mount *mnt, struct dentry *dentry)
 	return false;
 }
 
-bool has_locked_children(struct mount *mnt, struct dentry *dentry)
-{
-	guard(mount_locked_reader)();
-	return __has_locked_children(mnt, dentry);
-}
-
 /* locks: namespace_shared && pinned(mnt) || mount_locked_reader */
 static bool __has_children(struct mount *mnt, struct dentry *dentry)
 {
@@ -2441,7 +2435,7 @@ struct vfsmount *clone_private_mount(const struct path *path)
 	if (!ns_capable(old_mnt->mnt_ns->user_ns, CAP_SYS_ADMIN))
 		return ERR_PTR(-EPERM);
 
-	if (__has_locked_children(old_mnt, path->dentry))
+	if (has_locked_children(old_mnt, path->dentry))
 		return ERR_PTR(-EINVAL);
 
 	new_mnt = clone_mnt(old_mnt, path->dentry, CL_PRIVATE);
@@ -3056,7 +3050,7 @@ static struct mount *__do_loopback(const struct path *old_path,
 	if (recurse && !old->mnt_ns)
 		return ERR_PTR(-EINVAL);
 
-	if (!recurse && __has_locked_children(old, old_path->dentry))
+	if (!recurse && has_locked_children(old, old_path->dentry))
 		return ERR_PTR(-EINVAL);
 
 	if (recurse)
@@ -3544,7 +3538,7 @@ static int do_set_group(const struct path *from_path, const struct path *to_path
 		return -EINVAL;
 
 	/* From mount should not have locked children in place of To's root */
-	if (__has_locked_children(from, to->mnt.mnt_root))
+	if (has_locked_children(from, to->mnt.mnt_root))
 		return -EINVAL;
 
 	/* Setting sharing groups is only allowed on private mounts */

-- 
2.53.0


^ permalink raw reply	[flat|nested] 23+ messages in thread

* [PATCH 14/21] namespace: keep the private nullfs instance in knullfs
  2026-10-02 13:52 [PATCH 00/21] mount: more bugfixes, trapped in the Black Lodge edition Christian Brauner
                   ` (12 preceding siblings ...)
  2026-10-02 13:52 ` [PATCH 13/21] fhandle: decide the subtree check under mount_lock Christian Brauner
@ 2026-10-02 13:52 ` Christian Brauner
  2026-10-02 13:52 ` [PATCH 15/21] namespace: nothing is mounted on or written through knullfs Christian Brauner
                   ` (6 subsequent siblings)
  20 siblings, 0 replies; 23+ messages in thread
From: Christian Brauner @ 2026-10-02 13:52 UTC (permalink / raw)
  To: linux-fsdevel
  Cc: Alexander Viro, Jan Kara, linux-kernel, Jeff Layton, Jann Horn,
	Neil Brown, Amir Goldstein, Christian Brauner (Amutable)

The private nullfs instance that kernel threads are confined to is only
ever reachable through init_task's root and pwd. The following patches
point mounts at it as well, so keep it in a global.

No functional changes.

Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 fs/mount.h     |  1 +
 fs/namespace.c | 11 ++++++-----
 2 files changed, 7 insertions(+), 5 deletions(-)

diff --git a/fs/mount.h b/fs/mount.h
index 2223fb141499..4e68e5cbc254 100644
--- a/fs/mount.h
+++ b/fs/mount.h
@@ -6,6 +6,7 @@
 #include <linux/fs_pin.h>
 
 extern struct file_system_type nullfs_fs_type;
+extern struct vfsmount *knullfs;
 extern struct list_head notify_list;
 
 struct mnt_namespace {
diff --git a/fs/namespace.c b/fs/namespace.c
index e116894c8667..d1e83cda57e6 100644
--- a/fs/namespace.c
+++ b/fs/namespace.c
@@ -80,6 +80,7 @@ static u64 mnt_id_ctr = MNT_UNIQUE_ID_OFFSET;
 static struct hlist_head *mount_hashtable __ro_after_init;
 static struct hlist_head *mountpoint_hashtable __ro_after_init;
 static struct kmem_cache *mnt_cache __ro_after_init;
+struct vfsmount *knullfs __ro_after_init;	/* private nullfs instance */
 static DECLARE_RWSEM(namespace_sem);
 static HLIST_HEAD(unmounted);	/* protected by namespace_sem */
 static LIST_HEAD(ex_mountpoints); /* protected by namespace_sem */
@@ -6335,7 +6336,7 @@ static void __init init_mount_tree(void)
 	 *
 	 * (1) nullfs with mount id 1
 	 * (2) mutable rootfs with mount id 2
-	 * (3) private nullfs for kthreads (SB_KERNMOUNT)
+	 * (3) private nullfs for kthreads (SB_KERNMOUNT), kept in knullfs
 	 *
 	 * with (2) mounted on top of (1). The init_task's root and pwd
 	 * are pointed at (3) so all kthreads start isolated in nullfs.
@@ -6370,11 +6371,11 @@ static void __init init_mount_tree(void)
 		init_mnt_ns.nr_mounts++;
 	}
 
-	nullfs_mnt = kern_mount(&nullfs_fs_type);
-	if (IS_ERR(nullfs_mnt))
+	knullfs = kern_mount(&nullfs_fs_type);
+	if (IS_ERR(knullfs))
 		panic("VFS: Failed to create private nullfs instance");
-	root.mnt	= nullfs_mnt;
-	root.dentry	= nullfs_mnt->mnt_root;
+	root.mnt	= knullfs;
+	root.dentry	= knullfs->mnt_root;
 
 	init_task.nsproxy->mnt_ns = &init_mnt_ns;
 	get_mnt_ns(&init_mnt_ns);

-- 
2.53.0


^ permalink raw reply	[flat|nested] 23+ messages in thread

* [PATCH 15/21] namespace: nothing is mounted on or written through knullfs
  2026-10-02 13:52 [PATCH 00/21] mount: more bugfixes, trapped in the Black Lodge edition Christian Brauner
                   ` (13 preceding siblings ...)
  2026-10-02 13:52 ` [PATCH 14/21] namespace: keep the private nullfs instance in knullfs Christian Brauner
@ 2026-10-02 13:52 ` Christian Brauner
  2026-10-02 13:52 ` [PATCH 16/21] fsnotify: let a filesystem refuse marks on its objects Christian Brauner
                   ` (5 subsequent siblings)
  20 siblings, 0 replies; 23+ messages in thread
From: Christian Brauner @ 2026-10-02 13:52 UTC (permalink / raw)
  To: linux-fsdevel
  Cc: Alexander Viro, Jan Kara, linux-kernel, Jeff Layton, Jann Horn,
	Neil Brown, Amir Goldstein, Christian Brauner (Amutable)

Mark the root of knullfs with dont_mount(). Nothing is ever mounted on
the root of a kernel thread and the following patches make that root
reachable from userspace, so say it on the dentry where it doesn't
depend on the mount being in no namespace. Let do_lock_mount() refuse
such a target before it takes the inode lock and namespace_sem. The
flag is sticky so the check needs no lock. The one under the locks
stays for a mountpoint that is being removed.

Make the mount read-only as well. Its one inode is immutable so nothing
could be changed through it anyway, but MNT_READONLY makes that visible
the usual way: EROFS instead of EPERM and ST_RDONLY in statvfs().

Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 fs/namespace.c | 9 +++++++++
 1 file changed, 9 insertions(+)

diff --git a/fs/namespace.c b/fs/namespace.c
index d1e83cda57e6..e1b0ade95b0d 100644
--- a/fs/namespace.c
+++ b/fs/namespace.c
@@ -2814,6 +2814,11 @@ static void do_lock_mount(const struct path *path,
 
 		scoped_guard(mount_locked_reader) {
 			m = where_to_mount(path, &dentry, beneath);
+			/* sticky, so it takes no locks to refuse it */
+			if (unlikely(cant_mount(dentry))) {
+				res->parent = ERR_PTR(-ENOENT);
+				return;
+			}
 			if (&m->mnt != path->mnt) {
 				mntget(&m->mnt);
 				dget(dentry);
@@ -6374,6 +6379,10 @@ static void __init init_mount_tree(void)
 	knullfs = kern_mount(&nullfs_fs_type);
 	if (IS_ERR(knullfs))
 		panic("VFS: Failed to create private nullfs instance");
+	/* nothing is ever mounted on the root of a kernel thread */
+	dont_mount(knullfs->mnt_root);
+	/* and nothing is ever written through it */
+	knullfs->mnt_flags |= MNT_READONLY;
 	root.mnt	= knullfs;
 	root.dentry	= knullfs->mnt_root;
 

-- 
2.53.0


^ permalink raw reply	[flat|nested] 23+ messages in thread

* [PATCH 16/21] fsnotify: let a filesystem refuse marks on its objects
  2026-10-02 13:52 [PATCH 00/21] mount: more bugfixes, trapped in the Black Lodge edition Christian Brauner
                   ` (14 preceding siblings ...)
  2026-10-02 13:52 ` [PATCH 15/21] namespace: nothing is mounted on or written through knullfs Christian Brauner
@ 2026-10-02 13:52 ` Christian Brauner
  2026-10-02 14:26   ` Amir Goldstein
  2026-10-02 13:52 ` [PATCH 17/21] nullfs: refuse file locks Christian Brauner
                   ` (4 subsequent siblings)
  20 siblings, 1 reply; 23+ messages in thread
From: Christian Brauner @ 2026-10-02 13:52 UTC (permalink / raw)
  To: linux-fsdevel
  Cc: Alexander Viro, Jan Kara, linux-kernel, Jeff Layton, Jann Horn,
	Neil Brown, Amir Goldstein, Christian Brauner (Amutable)

Don't let nullfs be watched. fanotify refuses mount and filesystem marks
on SB_NOUSER superblocks but inode marks of inotify, fanotify and
dnotify go through. The one inode of knullfs is the root of every
kernel thread and the following patches make it reachable from
userspace as the directory that stands in for an unmounted mount. A
watch placed through one such directory would report the opens through
all the others, across users.

Add FS_DISALLOW_NOTIFY next to FS_DISALLOW_NOTIFY_PERM, refuse a mark on
any object of such a filesystem in fsnotify_add_mark_list() where every
backend ends up and set it for nullfs. There's nothing to watch on a
permanently empty and immutable filesystem.

Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 fs/notify/mark.c   | 4 ++++
 fs/nullfs.c        | 1 +
 include/linux/fs.h | 1 +
 3 files changed, 6 insertions(+)

diff --git a/fs/notify/mark.c b/fs/notify/mark.c
index b2640d836a71..d17628580a57 100644
--- a/fs/notify/mark.c
+++ b/fs/notify/mark.c
@@ -903,6 +903,10 @@ static int fsnotify_add_mark_list(struct fsnotify_mark *mark, void *obj,
 	if (WARN_ON(!fsnotify_valid_obj_type(obj_type)))
 		return -EINVAL;
 
+	/* the filesystem doesn't want its objects watched */
+	if (sb && (sb->s_type->fs_flags & FS_DISALLOW_NOTIFY))
+		return -EINVAL;
+
 	/*
 	 * Attach the sb info before attaching a connector to any object on sb.
 	 * The sb info will remain attached as long as sb lives.
diff --git a/fs/nullfs.c b/fs/nullfs.c
index 40aa228bd81a..55a04f2d7761 100644
--- a/fs/nullfs.c
+++ b/fs/nullfs.c
@@ -61,6 +61,7 @@ static int nullfs_init_fs_context(struct fs_context *fc)
 
 struct file_system_type nullfs_fs_type = {
 	.name			= "nullfs",
+	.fs_flags		= FS_DISALLOW_NOTIFY,
 	.init_fs_context	= nullfs_init_fs_context,
 	.kill_sb		= kill_anon_super,
 };
diff --git a/include/linux/fs.h b/include/linux/fs.h
index f9d1e05e8ae6..784fa20217c4 100644
--- a/include/linux/fs.h
+++ b/include/linux/fs.h
@@ -2296,6 +2296,7 @@ struct file_system_type {
 #define FS_POWER_FREEZE		256	/* Always freeze on suspend/hibernate */
 #define FS_USERNS_MOUNT_RESTRICTED 512	/* Restrict mount in userns if not already visible */
 #define FS_USERNS_DELEGATABLE	1024	/* Can be mounted inside userns from outside */
+#define FS_DISALLOW_NOTIFY	2048	/* No fsnotify marks on its objects */
 #define FS_RENAME_DOES_D_MOVE	32768	/* FS will handle d_move() during rename() internally. */
 	int (*init_fs_context)(struct fs_context *);
 	const struct fs_parameter_spec *parameters;

-- 
2.53.0


^ permalink raw reply	[flat|nested] 23+ messages in thread

* [PATCH 17/21] nullfs: refuse file locks
  2026-10-02 13:52 [PATCH 00/21] mount: more bugfixes, trapped in the Black Lodge edition Christian Brauner
                   ` (15 preceding siblings ...)
  2026-10-02 13:52 ` [PATCH 16/21] fsnotify: let a filesystem refuse marks on its objects Christian Brauner
@ 2026-10-02 13:52 ` Christian Brauner
  2026-10-02 13:52 ` [PATCH 18/21] nullfs: refuse leases and delegations Christian Brauner
                   ` (3 subsequent siblings)
  20 siblings, 0 replies; 23+ messages in thread
From: Christian Brauner @ 2026-10-02 13:52 UTC (permalink / raw)
  To: linux-fsdevel
  Cc: Alexander Viro, Jan Kara, linux-kernel, Jeff Layton, Jann Horn,
	Neil Brown, Amir Goldstein, Christian Brauner (Amutable)

Refuse flock() and POSIX locks on nullfs. Its one inode is the root of
every kernel thread and the following patches make it the directory
that stands in for an unmounted mount, so a lock taken through one such
directory would block the locks of every other holder and F_GETLK would
tell them the pid of the holder. Give the directory file operations of
its own: what libfs gives an empty directory plus ->lock and ->flock
that fail with ENOLCK, the way a filesystem without lock support does.

Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 fs/nullfs.c | 29 +++++++++++++++++++++++++++++
 1 file changed, 29 insertions(+)

diff --git a/fs/nullfs.c b/fs/nullfs.c
index 55a04f2d7761..d90c71f5eece 100644
--- a/fs/nullfs.c
+++ b/fs/nullfs.c
@@ -10,6 +10,34 @@ static const struct super_operations nullfs_super_operations = {
 	.statfs	= simple_statfs,
 };
 
+static loff_t nullfs_dir_llseek(struct file *file, loff_t offset, int whence)
+{
+	/* an empty directory has two entries . and .. at offsets 0 and 1 */
+	return generic_file_llseek_size(file, offset, whence, 2, 2);
+}
+
+static int nullfs_dir_readdir(struct file *file, struct dir_context *ctx)
+{
+	dir_emit_dots(file, ctx);
+	return 0;
+}
+
+/* the one inode of nullfs is shared by every holder, so no locks on it */
+static int nullfs_nolock(struct file *file, int cmd, struct file_lock *fl)
+{
+	return -ENOLCK;
+}
+
+/* what libfs gives an empty directory, plus the refusal of file locks */
+static const struct file_operations nullfs_dir_operations = {
+	.llseek		= nullfs_dir_llseek,
+	.read		= generic_read_dir,
+	.iterate_shared	= nullfs_dir_readdir,
+	.fsync		= noop_fsync,
+	.lock		= nullfs_nolock,
+	.flock		= nullfs_nolock,
+};
+
 static int nullfs_fs_fill_super(struct super_block *s, struct fs_context *fc)
 {
 	struct inode *inode;
@@ -30,6 +58,7 @@ static int nullfs_fs_fill_super(struct super_block *s, struct fs_context *fc)
 
 	/* nullfs is permanently empty... */
 	make_empty_dir_inode(inode);
+	inode->i_fop = &nullfs_dir_operations;
 	simple_inode_init_ts(inode);
 	inode->i_ino	= 1;
 	/* ... and immutable, reading it leaves no trace either. */

-- 
2.53.0


^ permalink raw reply	[flat|nested] 23+ messages in thread

* [PATCH 18/21] nullfs: refuse leases and delegations
  2026-10-02 13:52 [PATCH 00/21] mount: more bugfixes, trapped in the Black Lodge edition Christian Brauner
                   ` (16 preceding siblings ...)
  2026-10-02 13:52 ` [PATCH 17/21] nullfs: refuse file locks Christian Brauner
@ 2026-10-02 13:52 ` Christian Brauner
  2026-10-02 13:52 ` [PATCH 19/21] readdir: take no inode lock on an immutable directory Christian Brauner
                   ` (2 subsequent siblings)
  20 siblings, 0 replies; 23+ messages in thread
From: Christian Brauner @ 2026-10-02 13:52 UTC (permalink / raw)
  To: linux-fsdevel
  Cc: Alexander Viro, Jan Kara, linux-kernel, Jeff Layton, Jann Horn,
	Neil Brown, Amir Goldstein, Christian Brauner (Amutable)

Refuse leases and delegations on nullfs as well. generic_setlease()
grants read leases and directory delegations on directories, so the
owner of the shared inode, or anyone with CAP_LEASE, could put one on
the directory that stands in for every unmounted mount and F_GETLEASE
and F_GETDELEG would show it to every other holder. Nothing on nullfs
ever changes, so a lease on it would never be broken and never tell
anyone anything.

Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 fs/nullfs.c | 8 ++++++++
 1 file changed, 8 insertions(+)

diff --git a/fs/nullfs.c b/fs/nullfs.c
index d90c71f5eece..f76b87cf1841 100644
--- a/fs/nullfs.c
+++ b/fs/nullfs.c
@@ -28,6 +28,13 @@ static int nullfs_nolock(struct file *file, int cmd, struct file_lock *fl)
 	return -ENOLCK;
 }
 
+/* and no leases or delegations */
+static int nullfs_nolease(struct file *file, int arg, struct file_lease **flp,
+			  void **priv)
+{
+	return -EINVAL;
+}
+
 /* what libfs gives an empty directory, plus the refusal of file locks */
 static const struct file_operations nullfs_dir_operations = {
 	.llseek		= nullfs_dir_llseek,
@@ -36,6 +43,7 @@ static const struct file_operations nullfs_dir_operations = {
 	.fsync		= noop_fsync,
 	.lock		= nullfs_nolock,
 	.flock		= nullfs_nolock,
+	.setlease	= nullfs_nolease,
 };
 
 static int nullfs_fs_fill_super(struct super_block *s, struct fs_context *fc)

-- 
2.53.0


^ permalink raw reply	[flat|nested] 23+ messages in thread

* [PATCH 19/21] readdir: take no inode lock on an immutable directory
  2026-10-02 13:52 [PATCH 00/21] mount: more bugfixes, trapped in the Black Lodge edition Christian Brauner
                   ` (17 preceding siblings ...)
  2026-10-02 13:52 ` [PATCH 18/21] nullfs: refuse leases and delegations Christian Brauner
@ 2026-10-02 13:52 ` Christian Brauner
  2026-10-02 13:52 ` [PATCH 20/21] selftests/filesystems: add a helper that holds a readdir in a page fault Christian Brauner
  2026-10-02 13:52 ` [PATCH 21/21] selftests/filesystems: check that reading the root of an empty mount namespace stalls nobody Christian Brauner
  20 siblings, 0 replies; 23+ messages in thread
From: Christian Brauner @ 2026-10-02 13:52 UTC (permalink / raw)
  To: linux-fsdevel
  Cc: Alexander Viro, Jan Kara, linux-kernel, Jeff Layton, Jann Horn,
	Neil Brown, Amir Goldstein, Christian Brauner (Amutable),
	stable

iterate_dir() takes the directory's i_rwsem shared and holds it across
->iterate_shared(). For a directory that never has an entry and is
never removed the lock keeps nothing still, it only orders every reader
and every writer of that inode behind each other.

For the directory of a nullfs instance that matters. The instance of
the initial mount namespace is the root of every empty mount namespace
and the private instance is the root of every kernel thread, so one
inode is shared across users who have nothing else in common. And a
reader can hold the lock for as long as it likes: back the getdents()
buffer with a mapping of a file on a FUSE mount of your own, let the
copy of "." and ".." fault and let the server wait. Queue an exclusive
taker behind it, a mkdir() in that directory goes through start_dirop()
before the read-only mount is reported, and from then on every lookup
that misses the dcache in that directory, every create and every mount
on it waits until the server answers. One user of an empty mount
namespace stalls all the others.

Add FOP_IMMUTABLE for the file operations of a directory that never
changes and is never removed and let iterate_dir() skip the lock for
it. The flag never changes for a file, ->f_pos is protected by
f_pos_lock since directories are FMODE_ATOMIC_POS, IS_DEADDIR can't be
set on such a directory and neither touch_atime() nor fsnotify take
i_rwsem. Set it on the nullfs directory. The placeholder directories of
libfs never have an entry either but their owners remove them, so they
keep the lock.

Fixes: 9d4e752a24f7 ("namespace: allow creating empty mount namespaces")
Cc: stable@vger.kernel.org # v7.1+
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 fs/nullfs.c        |  1 +
 fs/readdir.c       | 13 +++++++++----
 include/linux/fs.h |  2 ++
 3 files changed, 12 insertions(+), 4 deletions(-)

diff --git a/fs/nullfs.c b/fs/nullfs.c
index f76b87cf1841..bfc04bca3940 100644
--- a/fs/nullfs.c
+++ b/fs/nullfs.c
@@ -44,6 +44,7 @@ static const struct file_operations nullfs_dir_operations = {
 	.lock		= nullfs_nolock,
 	.flock		= nullfs_nolock,
 	.setlease	= nullfs_nolease,
+	.fop_flags	= FOP_IMMUTABLE,
 };
 
 static int nullfs_fs_fill_super(struct super_block *s, struct fs_context *fc)
diff --git a/fs/readdir.c b/fs/readdir.c
index 76bb1ae3a450..f2288841337a 100644
--- a/fs/readdir.c
+++ b/fs/readdir.c
@@ -87,6 +87,8 @@ EXPORT_SYMBOL(wrap_directory_iterator);
 int iterate_dir(struct file *file, struct dir_context *ctx)
 {
 	struct inode *inode = file_inode(file);
+	/* never an entry, never removed: nothing for the lock to keep still */
+	bool locked = !(file->f_op->fop_flags & FOP_IMMUTABLE);
 	int res = -ENOTDIR;
 
 	if (!file->f_op->iterate_shared)
@@ -100,9 +102,11 @@ int iterate_dir(struct file *file, struct dir_context *ctx)
 	if (res)
 		goto out;
 
-	res = down_read_killable(&inode->i_rwsem);
-	if (res)
-		goto out;
+	if (locked) {
+		res = down_read_killable(&inode->i_rwsem);
+		if (res)
+			goto out;
+	}
 
 	res = -ENOENT;
 	if (!IS_DEADDIR(inode)) {
@@ -112,7 +116,8 @@ int iterate_dir(struct file *file, struct dir_context *ctx)
 		fsnotify_access(file);
 		file_accessed(file);
 	}
-	inode_unlock_shared(inode);
+	if (locked)
+		inode_unlock_shared(inode);
 out:
 	return res;
 }
diff --git a/include/linux/fs.h b/include/linux/fs.h
index 784fa20217c4..deb411e86661 100644
--- a/include/linux/fs.h
+++ b/include/linux/fs.h
@@ -1978,6 +1978,8 @@ struct file_operations {
 #define FOP_ASYNC_LOCK		((__force fop_flags_t)(1 << 6))
 /* File system supports uncached read/write buffered IO */
 #define FOP_DONTCACHE		((__force fop_flags_t)(1 << 7))
+/* Never changes and is never removed, readdir of a directory takes no lock */
+#define FOP_IMMUTABLE		((__force fop_flags_t)(1 << 8))
 
 /* Wrap a directory iterator that needs exclusive inode access */
 int wrap_directory_iterator(struct file *, struct dir_context *,

-- 
2.53.0


^ permalink raw reply	[flat|nested] 23+ messages in thread

* [PATCH 20/21] selftests/filesystems: add a helper that holds a readdir in a page fault
  2026-10-02 13:52 [PATCH 00/21] mount: more bugfixes, trapped in the Black Lodge edition Christian Brauner
                   ` (18 preceding siblings ...)
  2026-10-02 13:52 ` [PATCH 19/21] readdir: take no inode lock on an immutable directory Christian Brauner
@ 2026-10-02 13:52 ` Christian Brauner
  2026-10-02 13:52 ` [PATCH 21/21] selftests/filesystems: check that reading the root of an empty mount namespace stalls nobody Christian Brauner
  20 siblings, 0 replies; 23+ messages in thread
From: Christian Brauner @ 2026-10-02 13:52 UTC (permalink / raw)
  To: linux-fsdevel
  Cc: Alexander Viro, Jan Kara, linux-kernel, Jeff Layton, Jann Horn,
	Neil Brown, Amir Goldstein, Christian Brauner (Amutable)

A getdents64() whose buffer is a page registered with userfaultfd sits
in handle_userfault() with whatever iterate_dir() took before it copied
the entries. Add readdir_hold.h for tests that want to know what that
blocks: it opens the userfaultfd before the test enters a user
namespace, the fault happens in the kernel and needs CAP_SYS_PTRACE in
the initial one, starts the readdir in a thread, waits until that
thread is stuck, queues a create behind it and waits for it to settle,
then probes a lookup with a watchdog and says whether it came back.
The page is released afterwards so that everything drains on a kernel
that still holds the lock across the copy.

The two users follow.

Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 tools/testing/selftests/filesystems/readdir_hold.h | 224 +++++++++++++++++++++
 1 file changed, 224 insertions(+)

diff --git a/tools/testing/selftests/filesystems/readdir_hold.h b/tools/testing/selftests/filesystems/readdir_hold.h
new file mode 100644
index 000000000000..57eee1bef470
--- /dev/null
+++ b/tools/testing/selftests/filesystems/readdir_hold.h
@@ -0,0 +1,224 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * Hold a readdir of a directory in the page fault of its buffer and see
+ * whether a create and a lookup in that directory wait for it. For a
+ * directory that is permanently empty they must not.
+ */
+#ifndef __SELFTESTS_READDIR_HOLD_H
+#define __SELFTESTS_READDIR_HOLD_H
+
+#include <errno.h>
+#include <fcntl.h>
+#include <poll.h>
+#include <pthread.h>
+#include <stdbool.h>
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <unistd.h>
+#include <linux/userfaultfd.h>
+#include <sys/ioctl.h>
+#include <sys/mman.h>
+#include <sys/stat.h>
+#include <sys/syscall.h>
+
+#define HOLD_FAULT_MS	5000	/* for the reader to reach the fault */
+#define HOLD_QUEUE_MS	2000	/* for the create to queue up behind it */
+#define HOLD_LOOKUP_MS	5000	/* for the lookup to come back */
+
+struct readdir_hold {
+	int uffd;
+	int taskdir;		/* /proc/self/task, from before any namespace change */
+	long page_size;
+	char *page;		/* faults until released */
+	int dfd;
+	pid_t creator_tid;
+	int created;
+	int done[2];		/* the finder writes a byte when it is back */
+};
+
+/*
+ * The fault happens in the kernel, so the userfaultfd needs CAP_SYS_PTRACE
+ * in the initial user namespace or vm.unprivileged_userfaultfd. Call this
+ * before entering a user namespace.
+ */
+static inline int readdir_hold_init(struct readdir_hold *h)
+{
+	struct uffdio_api api = { .api = UFFD_API };
+
+	h->page = MAP_FAILED;
+	h->page_size = sysconf(_SC_PAGESIZE);
+	h->taskdir = open("/proc/self/task", O_RDONLY | O_DIRECTORY | O_CLOEXEC);
+	h->uffd = syscall(__NR_userfaultfd, O_CLOEXEC | O_NONBLOCK);
+	if (h->uffd < 0 || h->taskdir < 0 || ioctl(h->uffd, UFFDIO_API, &api)) {
+		if (h->uffd >= 0)
+			close(h->uffd);
+		if (h->taskdir >= 0)
+			close(h->taskdir);
+		h->uffd = h->taskdir = -1;
+		return -1;
+	}
+	return 0;
+}
+
+static inline void readdir_hold_destroy(struct readdir_hold *h)
+{
+	if (h->uffd >= 0)
+		close(h->uffd);
+	if (h->taskdir >= 0)
+		close(h->taskdir);
+	h->uffd = h->taskdir = -1;
+}
+
+static inline void *readdir_hold_reader(void *arg)
+{
+	struct readdir_hold *h = arg;
+
+	/* the first byte written to the buffer faults until released */
+	syscall(__NR_getdents64, h->dfd, h->page, h->page_size);
+	return NULL;
+}
+
+static inline void *readdir_hold_creator(void *arg)
+{
+	struct readdir_hold *h = arg;
+
+	h->creator_tid = syscall(__NR_gettid);
+	/* takes the directory lock exclusive before it fails */
+	mkdirat(h->dfd, "x", 0755);
+	__atomic_store_n(&h->created, 1, __ATOMIC_RELEASE);
+	return NULL;
+}
+
+static inline void *readdir_hold_finder(void *arg)
+{
+	struct readdir_hold *h = arg;
+	int fd;
+
+	/* a lookup that misses the dcache takes the lock shared */
+	fd = openat(h->dfd, "no_such_name", O_RDONLY | O_CLOEXEC);
+	if (fd >= 0)
+		close(fd);
+	if (write(h->done[1], "x", 1) != 1)
+		perror("readdir_hold: finder");
+	return NULL;
+}
+
+/* the creator is back, or waits in the kernel for the lock */
+static inline bool readdir_hold_creator_settled(struct readdir_hold *h)
+{
+	char path[32], buf[256], *p;
+	ssize_t n;
+	int fd;
+
+	if (__atomic_load_n(&h->created, __ATOMIC_ACQUIRE))
+		return true;
+	if (!h->creator_tid)
+		return false;
+	snprintf(path, sizeof(path), "%d/stat", h->creator_tid);
+	fd = openat(h->taskdir, path, O_RDONLY | O_CLOEXEC);
+	if (fd < 0)
+		return false;
+	n = read(fd, buf, sizeof(buf) - 1);
+	close(fd);
+	if (n <= 0)
+		return false;
+	buf[n] = 0;
+	/* "pid (comm) state ..." */
+	p = strrchr(buf, ')');
+	return p && p[1] == ' ' && p[2] == 'D';
+}
+
+static inline bool readdir_hold_faulted(struct readdir_hold *h)
+{
+	struct pollfd pfd = { .fd = h->uffd, .events = POLLIN };
+	struct uffd_msg msg;
+
+	if (poll(&pfd, 1, HOLD_FAULT_MS) != 1)
+		return false;
+	if (read(h->uffd, &msg, sizeof(msg)) != sizeof(msg))
+		return false;
+	return msg.event == UFFD_EVENT_PAGEFAULT;
+}
+
+/* let the reader go on */
+static inline void readdir_hold_release(struct readdir_hold *h)
+{
+	struct uffdio_copy cp = {
+		.dst = (unsigned long)h->page,
+		.len = h->page_size,
+	};
+	void *zero;
+
+	zero = calloc(1, h->page_size);
+	if (!zero)
+		return;
+	cp.src = (unsigned long)zero;
+	if (ioctl(h->uffd, UFFDIO_COPY, &cp) && errno != EEXIST)
+		perror("readdir_hold: UFFDIO_COPY");
+	free(zero);
+}
+
+/*
+ * A readdir of @dfd that sticks in the fault of its buffer, then a create
+ * and a lookup in @dfd. Whether the lookup came back while the readdir was
+ * still stuck goes to @stalled. Returns -1 when that could not be found
+ * out.
+ */
+static inline int readdir_hold_check(struct readdir_hold *h, int dfd,
+				     bool *stalled)
+{
+	pthread_t reader, creator, finder;
+	struct uffdio_register reg = {};
+	struct pollfd pfd;
+	int ret = -1, i;
+
+	h->dfd = dfd;
+	h->created = 0;
+	h->creator_tid = 0;
+	h->page = mmap(NULL, h->page_size, PROT_READ | PROT_WRITE,
+		       MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
+	if (h->page == MAP_FAILED)
+		return -1;
+	reg.range.start = (unsigned long)h->page;
+	reg.range.len = h->page_size;
+	reg.mode = UFFDIO_REGISTER_MODE_MISSING;
+	if (ioctl(h->uffd, UFFDIO_REGISTER, &reg) || pipe2(h->done, O_CLOEXEC))
+		goto out_page;
+
+	if (pthread_create(&reader, NULL, readdir_hold_reader, h))
+		goto out_pipe;
+	if (!readdir_hold_faulted(h))
+		goto out_reader;
+	if (pthread_create(&creator, NULL, readdir_hold_creator, h))
+		goto out_reader;
+	for (i = 0; i < HOLD_QUEUE_MS / 10 && !readdir_hold_creator_settled(h); i++)
+		usleep(10000);
+	if (pthread_create(&finder, NULL, readdir_hold_finder, h))
+		goto out_creator;
+
+	pfd.fd = h->done[0];
+	pfd.events = POLLIN;
+	*stalled = poll(&pfd, 1, HOLD_LOOKUP_MS) != 1;
+	ret = 0;
+
+	readdir_hold_release(h);
+	pthread_join(finder, NULL);
+out_creator:
+	if (ret)
+		readdir_hold_release(h);
+	pthread_join(creator, NULL);
+out_reader:
+	if (ret)
+		readdir_hold_release(h);
+	pthread_join(reader, NULL);
+out_pipe:
+	close(h->done[0]);
+	close(h->done[1]);
+out_page:
+	munmap(h->page, h->page_size);
+	h->page = MAP_FAILED;
+	return ret;
+}
+
+#endif /* __SELFTESTS_READDIR_HOLD_H */

-- 
2.53.0


^ permalink raw reply	[flat|nested] 23+ messages in thread

* [PATCH 21/21] selftests/filesystems: check that reading the root of an empty mount namespace stalls nobody
  2026-10-02 13:52 [PATCH 00/21] mount: more bugfixes, trapped in the Black Lodge edition Christian Brauner
                   ` (19 preceding siblings ...)
  2026-10-02 13:52 ` [PATCH 20/21] selftests/filesystems: add a helper that holds a readdir in a page fault Christian Brauner
@ 2026-10-02 13:52 ` Christian Brauner
  20 siblings, 0 replies; 23+ messages in thread
From: Christian Brauner @ 2026-10-02 13:52 UTC (permalink / raw)
  To: linux-fsdevel
  Cc: Alexander Viro, Jan Kara, linux-kernel, Jeff Layton, Jann Horn,
	Neil Brown, Amir Goldstein, Christian Brauner (Amutable)

A readdir of the root of an empty mount namespace stuck in the page
fault of its buffer must stall neither a create nor a lookup in that
directory. The test needs userfaultfd and skips without it.

Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 .../selftests/filesystems/empty_mntns/.gitignore   |  1 +
 .../selftests/filesystems/empty_mntns/Makefile     |  5 ++-
 .../filesystems/empty_mntns/root_readdir_test.c    | 45 ++++++++++++++++++++++
 3 files changed, 50 insertions(+), 1 deletion(-)

diff --git a/tools/testing/selftests/filesystems/empty_mntns/.gitignore b/tools/testing/selftests/filesystems/empty_mntns/.gitignore
index 8ea166068d6b..27bcbcaeb1d1 100644
--- a/tools/testing/selftests/filesystems/empty_mntns/.gitignore
+++ b/tools/testing/selftests/filesystems/empty_mntns/.gitignore
@@ -4,3 +4,4 @@ empty_mntns_test
 overmount_chroot_test
 internal_sb_reconfigure_test
 nullfs_atime_test
+root_readdir_test
diff --git a/tools/testing/selftests/filesystems/empty_mntns/Makefile b/tools/testing/selftests/filesystems/empty_mntns/Makefile
index 4b2ea6acec1c..22af2164c54e 100644
--- a/tools/testing/selftests/filesystems/empty_mntns/Makefile
+++ b/tools/testing/selftests/filesystems/empty_mntns/Makefile
@@ -4,7 +4,9 @@ CFLAGS += -Wall -O2 -g $(KHDR_INCLUDES) $(TOOLS_INCLUDES)
 LDLIBS += -lcap
 
 TEST_GEN_PROGS := empty_mntns_test overmount_chroot_test clone3_empty_mntns_test
-TEST_GEN_PROGS += internal_sb_reconfigure_test nullfs_atime_test
+TEST_GEN_PROGS += internal_sb_reconfigure_test nullfs_atime_test root_readdir_test
+
+LOCAL_HDRS += ../readdir_hold.h
 
 include ../../lib.mk
 
@@ -12,3 +14,4 @@ $(OUTPUT)/empty_mntns_test: ../utils.c
 $(OUTPUT)/overmount_chroot_test: ../utils.c
 $(OUTPUT)/clone3_empty_mntns_test: ../utils.c
 $(OUTPUT)/internal_sb_reconfigure_test: ../utils.c
+$(OUTPUT)/root_readdir_test: LDLIBS += -pthread
diff --git a/tools/testing/selftests/filesystems/empty_mntns/root_readdir_test.c b/tools/testing/selftests/filesystems/empty_mntns/root_readdir_test.c
new file mode 100644
index 000000000000..0ff544247027
--- /dev/null
+++ b/tools/testing/selftests/filesystems/empty_mntns/root_readdir_test.c
@@ -0,0 +1,45 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Reading the root of an empty mount namespace holds nothing that anybody
+ * else waits for: the directory never has an entry, so a readdir stuck in
+ * the page fault of its buffer stalls neither a create nor a lookup in it.
+ */
+#define _GNU_SOURCE
+#include <errno.h>
+#include <fcntl.h>
+#include <sched.h>
+#include <stdio.h>
+#include <stdlib.h>
+#include <unistd.h>
+
+#include "../../kselftest_harness.h"
+#include "../readdir_hold.h"
+#include "empty_mntns.h"
+
+TEST(readdir_blocks_nobody)
+{
+	struct readdir_hold hold;
+	bool stalled;
+	int dfd;
+
+	if (geteuid())
+		SKIP(return, "test requires root");
+	if (readdir_hold_init(&hold))
+		SKIP(return, "test requires userfaultfd");
+	if (unshare(UNSHARE_EMPTY_MNTNS)) {
+		readdir_hold_destroy(&hold);
+		if (errno == EINVAL)
+			SKIP(return, "UNSHARE_EMPTY_MNTNS not supported");
+		ASSERT_TRUE(false)
+			TH_LOG("unshare(UNSHARE_EMPTY_MNTNS): %m");
+	}
+	dfd = open("/", O_RDONLY | O_DIRECTORY | O_CLOEXEC);
+	ASSERT_GE(dfd, 0);
+	ASSERT_EQ(readdir_hold_check(&hold, dfd, &stalled), 0);
+	/* the lookup came back while the readdir was stuck in its fault */
+	EXPECT_FALSE(stalled);
+	close(dfd);
+	readdir_hold_destroy(&hold);
+}
+
+TEST_HARNESS_MAIN

-- 
2.53.0


^ permalink raw reply	[flat|nested] 23+ messages in thread

* Re: [PATCH 16/21] fsnotify: let a filesystem refuse marks on its objects
  2026-10-02 13:52 ` [PATCH 16/21] fsnotify: let a filesystem refuse marks on its objects Christian Brauner
@ 2026-10-02 14:26   ` Amir Goldstein
  0 siblings, 0 replies; 23+ messages in thread
From: Amir Goldstein @ 2026-10-02 14:26 UTC (permalink / raw)
  To: Christian Brauner
  Cc: linux-fsdevel, Alexander Viro, Jan Kara, linux-kernel,
	Jeff Layton, Jann Horn, Neil Brown

On Fri, Oct 2, 2026 at 3:54 PM Christian Brauner <brauner@kernel.org> wrote:
>
> Don't let nullfs be watched. fanotify refuses mount and filesystem marks
> on SB_NOUSER superblocks but inode marks of inotify, fanotify and
> dnotify go through. The one inode of knullfs is the root of every
> kernel thread and the following patches make it reachable from
> userspace as the directory that stands in for an unmounted mount. A
> watch placed through one such directory would report the opens through
> all the others, across users.
>
> Add FS_DISALLOW_NOTIFY next to FS_DISALLOW_NOTIFY_PERM, refuse a mark on
> any object of such a filesystem in fsnotify_add_mark_list() where every
> backend ends up and set it for nullfs. There's nothing to watch on a
> permanently empty and immutable filesystem.
>

I don't particularly mind this custom opt-out, but
shall we perhaps instead deny marks on SB_NOUSER directories?
IIRC, we already wanted to deny marks on any SB_NOUSER, but then
it turned out that some users set inotify watches on pipes, so we stepped
back and restricted SB_NOUSER only for fs/mount marks.
I think that the same logic extends well also for directories, because
I don't think that any existing SB_NOUSER fs is supposed to be
exposing them atm, so no regressions are to be expected.

Thanks,
Amir.

^ permalink raw reply	[flat|nested] 23+ messages in thread

end of thread, other threads:[~2026-10-02 14:26 UTC | newest]

Thread overview: 23+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-02 13:52 [PATCH 00/21] mount: more bugfixes, trapped in the Black Lodge edition Christian Brauner
2026-10-02 13:52 ` [PATCH 01/21] namespace: unhash a dentry before detaching the mounts on it Christian Brauner
2026-10-02 13:52 ` [PATCH 02/21] namei: don't reveal overmounted entries in refwalk Christian Brauner
2026-10-02 13:52 ` [PATCH 03/21] fcntl: refuse F_SET_RW_HINT on an immutable inode Christian Brauner
2026-10-02 13:52 ` [PATCH 04/21] selftests/filesystems: check that an immutable inode takes no write hint Christian Brauner
2026-10-02 13:52 ` [PATCH 05/21] namespace: refuse an automount below a mount that is in no namespace Christian Brauner
2026-10-02 13:52 ` [PATCH 06/21] namespace: handle mount locking for automounts correctly Christian Brauner
2026-10-02 13:52 ` [PATCH 07/21] nullfs: don't update the access time Christian Brauner
2026-10-02 13:52 ` [PATCH 08/21] namespace: never expire a locked mount Christian Brauner
2026-10-02 13:52 ` [PATCH 09/21] namespace: keep the lock on a mount that a propagated copy is moved beneath Christian Brauner
2026-10-02 13:52 ` [PATCH 10/21] selftests/filesystems: check that a lock lands on the right mount and stays Christian Brauner
2026-10-02 13:52 ` [PATCH 11/21] selftests/filesystems: check the atime of the empty mount namespace root Christian Brauner
2026-10-02 13:52 ` [PATCH 12/21] selftests/filesystems: check that an automount below an overlay layer is refused Christian Brauner
2026-10-02 13:52 ` [PATCH 13/21] fhandle: decide the subtree check under mount_lock Christian Brauner
2026-10-02 13:52 ` [PATCH 14/21] namespace: keep the private nullfs instance in knullfs Christian Brauner
2026-10-02 13:52 ` [PATCH 15/21] namespace: nothing is mounted on or written through knullfs Christian Brauner
2026-10-02 13:52 ` [PATCH 16/21] fsnotify: let a filesystem refuse marks on its objects Christian Brauner
2026-10-02 14:26   ` Amir Goldstein
2026-10-02 13:52 ` [PATCH 17/21] nullfs: refuse file locks Christian Brauner
2026-10-02 13:52 ` [PATCH 18/21] nullfs: refuse leases and delegations Christian Brauner
2026-10-02 13:52 ` [PATCH 19/21] readdir: take no inode lock on an immutable directory Christian Brauner
2026-10-02 13:52 ` [PATCH 20/21] selftests/filesystems: add a helper that holds a readdir in a page fault Christian Brauner
2026-10-02 13:52 ` [PATCH 21/21] selftests/filesystems: check that reading the root of an empty mount namespace stalls nobody Christian Brauner

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®