mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Justin Suess <utilityemal77@gmail.com>
To: "Christian Brauner" <brauner@kernel.org>,
	"Alexander Viro" <viro@zeniv.linux.org.uk>,
	"Jan Kara" <jack@suse.cz>, NeilBrown <neil@brown.name>,
	"Mickaël Salaün" <mic@digikod.net>,
	"Alexei Starovoitov" <ast@kernel.org>,
	"Daniel Borkmann" <daniel@iogearbox.net>,
	"Andrii Nakryiko" <andrii@kernel.org>,
	"Song Liu" <song@kernel.org>
Cc: linux-fsdevel@vger.kernel.org, bpf@vger.kernel.org,
	linux-security-module@vger.kernel.org,
	linux-kernel@vger.kernel.org, "Günther Noack" <gnoack@google.com>,
	"Paul Moore" <paul@paul-moore.com>,
	"James Morris" <jmorris@namei.org>,
	"Serge E . Hallyn" <serge@hallyn.com>,
	"Martin KaFai Lau" <martin.lau@linux.dev>,
	"Eduard Zingerman" <eddyz87@gmail.com>,
	"Yonghong Song" <yonghong.song@linux.dev>,
	"John Fastabend" <john.fastabend@gmail.com>,
	"Kumar Kartikeya Dwivedi" <memxor@gmail.com>,
	"Jiri Olsa" <jolsa@kernel.org>,
	"Jeff Layton" <jlayton@kernel.org>,
	"Amir Goldstein" <amir73il@gmail.com>,
	"Mateusz Guzik" <mjguzik@gmail.com>,
	"Shuah Khan" <shuah@kernel.org>, "Tingmao Wang" <m@maowtm.org>,
	"Justin Suess" <utilityemal77@gmail.com>
Subject: [RFC PATCH bpf-next 09/12] namei: add an rcu-walk mode to the ancestor walk
Date: Mon,  5 Oct 2026 20:20:16 -0400	[thread overview]
Message-ID: <20261006002020.2890858-10-utilityemal77@gmail.com> (raw)
In-Reply-To: <20261006002020.2890858-1-utilityemal77@gmail.com>

Let the stepwise engine walk lockless, entered with %VFS_WALK_RCU under
the caller's rcu_read_lock: stepping shares follow_dotdot_rcu()'s core,
factored into __path_walk_parent_rcu(), with per-step d_seq validation,
choose_mountpoint_rcu() for crossings and mount_lock revalidation
before concluding at the real root.  Any lost race surfaces as -ECHILD
and the caller retries in the referenced mode.

vfs_walk_handover() continues a lockless walk with references, so that
a caller can do at a position what no lockless walk can.  It mirrors
unlazy_walk(): the position is legitimized and becomes the starting
position of a second, reference-based walk that owns the references
acquired on it.

Both halves of that are ordered so that nothing has to be put: a failed
legitimization leaves no partial references, and the one case where
__legitimize_mnt() obliges a sleepable mntput() goes to
mnt_undo_legitimize().  Handing the reference straight over rather than
copying it out and dropping a second one is what keeps the whole
operation callable inside the caller's RCU read-side critical section,
where even a path_put() that provably cannot sleep is still a dput()
and so still a might_sleep().

Signed-off-by: Justin Suess <utilityemal77@gmail.com>
---
 fs/internal.h         |   7 ++
 fs/namei.c            | 196 ++++++++++++++++++++++++++++++++++++------
 include/linux/namei.h |   3 +
 3 files changed, 179 insertions(+), 27 deletions(-)

diff --git a/fs/internal.h b/fs/internal.h
index 3ce2220ff57f..8a5696157121 100644
--- a/fs/internal.h
+++ b/fs/internal.h
@@ -74,9 +74,14 @@ int lookup_noperm_common(struct qstr *qname, struct dentry *base);
 /*
  * The stepwise engine under vfs_walk_ancestors(); fs-internal so iterating
  * consumers (BPF) can drive it, with the walk invariants staying in namei.c.
+ * In rcu mode (%VFS_WALK_RCU) the caller holds rcu_read_lock() over the
+ * whole walk, no references are held, and vfs_walk_next() returning -ECHILD
+ * invalidates everything derived from the walk.
  */
 struct vfs_ancestor_walk {
 	struct path pos;
+	unsigned int seq;	/* pos.dentry->d_seq sample (rcu mode) */
+	unsigned int m_seq;	/* mount_lock sample (rcu mode) */
 	unsigned int pos_flags;	/* VFS_WALK_POS_* describing pos */
 	unsigned int flags;	/* VFS_WALK_* */
 };
@@ -85,6 +90,8 @@ void vfs_walk_start(struct vfs_ancestor_walk *aw, const struct path *path,
 		    unsigned int flags);
 int vfs_walk_next(struct vfs_ancestor_walk *aw);
 void vfs_walk_end(struct vfs_ancestor_walk *aw);
+bool vfs_walk_handover(struct vfs_ancestor_walk *to,
+		       struct vfs_ancestor_walk *from);
 
 void __init filename_init(void);
 
diff --git a/fs/namei.c b/fs/namei.c
index 73f25152d917..31f96602d4e9 100644
--- a/fs/namei.c
+++ b/fs/namei.c
@@ -2152,34 +2152,69 @@ static __always_inline const char *step_into(struct nameidata *nd, int flags,
 	return step_into_slowpath(nd, flags, dentry);
 }
 
-static struct dentry *follow_dotdot_rcu(struct nameidata *nd)
+/**
+ * __path_walk_parent_rcu - step towards the parent of the given struct path
+ * @path: position to step up from; updated in place on a mount crossing,
+ *        which is first if @path is the root of a mounted tree.  No
+ *        references are acquired; the callers layer their own bookkeeping
+ *        (path_connected(), nameidata updates, ...) on top
+ * @root: boundary as for choose_mountpoint_rcu(); if zero'ed, walk all the
+ *        way to the global root
+ * @flags: %LOOKUP_NO_XDEV fails a mount crossing with -ECHILD
+ * @m_seq: the walk's mount_lock sample
+ * @seqp: d_seq sample validating @path->dentry; updated to cover the new
+ *        @path->dentry when a mount is crossed
+ * @next_seqp: set to the returned parent's d_seq sample
+ *
+ * Returns: the parent dentry (which is @path->dentry itself if that is a
+ * disconnected root), NULL if @path is in the root with nothing to cross
+ * into, or ERR_PTR(-ECHILD) when a concurrent change was detected.
+ */
+static struct dentry *__path_walk_parent_rcu(struct path *path,
+					     const struct path *root, int flags,
+					     unsigned int m_seq, unsigned int *seqp,
+					     unsigned int *next_seqp)
 {
 	struct dentry *parent, *old;
 
-	if (path_equal(&nd->path, &nd->root))
-		goto in_root;
-	if (unlikely(nd->path.dentry == nd->path.mnt->mnt_root)) {
-		struct path path;
-		unsigned seq;
-		if (!choose_mountpoint_rcu(real_mount(nd->path.mnt),
-					   &nd->root, &path, &seq))
-			goto in_root;
-		if (unlikely(nd->flags & LOOKUP_NO_XDEV))
+	if (unlikely(path->dentry == path->mnt->mnt_root)) {
+		struct path mounted;
+		unsigned int seq;
+
+		if (!choose_mountpoint_rcu(real_mount(path->mnt),
+					   root, &mounted, &seq))
+			return NULL;
+		if (unlikely(flags & LOOKUP_NO_XDEV))
 			return ERR_PTR(-ECHILD);
-		nd->path = path;
-		nd->inode = path.dentry->d_inode;
-		nd->seq = seq;
+		*path = mounted;
+		*seqp = seq;
 		// makes sure that non-RCU pathwalk could reach this state
-		if (read_seqretry(&mount_lock, nd->m_seq))
+		if (read_seqretry(&mount_lock, m_seq))
 			return ERR_PTR(-ECHILD);
 		/* we know that mountpoint was pinned */
 	}
-	old = nd->path.dentry;
+	old = path->dentry;
 	parent = old->d_parent;
-	nd->next_seq = read_seqcount_begin(&parent->d_seq);
+	*next_seqp = read_seqcount_begin(&parent->d_seq);
 	// makes sure that non-RCU pathwalk could reach this state
-	if (read_seqcount_retry(&old->d_seq, nd->seq))
+	if (read_seqcount_retry(&old->d_seq, *seqp))
 		return ERR_PTR(-ECHILD);
+	return parent;
+}
+
+static struct dentry *follow_dotdot_rcu(struct nameidata *nd)
+{
+	struct dentry *parent;
+
+	if (path_equal(&nd->path, &nd->root))
+		goto in_root;
+	parent = __path_walk_parent_rcu(&nd->path, &nd->root, nd->flags,
+					nd->m_seq, &nd->seq, &nd->next_seq);
+	if (!parent)
+		goto in_root;
+	if (IS_ERR(parent))
+		return parent;
+	nd->inode = nd->path.dentry->d_inode;
 	if (unlikely(!path_connected(nd->path.mnt, parent)))
 		return ERR_PTR(-ECHILD);
 	return parent;
@@ -2247,7 +2282,9 @@ static const struct path vfs_walk_no_root;
  * vfs_walk_start - begin a stepwise ancestor walk
  * @aw: walk state, valid until vfs_walk_end()
  * @path: position to walk up from; never modified
- * @flags: %VFS_WALK_* flags; none defined yet, pass 0
+ * @flags: %VFS_WALK_RCU to walk lockless; the caller then holds
+ *         rcu_read_lock() from before vfs_walk_start() until after
+ *         vfs_walk_end(), and owns no references on yielded positions.
  */
 void vfs_walk_start(struct vfs_ancestor_walk *aw, const struct path *path,
 		    unsigned int flags)
@@ -2255,7 +2292,14 @@ void vfs_walk_start(struct vfs_ancestor_walk *aw, const struct path *path,
 	aw->pos = *path;
 	aw->flags = flags;
 	aw->pos_flags = 0;
-	path_get(&aw->pos);
+	if (flags & VFS_WALK_RCU) {
+		RCU_LOCKDEP_WARN(!rcu_read_lock_held(),
+				 "rcu-mode ancestor walk outside of RCU read-side critical section");
+		aw->m_seq = read_seqbegin(&mount_lock);
+		aw->seq = raw_seqcount_begin(&aw->pos.dentry->d_seq);
+	} else {
+		path_get(&aw->pos);
+	}
 }
 
 static int vfs_walk_step_ref(struct vfs_ancestor_walk *aw)
@@ -2286,6 +2330,35 @@ static int vfs_walk_step_ref(struct vfs_ancestor_walk *aw)
 	return 0;
 }
 
+static int vfs_walk_step_rcu(struct vfs_ancestor_walk *aw)
+{
+	struct dentry *parent;
+	unsigned int next_seq;
+
+	if (unlikely(aw->pos_flags & VFS_WALK_POS_DISCONNECTED)) {
+		/* Resume at the root of the disconnected position's mount. */
+		aw->pos.dentry = aw->pos.mnt->mnt_root;
+		aw->seq = raw_seqcount_begin(&aw->pos.dentry->d_seq);
+		aw->pos_flags = 0;
+		return read_seqretry(&mount_lock, aw->m_seq) ? -ECHILD : 0;
+	}
+
+	parent = __path_walk_parent_rcu(&aw->pos, &vfs_walk_no_root, 0,
+					aw->m_seq, &aw->seq, &next_seq);
+	if (!parent)
+		/* The real root, unless the mount tree moved. */
+		return read_seqretry(&mount_lock, aw->m_seq) ? -ECHILD : 1;
+	if (IS_ERR(parent))
+		return PTR_ERR(parent);
+	/* A crossing onto a disconnected root, as in vfs_walk_step_ref(). */
+	aw->pos_flags = parent == aw->pos.dentry ?
+		VFS_WALK_POS_DISCONNECTED | VFS_WALK_POS_MOUNTPOINT :
+		vfs_walk_pos_flags(aw->pos.mnt, parent);
+	aw->pos.dentry = parent;
+	aw->seq = next_seq;
+	return 0;
+}
+
 /**
  * vfs_walk_next - yield the walk's next position in @aw->pos
  * @aw: the walk
@@ -2295,12 +2368,14 @@ static int vfs_walk_step_ref(struct vfs_ancestor_walk *aw)
  * described at vfs_walk_ancestors().
  *
  * Returns: 0 with @aw->pos valid, 1 once the walk has passed the real
- * root.
+ * root, -ECHILD when an rcu-mode walk lost a race and must be retried
+ * (typically in the referenced mode).
  */
 int vfs_walk_next(struct vfs_ancestor_walk *aw)
 {
 	if (aw->flags & VFS_WALK_STARTED) {
-		int err = vfs_walk_step_ref(aw);
+		int err = (aw->flags & VFS_WALK_RCU) ?
+			vfs_walk_step_rcu(aw) : vfs_walk_step_ref(aw);
 
 		if (err)
 			return err;
@@ -2308,6 +2383,10 @@ int vfs_walk_next(struct vfs_ancestor_walk *aw)
 		aw->flags |= VFS_WALK_STARTED;
 		aw->pos_flags = vfs_walk_pos_flags(aw->pos.mnt, aw->pos.dentry);
 	}
+	/* The flags must describe the dentry the seq covers. */
+	if ((aw->flags & VFS_WALK_RCU) &&
+	    read_seqcount_retry(&aw->pos.dentry->d_seq, aw->seq))
+		return -ECHILD;
 	return 0;
 }
 
@@ -2317,7 +2396,59 @@ int vfs_walk_next(struct vfs_ancestor_walk *aw)
  */
 void vfs_walk_end(struct vfs_ancestor_walk *aw)
 {
-	path_put(&aw->pos);
+	if (!(aw->flags & VFS_WALK_RCU))
+		path_put(&aw->pos);
+}
+
+/**
+ * vfs_walk_handover - continue an rcu-mode walk with references
+ * @to: walk state to begin at @from's current position, owning the
+ *      references acquired on it; valid until vfs_walk_end() either way
+ * @from: an rcu-mode walk, positioned by a 0 return from vfs_walk_next()
+ *
+ * Mirrors unlazy_walk(): @from's current position is legitimized and
+ * becomes the starting position of the referenced walk @to, which the
+ * first vfs_walk_next() on it yields.  A failed legitimization leaves no
+ * partial references behind and a successful one moves straight into @to,
+ * so nothing is ever put here and the handover is safe within the caller's
+ * RCU read-side critical section - where a path_put() would not be, dput()
+ * being allowed to sleep.  @to itself is only usable once the caller has
+ * left it, its walk being reference-based.
+ *
+ * @from is untouched on success and may keep stepping, lockless, from
+ * where it stands.  The references do not conclude its walk: a concurrent
+ * rename may relocate the position the instant they are taken, as it may
+ * during any reference-based walk.
+ *
+ * Returns: false iff @from lost a race; it is then dead, as after -ECHILD
+ * from vfs_walk_next(), and @to is zeroed - safe to vfs_walk_end(), but
+ * not to step, so the caller has to remember it never started.
+ */
+bool vfs_walk_handover(struct vfs_ancestor_walk *to,
+		       struct vfs_ancestor_walk *from)
+{
+	struct path pos = from->pos;
+	int err;
+
+	err = __legitimize_mnt(pos.mnt, from->m_seq);
+	if (unlikely(err)) {
+		if (err < 0)
+			mnt_undo_legitimize(real_mount(pos.mnt));
+		goto dead;
+	}
+	if (unlikely(read_seqcount_retry(&pos.dentry->d_seq, from->seq) ||
+		     !lockref_get_not_dead(&pos.dentry->d_lockref))) {
+		mnt_undo_legitimize(real_mount(pos.mnt));
+		goto dead;
+	}
+	to->pos = pos;
+	to->flags = 0;
+	to->pos_flags = 0;
+	return true;
+
+dead:
+	memset(to, 0, sizeof(*to));
+	return false;
 }
 
 /**
@@ -2326,18 +2457,25 @@ void vfs_walk_end(struct vfs_ancestor_walk *aw)
  * @cb: callback invoked on @path, then on each ancestor up to the real
  *      root, crossing mount boundaries.  @cb must not sleep and returns
  *      %VFS_WALK_CONTINUE, %VFS_WALK_STOP or a negative errno to abort the
- *      walk.  @ancestor is only valid during the invocation; @cb must take
- *      its own references to keep a position.
+ *      walk; -ECHILD is reserved (see below).  @ancestor is only valid
+ *      during the invocation; @cb must take its own references to keep a
+ *      position.
  *      A position whose dentry is a disconnected root is flagged with
  *      %VFS_WALK_POS_DISCONNECTED (plus %VFS_WALK_POS_MOUNTPOINT when it
  *      is a mountpoint a mount crossing landed on rather than a parent);
  *      if @cb continues over it, the walk resumes at the root of that
  *      position's mount.
+ *      With %VFS_WALK_RCU, @cb accepts positions the walk holds no
+ *      references on: the walk then runs lockless (under rcu_read_lock)
+ *      and returns -ECHILD when it loses a race, or when @cb returns
+ *      -ECHILD.  The caller should then discard any state @cb accumulated
+ *      and retry without %VFS_WALK_RCU.
  * @data: opaque argument passed to @cb
- * @flags: %VFS_WALK_* flags; none defined yet, pass 0
+ * @flags: %VFS_WALK_RCU if @cb copes with unreferenced positions
  *
- * Returns: 0 once the real root was reached, 1 if @cb stopped the walk, or
- * the negative errno @cb aborted with.
+ * Returns: 0 once the real root was reached, 1 if @cb stopped the walk,
+ * -ECHILD if a lockless walk must be retried with references, or the
+ * negative errno @cb aborted with.
  */
 int vfs_walk_ancestors(const struct path *path,
 		       int (*cb)(const struct path *ancestor,
@@ -2347,6 +2485,8 @@ int vfs_walk_ancestors(const struct path *path,
 	struct vfs_ancestor_walk aw;
 	int ret;
 
+	if (flags & VFS_WALK_RCU)
+		rcu_read_lock();
 	vfs_walk_start(&aw, path, flags);
 	for (;;) {
 		ret = vfs_walk_next(&aw);
@@ -2365,6 +2505,8 @@ int vfs_walk_ancestors(const struct path *path,
 		}
 	}
 	vfs_walk_end(&aw);
+	if (flags & VFS_WALK_RCU)
+		rcu_read_unlock();
 	return ret;
 }
 
diff --git a/include/linux/namei.h b/include/linux/namei.h
index 3e198de7a0d3..8820a2833213 100644
--- a/include/linux/namei.h
+++ b/include/linux/namei.h
@@ -162,6 +162,9 @@ extern int follow_down_one(struct path *);
 extern int follow_down(struct path *path, unsigned int flags);
 extern int follow_up(struct path *);
 
+/* vfs_walk_ancestors() flags */
+#define VFS_WALK_RCU			BIT(0)
+
 /* per-position flags passed to the vfs_walk_ancestors() callback */
 #define VFS_WALK_POS_DISCONNECTED	BIT(0)
 /* the position is a mountpoint landed on by a mount crossing */
-- 
2.55.0


  parent reply	other threads:[~2026-10-06  0:20 UTC|newest]

Thread overview: 21+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-06  0:20 [RFC PATCH bpf-next 00/12] fs: unified VFS ancestor walk for Landlock and BPF Justin Suess
2026-10-06  0:20 ` [RFC PATCH bpf-next 01/12] namei: introduce __path_walk_parent() Justin Suess
2026-10-06  0:20 ` [RFC PATCH bpf-next 02/12] namei: add vfs_walk_ancestors() Justin Suess
2026-10-06  0:20 ` [RFC PATCH bpf-next 03/12] landlock: convert ancestor walk to vfs_walk_ancestors() Justin Suess
2026-10-06  1:10   ` bot+bpf-ci
2026-10-06  0:20 ` [RFC PATCH bpf-next 04/12] bpf: mark struct path trusted Justin Suess
2026-10-06  1:10   ` bot+bpf-ci
2026-10-06  0:20 ` [RFC PATCH bpf-next 05/12] namei: make vfs_walk_ancestors() stepwise Justin Suess
2026-10-06  0:20 ` [RFC PATCH bpf-next 06/12] bpf: add a path ancestor iterator Justin Suess
2026-10-06  1:11   ` bot+bpf-ci
2026-10-06  0:20 ` [RFC PATCH bpf-next 07/12] selftests/bpf: exercise the " Justin Suess
2026-10-06  1:10   ` bot+bpf-ci
2026-10-06  0:20 ` [RFC PATCH bpf-next 08/12] fs: add mnt_undo_legitimize() Justin Suess
2026-10-06  0:20 ` Justin Suess [this message]
2026-10-06  1:10   ` [RFC PATCH bpf-next 09/12] namei: add an rcu-walk mode to the ancestor walk bot+bpf-ci
2026-10-06  0:20 ` [RFC PATCH bpf-next 10/12] bpf: support "__uninit" iterator arguments in generic kfuncs Justin Suess
2026-10-06  0:20 ` [RFC PATCH bpf-next 11/12] bpf: add a lockless path ancestor iterator Justin Suess
2026-10-06  1:10   ` bot+bpf-ci
2026-10-06 14:44     ` Justin Suess
2026-10-06  0:20 ` [RFC PATCH bpf-next 12/12] selftests/bpf: exercise the " Justin Suess
2026-10-06  1:10   ` bot+bpf-ci

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261006002020.2890858-10-utilityemal77@gmail.com \
    --to=utilityemal77@gmail.com \
    --cc=amir73il@gmail.com \
    --cc=andrii@kernel.org \
    --cc=ast@kernel.org \
    --cc=bpf@vger.kernel.org \
    --cc=brauner@kernel.org \
    --cc=daniel@iogearbox.net \
    --cc=eddyz87@gmail.com \
    --cc=gnoack@google.com \
    --cc=jack@suse.cz \
    --cc=jlayton@kernel.org \
    --cc=jmorris@namei.org \
    --cc=john.fastabend@gmail.com \
    --cc=jolsa@kernel.org \
    --cc=linux-fsdevel@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-security-module@vger.kernel.org \
    --cc=m@maowtm.org \
    --cc=martin.lau@linux.dev \
    --cc=memxor@gmail.com \
    --cc=mic@digikod.net \
    --cc=mjguzik@gmail.com \
    --cc=neil@brown.name \
    --cc=paul@paul-moore.com \
    --cc=serge@hallyn.com \
    --cc=shuah@kernel.org \
    --cc=song@kernel.org \
    --cc=viro@zeniv.linux.org.uk \
    --cc=yonghong.song@linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®