mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v2 0/2] fsnotify: drop SRCU across inotify event delivery
@ 2026-10-08 12:02 Jia Zhu
  2026-10-08 12:02 ` [PATCH v2 1/2] inotify: drop fsnotify SRCU protection during " Jia Zhu
  2026-10-08 12:02 ` [PATCH v2 2/2] fsnotify: avoid unrelated mark reaper waits during group teardown Jia Zhu
  0 siblings, 2 replies; 3+ messages in thread
From: Jia Zhu @ 2026-10-08 12:02 UTC (permalink / raw)
  To: Jan Kara; +Cc: Jia Zhu, amir73il, linux-fsdevel, linux-kernel

We observed fsnotify teardown stalls in production. NMI backtraces
showed inotify_handle_inode_event() in memcg reclaim, and the mark
reaper, inotifywait and a systemd process were reported blocked for more
than 327 seconds.

We checked upstream and found the same issue: inotify allocates events
under fsnotify_mark_srcu, and the global mark reaper waits in
synchronize_srcu().

The production report shows multiple inotify release paths blocked
behind the global mark reaper. Because SRCU and the reaper are shared
across groups, a stalled delivery can delay other groups' close() and
exit paths.

Patch 1 pins the iterator marks and drops SRCU around inotify delivery.
It retains fanotify's pinning semantics and preserves event delivery to
a surviving inotify watch when another watch is removed. This leaves the
reclaim stall itself unchanged.

Patch 2 skips the global reaper flush when only the closing group
reference remains. As Jan noted on the earlier patch [1], this only
helps groups with no outstanding marks. It is optional and can be
dropped.

The teardown fast path was checked in QEMU with a test-only SRCU holder.

[1] https://lore.kernel.org/all/4467cbj7eivubtdwcawvr7mhrb7armgoc4xfqjjnncrbkxxag6@3mmou6nhviao/

Jia Zhu (2):
  inotify: drop fsnotify SRCU protection during event delivery
  fsnotify: avoid unrelated mark reaper waits during group teardown

 fs/notify/fsnotify.c             | 15 ++++++++++++---
 fs/notify/group.c                | 26 +++++++++++++++-----------
 fs/notify/inotify/inotify_user.c |  3 ++-
 fs/notify/mark.c                 | 26 ++++++++++++++++++++++----
 include/linux/fsnotify_backend.h |  2 ++
 5 files changed, 53 insertions(+), 19 deletions(-)


base-commit: df2908090cda368b01ff43709f51890076c56157
-- 
2.20.1

^ permalink raw reply	[flat|nested] 3+ messages in thread

* [PATCH v2 1/2] inotify: drop fsnotify SRCU protection during event delivery
  2026-10-08 12:02 [PATCH v2 0/2] fsnotify: drop SRCU across inotify event delivery Jia Zhu
@ 2026-10-08 12:02 ` Jia Zhu
  2026-10-08 12:02 ` [PATCH v2 2/2] fsnotify: avoid unrelated mark reaper waits during group teardown Jia Zhu
  1 sibling, 0 replies; 3+ messages in thread
From: Jia Zhu @ 2026-10-08 12:02 UTC (permalink / raw)
  To: Jan Kara; +Cc: Jia Zhu, amir73il, linux-fsdevel, linux-kernel

Inotify event allocation can enter reclaim while holding the global
fsnotify_mark_srcu read lock. This delays mark reclamation and can stall
close or task exit in unrelated groups with marks awaiting destruction.

Share the user-wait pinning implementation to pin iterator marks and drop
SRCU around inotify event delivery. Preserve the all-or-nothing behavior
of the existing fanotify user-wait interface.

Inotify watches on a directory and its child are independent. Skip marks
that can no longer be pinned and clear their report bits, so removing one
watch does not suppress delivery to the surviving watch. Keep all remaining
iterator heads pinned, including those belonging to other groups, until
SRCU is reacquired. If no reportable marks remain, roll back the pins and
return with SRCU still held.

This lets unrelated mark reclamation proceed while a callback sleeps.
Teardown still waits for pinned marks before flushing the group's event
queue.

Signed-off-by: Jia Zhu <zhujia.zj@bytedance.com>
---
 fs/notify/fsnotify.c             | 15 ++++++++++++---
 fs/notify/inotify/inotify_user.c |  3 ++-
 fs/notify/mark.c                 | 26 ++++++++++++++++++++++----
 include/linux/fsnotify_backend.h |  2 ++
 4 files changed, 38 insertions(+), 8 deletions(-)

diff --git a/fs/notify/fsnotify.c b/fs/notify/fsnotify.c
index 7e2f330fd2837..78d040399cbe7 100644
--- a/fs/notify/fsnotify.c
+++ b/fs/notify/fsnotify.c
@@ -341,7 +341,8 @@ static int send_to_group(__u32 mask, const void *data, int data_type,
 	__u32 marks_ignore_mask = 0;
 	bool is_dir = mask & FS_ISDIR;
 	struct fsnotify_mark *mark;
-	int type;
+	int type, ret;
+	bool pin_events;
 
 	if (!iter_info->report_mask)
 		return 0;
@@ -375,8 +376,16 @@ static int send_to_group(__u32 mask, const void *data, int data_type,
 						file_name, cookie, iter_info);
 	}
 
-	return fsnotify_handle_event(group, mask, data, data_type, dir,
-				     file_name, cookie, iter_info);
+	pin_events = group->flags & FSNOTIFY_GROUP_PIN_EVENTS;
+	if (pin_events && !fsnotify_prepare_inode_event(iter_info))
+		return 0;
+
+	ret = fsnotify_handle_event(group, mask, data, data_type, dir,
+				    file_name, cookie, iter_info);
+
+	if (pin_events)
+		fsnotify_finish_user_wait(iter_info);
+	return ret;
 }
 
 static struct fsnotify_mark *fsnotify_first_mark(struct fsnotify_mark_connector *const *connp)
diff --git a/fs/notify/inotify/inotify_user.c b/fs/notify/inotify/inotify_user.c
index 5f19c24ec187f..5955e4ae0816e 100644
--- a/fs/notify/inotify/inotify_user.c
+++ b/fs/notify/inotify/inotify_user.c
@@ -644,7 +644,8 @@ static struct fsnotify_group *inotify_new_group(unsigned int max_events)
 	struct inotify_event_info *oevent;
 
 	group = fsnotify_alloc_group(&inotify_fsnotify_ops,
-				     FSNOTIFY_GROUP_USER);
+				     FSNOTIFY_GROUP_USER |
+				     FSNOTIFY_GROUP_PIN_EVENTS);
 	if (IS_ERR(group))
 		return group;
 
diff --git a/fs/notify/mark.c b/fs/notify/mark.c
index b2640d836a712..d0deae0c1ce3a 100644
--- a/fs/notify/mark.c
+++ b/fs/notify/mark.c
@@ -553,7 +553,8 @@ static void fsnotify_put_mark_wake(struct fsnotify_mark *mark)
 	}
 }
 
-bool fsnotify_prepare_user_wait(struct fsnotify_iter_info *iter_info)
+static bool fsnotify_prepare_wait(struct fsnotify_iter_info *iter_info,
+				  bool report_partial)
 	__releases(&fsnotify_mark_srcu)
 {
 	int type;
@@ -564,15 +565,19 @@ bool fsnotify_prepare_user_wait(struct fsnotify_iter_info *iter_info)
 		/* This can fail if mark is being removed */
 		while (mark && !fsnotify_get_mark_safe(mark)) {
 			if (mark->group == iter_info->current_group) {
-				__release(&fsnotify_mark_srcu);
-				goto fail;
+				if (!report_partial)
+					goto fail;
+				/* The next cursor may belong to another group. */
+				iter_info->report_mask &= ~(1U << type);
 			}
-			/* This is a mark in an unrelated group, skip */
 			mark = fsnotify_next_mark(mark);
 			iter_info->marks[type] = mark;
 		}
 	}
 
+	if (report_partial && !iter_info->report_mask)
+		goto fail;
+
 	/*
 	 * Now that all marks are pinned by refcount in the inode / vfsmount / etc
 	 * lists, we can drop SRCU lock, and safely resume the list iteration
@@ -583,11 +588,24 @@ bool fsnotify_prepare_user_wait(struct fsnotify_iter_info *iter_info)
 	return true;
 
 fail:
+	__release(&fsnotify_mark_srcu);
 	for (type--; type >= 0; type--)
 		fsnotify_put_mark_wake(iter_info->marks[type]);
 	return false;
 }
 
+bool fsnotify_prepare_user_wait(struct fsnotify_iter_info *iter_info)
+	__releases(&fsnotify_mark_srcu)
+{
+	return fsnotify_prepare_wait(iter_info, false);
+}
+
+bool fsnotify_prepare_inode_event(struct fsnotify_iter_info *iter_info)
+	__releases(&fsnotify_mark_srcu)
+{
+	return fsnotify_prepare_wait(iter_info, true);
+}
+
 void fsnotify_finish_user_wait(struct fsnotify_iter_info *iter_info)
 	__acquires(&fsnotify_mark_srcu)
 {
diff --git a/include/linux/fsnotify_backend.h b/include/linux/fsnotify_backend.h
index 618eed4d6d724..b3132e98a0c3f 100644
--- a/include/linux/fsnotify_backend.h
+++ b/include/linux/fsnotify_backend.h
@@ -234,6 +234,7 @@ struct fsnotify_group {
 
 #define FSNOTIFY_GROUP_USER	0x01 /* user allocated group */
 #define FSNOTIFY_GROUP_DUPS	0x02 /* allow multiple marks per object */
+#define FSNOTIFY_GROUP_PIN_EVENTS 0x04 /* pin marks during events */
 	int flags;
 	unsigned int owner_flags;	/* stored flags of mark_mutex owner */
 
@@ -938,6 +939,7 @@ extern void fsnotify_put_mark(struct fsnotify_mark *mark);
 struct fsnotify_mark *fsnotify_next_mark(struct fsnotify_mark *mark);
 extern void fsnotify_finish_user_wait(struct fsnotify_iter_info *iter_info);
 extern bool fsnotify_prepare_user_wait(struct fsnotify_iter_info *iter_info);
+bool fsnotify_prepare_inode_event(struct fsnotify_iter_info *iter_info);
 extern void fsnotify_modify_mark_mask(struct fsnotify_mark *mark, u32 set, u32 clear);
 
 static inline void fsnotify_init_event(struct fsnotify_event *event)
-- 
2.20.1

^ permalink raw reply	[flat|nested] 3+ messages in thread

* [PATCH v2 2/2] fsnotify: avoid unrelated mark reaper waits during group teardown
  2026-10-08 12:02 [PATCH v2 0/2] fsnotify: drop SRCU across inotify event delivery Jia Zhu
  2026-10-08 12:02 ` [PATCH v2 1/2] inotify: drop fsnotify SRCU protection during " Jia Zhu
@ 2026-10-08 12:02 ` Jia Zhu
  1 sibling, 0 replies; 3+ messages in thread
From: Jia Zhu @ 2026-10-08 12:02 UTC (permalink / raw)
  To: Jan Kara; +Cc: Jia Zhu, amir73il, linux-fsdevel, linux-kernel

fsnotify_destroy_group() unconditionally flushes the global mark reaper,
so an empty inotify instance can hang on unrelated reclamation during
close or task exit:

  do_exit
    __fput
      inotify_release
        fsnotify_destroy_group
          fsnotify_wait_marks_destroyed
            __flush_work

Skip the flush when only the closing reference remains, allowing teardown
to finish without waiting for unrelated SRCU readers.

Verified in QEMU: empty groups and groups whose marks have been reclaimed
exit while an unrelated SRCU reader remains held; groups with pending
reclamation still wait.

Signed-off-by: Jia Zhu <zhujia.zj@bytedance.com>
---
 fs/notify/group.c | 26 +++++++++++++++-----------
 1 file changed, 15 insertions(+), 11 deletions(-)

diff --git a/fs/notify/group.c b/fs/notify/group.c
index b56d1c1d9644a..052a419ee52a6 100644
--- a/fs/notify/group.c
+++ b/fs/notify/group.c
@@ -46,6 +46,7 @@ void fsnotify_group_stop_queueing(struct fsnotify_group *group)
  * the group reference.
  * Note that another thread calling fsnotify_clear_marks_by_group() may still
  * hold a ref to the group.
+ * The caller must hold a reference and exclude new marks.
  */
 void fsnotify_destroy_group(struct fsnotify_group *group)
 {
@@ -67,19 +68,22 @@ void fsnotify_destroy_group(struct fsnotify_group *group)
 	 */
 	wait_event(group->notification_waitq, !atomic_read(&group->user_waits));
 
-	/*
-	 * Wait until all marks get really destroyed. We could actually destroy
-	 * them ourselves instead of waiting for worker to do it, however that
-	 * would be racy as worker can already be processing some marks before
-	 * we even entered fsnotify_destroy_group().
-	 */
-	fsnotify_wait_marks_destroyed();
+	/* Even detached marks hold a group reference until final destruction. */
+	if (refcount_read(&group->refcnt) == 1) {
+		/*
+		 * Pair the refcount read and this barrier with the release
+		 * decrement in fsnotify_put_group() (refcount_dec_and_test()),
+		 * ordering mark destruction before subsequent group teardown.
+		 */
+		smp_mb();
+	} else {
+		fsnotify_wait_marks_destroyed();
+	}
 
 	/*
-	 * Since we have waited for fsnotify_mark_srcu in
-	 * fsnotify_mark_destroy_list() there can be no outstanding event
-	 * notification against this group. So clearing the notification queue
-	 * of all events is reliable now.
+	 * Mark destruction waits for fsnotify_mark_srcu, so there can be no
+	 * outstanding event notification against this group. Clearing the
+	 * notification queue of all events is reliable now.
 	 */
 	fsnotify_flush_notify(group);
 
-- 
2.20.1

^ permalink raw reply	[flat|nested] 3+ messages in thread

end of thread, other threads:[~2026-10-08 12:03 UTC | newest]

Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-08 12:02 [PATCH v2 0/2] fsnotify: drop SRCU across inotify event delivery Jia Zhu
2026-10-08 12:02 ` [PATCH v2 1/2] inotify: drop fsnotify SRCU protection during " Jia Zhu
2026-10-08 12:02 ` [PATCH v2 2/2] fsnotify: avoid unrelated mark reaper waits during group teardown Jia Zhu

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®