From: Shakeel Butt <shakeel.butt@linux.dev>
To: Greg Kroah-Hartman <gregkh@linuxfoundation.org>,
Tejun Heo <tj@kernel.org>,
Christian Brauner <christian@brauner.io>
Cc: Sebastian Andrzej Siewior <bigeasy@linutronix.de>,
Meta kernel team <kernel-team@meta.com>,
linux-fsdevel@vger.kernel.org, driver-core@lists.linux.dev,
linux-kernel@vger.kernel.org
Subject: [PATCH 0/3] kernfs: don't hold kernfs_rwsem across dir_emit()
Date: Wed, 9 Sep 2026 17:36:47 -0700 [thread overview]
Message-ID: <20260910003650.1680854-1-shakeel.butt@linux.dev> (raw)
kernfs_fop_readdir() takes kernfs_rwsem for reading and holds it for the
whole listing, dir_emit() included. dir_emit() copies into a userspace
buffer, so it can fault, and under memory pressure that fault goes to
reclaim.
We are hitting this case on the Meta fleet very regularly. Many times it
is below[1], a monitoring daemon. It walks the cgroup tree and faults on
its own getdents(2) buffer with the lock held for reading.
below: page allocation stall for 120 secs: order:0,
mode:0x140dca(GFP_HIGHUSER_MOVABLE|__GFP_ZERO|__GFP_COMP)
nodemask=(null),cpuset=hostcritical.slice,mems_allowed=0
Call Trace:
<TASK>
dump_stack_lvl+0x5d/0x80
__alloc_frozen_pages_noprof+0x5f4d/0x6300
? memcg_list_lru_alloc+0x73/0x320
? ima_file_check+0xd0/0x7d0
vma_alloc_folio_noprof+0x145/0x560
handle_mm_fault+0x17c9/0x2720
? find_vma+0x27/0x30
do_user_addr_fault+0x39f/0x6e0
exc_page_fault+0x8f/0x110
asm_exc_page_fault+0x22/0x30
RIP: 0010:filldir64+0xd7/0x1a0
[Code:/RSP:/RAX:..R15: register block elided]
kernfs_fop_readdir+0x2de/0x420
iterate_dir+0x8c/0x1f0
__se_sys_getdents64+0x61/0xe0
? copy_page_from_iter+0x860/0x860
do_syscall_64+0x6a/0x250
entry_SYSCALL_64_after_hwframe+0x4b/0x53
</TASK>
kernfs_rwsem is the lock every create, remove and rename in the hierarchy
needs, and sysfs and cgroupfs have one per machine. More importantly, the
userspace OOM killers such as systemd-oomd also traverse the cgroupfs
hierarchy and can get stuck behind below, which is stuck in reclaim. They
are supposed to relieve memory pressure on the system, but can get stuck
themselves.
The fix is to not hold the lock across dir_emit(). Copy the name while the
lock is held, drop the lock, emit, and take it again.
That opens a window. With the lock dropped, the entry the listing stopped
on can be removed before the listing picks up again. readdir remembers its
place as a name hash, and the search for a hash that is no longer there can
land on the entry after the missing one. That entry is then stepped over
and never reported.
So patch 1 closes that race, patch 2 is the fix, and patch 3 adds tests.
Patch 1 is worth having on its own. Without patch 2 the same search can
land on the entry before the missing one instead, and report it a second
time, between two getdents(2) calls.
Patches 1 and 2 apply to the vfs-7.4.kernfs branch of the vfs tree.
Patch 3 applies on top of "kernfs: three standalone fixes" [2], which is
still pending on the same branch.
[1] https://github.com/facebookincubator/below
[2] https://lore.kernel.org/20260905191613.3143937-1-shakeel.butt@linux.dev/
Shakeel Butt (3):
kernfs: don't repeat or skip an entry when readdir resumes
kernfs: don't hold kernfs_rwsem across dir_emit()
selftests: cover readdir resuming at a removed entry
fs/kernfs/dir.c | 77 +++++--
.../selftests/filesystems/kernfs_test.c | 200 ++++++++++++++++++
2 files changed, 256 insertions(+), 21 deletions(-)
--
2.53.0-Meta
next reply other threads:[~2026-09-10 0:37 UTC|newest]
Thread overview: 4+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-10 0:36 Shakeel Butt [this message]
2026-09-10 0:36 ` [PATCH 1/3] kernfs: don't repeat or skip an entry when readdir resumes Shakeel Butt
2026-09-10 0:36 ` [PATCH 2/3] kernfs: don't hold kernfs_rwsem across dir_emit() Shakeel Butt
2026-09-10 0:36 ` [PATCH 3/3] selftests: cover readdir resuming at a removed entry Shakeel Butt
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260910003650.1680854-1-shakeel.butt@linux.dev \
--to=shakeel.butt@linux.dev \
--cc=bigeasy@linutronix.de \
--cc=christian@brauner.io \
--cc=driver-core@lists.linux.dev \
--cc=gregkh@linuxfoundation.org \
--cc=kernel-team@meta.com \
--cc=linux-fsdevel@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=tj@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®