From: Zhihao Cheng <chengzhihao1@huawei.com>
To: Nikolaos Karaolidis <nick@karaolidis.com>,
Christian Brauner <brauner@kernel.org>,
Alexander Viro <viro@zeniv.linux.org.uk>
Cc: Jan Kara <jack@suse.cz>, Mateusz Guzik <mjguzik@gmail.com>,
<linux-fsdevel@vger.kernel.org>, <linux-mm@kvack.org>,
<linux-kernel@vger.kernel.org>, <stable@vger.kernel.org>
Subject: Re: [PATCH] fs: don't lose inodes that another LRU walker is isolating
Date: Sat, 10 Oct 2026 11:17:46 +0800 [thread overview]
Message-ID: <a2f5f4f7-b1ec-bf23-d990-aebbdf6de177@huawei.com> (raw)
In-Reply-To: <20261009142431.713875-1-nick@karaolidis.com>
在 2026/10/9 22:24, Nikolaos Karaolidis 写道:
> When inode_lru_isolate() finds an unused inode that still has page
> cache, it sets I_LRU_ISOLATING, drops i_lock and the list_lru lock,
> invalidates the page cache, clears the flag again and returns
> LRU_RETRY so that the restarted walk can free the inode.
>
> Another walker of the same list can reach the inode in that window.
> It sees a non-zero i_state and takes the inode off the LRU, as it
> would an inode that is in use. Page cache deletion requeues the inode
> once the mapping is shrinkable, so this is harmless if it happens
> before the last deletion. If it happens after, nothing puts the inode
> back: it stays in memory, unused, clean and without page cache, out
> of reach of the shrinker and of drop_caches, until it is looked up
> again or the filesystem is unmounted.
>
> Before commit 2a0629834cd8 ("vfs: Don't evict inode under the inode
> lru traversing context") the pin was an i_count reference, and the
> final iput() put the inode back on the LRU.
>
> Skip inodes that are being isolated and leave them to the walker that
> pinned them, which restarts its walk once it has unpinned them.
>
> The window is reached by any unused inode with page cache on
> CONFIG_HIGHMEM kernels, and otherwise by inodes whose page cache is a
> single shadow entry at index 0. On an i386 HIGHMEM guest, eight
> concurrent "echo 2 > /proc/sys/vm/drop_caches" over 2000 freshly read
> one-page ext4 files leave roughly 400 of the inodes off the LRU per
> round; on x86_64, with the files' pages first evicted through
> memory.reclaim, 3 to 4 per round. With this patch, none in either case.
>
> Fixes: 2a0629834cd8 ("vfs: Don't evict inode under the inode lru traversing context")
> Cc: stable@vger.kernel.org
> Assisted-by: LLM
> Signed-off-by: Nikolaos Karaolidis <nick@karaolidis.com>
> ---
>
> Found while working on reclaim of inodes pinned by page cache in
> ZONE_MOVABLE.
>
> Tested on master af32da41b032 (v7.3-rc6-232) in QEMU/KVM with a small
> stress program: mount an ext4 image, read 2000 one-page files, run 8
> concurrent "echo 2 > /proc/sys/vm/drop_caches", drop caches again
> single-threaded (2, then 1, then 2), and count the filesystem's inodes
> still in memory via how many the umount frees:
>
> i386 HIGHMEM: ~390 of 2000 lost per round without the patch
> (3 runs x 20 rounds), 0 with it
> x86_64, pages evicted through memory.reclaim first (single shadow
> entry left): ~4 per round without, 0 with
> x86_64 with KASAN, PROVE_LOCKING, DEBUG_VM, DEBUG_LIST:
> 6109 lost in 50 rounds without, 0 with, no splats
>
> The reproducer is a single C file run as init; I can post it if useful.
>
> I went with skipping the inode rather than requeueing it on unpin
> (which would be closer to what the final iput() did before
> 2a0629834cd8) because it avoids the remove/re-add and leaves the inode
> where the pinning walker finds it again. Happy to switch if preferred.
>
> Jan, this applies cleanly before or after "fs: Basic infrastructure for
> offloading inode reclaim" (v2 3/5); happy to rebase onto that series if
> you'd rather take it there.
>
> For stable: 6.18 and older need inode->i_state instead of
> inode_state_read(). I have only tested master.
>
> An LLM was used to check the analysis against the code, write the
> reproducer and draft the changelog; I have reviewed all of it.
>
> fs/inode.c | 6 ++++++
> 1 file changed, 6 insertions(+)
>
Good catch.
Reviewed-by: Zhihao Cheng <chengzhihao1@huawei.com>
> diff --git a/fs/inode.c b/fs/inode.c
> index a9d37be390a1..74dfc379b377 100644
> --- a/fs/inode.c
> +++ b/fs/inode.c
> @@ -942,6 +942,12 @@ static enum lru_status inode_lru_isolate(struct list_head *item,
> if (!spin_trylock(&inode->i_lock))
> return LRU_SKIP;
>
> + /* Another walker is dropping its page cache, leave it to that one */
> + if (inode_state_read(inode) & I_LRU_ISOLATING) {
> + spin_unlock(&inode->i_lock);
> + return LRU_SKIP;
> + }
> +
> /*
> * Inodes can get referenced, redirtied, or repopulated while
> * they're already on the LRU, and this can make them
>
> base-commit: af32da41b0327b9c6a37856ba82b6760d6c8d10e
> --
> 2.55.0
> .
>
prev parent reply other threads:[~2026-10-10 3:17 UTC|newest]
Thread overview: 2+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-09 14:24 Nikolaos Karaolidis
2026-10-10 3:17 ` Zhihao Cheng [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=a2f5f4f7-b1ec-bf23-d990-aebbdf6de177@huawei.com \
--to=chengzhihao1@huawei.com \
--cc=brauner@kernel.org \
--cc=jack@suse.cz \
--cc=linux-fsdevel@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=mjguzik@gmail.com \
--cc=nick@karaolidis.com \
--cc=stable@vger.kernel.org \
--cc=viro@zeniv.linux.org.uk \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®