From: Andrew Morton <akpm@linux-foundation.org>
To: Ayush Ranjan <ayushr@modal.com>
Cc: Pedro Falcato <pfalcato@suse.de>, Hugh Dickins <hughd@google.com>,
Matthew Wilcox <willy@infradead.org>, Jan Kara <jack@suse.cz>,
Baolin Wang <baolin.wang@linux.alibaba.com>,
David Hildenbrand <david@kernel.org>,
Gregory Price <gourry@gourry.net>,
linux-mm@kvack.org, linux-fsdevel@vger.kernel.org,
linux-kernel@vger.kernel.org
Subject: Re: [BUG] shmem: FALLOC_FL_PUNCH_HOLE vs fault-around race corrupts page cache / rss counters
Date: Sun, 4 Oct 2026 22:28:31 -0700 [thread overview]
Message-ID: <20261004222831.bdd648a2b406f02747dd9e40@linux-foundation.org> (raw)
In-Reply-To: <20261003033107.1488699-1-ayushr@modal.com>
On Sat, 3 Oct 2026 03:31:03 +0000 Ayush Ranjan <ayushr@modal.com> wrote:
> Hi all,
>
> Gentle ping on this. The reproducer in my previous mail [1] triggers
> "Bad page cache ... still mapped when deleted" on 6.18.46 within a
> couple of minutes on a 128-CPU bare-metal box, with no fork() and no
> gVisor involved.
>
> We continue to hit this in production at low frequency, so I am happy
> to test patches or collect more data if that would help.
>
fwiw I made gpt and gemini argue about this for a while and ended up
with the below.
From: Andrew Morton <akpm@linux-foundation.org>
Subject: mm: shmem: serialize fault-around against hole punching
Date: Sun Oct 4 09:05:31 PM PDT 2026
shmem uses filemap_map_pages() for fault-around. Unlike shmem_fault(),
filemap_map_pages() does not participate in shmem's fallocate exclusion
protocol.
During a partial hole punch of a large shmem folio, the folio can be split
and the resulting smaller folios can remain temporarily visible in the page
cache while shmem_undo_range() restarts its walk. Fault-around can then map
one of those folios again before the hole-punch path removes it.
shmem_undo_range() may subsequently delete that folio from the page cache
despite the new userspace mapping. This can trigger "still mapped when
deleted" warnings and leave stale mappings or inconsistent RSS/page-table
accounting behind.
The race dates back to d7c1755179b8 ("mm: implement ->map_pages for
shmem/tmpfs"), which enabled generic fault-around for shmem without making
it participate in shmem's hole-punch exclusion protocol.
Serialize shmem fault-around against the hole-punch unmap/truncate sequence
with mapping->invalidate_lock. The hole-punch side holds the lock
exclusively while unmapping and removing pages.
do_fault_around() invokes ->map_pages() under rcu_read_lock(), so the
fault-around side cannot block on invalidate_lock. Use the shared trylock
instead. If the trylock fails, skip fault-around and let the normal shmem
fault path handle the fault. Otherwise hold the shared lock while
filemap_map_pages() installs mappings.
This preserves fault-around in the uncontended case while ensuring that
pages in a punched range cannot be remapped between unmap_mapping_range()
and shmem_truncate_range().
Fixes: d7c1755179b8 ("mm: implement ->map_pages for shmem/tmpfs")
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
---
mm/shmem.c | 27 +++++++++++++++++++++++++--
1 file changed, 25 insertions(+), 2 deletions(-)
--- a/mm/shmem.c~a
+++ a/mm/shmem.c
@@ -2938,6 +2938,27 @@ static vm_fault_t shmem_fault(struct vm_
return ret;
}
+/*
+ * A hole punch can temporarily leave split folios visible in the page cache
+ * after it has unmapped the range. Do not let fault-around map them again
+ * before shmem_truncate_range() removes them. ->map_pages() runs under
+ * rcu_read_lock(), so this exclusion must be non-blocking.
+ */
+static vm_fault_t shmem_map_pages(struct vm_fault *vmf,
+ pgoff_t start_pgoff, pgoff_t end_pgoff)
+{
+ struct address_space *mapping = vmf->vma->vm_file->f_mapping;
+ vm_fault_t ret;
+
+ if (!filemap_invalidate_trylock_shared(mapping))
+ return 0;
+
+ ret = filemap_map_pages(vmf, start_pgoff, end_pgoff);
+ filemap_invalidate_unlock_shared(mapping);
+
+ return ret;
+}
+
unsigned long shmem_get_unmapped_area(struct file *file,
unsigned long uaddr, unsigned long len,
unsigned long pgoff, unsigned long flags)
@@ -3865,10 +3886,12 @@ static long shmem_fallocate(struct file
WRITE_ONCE(inode->i_private, &shmem_falloc);
spin_unlock(&inode->i_lock);
+ filemap_invalidate_lock(mapping);
if ((u64)unmap_end > (u64)unmap_start)
unmap_mapping_range(mapping, unmap_start,
1 + unmap_end - unmap_start, 0);
shmem_truncate_range(inode, offset, offset + len - 1);
+ filemap_invalidate_unlock(mapping);
/* No need to unmap again: hole-punching leaves COWed pages */
spin_lock(&inode->i_lock);
@@ -5467,7 +5490,7 @@ static const struct super_operations shm
static const struct vm_operations_struct shmem_vm_ops = {
.fault = shmem_fault,
- .map_pages = filemap_map_pages,
+ .map_pages = shmem_map_pages,
#ifdef CONFIG_NUMA
.set_policy = shmem_set_policy,
.get_policy = shmem_get_policy,
@@ -5479,7 +5502,7 @@ static const struct vm_operations_struct
static const struct vm_operations_struct shmem_anon_vm_ops = {
.fault = shmem_fault,
- .map_pages = filemap_map_pages,
+ .map_pages = shmem_map_pages,
#ifdef CONFIG_NUMA
.set_policy = shmem_set_policy,
.get_policy = shmem_get_policy,
_
prev parent reply other threads:[~2026-10-05 5:28 UTC|newest]
Thread overview: 11+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-24 6:16 Ayush Ranjan
2026-09-24 7:12 ` David Hildenbrand (Arm)
2026-09-24 8:34 ` Pedro Falcato
2026-09-24 9:15 ` Jan Kara
2026-09-25 5:32 ` Ayush Ranjan
2026-09-24 9:30 ` Baolin Wang
2026-09-25 5:33 ` Ayush Ranjan
2026-09-25 5:30 ` Ayush Ranjan
2026-09-25 6:50 ` Ayush Ranjan
2026-10-03 3:31 ` Ayush Ranjan
2026-10-05 5:28 ` Andrew Morton [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261004222831.bdd648a2b406f02747dd9e40@linux-foundation.org \
--to=akpm@linux-foundation.org \
--cc=ayushr@modal.com \
--cc=baolin.wang@linux.alibaba.com \
--cc=david@kernel.org \
--cc=gourry@gourry.net \
--cc=hughd@google.com \
--cc=jack@suse.cz \
--cc=linux-fsdevel@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=pfalcato@suse.de \
--cc=willy@infradead.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®