* [PATCH] mm/fadvise: skip remote LRU drains for ineligible folios
@ 2026-10-05 22:27 Serapheim Dimitropoulos
2026-10-07 20:18 ` Andrew Morton
0 siblings, 1 reply; 2+ messages in thread
From: Serapheim Dimitropoulos @ 2026-10-05 22:27 UTC (permalink / raw)
To: Andrew Morton, linux-mm
Cc: David Hildenbrand, Matthew Wilcox (Oracle),
Jan Kara, fujunjie, Lorenzo Stoakes, Liam R. Howlett,
Vlastimil Babka, Mike Rapoport, Suren Baghdasaryan, Michal Hocko,
linux-fsdevel, linux-kernel, Serapheim Dimitropoulos
From: Serapheim Dimitropoulos <sdimitropoulos@coreweave.com>
POSIX_FADV_DONTNEED retries invalidation after a global LRU drain whenever
mapping_try_invalidate() reports a failed eviction. This includes failures
for mapped, dirty or writeback folios, which dropping LRU batch references
cannot make evictable while those conditions persist.
Filter those failures in mapping_try_invalidate(), while the folio is still
locked and before deactivation can enqueue another batch reference. Leave
mapping_evict_folio() and its eviction safety checks unchanged.
Use a boolean retry flag instead of a failure count: generic_fadvise() only
needs to decide whether to drain and retry once. This remains a heuristic,
not a test for remote LRU references. Preserve the retry for other failures
on clean, unmapped folios, including failures from filemap_release_folio()
and remove_mapping(), rather than limiting it to the early refcount check.
This follows the problem identified in fujunjie's earlier proposal, with
the filtering kept in mapping_try_invalidate() and a conservative fallback
for other eviction failures.
Link: https://lkml.rescloud.iu.edu/2605.0/07547.html
Link: https://lkml.iu.edu/2605.1/03284.html
Link: https://lkml.iu.edu/2605.1/03732.html
Signed-off-by: Serapheim Dimitropoulos <sdimitropoulos@coreweave.com>
---
Tested baseline and patched kernels in QEMU on ext4, XFS and OverlayFS.
For clean mapped files on each filesystem, 32 POSIX_FADV_DONTNEED calls
produced 32 lru_add_drain_all() calls before the patch and none afterwards.
A separate dirty, unmapped test using ext4 data=journal showed the same
reduction.
Cross-CPU and mixed-range tests confirmed that clean, unmapped folios were
still evicted and the global-drain fallback remained available. Writeback
and concurrent pwrite/mmap/fadvise tests also passed, with file contents
verified after fsync, eviction and refault.
Skipping the retry can miss opportunistic eviction if a folio becomes
eligible immediately afterwards; POSIX_FADV_DONTNEED remains advisory.
---
mm/fadvise.c | 12 ++++++------
mm/internal.h | 2 +-
mm/truncate.c | 22 +++++++++++++---------
3 files changed, 20 insertions(+), 16 deletions(-)
diff --git a/mm/fadvise.c b/mm/fadvise.c
index b63fe2141..daf655846 100644
--- a/mm/fadvise.c
+++ b/mm/fadvise.c
@@ -141,7 +141,7 @@ int generic_fadvise(struct file *file, loff_t offset, loff_t len, int advice)
}
if (end_index >= start_index) {
- unsigned long nr_failed = 0;
+ bool need_drain = false;
/*
* It's common to FADV_DONTNEED right after
@@ -155,14 +155,14 @@ int generic_fadvise(struct file *file, loff_t offset, loff_t len, int advice)
lru_add_drain();
mapping_try_invalidate(mapping, start_index, end_index,
- &nr_failed);
+ &need_drain);
/*
- * The failures may be due to the folio being
- * in the LRU cache of a remote CPU. Drain all
- * caches and try again.
+ * Clean, unmapped folios may still have references in
+ * remote LRU batches. Drain and retry only if a failure
+ * might be resolved by dropping those references.
*/
- if (nr_failed) {
+ if (need_drain) {
lru_add_drain_all();
invalidate_mapping_pages(mapping, start_index,
end_index);
diff --git a/mm/internal.h b/mm/internal.h
index 0ca863f26..1637d4275 100644
--- a/mm/internal.h
+++ b/mm/internal.h
@@ -628,7 +628,7 @@ bool truncate_inode_partial_folio(struct folio *folio, loff_t lstart,
loff_t lend, pgoff_t *pstart, pgoff_t *pend);
long mapping_evict_folio(struct address_space *mapping, struct folio *folio);
unsigned long mapping_try_invalidate(struct address_space *mapping,
- pgoff_t start, pgoff_t end, unsigned long *nr_failed);
+ pgoff_t start, pgoff_t end, bool *need_drain);
/**
* folio_evictable - Test whether a folio is evictable.
diff --git a/mm/truncate.c b/mm/truncate.c
index 5b1b13cf8..42b6a1b63 100644
--- a/mm/truncate.c
+++ b/mm/truncate.c
@@ -567,13 +567,14 @@ EXPORT_SYMBOL(truncate_inode_pages_final);
* @mapping: the address_space which holds the folios to invalidate
* @start: the offset 'from' which to invalidate
* @end: the offset 'to' which to invalidate (inclusive)
- * @nr_failed: How many folio invalidations failed
+ * @need_drain: Optional flag, set if a remote LRU drain may help eviction
*
- * This function is similar to invalidate_mapping_pages(), except that it
- * returns the number of folios which could not be evicted in @nr_failed.
+ * This function is similar to invalidate_mapping_pages(), except that it can
+ * indicate whether a remote LRU drain may allow a failed eviction to succeed.
+ * Callers using @need_drain must initialize it to false before the first call.
*/
unsigned long mapping_try_invalidate(struct address_space *mapping,
- pgoff_t start, pgoff_t end, unsigned long *nr_failed)
+ pgoff_t start, pgoff_t end, bool *need_drain)
{
pgoff_t indices[FOLIO_BATCH_SIZE];
struct folio_batch fbatch;
@@ -599,17 +600,20 @@ unsigned long mapping_try_invalidate(struct address_space *mapping,
}
ret = mapping_evict_folio(mapping, folio);
+ /*
+ * Draining LRU batches cannot unmap a folio or clean it.
+ * Other failures may be due to remote batch references.
+ */
+ if (!ret && need_drain && !folio_mapped(folio) &&
+ !folio_test_dirty(folio) && !folio_test_writeback(folio))
+ *need_drain = true;
folio_unlock(folio);
/*
* Invalidation is a hint that the folio is no longer
* of interest and try to speed up its reclaim.
*/
- if (!ret) {
+ if (!ret)
deactivate_file_folio(folio);
- /* Likely in the lru cache of a remote CPU */
- if (nr_failed)
- (*nr_failed)++;
- }
count += ret;
}
---
base-commit: 0aaec43576cd41474a8e90c919631c0c9fc88417
change-id: 20261005-fadvise-lru-drain-d5b75785c9fd
Best regards,
--
Serapheim Dimitropoulos <sdimitropoulos@coreweave.com>
^ permalink raw reply [flat|nested] 2+ messages in thread
* Re: [PATCH] mm/fadvise: skip remote LRU drains for ineligible folios
2026-10-05 22:27 [PATCH] mm/fadvise: skip remote LRU drains for ineligible folios Serapheim Dimitropoulos
@ 2026-10-07 20:18 ` Andrew Morton
0 siblings, 0 replies; 2+ messages in thread
From: Andrew Morton @ 2026-10-07 20:18 UTC (permalink / raw)
To: Serapheim Dimitropoulos
Cc: linux-mm, David Hildenbrand, Matthew Wilcox (Oracle),
Jan Kara, fujunjie, Lorenzo Stoakes, Liam R. Howlett,
Vlastimil Babka, Mike Rapoport, Suren Baghdasaryan, Michal Hocko,
linux-fsdevel, linux-kernel, Serapheim Dimitropoulos
On Tue, 06 Oct 2026 00:27:19 +0200 Serapheim Dimitropoulos <serapheimd@gmail.com> wrote:
> From: Serapheim Dimitropoulos <sdimitropoulos@coreweave.com>
>
> POSIX_FADV_DONTNEED retries invalidation after a global LRU drain whenever
> mapping_try_invalidate() reports a failed eviction. This includes failures
> for mapped, dirty or writeback folios, which dropping LRU batch references
> cannot make evictable while those conditions persist.
>
> Filter those failures in mapping_try_invalidate(), while the folio is still
> locked and before deactivation can enqueue another batch reference. Leave
> mapping_evict_folio() and its eviction safety checks unchanged.
>
> Use a boolean retry flag instead of a failure count: generic_fadvise() only
> needs to decide whether to drain and retry once. This remains a heuristic,
> not a test for remote LRU references. Preserve the retry for other failures
> on clean, unmapped folios, including failures from filemap_release_folio()
> and remove_mapping(), rather than limiting it to the early refcount check.
>
> This follows the problem identified in fujunjie's earlier proposal, with
> the filtering kept in mapping_try_invalidate() and a conservative fallback
> for other eviction failures.
It's not particularly clear from the above, but this appears to be a
performance optimization. No increase in POSIX_FADV_DONTNEED's success
rate is expected?
> Link: https://lkml.rescloud.iu.edu/2605.0/07547.html
> Link: https://lkml.iu.edu/2605.1/03284.html
> Link: https://lkml.iu.edu/2605.1/03732.html
> Signed-off-by: Serapheim Dimitropoulos <sdimitropoulos@coreweave.com>
> ---
> Tested baseline and patched kernels in QEMU on ext4, XFS and OverlayFS.
> For clean mapped files on each filesystem, 32 POSIX_FADV_DONTNEED calls
> produced 32 lru_add_drain_all() calls before the patch and none afterwards.
> A separate dirty, unmapped test using ext4 data=journal showed the same
> reduction.
The whole point of the patch is a performance optimization, so it lives
or dies by measurements. This vital info shouldn't be below the ---
throwaway line!
Also, fujunjie's original had timing measurements, which are nice to
see.
Anyway, I totally believe that this makes things faster so there's no
need to do much work on this. Just saying.
We're in lockdown mode for this -rc cycle so please await review and
plan to respin/resend after next -rc1, thanks.
^ permalink raw reply [flat|nested] 2+ messages in thread
end of thread, other threads:[~2026-10-07 20:18 UTC | newest]
Thread overview: 2+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-05 22:27 [PATCH] mm/fadvise: skip remote LRU drains for ineligible folios Serapheim Dimitropoulos
2026-10-07 20:18 ` Andrew Morton
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®