* [PATCH RFC v3 00/17] mm/mglru: frequency guided promotion (MGLRU-FG) and flag cleanup
@ 2026-10-03 12:55 Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 01/17] mm/memcontrol: allow update of LRU statistic without holding LRU lock Kairui Song via B4 Relay
` (16 more replies)
0 siblings, 17 replies; 21+ messages in thread
From: Kairui Song via B4 Relay @ 2026-10-03 12:55 UTC (permalink / raw)
To: linux-mm
Cc: Andrew Morton, Johannes Weiner, Muchun Song, Qi Zheng,
Ying Huang, Chris Li, Baoquan He, Nico Pache, Usama Arif,
Michal Hocko, Roman Gushchin, Shakeel Butt, David Hildenbrand,
Lorenzo Stoakes, Barry Song, Axel Rasmussen, Yuanchu Xie, Wei Xu,
Vlastimil Babka, Suren Baghdasaryan, Kemeng Shi, Nhat Pham,
Youngjun Park, Zi Yan, Gregory Price, Matthew Wilcox (Oracle),
Baolin Wang, Ryan Roberts, Dev Jain, Lance Yang, Hugh Dickins,
SeongJae Park, David Rientjes, Yu Zhao, Vernon Yang,
Zicheng Wang, Chen Ridong, Tal Zussman, Kairui Song, Xiang Liu,
Ehab Ababneh, zhaozhengzhuo, Lian Wang, Kunwu Chan, linux-kernel,
cgroups, Kairui Song
Hi all,
This is the updated RFC following the idea proposed at LSF/MM/BPF [1],
based on current mm-new. Slightly adjusted the refs increasement part,
so Android test and build kernel test are given even better result.
Retested most other result and they pretty much just match V2's
improvement. Also, dropped the usage of folio_activate in a cold path
to avoid a potential jitter.
In summary, we can see a 10% - 40% higher performance or lower refault in
various different tests, certain workload gets a dramatically reduce of
runtime, while reducing the page flags usage by 1 bit. The gain here is
mostly from real improvement of LRU's ability to distinguish the hotter
workingset. Tested across multiple servers of different archs, desktops,
and Android, all shows very promising results. Compared to V2, V3
is more effected by anon over-reclaim, but provides over-all better
results, that is a problem that should be fixed later or seperately.
It's already very usable, stable, and performing well, but I'll keep
it RFC as this is a major change to LRU, including changing the
Active/Inactive reading, in a good way I think.
It also fixes several long-standing issues including under-accounted PSI
and poor workingset tracking (especially for page cache).
Test results (CLRU means classical LRU):
Build kernel test:
==================
Running make -j48 in a 3G memcg, using ZRAM as swap (256G, lzo-rle) and
holding the kernel and build output on a NVMe drive, 3 runs of 5 swappiness
configurations (15 builds per kernel) [2]; the patched version is better
than mainline at every swappiness value, measuring the total average:
real sys pgpgin pswpin pswpout refault_file refault_anon
CLRU 1m37s 13m54s 19.21M 2.43M 10.65M 2.21M 2.49M
Before 1m32s 9m12s 12.44M 2.05M 8.86M 472k 2.03M
After 1m31s 8m30s 11.21M 1.90M 8.99M 391k 1.88M
delta -1s -42s -9.9% -7.2% +1.5% -17.1% -7.1%
Same test with all 96 threads busy (-j96), which quadruples the reclaim
pressure (pgfault 109M vs 2M per build), 3-4 runs each:
real sys pgpgin refault_file refault_anon
Before 1m54s 72m12s 111.3M 488k 18.7M
After 1m45s 57m10s 100.0M 362k 17.8M
delta -9.5s -15m(-21%) -10.1% -25.8% -4.9%
Same -j48 test, disk swap instead of ZRAM (SSD-backed, the kernel and build
output stay on NVMe). There is a slight regression vs mainline, mostly
from increased swap-out (+17.5% pswpout): while the file working set is
better protected, the 3G memcg pushes more anon out. It still beats CLRU
by a lot, and file refault is lower. Could be related to recent
upstream changes or over-reclaim of anon.
real sys pgpgin pswpin pswpout refault_file refault_anon
CLRU 4m20s 21m39s 37.54M 3.00M 10.51M 6.22M 3.05M
Before 2m52s 10m11s 10.08M 1.58M 5.04M 422k 1.06M
After 3m04s 11m06s 10.19M 1.68M 5.92M 380k 1.11M
delta +12s +9% +1.1% +6.5% +17.5% -10.1% +4.3%
For reference, test result from V2:
real sys pgpgin pswpin pswpout refault_file refault_anon
CLRU 6m06s 31m01s 50.3M 3.20M 13.8M 10.3M 3.35M
Before 2m58s 10m58s 10.30M 1.60M 5.25M 434k 1.06M
After 2m50s 10m38s 8.79M 1.34M 4.82M 377k 844k
delta -8s -20s -15% -16% -8% -13% -20%
MongoDB YCSB workloadb [3]
==========================
With recordcount:20000000 operationcount:6000000, threads:48,
in a 16G memcg, 3 runs:
CLRU: 98389.94 ops/s
MGLRU Before: 83700.34 ops/s
MGLRU After: 94951.21 ops/s (+13.4%)
There is still a little gap to CLRU, and this is the only test behind
CLRU, which I believe is related to writeback threshold (64 vs 32) which
we can tune later. Test from community didn't show such gap [8].
Chromium & Node.js test [4]
===========================
Using ZRAM as swap, on a 48c96t machine with 128G memory, 64 workers, run
for 1 hour:
Total requests:
CLRU: 63822
MGLRU Before: 132763
MGLRU After: 233774 (+76.0%)
(NOTE: It seems some recent change broken MGLRU's fairness guarteen and
also made this test dramatically faster than a few months ago, which isn't
related to this series and reading are even better now, but I'll take a
deeper look later.)
FIO cached with zipf
====================
Using an NVMe disk, in a 16G cgroup, total file size 40G, this measures
the LRU's theoretical ability to distinguish the hotter portion, 3 test
run each config:
fio --name=fg --numjobs=16 --nrfiles=1 \
--filename_format="$testdir/rnvmedk.\$jobnum.img" \
--size=${FILE_MIB}M \
--buffered=1 --ioengine=sync --rw=randread \
--random_distribution=zipf:$ZIPF --bs=4k --time_based \
--ramp_time=45s --runtime=600s --group_reporting
Avg IOPS (higher is better):
+----------+----------+----------+----------+----------+-----------+
| Config | 0.8 | 0.9 | 1.1 | 1.2 | avg delta |
+----------+----------+----------+----------+----------+-----------+
| CLRU | 400,000 | 601,667 | 2,074,333| 4,354,333| +1.2% |
| Before | 382,667 | 604,667 | 2,071,000| 4,334,667| -- |
| After | 434,667 | 669,000 | 2,301,333| 4,800,000| +11.5% |
+----------+----------+----------+----------+----------+-----------+
Delta per zipf (higher is better): +13.6 / +10.6 / +11.1 / +10.7 %
Throughput-normalized file miss (refault/read, lower is better):
+----------+----------+----------+----------+----------+-----------+
| Config | 0.8 | 0.9 | 1.1 | 1.2 | avg delta |
+----------+----------+----------+----------+----------+-----------+
| CLRU | 0.17868 | 0.11840 | 0.03216 | 0.01274 | -2.8% |
| Before | 0.18237 | 0.12232 | 0.03296 | 0.01321 | -- |
| After | 0.16680 | 0.10994 | 0.02907 | 0.01142 | -11.0% |
+----------+----------+----------+----------+----------+-----------+
Delta per zipf (lower is better): -8.5 / -10.1 / -11.8 / -13.6 %
(MB/s ~ IOPS x 4 KiB; e.g. 4.80M IOPS ~ 18.75 GB/s.)
On the throughput-normalized file miss rate (refaults per read, the
metric that reflects LRU workingset-detection accuracy), unpatched MGLRU
is ~2–3% worse then CLRU across every zipfian access pattern on
this page-cache read workload which the standard model of real
cache-locality skew. This matches the long complained MGLRU cache issue
from community. However the comparable IOPS largely reflects MGLRU's lower
internal LRU/bookkeeping overhead masking the higher miss rate.
And, the patched MGLRU-FG, lowers the miss rate ~8–14% versus default
MGLRU and ~6–10% versus CLRU (best of all three) while raising
IOPS ~10-11% versus both. It detects the workingset more accurately
than all others while retaining MGLRU's lower overhead than CLRU.
So in summary: MGLFU-FG provides a ~10% gain on zipf access on real
high performance disks compared to CLRU, while unpatched MGLRU is ~1-3%
worse than CLRU.
LevelDB Scan/Get
================
I also retested the LevelDB benchmark from the cache_ext paper [5].
Interestingly, mainline MGLRU already beats CLRU on this one after a
recent change in lru_gen_folio_seq that bumps new folios with refs == 1
to the second-oldest generation. That change accidentally gave random
reads a higher hotness level while making sequential reads much colder:
sequential reads involve many readahead hits, and readahead folios start
with refs == 0, so they're already in the oldest generation, and
folio_mark_accessed() on a readahead hit has almost no effect on generations
in mainline MGLRU. Meanwhile, all direct-hit (random read) folios start
with refs == 1 in the second-oldest generation. As a result, the scan-get
test natively protects the "get" part and sacrifices the "scan" part.
That's not the best solution though. It's unreliable because it depends
on LRU drain timing, and it hurts workloads where the sequential part is
actually hotter (any workload involving a hotter large file and many
small cold files will be affected).
This series improves on that base: it covers ordinary workloads without
hurting the scan-get workload and without relying on that initial bump.
LevelDB Scan / Get, Throughput Total:
CLRU: 4668.8 ops/s
MGLRU: 5026.9 ops/s (faster than CLRU, but hurts other workloads)
MGLRU After: 5029.7 ops/s (fastest in all cases, and no regression)
The hot-sequential and cold-random workload can be easily reproduced with
SQLite and grep. SQLite continuously scans and looks up a small hot
portion of a DB file, while grep iterates over a set of small files much
larger than RAM [6]:
SQLite scan & lookup time: Grep iterate time:
CLRU: 14.51ms 13281.37ms
MGLRU mainline: 567.05ms 13694.47ms
MGLRU After this series: 10.58ms 12930.43ms
The grep cold portion is larger than RAM and accessed only once per
iteration, so there's no promotion of any of it. CLRU handles
this reasonably; mainline MGLRU has a clear regression; MGLRU-FG now
not only recovers but is able to catch some hot parts from the cold grep
workload. This test is somewhat subjective, but the signal is clear.
Android (For reference)
=======================
Testing on Android is quite difficult: lack of mainline support, and it is
very noisy to get a stable result. I did an informal backport of the latest
MGLRU-FG patches onto the 6.1 GKI tree of a Pixel 9 Pro (16 GB RAM, 8 GB
ZRAM, Android 17), preserving the frozen kABI layouts so the implementation
is limited. The "before" and "after" kernels share the same base and differ
only by the FG series: the "before" tree has MGLRU aligned to the current
upstream code, so that unrelated backport deltas cancel out. (And BTW that
alignment also improved the performance by a lot, which matches the report
from previous MGLRU reclaim optimization series [12]).
The workload is a churn loop driven over adb: 34 common apps and 20 Chrome
tabs are launched and cycled, with a brief scroll per app during the cold
build of each iteration, 16 iterations per run, about one hour per run. The
loop exhausts memory in every iteration (8 GB ZRAM full, MemFree down to
~80 MB). The two kernels are run in an interleaved rotation and compared on
matched per-iteration samples (same period, same iteration index), which
cancels most of the run-to-run drift. 6 runs per kernel, roughly more
than 18 hours of device time in total.
The memory-management stack in Android is built around the non-FG MGLRU
behaviour, and the 6.1 base lacks some of the upstream infrastructure. Even
so, FG holds its own, it lowers reclaim traffic on both the file and the anon
side:
Before After
------------------------------------------------------
workingset_refault_anon 2.69M 2.38M (-11%)
workingset_refault_file 3.65M 3.16M (-14%)
pgscan_anon 8.91M 7.50M (-16%)
pgsteal_anon 3.42M 2.97M (-13%)
pgscan_file 13.01M 11.07M (-15%)
pgsteal_file 11.20M 9.38M (-16%)
pswpout 3.82M 3.35M (-12%)
pgpgin 40.78M 33.58M (-18%)
("Before" is the aligned tree without the FG series, "After" adds FG. The
numbers are trimmed means of the 6 interleaved runs per kernel: for each
counter the highest and lowest run are dropped and the middle four averaged;
each counter is the vmstat delta of one full run.)
On the 96 matched per-iteration samples the reductions are significant for
every counter in the table, and also for pgmajfault (-12%); only pgpgout is
neutral. An earlier revision of the series was statistically neutral on the
anon side and reduced the file side only, while this version improves both.
I also ran the Android Jank test from Zicheng [7] on this device (chrome
scroll, FrameTimeline), plus a 34-app keepalive run. No obvious difference
was observed: the run-mean chrome scroll fps of every kernel sits in a
100.6 - 103.6 band on a 120 Hz panel, and the jank ratios are all below 8%
per scroll round (n=3-8 gated jank frames per run, so Poisson noise dominates)
with no consistent ordering across kernels. The keepalive test does not
discriminate on this build: 31-32 of 34 apps stay alive with every kernel.
On the previous build a run kept slightly more apps alive with FG (average
4.8 vs 4.5, peak 23 vs 20), that could be noise or a slight improvement.
Both kernels saturate all 8 cores with no obvious CPU usage difference
(99.8-99.9% utilization, ~518s busy per 65s window).
I also did a test on another Android phone with a 5.15 kernel (Xperia 1 V),
which has all apps in one global memcg (this Pixel 9 Pro has each app in a
separate memcg). The result looks much better there, either due to the memcg
layout or the non-reclaiming anon shadow (the 6.1 kernel reclaims anon
shadow). But in either case, the performance is a positive reading.
So in summary: with the final revision we get lower file refaults, lower
pgpgin, and lower anon reclaim and swap traffic, with no measured
user-visible cost.
Others
======
Additionally, PSI, smaps, and readahead should all benefit from better
accuracy since this series unifies the flag usage between classical
LRU and MGLRU.
Other tests such as MySQL are looking fine, with no regressions. There is
also community test report [8].
Refault distance is not included yet, so MGLRU may respond more slowly to
workingset shifts. That can be added later, as previously demonstrated
[9], [10].
Extra note about future development: this series is actually highly
compatible with ideas like workingset reporting [11]. The "gen climbing
folio" design may appear to conflict with workingset reporting's idea,
but it doesn't. The solution is simple and straightforward: once we can
extend the generation number to a larger value (e.g. 64 or 128), the
refs-driven promotion can stop at a lower gen (e.g. oldest_gen + 16),
leaving the remaining newer generations as perfectly time-gap-separated
bins.
The tier count is not fixed either; we'll need to find a way to tune it
if tiers go beyond 4 though, but that shouldn't be hard, not a blocker
either.
More details are in the individual commit messages. LLM is used to help
improve the comments and tests as I'm really not good at that :)
Link: https://lore.kernel.org/linux-mm/CAMgjq7BoekNjg-Ra3C8M7=8=75su38w=HD782T5E_cxyeCeH_g@mail.gmail.com/ [1]
Link: https://lore.kernel.org/linux-mm/CAGsJ_4xre-x0e+qNVm=KLFnO1dbPkPX5RuecqwvTZu-vS+o8yQ@mail.gmail.com/ [2]
Link: https://github.com/brianfrankcooper/YCSB/blob/master/workloads/workloadb [3]
Link: https://lore.kernel.org/all/20221220214923.1229538-1-yuzhao@google.com/ [4]
Link: https://dl.acm.org/doi/10.1145/3731569.3764820 [5]
Link: https://github.com/ryncsn/emm-test-project/tree/master/sqlite-grep [6]
Link: https://github.com/purplewall1206/android-perf-bench [7]
Link: https://lore.kernel.org/linux-mm/92DCEFFD13221261+20260828102344.1537874-1-zhaozhengzhuo@uniontech.com/ [8]
Link: https://lwn.net/Articles/945266/ [9]
Link: https://lore.kernel.org/linux-mm/20260502-mglru-fg-v1-0-913619b014d9@tencent.com/ [10]
Link: https://lwn.net/Articles/976985/ [11]
Link: https://lore.kernel.org/linux-mm/20260417025123.2971253-1-wxy2009nrrr@163.com/ [12]
Assisted-by: LLM
Signed-off-by: Kairui Song <kasong@tencent.com>
Tested-by: zhaozhengzhuo <zhaozhengzhuo@uniontech.com>
---
Changes in v3:
- Micro-optimization for page flags operations: if the flags are
unchanged after calculation, skip the cmpxchg.
- Avoid touching folio_activate even in the cold path. I tried multiple
ways for that lockless promotion, PG_lru in the earlier RFC, and
previous folio_activate, the new speculative barrier + retry seem the
best solution.
- Apply tier cap for MADV_PAGEOUT.
- Only promote folio during page table walk or rmap if the folio
is in the min gen. May worth trying to enlarge the range to
(max_seq - MIN_NR_GENS) in next version.
- Fix a potential folio leak caused by reparenting.
- Retest shows great result especially for Android case.
- Link to v2: https://patch.msgid.link/20260911-mglru-fg-v2-0-f26e5cb26da7@tencent.com
Changes in v2:
- Rebased; dropped v1 03/05/06 (already upstream), folded v1 01 and 07
into patches 01 and 04; new patches 06, 07, 10, 12, 13.
- Make folio_test_workingset() itself arbitrate, dropping the parallel
folio_is_* helpers; convert the last raw PageWorkingset() user
(erofs). (Johannes)
- Use LRU_REF_MAPPED/LRU_REF_EXEC flags instead of is_fault/is_exec
booleans. (Barry)
- Fix syzbot "WARNING in folio_inc_lru_refs" on off-LRU folios.
- Account active/inactive per folio from refs, not the gen window:
/proc/vmstat and memory.stat no longer jump on aging or reverse on
swapless machines.
- Make folio_inc_lru_refs() lockless; add folio_inc_lru_refs_fast()
for the gup fast paths.
- Convert DAMON and khugepaged to the refs-based operations.
- Drop the lru_size WARN_ON_ONCE() and lockdep_assert_held(): the
counter is lockless now, so transient negatives are expected.
- Link to v1: https://patch.msgid.link/20260804-mglru-fg-v1-0-4d8dad39dad6@tencent.com
---
Kairui Song (17):
mm/memcontrol: allow update of LRU statistic without holding LRU lock
mm/mglru: make generation page counters atomic
mm/memcg: add folio-based lruvec live helper
mm/mglru: frequency guided workingset promotion (MGLRU-FG)
mm/mglru: make folio lru referenced times count a generic API
mm/mglru: move add/del LRU size accounting out of lru_gen_update_size()
mm/mglru, gup: mark folios referenced via a fast helper
mm/smap: convert to LRU refs based operations
mm/madvise: adapt for LRU refs based operations in MGLRU
mm/damon: convert to LRU refs based operations
mm/huge_memory: mark file folio as accessed more accurately on split
mm/mglru: folio LRU refs based active/inactive number accounting
mm/mglru: make folio_inc_lru_refs lruvec lockless
mm/mglru: reparent folios from all generations
mm/khugepaged: check folio referenced state via LRU refs under MGLRU
mm/mglru: make folio_test_workingset() work based on folio LRU refs
Documentation/mm: multi-gen LRU: update for frequency guided promotion
Documentation/mm/multigen_lru.rst | 56 +++--
fs/btrfs/compression.c | 1 +
fs/erofs/zdata.c | 3 +-
fs/proc/task_mmu.c | 22 +-
include/linux/memcontrol.h | 45 +++-
include/linux/mm_inline.h | 370 ++++++++++++++++-----------
include/linux/mmzone.h | 168 +++++++++----
include/linux/page-flags.h | 2 -
kernel/bounds.c | 2 +-
mm/damon/paddr.c | 9 +-
mm/filemap.c | 1 +
mm/folio.c | 85 +------
mm/gup.c | 6 +-
mm/huge_memory.c | 8 +-
mm/khugepaged.c | 4 +-
mm/madvise.c | 57 +++--
mm/memcontrol.c | 6 +-
mm/migrate.c | 2 -
mm/page_io.c | 1 +
mm/vmscan.c | 515 +++++++++++++++++++++++++-------------
mm/workingset.c | 45 ++--
21 files changed, 882 insertions(+), 526 deletions(-)
---
base-commit: 763ad0211c7b587344f03bc4d1299810aeb736f4
change-id: 20260722-mglru-fg-3a2c8574725b
Best regards,
--
Kairui Song <kasong@tencent.com>
^ permalink raw reply [flat|nested] 21+ messages in thread
* [PATCH RFC v3 01/17] mm/memcontrol: allow update of LRU statistic without holding LRU lock
2026-10-03 12:55 [PATCH RFC v3 00/17] mm/mglru: frequency guided promotion (MGLRU-FG) and flag cleanup Kairui Song via B4 Relay
@ 2026-10-03 12:55 ` Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 02/17] mm/mglru: make generation page counters atomic Kairui Song via B4 Relay
` (15 subsequent siblings)
16 siblings, 0 replies; 21+ messages in thread
From: Kairui Song via B4 Relay @ 2026-10-03 12:55 UTC (permalink / raw)
To: linux-mm
Cc: Andrew Morton, Johannes Weiner, Muchun Song, Qi Zheng,
Ying Huang, Chris Li, Baoquan He, Nico Pache, Usama Arif,
Michal Hocko, Roman Gushchin, Shakeel Butt, David Hildenbrand,
Lorenzo Stoakes, Barry Song, Axel Rasmussen, Yuanchu Xie, Wei Xu,
Vlastimil Babka, Suren Baghdasaryan, Kemeng Shi, Nhat Pham,
Youngjun Park, Zi Yan, Gregory Price, Matthew Wilcox (Oracle),
Baolin Wang, Ryan Roberts, Dev Jain, Lance Yang, Hugh Dickins,
SeongJae Park, David Rientjes, Yu Zhao, Vernon Yang,
Zicheng Wang, Chen Ridong, Tal Zussman, Kairui Song, Xiang Liu,
Ehab Ababneh, zhaozhengzhuo, Lian Wang, Kunwu Chan, linux-kernel,
cgroups, Kairui Song
From: Kairui Song <kasong@tencent.com>
To enable moving file pages in folio_mark_accessed directly and lazily
for MGLRU, allow updating the LRU statistic atomically without holding a
lock. It may cause temporary counter underflow, which should be fine as
we still follow final consistency of the counter, and it only serves as
a factor for calculating the reclaim budget in vmscan. A little
inaccuracy has no visible effect.
Signed-off-by: Kairui Song <kasong@tencent.com>
---
include/linux/memcontrol.h | 6 +++---
include/linux/mm_inline.h | 3 +--
mm/memcontrol.c | 6 +++---
3 files changed, 7 insertions(+), 8 deletions(-)
diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h
index 74110a324f9e..bdc5dc925c1d 100644
--- a/include/linux/memcontrol.h
+++ b/include/linux/memcontrol.h
@@ -106,7 +106,7 @@ struct mem_cgroup_per_node {
/* Written on every LRU update and on every reclaim iteration. */
__cacheline_group_begin_aligned(memcg_pn_write_hot);
- long lru_zone_size[MAX_NR_ZONES][NR_LRU_LISTS];
+ atomic_long_t lru_zone_size[MAX_NR_ZONES][NR_LRU_LISTS];
struct mem_cgroup_reclaim_iter iter;
#ifdef CONFIG_MEMCG_NMI_SAFETY_REQUIRES_ATOMIC
/* slab stats for nmi context */
@@ -927,8 +927,8 @@ unsigned long mem_cgroup_get_zone_lru_size(const struct lruvec *lruvec,
const struct mem_cgroup_per_node *mz;
mz = container_of_const(lruvec, struct mem_cgroup_per_node, lruvec);
- val = READ_ONCE(mz->lru_zone_size[zone_idx][lru]);
- if (WARN_ON_ONCE(val < 0))
+ val = atomic_long_read(&mz->lru_zone_size[zone_idx][lru]);
+ if (val < 0)
return 0;
return val;
diff --git a/include/linux/mm_inline.h b/include/linux/mm_inline.h
index ab69b9930893..597f013c8e04 100644
--- a/include/linux/mm_inline.h
+++ b/include/linux/mm_inline.h
@@ -47,11 +47,10 @@ static __always_inline void __update_lru_size(struct lruvec *lruvec,
{
struct pglist_data *pgdat = lruvec_pgdat(lruvec);
- lockdep_assert_held(&lruvec->lru_lock);
WARN_ON_ONCE(nr_pages != (int)nr_pages);
mod_lruvec_state(lruvec, NR_LRU_BASE + lru, nr_pages);
- __mod_zone_page_state(&pgdat->node_zones[zid],
+ mod_zone_page_state(&pgdat->node_zones[zid],
NR_ZONE_LRU_BASE + lru, nr_pages);
}
diff --git a/mm/memcontrol.c b/mm/memcontrol.c
index aad0498a7bd6..2f712b5e17d4 100644
--- a/mm/memcontrol.c
+++ b/mm/memcontrol.c
@@ -1557,8 +1557,8 @@ struct lruvec *folio_lruvec_lock_irqsave(const struct folio *folio,
* @zid: zone id of the accounted pages
* @nr_pages: positive when adding or negative when removing
*
- * This function must be called under lru_lock, just before a page is added
- * to or just after a page is removed from an lru list.
+ * This function must be called when a page is added to or removed from
+ * an lru list. Caller need to protect the lruvec from being freed.
*/
void mem_cgroup_update_lru_size(struct lruvec *lruvec, enum lru_list lru,
int zid, long nr_pages)
@@ -1569,7 +1569,7 @@ void mem_cgroup_update_lru_size(struct lruvec *lruvec, enum lru_list lru,
return;
mz = container_of(lruvec, struct mem_cgroup_per_node, lruvec);
- mz->lru_zone_size[zid][lru] += nr_pages;
+ atomic_long_add(nr_pages, &mz->lru_zone_size[zid][lru]);
}
/**
--
2.55.0
^ permalink raw reply [flat|nested] 21+ messages in thread
* [PATCH RFC v3 02/17] mm/mglru: make generation page counters atomic
2026-10-03 12:55 [PATCH RFC v3 00/17] mm/mglru: frequency guided promotion (MGLRU-FG) and flag cleanup Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 01/17] mm/memcontrol: allow update of LRU statistic without holding LRU lock Kairui Song via B4 Relay
@ 2026-10-03 12:55 ` Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 03/17] mm/memcg: add folio-based lruvec live helper Kairui Song via B4 Relay
` (14 subsequent siblings)
16 siblings, 0 replies; 21+ messages in thread
From: Kairui Song via B4 Relay @ 2026-10-03 12:55 UTC (permalink / raw)
To: linux-mm
Cc: Andrew Morton, Johannes Weiner, Muchun Song, Qi Zheng,
Ying Huang, Chris Li, Baoquan He, Nico Pache, Usama Arif,
Michal Hocko, Roman Gushchin, Shakeel Butt, David Hildenbrand,
Lorenzo Stoakes, Barry Song, Axel Rasmussen, Yuanchu Xie, Wei Xu,
Vlastimil Babka, Suren Baghdasaryan, Kemeng Shi, Nhat Pham,
Youngjun Park, Zi Yan, Gregory Price, Matthew Wilcox (Oracle),
Baolin Wang, Ryan Roberts, Dev Jain, Lance Yang, Hugh Dickins,
SeongJae Park, David Rientjes, Yu Zhao, Vernon Yang,
Zicheng Wang, Chen Ridong, Tal Zussman, Kairui Song, Xiang Liu,
Ehab Ababneh, zhaozhengzhuo, Lian Wang, Kunwu Chan, linux-kernel,
cgroups, Kairui Song
From: Kairui Song <kasong@tencent.com>
No feature change, convert them to atomic so we can update them without
holding the LRU lock. There is no risk of overflow. The reader always
compares and uses zero instead if the counter values are negative. It
follows final consistency.
Signed-off-by: Kairui Song <kasong@tencent.com>
---
include/linux/mm_inline.h | 6 ++----
include/linux/mmzone.h | 2 +-
mm/vmscan.c | 27 ++++++++++++---------------
3 files changed, 15 insertions(+), 20 deletions(-)
diff --git a/include/linux/mm_inline.h b/include/linux/mm_inline.h
index 597f013c8e04..f52f02e8e5be 100644
--- a/include/linux/mm_inline.h
+++ b/include/linux/mm_inline.h
@@ -265,11 +265,9 @@ static inline void lru_gen_update_size(struct lruvec *lruvec, struct folio *foli
VM_WARN_ON_ONCE(old_gen == -1 && new_gen == -1);
if (old_gen >= 0)
- WRITE_ONCE(lrugen->nr_pages[old_gen][type][zone],
- lrugen->nr_pages[old_gen][type][zone] - delta);
+ atomic_long_sub(delta, &lrugen->nr_pages[old_gen][type][zone]);
if (new_gen >= 0)
- WRITE_ONCE(lrugen->nr_pages[new_gen][type][zone],
- lrugen->nr_pages[new_gen][type][zone] + delta);
+ atomic_long_add(delta, &lrugen->nr_pages[new_gen][type][zone]);
/* addition */
if (old_gen < 0) {
diff --git a/include/linux/mmzone.h b/include/linux/mmzone.h
index 3b96d6c7123b..09ce82fa841a 100644
--- a/include/linux/mmzone.h
+++ b/include/linux/mmzone.h
@@ -554,7 +554,7 @@ struct lru_gen_folio {
/* the multi-gen LRU lists, lazily sorted on eviction */
struct list_head folios[MAX_NR_GENS][ANON_AND_FILE][MAX_NR_ZONES];
/* the multi-gen LRU sizes, eventually consistent */
- long nr_pages[MAX_NR_GENS][ANON_AND_FILE][MAX_NR_ZONES];
+ atomic_long_t nr_pages[MAX_NR_GENS][ANON_AND_FILE][MAX_NR_ZONES];
/* the exponential moving average of refaulted */
unsigned long avg_refaulted[ANON_AND_FILE][MAX_NR_TIERS];
/* the exponential moving average of evicted+protected */
diff --git a/mm/vmscan.c b/mm/vmscan.c
index 91295070ca33..0a29dbf9fa3f 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -3495,8 +3495,7 @@ static void reset_batch_size(struct lru_gen_mm_walk *walk)
continue;
walk->nr_pages[gen][type][zone] = 0;
- WRITE_ONCE(lrugen->nr_pages[gen][type][zone],
- lrugen->nr_pages[gen][type][zone] + delta);
+ atomic_long_add(delta, &lrugen->nr_pages[gen][type][zone]);
if (lru_gen_is_active(lruvec, gen))
lru += LRU_ACTIVE;
@@ -4117,11 +4116,8 @@ static bool inc_min_seq(struct lruvec *lruvec, int type, int swappiness)
break;
}
flush_lru_batch(head, &batch_end, target_list);
-
- WRITE_ONCE(lrugen->nr_pages[old_gen][type][zone],
- lrugen->nr_pages[old_gen][type][zone] - delta);
- WRITE_ONCE(lrugen->nr_pages[target_gen][type][zone],
- lrugen->nr_pages[target_gen][type][zone] + delta);
+ atomic_long_sub(delta, &lrugen->nr_pages[old_gen][type][zone]);
+ atomic_long_add(delta, &lrugen->nr_pages[target_gen][type][zone]);
if (!remaining)
return false;
}
@@ -4226,8 +4222,8 @@ static bool inc_max_seq(struct lruvec *lruvec, unsigned long seq, int swappiness
for (type = 0; type < ANON_AND_FILE; type++) {
for (zone = 0; zone < MAX_NR_ZONES; zone++) {
enum lru_list lru = type * LRU_INACTIVE_FILE;
- long delta = lrugen->nr_pages[prev][type][zone] -
- lrugen->nr_pages[next][type][zone];
+ long delta = atomic_long_read(&lrugen->nr_pages[prev][type][zone]) -
+ atomic_long_read(&lrugen->nr_pages[next][type][zone]);
if (!delta)
continue;
@@ -4345,7 +4341,8 @@ static unsigned long lruvec_evictable_size(struct lruvec *lruvec, int swappiness
for (seq = min_seq[type]; seq <= max_seq; seq++) {
gen = lru_gen_from_seq(seq);
for (zone = 0; zone < MAX_NR_ZONES; zone++)
- total += max(READ_ONCE(lrugen->nr_pages[gen][type][zone]), 0L);
+ total += max(atomic_long_read(&lrugen->nr_pages[gen][type][zone]),
+ 0L);
}
}
@@ -4765,7 +4762,7 @@ static void __lru_gen_reparent_memcg(struct lruvec *child_lruvec, struct lruvec
for (i = 0; i < get_nr_gens(child_lruvec, type); i++) {
int gen = lru_gen_from_seq(child_lrugen->max_seq - i);
- long nr_pages = child_lrugen->nr_pages[gen][type][zone];
+ long nr_pages = atomic_long_read(&child_lrugen->nr_pages[gen][type][zone]);
int child_lru_active = lru_gen_is_active(child_lruvec, gen) ? LRU_ACTIVE : 0;
int parent_lru_active = lru_gen_is_active(parent_lruvec, gen) ? LRU_ACTIVE : 0;
@@ -4773,9 +4770,8 @@ static void __lru_gen_reparent_memcg(struct lruvec *child_lruvec, struct lruvec
list_splice_tail_init(&child_lrugen->folios[gen][type][zone],
&parent_lrugen->folios[gen][type][zone]);
- WRITE_ONCE(child_lrugen->nr_pages[gen][type][zone], 0);
- WRITE_ONCE(parent_lrugen->nr_pages[gen][type][zone],
- parent_lrugen->nr_pages[gen][type][zone] + nr_pages);
+ atomic_long_set(&child_lrugen->nr_pages[gen][type][zone], 0);
+ atomic_long_add(nr_pages, &parent_lrugen->nr_pages[gen][type][zone]);
if (lru_gen_is_active(child_lruvec, gen) != lru_gen_is_active(parent_lruvec, gen)) {
__update_lru_size(child_lruvec, lru + child_lru_active, zone, -nr_pages);
@@ -5855,7 +5851,8 @@ static int lru_gen_seq_show(struct seq_file *m, void *v)
char mark = full && seq < min_seq[type] ? 'x' : ' ';
for (zone = 0; zone < MAX_NR_ZONES; zone++)
- size += max(READ_ONCE(lrugen->nr_pages[gen][type][zone]), 0L);
+ size += max(atomic_long_read(&lrugen->nr_pages[gen][type][zone]),
+ 0L);
seq_printf(m, " %10lu%c", size, mark);
}
--
2.55.0
^ permalink raw reply [flat|nested] 21+ messages in thread
* [PATCH RFC v3 03/17] mm/memcg: add folio-based lruvec live helper
2026-10-03 12:55 [PATCH RFC v3 00/17] mm/mglru: frequency guided promotion (MGLRU-FG) and flag cleanup Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 01/17] mm/memcontrol: allow update of LRU statistic without holding LRU lock Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 02/17] mm/mglru: make generation page counters atomic Kairui Song via B4 Relay
@ 2026-10-03 12:55 ` Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 04/17] mm/mglru: frequency guided workingset promotion (MGLRU-FG) Kairui Song via B4 Relay
` (13 subsequent siblings)
16 siblings, 0 replies; 21+ messages in thread
From: Kairui Song via B4 Relay @ 2026-10-03 12:55 UTC (permalink / raw)
To: linux-mm
Cc: Andrew Morton, Johannes Weiner, Muchun Song, Qi Zheng,
Ying Huang, Chris Li, Baoquan He, Nico Pache, Usama Arif,
Michal Hocko, Roman Gushchin, Shakeel Butt, David Hildenbrand,
Lorenzo Stoakes, Barry Song, Axel Rasmussen, Yuanchu Xie, Wei Xu,
Vlastimil Babka, Suren Baghdasaryan, Kemeng Shi, Nhat Pham,
Youngjun Park, Zi Yan, Gregory Price, Matthew Wilcox (Oracle),
Baolin Wang, Ryan Roberts, Dev Jain, Lance Yang, Hugh Dickins,
SeongJae Park, David Rientjes, Yu Zhao, Vernon Yang,
Zicheng Wang, Chen Ridong, Tal Zussman, Kairui Song, Xiang Liu,
Ehab Ababneh, zhaozhengzhuo, Lian Wang, Kunwu Chan, linux-kernel,
cgroups, Kairui Song
From: Kairui Song <kasong@tencent.com>
Add a helper that resolves a stable lruvec for a folio under RCU
without taking the lruvec lock. It takes a folio directly so the
lruvec lookup happens inside the RCU read-side critical section,
which a lruvec-based interface cannot guarantee.
No functional change.
Signed-off-by: Kairui Song <kasong@tencent.com>
---
include/linux/memcontrol.h | 39 +++++++++++++++++++++++++++++++++++++++
1 file changed, 39 insertions(+)
diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h
index bdc5dc925c1d..d0a9229e8cd7 100644
--- a/include/linux/memcontrol.h
+++ b/include/linux/memcontrol.h
@@ -1535,6 +1535,45 @@ static inline void lruvec_lock_irq(struct lruvec *lruvec)
spin_lock_irq(&lruvec->lru_lock);
}
+/**
+ * folio_lruvec_live_get - get a live lruvec for a folio under RCU
+ * @folio: the folio
+ *
+ * Computes @folio's lruvec and walks up to the nearest live ancestor
+ * if the folio's memcg is dying. Paired with folio_lruvec_live_put().
+ * The result may be stale: RCU keeps it alive but does not pin @folio
+ * to it. That is fine as the counters are fixed up on reparenting.
+ *
+ * Return: the live lruvec, with rcu_read_lock held.
+ */
+static inline struct lruvec *folio_lruvec_live_get(struct folio *folio)
+{
+#ifdef CONFIG_MEMCG
+ struct lruvec *lruvec;
+ struct pglist_data *pgdat;
+ struct mem_cgroup *memcg;
+
+ rcu_read_lock();
+ lruvec = folio_lruvec(folio);
+ pgdat = lruvec_pgdat(lruvec);
+ memcg = lruvec_memcg(lruvec);
+ while (unlikely(memcg && css_is_dying(&memcg->css))) {
+ memcg = parent_mem_cgroup(memcg);
+ lruvec = mem_cgroup_lruvec(memcg, pgdat);
+ }
+ return lruvec;
+#else
+ return folio_lruvec(folio);
+#endif
+}
+
+static inline void folio_lruvec_live_put(struct lruvec *lruvec)
+{
+#ifdef CONFIG_MEMCG
+ rcu_read_unlock();
+#endif
+}
+
static inline struct lruvec *lruvec_live_lock_irq(struct lruvec *lruvec)
{
#ifdef CONFIG_MEMCG
--
2.55.0
^ permalink raw reply [flat|nested] 21+ messages in thread
* [PATCH RFC v3 04/17] mm/mglru: frequency guided workingset promotion (MGLRU-FG)
2026-10-03 12:55 [PATCH RFC v3 00/17] mm/mglru: frequency guided promotion (MGLRU-FG) and flag cleanup Kairui Song via B4 Relay
` (2 preceding siblings ...)
2026-10-03 12:55 ` [PATCH RFC v3 03/17] mm/memcg: add folio-based lruvec live helper Kairui Song via B4 Relay
@ 2026-10-03 12:55 ` Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 05/17] mm/mglru: make folio lru referenced times count a generic API Kairui Song via B4 Relay
` (12 subsequent siblings)
16 siblings, 0 replies; 21+ messages in thread
From: Kairui Song via B4 Relay @ 2026-10-03 12:55 UTC (permalink / raw)
To: linux-mm
Cc: Andrew Morton, Johannes Weiner, Muchun Song, Qi Zheng,
Ying Huang, Chris Li, Baoquan He, Nico Pache, Usama Arif,
Michal Hocko, Roman Gushchin, Shakeel Butt, David Hildenbrand,
Lorenzo Stoakes, Barry Song, Axel Rasmussen, Yuanchu Xie, Wei Xu,
Vlastimil Babka, Suren Baghdasaryan, Kemeng Shi, Nhat Pham,
Youngjun Park, Zi Yan, Gregory Price, Matthew Wilcox (Oracle),
Baolin Wang, Ryan Roberts, Dev Jain, Lance Yang, Hugh Dickins,
SeongJae Park, David Rientjes, Yu Zhao, Vernon Yang,
Zicheng Wang, Chen Ridong, Tal Zussman, Kairui Song, Xiang Liu,
Ehab Ababneh, zhaozhengzhuo, Lian Wang, Kunwu Chan, linux-kernel,
cgroups, Kairui Song
From: Kairui Song <kasong@tencent.com>
Complement MGLRU's eviction-time tier-PID protection with access-time
frequency-guided promotion. Introduce a unified set of helpers built based
on referenced (access) count of a folio.
Each access increments a folio's referenced count stored in folio flags
(refs), refs still maps to a logarithmic tier just like before, but with
more formal bit definitions, a few special thresholds are introduced:
LRU_REFS_REFERENCED (1), LRU_REFS_WORKINGSET (2), LRU_REFS_PROTECTED (3),
and LRU_REFS_MAX (7). When refs reaches a certain threshold, the folio is
promoted proactively instead of waiting for the PID controller to kick in.
Also simplify MGLRU's usage of PG_workingset and PG_referenced: they
become the low two bits of the refs count, with the higher bits
provided by LRU_REFS_MASK. This reduces MGLRU's original refs count
bit usage by one, since only one extra bit is now needed to record a
max referenced count of 7, and makes MGLRU's refs accounting more
accurate.
This doesn't affect classical LRU in any way, and it addresses several
shortcomings of MGLRU's old tier-only cache protection model:
- Long feedback loop: protection only activated after enough re-faults,
by which time the hot folios are already evicted, or no longer hot.
- Limited tier resolution: once referenced count exceeded the bits
limit (8 previously), MGLRU could no longer distinguish hotter folios as
they are capped by the tier. And what's worse, PG_workingset forces
a folio to stay on tier 3.
- Eviction hotness reversion: because PID protection activates upon
eviction and always targets the LRU tail, it tends to protect cold tail
folios at the expense of hotter head folios. Once the tail folios
consume the PID protection budget, head folios lose their protection.
- Additionally, the PID cannot distinguish the access time of folios
that share the same reference count, and there are only 4 tiers.
To achieve a frequency-guided framework, this commit introduces and reworks
the LRU_REFS related helpers and definitions; most of the work is done by
the helpers below, and their inline comments describe the details.
- folio_inc_lru_refs(): Called on any cache access (folio_mark_accessed)
or page table access. This is the main helper: it promotes folios
according to their access frequency.
Promotion is lazy: the gen bits and size counters are updated eagerly,
while the list move is deferred to the next isolation. NOTE: For now,
the lruvec lock is unconditionally taken on every on-list access to
block concurrent aging; a lockless fast path will be implemented very
soon in a following commit.
- folio_inc_lru_refs_walk(): Used by the PTE walk path during aging and
by the batched look-around, both without the LRU lock, so max_seq may
be stale. A first page table access only advances a folio out of the
oldest generation (second chance), a second access promotes it to the
newest generation; this also performs lazy promotion.
- folio_inc_lru_refs_isolated(): Used by the rmap check before
eviction. The folio is isolated and hence this doesn't perform
promotion by itself; the folio will be added back to the right gen
upon return according to the access frequency. This path also has a
higher promotion bias.
The eviction-time folio_inc_gen() still handles PID protection, but the
protection ratio is softer than before since proactive promotion is
mostly good enough already. The PID gain factors are relaxed from
(2:3) to (1:2) and the setpoint now spans the cumulative mass of the
tiers below the candidate instead of tier 0 alone. folio_inc_gen()
caps refs at WORKINGSET so the folio retains enough history to stay
above the cold tier. Tier 1 is the fallback tier for PID, and tier 2
is the fallback tier for frequency-guided promotion. The forced
protection for full-refs folios is removed, obsoleted by the proactive
promotion.
Refaults are now activated purely according to access frequency:
the old fault bias applied in folio_add_lru() is simplified, since a
folio's access history is a more consistent signal than the context of
the faulting task. Page table access is still considered a slightly
stronger signal.
This also redefines PG_workingset and PG_referenced as the low two
bits of the refs count, eliminating the old restriction where
LRU_REFS_MASK was only valid when PG_referenced was set, and allows
all paths to use the same encoding consistently. Following this idea,
a workingset folio is now defined as refs >= LRU_REFS_WORKINGSET (2),
matching the active/inactive LRU's definition and giving in-kernel
consumers (PSI, readahead) consistent behavior on MGLRU, fixing the
longstanding issue that these users don't work well with MGLRU.
PG_workingset and PG_referenced are no longer independent flags under
MGLRU. Adjusting existing raw folio_test_*() callers to the new
semantics is left as follow-ups.
Signed-off-by: Kairui Song <kasong@tencent.com>
---
include/linux/mm_inline.h | 83 ++++++------
include/linux/mmzone.h | 144 +++++++++++++++------
kernel/bounds.c | 2 +-
mm/folio.c | 50 +------
mm/vmscan.c | 322 +++++++++++++++++++++++++++++++---------------
mm/workingset.c | 45 ++++---
6 files changed, 397 insertions(+), 249 deletions(-)
diff --git a/include/linux/mm_inline.h b/include/linux/mm_inline.h
index f52f02e8e5be..f778e4056fbe 100644
--- a/include/linux/mm_inline.h
+++ b/include/linux/mm_inline.h
@@ -144,12 +144,13 @@ static inline int lru_hist_from_seq(unsigned long seq)
return seq % NR_HIST_GENS;
}
-static inline int lru_tier_from_refs(int refs, bool workingset)
+static inline int lru_tier_from_refs(unsigned int refs)
{
- VM_WARN_ON_ONCE(refs > BIT(LRU_REFS_WIDTH));
-
- /* see the comment on MAX_NR_TIERS */
- return workingset ? MAX_NR_TIERS - 1 : order_base_2(refs);
+ BUILD_BUG_ON(fls(LRU_REFS_MAX - 1) > MAX_NR_TIERS - 1);
+ VM_WARN_ON_ONCE(refs > LRU_REFS_MAX);
+ if (refs < LRU_REFS_WORKINGSET)
+ return 0;
+ return fls(refs - 1);
}
/**
@@ -187,21 +188,24 @@ static inline int lru_get_gen_flags(unsigned long flags)
* @flags: pointer to the folio flags
* @refs: referenced / access count number, between 0 and LRU_REFS_MAX, inclusive.
*
- * For MGLRU, PG_referenced holds the first ref, and the extra bits hold the
- * remaining refs. For classical LRU the extra bits are not used, so it can
- * also be seen as the refs count never exceeds 1. In both cases, refs == 1
- * means PG_referenced is set and the extra bits are zero, and refs == 0 means
- * PG_referenced and the extra bits are all unset.
+ * For MGLRU, PG_referenced, PG_workingset are used as the lower two bits of
+ * refs counter, and extra bits hold the remaining higher bits. For classical
+ * LRU the extra bits are not used, and the two flags has no direct
+ * relationship with each other, but this helper can still be used to sync
+ * them. For both cases, refs == 0 means these two flags and the extra bits
+ * are all unset. And refs == 1 / 2 / 3 means PG_referenced and PG_workingset
+ * are set in an bit order way, which is more meaningful for MGLRU though.
*/
static inline void lru_set_refs_flags(unsigned long *flags, unsigned int refs)
{
VM_WARN_ON_ONCE(refs > LRU_REFS_MAX);
- BUILD_BUG_ON(LRU_REFS_MAX != (LRU_REFS_MASK >> LRU_REFS_PGOFF) + 1);
-
+ BUILD_BUG_ON(LRU_REFS_MASK & (BIT(PG_referenced) | BIT(PG_workingset)));
*flags &= ~LRU_REFS_FLAGS;
- if (!refs)
- return;
- *flags |= (BIT(PG_referenced) | ((refs - 1UL) << LRU_REFS_PGOFF));
+ if (refs & BIT(0))
+ *flags |= BIT(PG_referenced);
+ if (refs & BIT(1))
+ *flags |= BIT(PG_workingset);
+ *flags |= ((unsigned long)refs >> 2) << LRU_REFS_PGOFF;
}
/**
@@ -212,13 +216,13 @@ static inline void lru_set_refs_flags(unsigned long *flags, unsigned int refs)
*/
static inline int lru_get_refs_flags(unsigned long flags)
{
- if (!(flags & BIT(PG_referenced)))
- return 0;
- /*
- * Return the total number of accesses including PG_referenced. Also see
- * the comment on LRU_REFS_FLAGS.
- */
- return ((flags & LRU_REFS_MASK) >> LRU_REFS_PGOFF) + 1;
+ int refs;
+
+ /* Return the total number of accesses. See the comment above MAX_NR_TIERS. */
+ refs = (flags & BIT(PG_referenced)) ? BIT(0) : 0;
+ refs += (flags & BIT(PG_workingset)) ? BIT(1) : 0;
+ refs += ((flags & LRU_REFS_MASK) >> LRU_REFS_PGOFF) << 2;
+ return refs;
}
static inline int folio_lru_refs(const struct folio *folio)
@@ -236,6 +240,8 @@ static inline void folio_set_lru_refs(struct folio *folio, unsigned int refs)
} while (!try_cmpxchg(folio_flags(folio, 0), &old_flags, new_flags));
}
+void folio_inc_lru_refs(struct folio *folio, unsigned int flags);
+
static inline int folio_lru_gen(const struct folio *folio)
{
return lru_get_gen_flags(READ_ONCE(*const_folio_flags(folio, 0)));
@@ -243,7 +249,7 @@ static inline int folio_lru_gen(const struct folio *folio)
static inline bool lru_gen_is_active(const struct lruvec *lruvec, int gen)
{
- unsigned long max_seq = lruvec->lrugen.max_seq;
+ unsigned long max_seq = READ_ONCE(lruvec->lrugen.max_seq);
VM_WARN_ON_ONCE(gen >= MAX_NR_GENS);
@@ -300,23 +306,24 @@ static inline unsigned long lru_gen_folio_seq(const struct lruvec *lruvec,
bool reclaiming)
{
int gen;
+ int refs = folio_lru_refs(folio);
int type = folio_is_file_lru(folio);
const struct lru_gen_folio *lrugen = &lruvec->lrugen;
/*
- * +-----------------------------------+-----------------------------------+
- * | Accessed through page tables and | Accessed through file descriptors |
- * | promoted by folio_update_gen() | and protected by folio_inc_gen() |
- * +-----------------------------------+-----------------------------------+
- * | PG_active (set while isolated) | |
- * +-----------------+-----------------+-----------------+-----------------+
- * | PG_workingset | PG_referenced | PG_workingset | LRU_REFS_FLAGS |
- * +-----------------------------------+-----------------------------------+
- * |<---------- MIN_NR_GENS ---------->| |
- * |<---------------------------- MAX_NR_GENS ---------------------------->|
+ * +------------------------------------------+------------------------------------------+
+ * | Accessed through page tables and | Accessed through file descriptors |
+ * | promoted by folio_inc_lru_refs_walk() | protected by folio_inc_lru_refs/inc_gen |
+ * +------------------------------------------+------------------------------------------+
+ * | PG_active (set at isolation or refault) | |
+ * +--------------------+---------------------+--------------------+---------------------+
+ * | LRU_REFS_MAX | LRU_REFS_WORKINGSET | LRU_REFS_MAX | LRU_REFS_WORKINGSET |
+ * +------------------------------------------+------------------------------------------+
+ * |<-------------- MIN_NR_GENS ------------->| |
+ * |<----------------------------------- MAX_NR_GENS ----------------------------------->|
*/
if (folio_test_active(folio))
- gen = MIN_NR_GENS - folio_test_workingset(folio);
+ gen = MIN_NR_GENS - (refs >= LRU_REFS_WORKINGSET);
else if (reclaiming)
gen = MAX_NR_GENS;
else if ((!folio_is_file_lru(folio) && !folio_test_swapcache(folio)) ||
@@ -324,7 +331,7 @@ static inline unsigned long lru_gen_folio_seq(const struct lruvec *lruvec,
(folio_test_dirty(folio) || folio_test_writeback(folio))))
gen = MIN_NR_GENS;
else
- gen = MAX_NR_GENS - (folio_test_workingset(folio) || folio_test_referenced(folio));
+ gen = MAX_NR_GENS - (refs >= LRU_REFS_WORKINGSET);
return max(READ_ONCE(lrugen->max_seq) - gen + 1, READ_ONCE(lrugen->min_seq[type]));
}
@@ -338,6 +345,7 @@ static inline bool lru_gen_add_folio(struct lruvec *lruvec, struct folio *folio,
int zone = folio_zonenum(folio);
struct lru_gen_folio *lrugen = &lruvec->lrugen;
+ BUILD_BUG_ON(BIT(LRU_GEN_WIDTH - 1) != MAX_NR_GENS);
VM_WARN_ON_ONCE_FOLIO(gen != -1, folio);
if (folio_test_unevictable(folio) || !lrugen->enabled)
@@ -392,7 +400,6 @@ static inline bool lru_gen_del_folio(struct lruvec *lruvec, struct folio *folio,
*/
static inline void folio_migrate_lru_refs(struct folio *new, const struct folio *old)
{
- BUILD_BUG_ON(LRU_REFS_MASK & BIT(PG_referenced));
folio_set_lru_refs(new, folio_lru_refs(old));
}
#else /* !CONFIG_LRU_GEN */
@@ -422,6 +429,10 @@ static inline bool lru_gen_del_folio(struct lruvec *lruvec, struct folio *folio,
return false;
}
+static inline void folio_inc_lru_refs(struct folio *folio, unsigned int flags)
+{
+}
+
static inline void folio_migrate_lru_refs(struct folio *new, const struct folio *old)
{
if (folio_test_referenced(old))
diff --git a/include/linux/mmzone.h b/include/linux/mmzone.h
index 09ce82fa841a..e3a438bc2719 100644
--- a/include/linux/mmzone.h
+++ b/include/linux/mmzone.h
@@ -453,57 +453,121 @@ enum lruvec_flags {
#define MAX_NR_GENS 4U
/*
- * Each generation is divided into multiple tiers. A folio accessed N times
- * through file descriptors is in tier order_base_2(N). A folio in the first
- * tier (N=0,1) is marked by PG_referenced unless it was faulted in through page
- * tables or read ahead. A folio in the last tier (MAX_NR_TIERS-1) is marked by
- * PG_workingset. A folio in any other tier (1<N<5) between the first and last
- * is marked by additional bits of LRU_REFS_WIDTH in folio->flags.
+ * Each generation is divided into multiple tiers. A folio's referenced
+ * count maps to a tier as shown below:
*
- * In contrast to moving across generations which requires the LRU lock, moving
- * across tiers only involves atomic operations on folio->flags and therefore
- * has a negligible cost in the buffered access path. In the eviction path,
- * comparisons of refaulted/(evicted+protected) from the first tier and the rest
- * infer whether folios accessed multiple times through file descriptors are
- * statistically hot and thus worth protecting.
+ * MGLRU (frequency guidance)
+ * Refs Tier |- Refs: how many times (at least) a folio has been referenced.
+ * 0 0 |- Mostly cold pages, readahead, etc. [1]
+ * 1 0 |= LRU_REFS_REFERENCED: Used at least once. [2]
+ * -WORKINGSET-+|- Pages beyond are workingset and never fall below this floor. [3]
+ * 2 1<-+|= LRU_REFS_WORKINGSET: Classical workingset, accessed twice, protected. [4]
+ * 3 2 |- LRU_REFS_PROTECTED: Protected workingset, promoted pages capped at here. [5]
+ * 4 2 |
+ * 5 3 |- The tier here is MAX_NR_TIERS - 1
+ * 6 3 |
+ * 7 3 |= LRU_REFS_MAX: Promotion candidate. [6]
+ * -PROMOTION->-/
*
- * MAX_NR_TIERS is set to 4 so that the multi-gen LRU can support twice the
- * number of categories of the active/inactive LRU when keeping track of
- * accesses through file descriptors. This uses MAX_NR_TIERS-2 spare bits in
- * folio->flags, masked by LRU_REFS_MASK.
+ * Ideally each tier holds folios of similar access patterns: lower tiers
+ * are less important and evicted faster. A page's reference count and
+ * tier are capped when it changes generation, preventing it from
+ * dominating the new generation based on old-generation access history.
+ * Generation ordering already ensures a newer-gen page is hotter than an
+ * older-gen one regardless of tier.
+ *
+ * Refs tracks accesses from two sources: page table (lazily collected by
+ * the page table aging walk or rmap eviction lookup) and file descriptors
+ * (by folio_mark_accessed). Page table accesses are weighted heavier
+ * because the accessed bit is sticky (undercounts repeated accesses),
+ * passively collected, and page faults are generally more important as
+ * userspace does not expect a memory access to block on reclaim. Both
+ * access types increment refs by one; the result is capped at
+ * LRU_REFS_PROTECTED on promotion or deferral, or LRU_REFS_MAX otherwise.
+ *
+ * 1. Tier is fls(N-1) for N > 1, 0 otherwise. Folios with zero
+ * accesses (refs == 0) are generally cold, e.g. readahead folios.
+ *
+ * Freshly allocated folios start with refs == 0; faulted and mapped
+ * folios have their page table access bit set, so the first page table
+ * access check always sets LRU_REFS_REFERENCED. That is the second
+ * chance described above MIN_NR_GENS: it only moves the folio one
+ * generation forward if it is in the oldest generation.
+ *
+ * 2. Folios accessed once stay on tier 0: one-time usage does not
+ * qualify for protection. A second access advances the folio,
+ * aligning with classical LRU's use-twice threshold. A second page
+ * table access promotes to the latest gen; file access only defers
+ * eviction from the oldest gen.
+ *
+ * 3. Folios accessed at least twice are considered workingset. This
+ * mostly aligns with classical LRU: at least one I/O is saved by
+ * keeping them in memory. Folios at or above this level never fall
+ * below tier 1 (the workingset floor), so tier 0 stays a clean tier
+ * for cold cache while tier 1 serves as the fallback line for
+ * actually reused or historically hot folios.
+ *
+ * Folios refaulted through a page fault at refs 1 will enter the second
+ * newest gen, so faulting will be protected better.
+ *
+ * 4. Starting from tier 1, PID protection sacrifices lower tiers to
+ * protect higher tiers by comparing refault rates for long-term
+ * accuracy, and caps higher refs to this value. Since PID protection
+ * bypasses page table lookup and clearing, when a further eviction
+ * attempt occurs after PID loosens, the folio's page table access is
+ * rechecked and a referenced folio is put back with refs capped to
+ * this value as well. This also gives folios a fair opportunity to be
+ * promoted by file access again.
+ *
+ * Folios refaulted through a page fault at tier 1 or above are activated
+ * and enter the newest gen. Non fault page will enter second oldest gen,
+ * driving aging and workingset shifting.
+ *
+ * 5. Pages beyond the ordinary workingset tier form new tiers for the
+ * PID controller to protect differently. Folios at or above this
+ * level are capped at LRU_REFS_PROTECTED on promotion or deferral,
+ * and at LRU_REFS_WORKINGSET under PID protection in the oldest
+ * generation, where they represent a historical workingset.
+ *
+ * 6. Folios that reach LRU_REFS_MAX are advanced to the next generation
+ * on further access, with refs capped to LRU_REFS_PROTECTED. This
+ * gives them a fair start for advancement to an even newer generation
+ * while keeping hot folios distinguishable.
+ *
+ * Tiering uses PG_referenced and PG_workingset as the lower two bits,
+ * and the bits masked by LRU_REFS_MASK as the higher bits, so the refs
+ * count ranges from 0 to LRU_REFS_MAX. A folio is on the workingset
+ * tier once accessed at least twice, which is more consistent with the
+ * classical LRU.
+ *
+ * A folio's referenced count never goes backwards except upon gen
+ * increase as described above, or when explicitly reset by
+ * lru_gen_clear_refs(). Refault of a reclaimed folio restores
+ * its referenced count, capped at LRU_REFS_PROTECTED, which aligns with
+ * promotion. Page table refaults of previous workingset folios send
+ * them to the latest gen, driving aging faster.
+ *
+ * MAX_NR_TIERS is set to 4 so that the multi-gen LRU can support twice
+ * the number of categories of the active/inactive LRU.
*/
#define MAX_NR_TIERS 4U
#define LRU_TIER_MIN 0U
#define LRU_TIER_MAX (MAX_NR_TIERS - 1)
+/* Access source flags for folio_inc_lru_refs() */
+#define LRU_REF_MAPPED 0x1U
+#define LRU_REF_EXEC 0x2U
+
+#define LRU_REFS_REFERENCED 0x1
+#define LRU_REFS_WORKINGSET 0x2
+#define LRU_REFS_PROTECTED 0x3
+
#ifndef __GENERATING_BOUNDS_H
#define LRU_GEN_MASK ((BIT(LRU_GEN_WIDTH) - 1) << LRU_GEN_PGOFF)
#define LRU_REFS_MASK ((BIT(LRU_REFS_WIDTH) - 1) << LRU_REFS_PGOFF)
-#define LRU_REFS_MAX BIT(LRU_REFS_WIDTH)
-
-/*
- * For folios accessed multiple times through file descriptors,
- * lru_gen_inc_refs() sets additional bits of LRU_REFS_WIDTH in folio->flags
- * after PG_referenced, then PG_workingset after LRU_REFS_WIDTH. After all its
- * bits are set, i.e., LRU_REFS_FLAGS|BIT(PG_workingset), a folio is lazily
- * promoted into the second oldest generation in the eviction path. And when
- * folio_inc_gen() does that, it clears LRU_REFS_FLAGS so that
- * lru_gen_inc_refs() can start over. Note that for this case, LRU_REFS_MASK is
- * only valid when PG_referenced is set.
- *
- * For folios accessed multiple times through page tables, folio_update_gen()
- * from a page table walk or lru_gen_set_refs() from a rmap walk sets
- * PG_referenced after the accessed bit is cleared for the first time.
- * Thereafter, those two paths set PG_workingset and promote folios to the
- * youngest generation. Like folio_inc_gen(), folio_update_gen() also clears
- * PG_referenced. Note that for this case, LRU_REFS_MASK is not used.
- *
- * For both cases above, after PG_workingset is set on a folio, it remains until
- * this folio is either reclaimed, or "deactivated" by lru_gen_clear_refs(). It
- * can be set again if lru_gen_test_recent() returns true upon a refault.
- */
-#define LRU_REFS_FLAGS (LRU_REFS_MASK | BIT(PG_referenced))
+#define LRU_REFS_FLAGS (LRU_REFS_MASK | BIT(PG_referenced) | BIT(PG_workingset))
+#define LRU_REFS_MAX (BIT(LRU_REFS_WIDTH + 2) - 1)
struct lruvec;
struct page_vma_mapped_walk;
diff --git a/kernel/bounds.c b/kernel/bounds.c
index 02b619eb6106..06a034713b5d 100644
--- a/kernel/bounds.c
+++ b/kernel/bounds.c
@@ -25,7 +25,7 @@ int main(void)
DEFINE(SPINLOCK_SIZE, sizeof(spinlock_t));
#ifdef CONFIG_LRU_GEN
DEFINE(LRU_GEN_WIDTH, order_base_2(MAX_NR_GENS + 1));
- DEFINE(__LRU_REFS_WIDTH, MAX_NR_TIERS - 2);
+ DEFINE(__LRU_REFS_WIDTH, MAX_NR_TIERS - 3);
#else
DEFINE(LRU_GEN_WIDTH, 0);
DEFINE(__LRU_REFS_WIDTH, 0);
diff --git a/mm/folio.c b/mm/folio.c
index 35e242b48870..e200efab814f 100644
--- a/mm/folio.c
+++ b/mm/folio.c
@@ -273,7 +273,6 @@ static void lru_activate(struct lruvec *lruvec, struct folio *folio)
if (folio_test_active(folio) || folio_test_unevictable(folio))
return;
-
lruvec_del_folio(lruvec, folio);
folio_set_active(folio);
lruvec_add_folio(lruvec, folio);
@@ -352,32 +351,6 @@ static void __lru_cache_activate_folio(struct folio *folio)
#ifdef CONFIG_LRU_GEN
-static void lru_gen_inc_refs(struct folio *folio)
-{
- unsigned long new_flags, old_flags = READ_ONCE(*folio_flags(folio, 0));
- int refs;
-
- if (folio_test_unevictable(folio))
- return;
-
- /* see the comment on LRU_REFS_FLAGS */
- if (!folio_lru_refs(folio)) {
- folio_set_lru_refs(folio, 1);
- return;
- }
-
- do {
- new_flags = old_flags;
- refs = lru_get_refs_flags(old_flags);
- if (refs == LRU_REFS_MAX) {
- if (!folio_test_workingset(folio))
- folio_set_workingset(folio);
- return;
- }
- lru_set_refs_flags(&new_flags, refs + 1);
- } while (!try_cmpxchg(folio_flags(folio, 0), &old_flags, new_flags));
-}
-
static bool lru_gen_clear_refs(struct folio *folio)
{
int gen = folio_lru_gen(folio);
@@ -388,7 +361,6 @@ static bool lru_gen_clear_refs(struct folio *folio)
return true;
folio_set_lru_refs(folio, 0);
- folio_clear_workingset(folio);
rcu_read_lock();
seq = READ_ONCE(folio_lruvec(folio)->lrugen.min_seq[type]);
@@ -399,10 +371,6 @@ static bool lru_gen_clear_refs(struct folio *folio)
#else /* !CONFIG_LRU_GEN */
-static void lru_gen_inc_refs(struct folio *folio)
-{
-}
-
static bool lru_gen_clear_refs(struct folio *folio)
{
return false;
@@ -428,7 +396,8 @@ void folio_mark_accessed(struct folio *folio)
if (folio_test_dropbehind(folio))
return;
if (lru_gen_enabled()) {
- lru_gen_inc_refs(folio);
+ if (!folio_test_unevictable(folio))
+ folio_inc_lru_refs(folio, 0);
return;
}
@@ -474,21 +443,6 @@ void folio_add_lru(struct folio *folio)
folio_test_unevictable(folio), folio);
VM_BUG_ON_FOLIO(folio_test_lru(folio), folio);
- /*
- * For refaulted workingset folios, set PG_active so they
- * can be added to active generations.
- * For prefaulted file folios, folio_mark_accessed() sets
- * PG_referenced so lru_gen_folio_seq() places them into
- * the second oldest generation.
- */
- if (lru_gen_enabled() && !folio_test_unevictable(folio) &&
- lru_gen_in_fault() && !(current->flags & PF_MEMALLOC)) {
- if (folio_test_workingset(folio))
- folio_set_active(folio);
- else if (!folio_test_referenced(folio))
- folio_mark_accessed(folio);
- }
-
folio_batch_add_and_move(folio, lru_add);
}
EXPORT_SYMBOL(folio_add_lru);
diff --git a/mm/vmscan.c b/mm/vmscan.c
index 0a29dbf9fa3f..a8c49262d5b3 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -931,38 +931,188 @@ enum folio_references {
};
#ifdef CONFIG_LRU_GEN
+/******************************************************************************
+ * Referenced count feedback
+ ******************************************************************************/
+
/*
- * Only used on a mapped folio in the eviction (rmap walk) path, where promotion
- * needs to be done by taking the folio off the LRU list and then adding it back
- * with PG_active set. In contrast, the aging (page table walk) path uses
- * folio_update_gen().
+ * The folio_inc_lru_refs{_*} helpers below collect the referenced info
+ * (hotness) from other parts, including the page table walker, the rmap walk
+ * upon eviction, the rmap lookaround, and file descriptors
+ * (folio_mark_accessed).
+ *
+ * Page table accesses use refs from either source. The first reference
+ * advances only oldest-gen folios by one generation; a second promotes to
+ * the newest. Executable file folios promote on their first access to avoid
+ * I/O thrashing. Isolated rmap accesses use evict_folios() for putback.
+ *
+ * Promotion changes folio->flags; sort_folio() or inc_min_seq() moves the
+ * folio later. Young PTEs found during unmapping use folio_mark_accessed()
+ * and follow the file descriptor rules below.
+ *
+ * File descriptor accesses do not promote. They only defer eviction from
+ * the oldest generation, and only once the folio is a workingset folio
+ * (LRU_REFS_WORKINGSET), leaving the rest to PID protection. Page table
+ * accesses are treated more generously because the accessed bit is sticky
+ * (it under-counts repeated accesses) and because a page fault is more
+ * costly than file descriptor I/O.
+ *
+ * PID protection operates on tier > 0 folios. The one proactive promotion
+ * outside of it and the page table path is the overflow case where the
+ * referenced count exceeds LRU_REFS_MAX, which means the folio is hotter
+ * than everything else in its generation.
+ *
+ * Whenever a folio changes generation here its referenced count is capped at
+ * LRU_REFS_PROTECTED, so it starts at or below the protected tier regardless
+ * of its old-generation access history. PID protection (folio_inc_gen) caps
+ * at LRU_REFS_WORKINGSET independently.
*/
-static bool lru_gen_set_refs(struct folio *folio, const vma_flags_t *vma_flags)
-{
- /* see the comment on LRU_REFS_FLAGS */
- if (!folio_test_referenced(folio) && !folio_test_workingset(folio)) {
- /* Activate file-backed executable folios after first usage. */
- if (is_exec_file_folio(folio, vma_flags)) {
- folio_set_workingset(folio);
- folio_set_lru_refs(folio, 0);
- return true;
+
+/*
+ * Update the folio's lru refs indicator. The caller doesn't need to hold
+ * the folio lock, isolate the folio, or hold the lruvec lock. Used by both
+ * cache access (flags == 0) and page table access (LRU_REF_MAPPED,
+ * optionally with LRU_REF_EXEC).
+ */
+void folio_inc_lru_refs(struct folio *folio, unsigned int flags)
+{
+ int max_gen, min_gen;
+ int type, refs, old_gen, gen;
+ unsigned long new_flags, old_flags, max_seq;
+ struct lru_gen_folio *lrugen;
+ struct lruvec *lruvec = NULL;
+
+ type = folio_is_file_lru(folio);
+ old_flags = READ_ONCE(*folio_flags(folio, 0));
+ do {
+ new_flags = old_flags;
+ old_gen = lru_get_gen_flags(old_flags);
+ refs = lru_get_refs_flags(old_flags) + 1;
+ gen = old_gen;
+ if (old_gen < 0)
+ goto out;
+ /*
+ * Lock the lruvec if the folio is on-list. We are already
+ * doing lazy promotion so in theory we don't need this,
+ * but for now, concurrent aging would still corrupt the
+ * size counters. This is a temporary limitation and
+ * will be lifted very soon, so the lock here is not a
+ * performance concern.
+ */
+ if (!lruvec) {
+ lruvec = lruvec_live_lock_irq(folio_lruvec(folio));
+ lrugen = &lruvec->lrugen;
}
+ max_seq = READ_ONCE(lrugen->max_seq);
+ max_gen = lru_gen_from_seq(max_seq);
+ min_gen = lru_gen_from_seq(READ_ONCE(lrugen->min_seq[type]));
+ if (old_gen == max_gen)
+ goto out;
- folio_set_lru_refs(folio, 1);
- return false;
- }
+ if (flags & (LRU_REF_MAPPED | LRU_REF_EXEC)) {
+ /* Promote second page table access or executable */
+ if (refs > LRU_REFS_REFERENCED || flags & LRU_REF_EXEC)
+ gen = max_gen;
+ /* First access only defers eviction from the oldest gen */
+ else if (old_gen == min_gen)
+ gen = (old_gen + 1) % MAX_NR_GENS;
+ refs = min(refs, LRU_REFS_PROTECTED);
+ } else if (refs > LRU_REFS_MAX) {
+ /* LRU refs counting overflow, bump the gen */
+ gen = (old_gen + 1) % MAX_NR_GENS;
+ refs = LRU_REFS_PROTECTED;
+ } else if (old_gen == min_gen && refs >= LRU_REFS_WORKINGSET) {
+ /* Defer eviction of just accessed workingset */
+ gen = (old_gen + 1) % MAX_NR_GENS;
+ refs = min(refs, LRU_REFS_PROTECTED);
+ }
+out:
+ refs = min(refs, LRU_REFS_MAX);
+ lru_set_refs_flags(&new_flags, refs);
+ if (gen != old_gen)
+ lru_set_gen_flags(&new_flags, gen);
+ if (new_flags == old_flags)
+ break;
+ } while (!try_cmpxchg(folio_flags(folio, 0), &old_flags, new_flags));
- /* Promote on second access */
- if (folio_lru_refs(folio) > 1) {
- folio_set_workingset(folio);
- folio_set_lru_refs(folio, 0);
- } else {
- folio_mark_accessed(folio);
- }
- return true;
+ if (gen != old_gen)
+ lru_gen_update_size(lruvec, folio, old_gen, gen);
+ if (lruvec)
+ lruvec_unlock_irq(lruvec);
+}
+
+/*
+ * Update the folio's lru refs indicator during a page table walk or the
+ * look-around. max_seq can be stale as neither holds the LRU lock.
+ *
+ * Returns the old generation and stores the new generation in @new_gen if
+ * the folio is on the LRU and not in the newest generation, or -1 otherwise.
+ */
+static int folio_inc_lru_refs_walk(struct folio *folio, struct lruvec *lruvec,
+ const vma_flags_t *vma_flags,
+ int *new_gen, int *type)
+{
+ unsigned long new_flags, old_flags = READ_ONCE(*folio_flags(folio, 0));
+ unsigned long max_seq = READ_ONCE(lruvec->lrugen.max_seq);
+ int refs, gen, min_gen, max_gen, ret;
+
+ max_gen = lru_gen_from_seq(max_seq);
+
+ do {
+ gen = lru_get_gen_flags(old_flags);
+ refs = lru_get_refs_flags(old_flags) + 1;
+ *type = folio_flags_is_file_lru(&old_flags);
+ min_gen = lru_gen_from_seq(READ_ONCE(lruvec->lrugen.min_seq[*type]));
+ new_flags = old_flags;
+
+ if (gen >= 0 && gen != max_gen) {
+ ret = gen;
+ /* Promote second page table access or executable */
+ if (refs > LRU_REFS_REFERENCED || is_exec_file_folio(folio, vma_flags))
+ *new_gen = max_gen;
+ /* First access only defers eviction from the oldest gen */
+ else if (gen == min_gen)
+ *new_gen = (gen + 1) % MAX_NR_GENS;
+ else
+ *new_gen = gen;
+ lru_set_gen_flags(&new_flags, *new_gen);
+ lru_set_refs_flags(&new_flags, min(refs, LRU_REFS_PROTECTED));
+ } else {
+ ret = -1;
+ lru_set_refs_flags(&new_flags, min(refs, LRU_REFS_MAX));
+ }
+ if (new_flags == old_flags)
+ break;
+ } while (!try_cmpxchg(folio_flags(folio, 0), &old_flags, new_flags));
+
+ return ret;
+}
+
+/*
+ * Update the lru refs indicator of an isolated folio, only used on
+ * mapped folios upon the final eviction.
+ *
+ * Increments the refs count (capped at LRU_REFS_PROTECTED, evict_folios()
+ * then caps it at LRU_REFS_WORKINGSET). Returns true if the caller should
+ * activate the folio (second access or executable), false to put it back
+ * for a second chance.
+ */
+static bool folio_inc_lru_refs_isolated(struct folio *folio, const vma_flags_t *vma_flags)
+{
+ unsigned long new_flags, old_flags = READ_ONCE(*folio_flags(folio, 0));
+ int refs;
+
+ do {
+ new_flags = old_flags;
+ refs = lru_get_refs_flags(old_flags) + 1;
+ lru_set_refs_flags(&new_flags, min(refs, LRU_REFS_PROTECTED));
+ } while (!try_cmpxchg(folio_flags(folio, 0), &old_flags, new_flags));
+
+ /* Promote second page table access or executable */
+ return refs > LRU_REFS_REFERENCED || is_exec_file_folio(folio, vma_flags);
}
#else
-static bool lru_gen_set_refs(struct folio *folio, const vma_flags_t *vma_flags)
+static bool folio_inc_lru_refs_isolated(struct folio *folio, const vma_flags_t *vma_flags)
{
return false;
}
@@ -997,7 +1147,8 @@ static enum folio_references folio_check_references(struct folio *folio,
if (!referenced_ptes)
return FOLIOREF_RECLAIM;
- return lru_gen_set_refs(folio, &vma_flags) ? FOLIOREF_ACTIVATE : FOLIOREF_KEEP;
+ return folio_inc_lru_refs_isolated(folio, &vma_flags) ?
+ FOLIOREF_ACTIVATE : FOLIOREF_KEEP;
}
referenced_folio = folio_test_clear_referenced(folio);
@@ -3298,9 +3449,9 @@ static bool iterate_mm_list_nowalk(struct lruvec *lruvec, unsigned long seq)
* P term over the generations previously evicted, using the smoothing factor
* 1/2; the D term isn't supported.
*
- * The setpoint (SP) is always the first tier of one type; the process variable
- * (PV) is either any tier of the other type or any other tier of the same
- * type.
+ * To select a type, the setpoint (SP) is all tiers of one type and the process
+ * variable (PV) is all tiers of the other type. To select a tier to protect,
+ * the SP is the tiers below it and the PV is the tier itself.
*
* The error is the difference between the SP and the PV; the correction is to
* turn off protection when SP>PV or turn on protection when SP<PV.
@@ -3388,50 +3539,15 @@ static bool positive_ctrl_err(struct ctrl_pos *sp, struct ctrl_pos *pv)
* the aging
******************************************************************************/
-/* promote pages accessed through page tables */
-static int folio_update_gen(struct folio *folio, int new_gen, int *type,
- const vma_flags_t *vma_flags)
-{
- unsigned long new_flags, old_flags = READ_ONCE(*folio_flags(folio, 0));
- int old_gen;
-
- /*
- * See the comment on LRU_REFS_FLAGS, and activate file-backed
- * executable folios after first usage to avoid typical IO
- * thrashing from reclaiming.
- */
- if (!folio_test_referenced(folio) && !folio_test_workingset(folio) &&
- !is_exec_file_folio(folio, vma_flags)) {
- folio_set_lru_refs(folio, 1);
- return -1;
- }
-
- do {
- old_gen = lru_get_gen_flags(old_flags);
- new_flags = old_flags;
-
- /* lru_gen_del_folio() has isolated this page? */
- if (old_gen < 0)
- break;
-
- lru_set_gen_flags(&new_flags, new_gen);
- lru_set_refs_flags(&new_flags, 0);
- new_flags |= BIT(PG_workingset);
- } while (!try_cmpxchg(folio_flags(folio, 0), &old_flags, new_flags));
-
- *type = folio_flags_is_file_lru(&old_flags);
- return old_gen;
-}
-
static int __folio_inc_gen(struct folio *folio, int old_gen, bool *increased)
{
unsigned long new_flags, old_flags = READ_ONCE(*folio_flags(folio, 0));
- int new_gen;
+ int refs, new_gen;
do {
new_gen = lru_get_gen_flags(old_flags);
- /* folio_update_gen() has promoted this page? */
+ /* folio_inc_lru_refs() has promoted this page? */
if (new_gen >= 0 && new_gen != old_gen) {
if (increased)
*increased = false;
@@ -3440,9 +3556,9 @@ static int __folio_inc_gen(struct folio *folio, int old_gen, bool *increased)
new_flags = old_flags;
new_gen = (old_gen + 1) % MAX_NR_GENS;
-
+ refs = lru_get_refs_flags(old_flags);
lru_set_gen_flags(&new_flags, new_gen);
- lru_set_refs_flags(&new_flags, 0);
+ lru_set_refs_flags(&new_flags, min(refs, LRU_REFS_WORKINGSET));
} while (!try_cmpxchg(folio_flags(folio, 0), &old_flags, new_flags));
if (increased)
@@ -3450,17 +3566,21 @@ static int __folio_inc_gen(struct folio *folio, int old_gen, bool *increased)
return new_gen;
}
-/* protect pages accessed multiple times through file descriptors */
+/*
+ * Force bump a folio's generation. Used for PID protection or to skip a
+ * folio from a zone ineligible for the current reclaim.
+ */
static int folio_inc_gen(struct lruvec *lruvec, struct folio *folio)
{
+ bool gen_increased;
int type = folio_is_file_lru(folio);
struct lru_gen_folio *lrugen = &lruvec->lrugen;
int new_gen, old_gen = lru_gen_from_seq(lrugen->min_seq[type]);
- bool gen_increased;
new_gen = __folio_inc_gen(folio, old_gen, &gen_increased);
if (gen_increased)
lru_gen_update_size(lruvec, folio, old_gen, new_gen);
+
return new_gen;
}
@@ -3656,25 +3776,25 @@ static void walk_update_folio(struct lru_gen_mm_walk *walk, struct vm_area_struc
struct lruvec *lruvec, struct folio *folio, bool dirty)
{
int new_gen, old_gen, type;
+ unsigned int flags = LRU_REF_MAPPED;
if (!folio)
return;
- new_gen = lru_gen_from_seq(READ_ONCE(lruvec->lrugen.max_seq));
-
if (dirty && !folio_test_dirty(folio) &&
!(folio_test_anon(folio) && folio_test_swapbacked(folio) &&
!folio_test_swapcache(folio)))
folio_mark_dirty(folio);
if (walk) {
- old_gen = folio_update_gen(folio, new_gen, &type, &vma->flags);
+ old_gen = folio_inc_lru_refs_walk(folio, lruvec, &vma->flags,
+ &new_gen, &type);
if (old_gen >= 0 && old_gen != new_gen)
update_batch_size(walk, folio, old_gen, new_gen, type);
- } else if (lru_gen_set_refs(folio, &vma->flags)) {
- old_gen = folio_lru_gen(folio);
- if (old_gen >= 0 && old_gen != new_gen)
- folio_activate(folio);
+ } else {
+ if (is_exec_file_folio(folio, &vma->flags))
+ flags |= LRU_REF_EXEC;
+ folio_inc_lru_refs(folio, flags);
}
}
@@ -4081,7 +4201,7 @@ static bool inc_min_seq(struct lruvec *lruvec, int type, int swappiness)
struct folio *folio = list_entry(pos, struct folio, lru);
long nr_pages = folio_nr_pages(folio);
int refs = folio_lru_refs(folio);
- bool workingset = folio_test_workingset(folio);
+ int tier = lru_tier_from_refs(refs);
bool gen_increased;
VM_WARN_ON_ONCE_FOLIO(folio_test_unevictable(folio), folio);
@@ -4101,13 +4221,8 @@ static bool inc_min_seq(struct lruvec *lruvec, int type, int swappiness)
delta += nr_pages;
batch_end = &folio->lru;
- /* don't count the workingset being lazily promoted */
- if (refs + workingset != BIT(LRU_REFS_WIDTH) + 1) {
- int tier = lru_tier_from_refs(refs, workingset);
-
- WRITE_ONCE(lrugen->protected[hist][type][tier],
- lrugen->protected[hist][type][tier] + nr_pages);
- }
+ WRITE_ONCE(lrugen->protected[hist][type][tier],
+ lrugen->protected[hist][type][tier] + nr_pages);
} else {
flush_lru_batch(head, &batch_end, target_list);
list_move(&folio->lru, &lrugen->folios[new_gen][type][zone]);
@@ -4822,8 +4937,7 @@ static bool sort_folio(struct lruvec *lruvec, struct folio *folio, struct scan_c
int zone = folio_zonenum(folio);
int delta = folio_nr_pages(folio);
int refs = folio_lru_refs(folio);
- bool workingset = folio_test_workingset(folio);
- int tier = lru_tier_from_refs(refs, workingset);
+ int tier = lru_tier_from_refs(refs);
struct lru_gen_folio *lrugen = &lruvec->lrugen;
VM_WARN_ON_ONCE_FOLIO(gen >= MAX_NR_GENS, folio);
@@ -4839,17 +4953,15 @@ static bool sort_folio(struct lruvec *lruvec, struct folio *folio, struct scan_c
}
/* protected */
- if (tier > tier_idx || refs + workingset == BIT(LRU_REFS_WIDTH) + 1) {
+ if (tier > tier_idx) {
+ int hist = lru_hist_from_seq(lrugen->min_seq[type]);
+
gen = folio_inc_gen(lruvec, folio);
list_move(&folio->lru, &lrugen->folios[gen][type][zone]);
- /* don't count the workingset being lazily promoted */
- if (refs + workingset != BIT(LRU_REFS_WIDTH) + 1) {
- int hist = lru_hist_from_seq(lrugen->min_seq[type]);
+ WRITE_ONCE(lrugen->protected[hist][type][tier],
+ lrugen->protected[hist][type][tier] + delta);
- WRITE_ONCE(lrugen->protected[hist][type][tier],
- lrugen->protected[hist][type][tier] + delta);
- }
return true;
}
@@ -4877,10 +4989,6 @@ static bool isolate_folio(struct lruvec *lruvec, struct folio *folio, struct sca
return false;
}
- /* see the comment on LRU_REFS_FLAGS */
- if (!folio_test_referenced(folio))
- folio_set_lru_refs(folio, 0);
-
success = lru_gen_del_folio(lruvec, folio, true);
VM_WARN_ON_ONCE_FOLIO(!success, folio);
@@ -4968,13 +5076,14 @@ static int get_tier_idx(struct lruvec *lruvec, int type)
struct ctrl_pos sp, pv;
/*
- * To leave a margin for fluctuations, use a larger gain factor (2:3).
- * This value is chosen because any other tier would have at least twice
- * as many refaults as the first tier.
+ * To leave a margin for fluctuations, use a larger gain factor (1:2).
+ * Stop at the first tier whose refault rate is clearly worse than
+ * that of the cumulative mass of the tiers below it; the PID
+ * protects the tiers above it.
*/
- read_ctrl_pos(lruvec, type, LRU_TIER_MIN, LRU_TIER_MIN, 2, &sp);
for (tier = LRU_TIER_MIN + 1; tier <= LRU_TIER_MAX; tier++) {
- read_ctrl_pos(lruvec, type, tier, tier, 3, &pv);
+ read_ctrl_pos(lruvec, type, LRU_TIER_MIN, tier - 1, 1, &sp);
+ read_ctrl_pos(lruvec, type, tier, tier, 2, &pv);
if (!positive_ctrl_err(&sp, &pv))
break;
}
@@ -5113,12 +5222,13 @@ static int evict_folios(unsigned long nr_to_scan, struct lruvec *lruvec,
}
/*
- * See the comments on LRU_REFS_FLAGS.
- *
* The rejected folios are never added to the oldest generation,
* so this effectively promotes them by at least one generation.
+ * PG_active is the only placement hint here, so a folio kept on
+ * its first page table access lands in the second newest one.
+ * See "Referenced count feedback" above.
*/
- folio_set_lru_refs(folio, 0);
+ folio_set_lru_refs(folio, min(folio_lru_refs(folio), LRU_REFS_WORKINGSET));
if (lru_gen_folio_seq(lruvec, folio, false) == min_seq[type])
folio_set_active(folio);
}
diff --git a/mm/workingset.c b/mm/workingset.c
index 1504f91cdca5..16cdd88e275e 100644
--- a/mm/workingset.c
+++ b/mm/workingset.c
@@ -188,6 +188,12 @@
#define EVICTION_SHIFT_ANON (EVICTION_SHIFT + SWAP_COUNT_SHIFT)
#define EVICTION_MASK (~0UL >> EVICTION_SHIFT)
#define EVICTION_MASK_ANON (~0UL >> EVICTION_SHIFT_ANON)
+/*
+ * LRU refs uses LRU_REFS_WIDTH + 2 bits, the 2 bits being PG_workingset
+ * and PG_referenced. The lowest bit is recorded in the workingset field
+ * of the shadow entry (to reuse pack_shadow()).
+ */
+#define LRU_REFS_BITS ((LRU_REFS_WIDTH + 2) - 1)
/*
* Eviction timestamps need to be able to cover the full range of
@@ -242,13 +248,12 @@ static void *lru_gen_eviction(struct folio *folio)
int type = folio_is_file_lru(folio);
int delta = folio_nr_pages(folio);
int refs = folio_lru_refs(folio);
- bool workingset = folio_test_workingset(folio);
- int tier = lru_tier_from_refs(refs, workingset);
+ int tier = lru_tier_from_refs(refs);
struct mem_cgroup *memcg;
struct pglist_data *pgdat = folio_pgdat(folio);
unsigned short memcg_id;
- BUILD_BUG_ON(LRU_GEN_WIDTH + LRU_REFS_WIDTH >
+ BUILD_BUG_ON(LRU_GEN_WIDTH + LRU_REFS_BITS >
BITS_PER_LONG - max(EVICTION_SHIFT, EVICTION_SHIFT_ANON));
rcu_read_lock();
@@ -256,14 +261,14 @@ static void *lru_gen_eviction(struct folio *folio)
lruvec = mem_cgroup_lruvec(memcg, pgdat);
lrugen = &lruvec->lrugen;
min_seq = READ_ONCE(lrugen->min_seq[type]);
- token = (min_seq << LRU_REFS_WIDTH) | max(refs - 1, 0);
+ token = (min_seq << LRU_REFS_BITS) | refs >> 1;
hist = lru_hist_from_seq(min_seq);
atomic_long_add(delta, &lrugen->evicted[hist][type][tier]);
memcg_id = mem_cgroup_private_id(memcg);
rcu_read_unlock();
- return pack_shadow(memcg_id, pgdat, token, workingset, type);
+ return pack_shadow(memcg_id, pgdat, token, refs & 1, type);
}
/*
@@ -284,9 +289,9 @@ static bool lru_gen_test_recent(void *shadow, struct lruvec **lruvec,
*lruvec = mem_cgroup_lruvec(memcg, pgdat);
max_seq = READ_ONCE((*lruvec)->lrugen.max_seq);
- max_seq &= (file ? EVICTION_MASK : EVICTION_MASK_ANON) >> LRU_REFS_WIDTH;
+ max_seq &= (file ? EVICTION_MASK : EVICTION_MASK_ANON) >> LRU_REFS_BITS;
- return abs_diff(max_seq, *token >> LRU_REFS_WIDTH) < MAX_NR_GENS;
+ return abs_diff(max_seq, *token >> LRU_REFS_BITS) < MAX_NR_GENS;
}
static void lru_gen_refault(struct folio *folio, void *shadow)
@@ -314,22 +319,26 @@ static void lru_gen_refault(struct folio *folio, void *shadow)
lrugen = &lruvec->lrugen;
hist = lru_hist_from_seq(READ_ONCE(lrugen->min_seq[type]));
- refs = (token & (BIT(LRU_REFS_WIDTH) - 1)) + 1;
- tier = lru_tier_from_refs(refs, workingset);
+ refs = ((token & (BIT(LRU_REFS_BITS) - 1)) << 1) + workingset;
+ tier = lru_tier_from_refs(refs);
atomic_long_add(delta, &lrugen->refaulted[hist][type][tier]);
- if (workingset) {
- /*
- * see folio_add_lru(), where folio_set_active() is
- * called for workingset folios
- */
- if (lru_gen_in_fault())
+ /*
+ * Activate a fault-driven refault folio, which would have been
+ * promoted had it stayed in memory.
+ */
+ if (refs >= LRU_REFS_REFERENCED) {
+ if (lru_gen_in_fault()) {
+ folio_set_active(folio);
mod_lruvec_state(lruvec, WORKINGSET_ACTIVATE_BASE + type, delta);
- folio_set_workingset(folio);
+ }
+ /* Refault is also promotion, cap the refs like folio_inc_lru_refs */
+ folio_set_lru_refs(folio, min(refs, LRU_REFS_PROTECTED));
+ }
+
+ if (refs >= LRU_REFS_WORKINGSET)
mod_lruvec_state(lruvec, WORKINGSET_RESTORE_BASE + type, delta);
- } else
- set_mask_bits(&folio->flags.f, LRU_REFS_MASK, (refs - 1UL) << LRU_REFS_PGOFF);
unlock:
rcu_read_unlock();
}
--
2.55.0
^ permalink raw reply [flat|nested] 21+ messages in thread
* [PATCH RFC v3 05/17] mm/mglru: make folio lru referenced times count a generic API
2026-10-03 12:55 [PATCH RFC v3 00/17] mm/mglru: frequency guided promotion (MGLRU-FG) and flag cleanup Kairui Song via B4 Relay
` (3 preceding siblings ...)
2026-10-03 12:55 ` [PATCH RFC v3 04/17] mm/mglru: frequency guided workingset promotion (MGLRU-FG) Kairui Song via B4 Relay
@ 2026-10-03 12:55 ` Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 06/17] mm/mglru: move add/del LRU size accounting out of lru_gen_update_size() Kairui Song via B4 Relay
` (11 subsequent siblings)
16 siblings, 0 replies; 21+ messages in thread
From: Kairui Song via B4 Relay @ 2026-10-03 12:55 UTC (permalink / raw)
To: linux-mm
Cc: Andrew Morton, Johannes Weiner, Muchun Song, Qi Zheng,
Ying Huang, Chris Li, Baoquan He, Nico Pache, Usama Arif,
Michal Hocko, Roman Gushchin, Shakeel Butt, David Hildenbrand,
Lorenzo Stoakes, Barry Song, Axel Rasmussen, Yuanchu Xie, Wei Xu,
Vlastimil Babka, Suren Baghdasaryan, Kemeng Shi, Nhat Pham,
Youngjun Park, Zi Yan, Gregory Price, Matthew Wilcox (Oracle),
Baolin Wang, Ryan Roberts, Dev Jain, Lance Yang, Hugh Dickins,
SeongJae Park, David Rientjes, Yu Zhao, Vernon Yang,
Zicheng Wang, Chen Ridong, Tal Zussman, Kairui Song, Xiang Liu,
Ehab Ababneh, zhaozhengzhuo, Lian Wang, Kunwu Chan, linux-kernel,
cgroups, Kairui Song
From: Kairui Song <kasong@tencent.com>
Pure code shuffle, no behavior change. Make the folio LRU refs helpers
available regardless of CONFIG_LRU_GEN; without MGLRU the refs bits are
never set, so they are inert. Move the lru_gen_* helpers into the
CONFIG_LRU_GEN section unchanged. This prepares for unifying the API
for checking folio referenced and workingset status.
folio_migrate_lru_refs() now migrates the complete refs count,
including PG_workingset (bit 1 of the encoding), bitwise identical to
the direct copy dropped from folio_migrate_flags(). It also gains
off-LRU VM_WARN_ON_ONCE checks, compiled out unless CONFIG_DEBUG_VM.
lru_set_refs_flags() skips the nonexistent mask bits when
LRU_REFS_WIDTH collapses to 0, keeping the generic helpers inert for
!CONFIG_LRU_GEN builds.
Signed-off-by: Kairui Song <kasong@tencent.com>
---
include/linux/mm_inline.h | 145 +++++++++++++++++++++++-----------------------
mm/migrate.c | 2 -
2 files changed, 72 insertions(+), 75 deletions(-)
diff --git a/include/linux/mm_inline.h b/include/linux/mm_inline.h
index f778e4056fbe..a3a302e22a47 100644
--- a/include/linux/mm_inline.h
+++ b/include/linux/mm_inline.h
@@ -105,54 +105,6 @@ static __always_inline enum lru_list folio_lru_list(const struct folio *folio)
return lru;
}
-#ifdef CONFIG_LRU_GEN
-
-static inline bool lru_gen_switching(void)
-{
- DECLARE_STATIC_KEY_FALSE(lru_switch);
-
- return static_branch_unlikely(&lru_switch);
-}
-#ifdef CONFIG_LRU_GEN_ENABLED
-static inline bool lru_gen_enabled(void)
-{
- DECLARE_STATIC_KEY_TRUE(lru_gen_caps[NR_LRU_GEN_CAPS]);
-
- return static_branch_likely(&lru_gen_caps[LRU_GEN_CORE]);
-}
-#else
-static inline bool lru_gen_enabled(void)
-{
- DECLARE_STATIC_KEY_FALSE(lru_gen_caps[NR_LRU_GEN_CAPS]);
-
- return static_branch_unlikely(&lru_gen_caps[LRU_GEN_CORE]);
-}
-#endif
-
-static inline bool lru_gen_in_fault(void)
-{
- return current->in_lru_fault;
-}
-
-static inline int lru_gen_from_seq(unsigned long seq)
-{
- return seq % MAX_NR_GENS;
-}
-
-static inline int lru_hist_from_seq(unsigned long seq)
-{
- return seq % NR_HIST_GENS;
-}
-
-static inline int lru_tier_from_refs(unsigned int refs)
-{
- BUILD_BUG_ON(fls(LRU_REFS_MAX - 1) > MAX_NR_TIERS - 1);
- VM_WARN_ON_ONCE(refs > LRU_REFS_MAX);
- if (refs < LRU_REFS_WORKINGSET)
- return 0;
- return fls(refs - 1);
-}
-
/**
* lru_set_gen_flags - Set the LRU generation number to specified folio flags.
* @flags: pointer to the folio flags
@@ -205,7 +157,8 @@ static inline void lru_set_refs_flags(unsigned long *flags, unsigned int refs)
*flags |= BIT(PG_referenced);
if (refs & BIT(1))
*flags |= BIT(PG_workingset);
- *flags |= ((unsigned long)refs >> 2) << LRU_REFS_PGOFF;
+ if (LRU_REFS_WIDTH)
+ *flags |= ((unsigned long)refs >> 2) << LRU_REFS_PGOFF;
}
/**
@@ -240,7 +193,77 @@ static inline void folio_set_lru_refs(struct folio *folio, unsigned int refs)
} while (!try_cmpxchg(folio_flags(folio, 0), &old_flags, new_flags));
}
+#ifdef CONFIG_LRU_GEN
void folio_inc_lru_refs(struct folio *folio, unsigned int flags);
+#else
+static inline void folio_inc_lru_refs(struct folio *folio, unsigned int flags)
+{
+ /* Should not be called with !CONFIG_LRU_GEN */
+ WARN_ON_ONCE(1);
+}
+#endif
+
+/**
+ * folio_migrate_lru_refs - copy the reference state to a new folio
+ * @new: the destination folio
+ * @old: the source folio
+ *
+ * Transfer the reference state to @new during migration: the MGLRU
+ * refs count, or PG_referenced and PG_workingset (bits 0-1 of the
+ * encoding) for the active/inactive LRU.
+ */
+static inline void folio_migrate_lru_refs(struct folio *new, const struct folio *old)
+{
+ folio_set_lru_refs(new, folio_lru_refs(old));
+}
+
+#ifdef CONFIG_LRU_GEN
+
+static inline bool lru_gen_switching(void)
+{
+ DECLARE_STATIC_KEY_FALSE(lru_switch);
+
+ return static_branch_unlikely(&lru_switch);
+}
+#ifdef CONFIG_LRU_GEN_ENABLED
+static inline bool lru_gen_enabled(void)
+{
+ DECLARE_STATIC_KEY_TRUE(lru_gen_caps[NR_LRU_GEN_CAPS]);
+
+ return static_branch_likely(&lru_gen_caps[LRU_GEN_CORE]);
+}
+#else
+static inline bool lru_gen_enabled(void)
+{
+ DECLARE_STATIC_KEY_FALSE(lru_gen_caps[NR_LRU_GEN_CAPS]);
+
+ return static_branch_unlikely(&lru_gen_caps[LRU_GEN_CORE]);
+}
+#endif
+
+static inline bool lru_gen_in_fault(void)
+{
+ return current->in_lru_fault;
+}
+
+static inline int lru_gen_from_seq(unsigned long seq)
+{
+ return seq % MAX_NR_GENS;
+}
+
+static inline int lru_hist_from_seq(unsigned long seq)
+{
+ return seq % NR_HIST_GENS;
+}
+
+static inline int lru_tier_from_refs(unsigned int refs)
+{
+ BUILD_BUG_ON(fls(LRU_REFS_MAX - 1) > MAX_NR_TIERS - 1);
+ VM_WARN_ON_ONCE(refs > LRU_REFS_MAX);
+ if (refs < LRU_REFS_WORKINGSET)
+ return 0;
+ return fls(refs - 1);
+}
static inline int folio_lru_gen(const struct folio *folio)
{
@@ -388,20 +411,6 @@ static inline bool lru_gen_del_folio(struct lruvec *lruvec, struct folio *folio,
return true;
}
-
-/**
- * folio_migrate_lru_refs - copy the reference state to a new folio
- * @new: the destination folio
- * @old: the source folio
- *
- * Transfer the reference state to @new during migration: the MGLRU
- * refs count, including PG_referenced, or just PG_referenced for the
- * active/inactive LRU.
- */
-static inline void folio_migrate_lru_refs(struct folio *new, const struct folio *old)
-{
- folio_set_lru_refs(new, folio_lru_refs(old));
-}
#else /* !CONFIG_LRU_GEN */
static inline bool lru_gen_enabled(void)
@@ -428,16 +437,6 @@ static inline bool lru_gen_del_folio(struct lruvec *lruvec, struct folio *folio,
{
return false;
}
-
-static inline void folio_inc_lru_refs(struct folio *folio, unsigned int flags)
-{
-}
-
-static inline void folio_migrate_lru_refs(struct folio *new, const struct folio *old)
-{
- if (folio_test_referenced(old))
- folio_set_referenced(new);
-}
#endif /* CONFIG_LRU_GEN */
static __always_inline
diff --git a/mm/migrate.c b/mm/migrate.c
index 7bdcdb57652f..5eefd8547f70 100644
--- a/mm/migrate.c
+++ b/mm/migrate.c
@@ -786,8 +786,6 @@ void folio_migrate_flags(struct folio *newfolio, struct folio *folio)
folio_set_active(newfolio);
} else if (folio_test_clear_unevictable(folio))
folio_set_unevictable(newfolio);
- if (folio_test_workingset(folio))
- folio_set_workingset(newfolio);
if (folio_test_checked(folio))
folio_set_checked(newfolio);
/*
--
2.55.0
^ permalink raw reply [flat|nested] 21+ messages in thread
* [PATCH RFC v3 06/17] mm/mglru: move add/del LRU size accounting out of lru_gen_update_size()
2026-10-03 12:55 [PATCH RFC v3 00/17] mm/mglru: frequency guided promotion (MGLRU-FG) and flag cleanup Kairui Song via B4 Relay
` (4 preceding siblings ...)
2026-10-03 12:55 ` [PATCH RFC v3 05/17] mm/mglru: make folio lru referenced times count a generic API Kairui Song via B4 Relay
@ 2026-10-03 12:55 ` Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 07/17] mm/mglru, gup: mark folios referenced via a fast helper Kairui Song via B4 Relay
` (10 subsequent siblings)
16 siblings, 0 replies; 21+ messages in thread
From: Kairui Song via B4 Relay @ 2026-10-03 12:55 UTC (permalink / raw)
To: linux-mm
Cc: Andrew Morton, Johannes Weiner, Muchun Song, Qi Zheng,
Ying Huang, Chris Li, Baoquan He, Nico Pache, Usama Arif,
Michal Hocko, Roman Gushchin, Shakeel Butt, David Hildenbrand,
Lorenzo Stoakes, Barry Song, Axel Rasmussen, Yuanchu Xie, Wei Xu,
Vlastimil Babka, Suren Baghdasaryan, Kemeng Shi, Nhat Pham,
Youngjun Park, Zi Yan, Gregory Price, Matthew Wilcox (Oracle),
Baolin Wang, Ryan Roberts, Dev Jain, Lance Yang, Hugh Dickins,
SeongJae Park, David Rientjes, Yu Zhao, Vernon Yang,
Zicheng Wang, Chen Ridong, Tal Zussman, Kairui Song, Xiang Liu,
Ehab Ababneh, zhaozhengzhuo, Lian Wang, Kunwu Chan, linux-kernel,
cgroups, Kairui Song
From: Kairui Song <kasong@tencent.com>
Pure code shuffle, no behavior change. lru_gen_update_size() now only
updates the per-generation counters plus the promotion move; the
addition and deletion active/inactive accounting move to
lru_gen_add_folio() and lru_gen_del_folio(), which know the generation
and can apply the active-window test directly.
The unevictable case is untouched: unevictable folios never enter
lru_gen_add_folio() and keep being accounted by lruvec_add_folio() as
before.
Signed-off-by: Kairui Song <kasong@tencent.com>
---
include/linux/mm_inline.h | 28 ++++++++++++++--------------
1 file changed, 14 insertions(+), 14 deletions(-)
diff --git a/include/linux/mm_inline.h b/include/linux/mm_inline.h
index a3a302e22a47..48a945c2cbaf 100644
--- a/include/linux/mm_inline.h
+++ b/include/linux/mm_inline.h
@@ -298,21 +298,9 @@ static inline void lru_gen_update_size(struct lruvec *lruvec, struct folio *foli
if (new_gen >= 0)
atomic_long_add(delta, &lrugen->nr_pages[new_gen][type][zone]);
- /* addition */
- if (old_gen < 0) {
- if (lru_gen_is_active(lruvec, new_gen))
- lru += LRU_ACTIVE;
- __update_lru_size(lruvec, lru, zone, delta);
+ /* return now if not a promotion */
+ if (old_gen < 0 || new_gen < 0)
return;
- }
-
- /* deletion */
- if (new_gen < 0) {
- if (lru_gen_is_active(lruvec, old_gen))
- lru += LRU_ACTIVE;
- __update_lru_size(lruvec, lru, zone, -delta);
- return;
- }
/* promotion */
if (!lru_gen_is_active(lruvec, old_gen) && lru_gen_is_active(lruvec, new_gen)) {
@@ -366,6 +354,8 @@ static inline bool lru_gen_add_folio(struct lruvec *lruvec, struct folio *folio,
int gen = folio_lru_gen(folio);
int type = folio_is_file_lru(folio);
int zone = folio_zonenum(folio);
+ int delta = folio_nr_pages(folio);
+ enum lru_list lru = type * LRU_INACTIVE_FILE;
struct lru_gen_folio *lrugen = &lruvec->lrugen;
BUILD_BUG_ON(BIT(LRU_GEN_WIDTH - 1) != MAX_NR_GENS);
@@ -381,6 +371,10 @@ static inline bool lru_gen_add_folio(struct lruvec *lruvec, struct folio *folio,
set_mask_bits(folio_flags(folio, 0), LRU_GEN_MASK | BIT(PG_active), flags);
lru_gen_update_size(lruvec, folio, -1, gen);
+ if (lru_gen_is_active(lruvec, gen))
+ lru += LRU_ACTIVE;
+ __update_lru_size(lruvec, lru, zone, delta);
+
/* for folio_rotate_reclaimable() */
if (reclaiming)
list_add_tail(&folio->lru, &lrugen->folios[gen][type][zone]);
@@ -394,6 +388,9 @@ static inline bool lru_gen_del_folio(struct lruvec *lruvec, struct folio *folio,
{
unsigned long flags;
int gen = folio_lru_gen(folio);
+ int zone = folio_zonenum(folio);
+ int delta = folio_nr_pages(folio);
+ enum lru_list lru = folio_is_file_lru(folio) * LRU_INACTIVE_FILE;
if (gen < 0)
return false;
@@ -407,6 +404,9 @@ static inline bool lru_gen_del_folio(struct lruvec *lruvec, struct folio *folio,
gen = ((flags & LRU_GEN_MASK) >> LRU_GEN_PGOFF) - 1;
lru_gen_update_size(lruvec, folio, gen, -1);
+ if (lru_gen_is_active(lruvec, gen))
+ lru += LRU_ACTIVE;
+ __update_lru_size(lruvec, lru, zone, -delta);
list_del(&folio->lru);
return true;
--
2.55.0
^ permalink raw reply [flat|nested] 21+ messages in thread
* [PATCH RFC v3 07/17] mm/mglru, gup: mark folios referenced via a fast helper
2026-10-03 12:55 [PATCH RFC v3 00/17] mm/mglru: frequency guided promotion (MGLRU-FG) and flag cleanup Kairui Song via B4 Relay
` (5 preceding siblings ...)
2026-10-03 12:55 ` [PATCH RFC v3 06/17] mm/mglru: move add/del LRU size accounting out of lru_gen_update_size() Kairui Song via B4 Relay
@ 2026-10-03 12:55 ` Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 08/17] mm/smap: convert to LRU refs based operations Kairui Song via B4 Relay
` (9 subsequent siblings)
16 siblings, 0 replies; 21+ messages in thread
From: Kairui Song via B4 Relay @ 2026-10-03 12:55 UTC (permalink / raw)
To: linux-mm
Cc: Andrew Morton, Johannes Weiner, Muchun Song, Qi Zheng,
Ying Huang, Chris Li, Baoquan He, Nico Pache, Usama Arif,
Michal Hocko, Roman Gushchin, Shakeel Butt, David Hildenbrand,
Lorenzo Stoakes, Barry Song, Axel Rasmussen, Yuanchu Xie, Wei Xu,
Vlastimil Babka, Suren Baghdasaryan, Kemeng Shi, Nhat Pham,
Youngjun Park, Zi Yan, Gregory Price, Matthew Wilcox (Oracle),
Baolin Wang, Ryan Roberts, Dev Jain, Lance Yang, Hugh Dickins,
SeongJae Park, David Rientjes, Yu Zhao, Vernon Yang,
Zicheng Wang, Chen Ridong, Tal Zussman, Kairui Song, Xiang Liu,
Ehab Ababneh, zhaozhengzhuo, Lian Wang, Kunwu Chan, linux-kernel,
cgroups, Kairui Song
From: Kairui Song <kasong@tencent.com>
Introduce folio_inc_lru_refs_fast(): raise refs from 0 to
LRU_REFS_REFERENCED via cmpxchg, or just set PG_referenced for the
classical LRU as before. No counter update, promotion, and guarteens
no other complex operations in following commits.
Convert the gup fast-path sites. No behavior change for classical LRU.
Signed-off-by: Kairui Song <kasong@tencent.com>
---
include/linux/mm_inline.h | 25 +++++++++++++++++++++++++
mm/gup.c | 6 +++---
2 files changed, 28 insertions(+), 3 deletions(-)
diff --git a/include/linux/mm_inline.h b/include/linux/mm_inline.h
index 48a945c2cbaf..cb54913aa184 100644
--- a/include/linux/mm_inline.h
+++ b/include/linux/mm_inline.h
@@ -439,6 +439,31 @@ static inline bool lru_gen_del_folio(struct lruvec *lruvec, struct folio *folio,
}
#endif /* CONFIG_LRU_GEN */
+/**
+ * folio_inc_lru_refs_fast - Bump folio refs count without promotion.
+ * @folio: the folio
+ *
+ * Raise refs from 0 to LRU_REFS_REFERENCED and leave hotter folios
+ * untouched. For the classical LRU, a plain PG_referenced set.
+ */
+static __always_inline void folio_inc_lru_refs_fast(struct folio *folio)
+{
+ unsigned long new_flags, old_flags;
+
+ if (!lru_gen_enabled()) {
+ folio_set_referenced(folio);
+ return;
+ }
+
+ old_flags = READ_ONCE(*folio_flags(folio, 0));
+ do {
+ if (lru_get_refs_flags(old_flags))
+ break;
+ new_flags = old_flags;
+ lru_set_refs_flags(&new_flags, LRU_REFS_REFERENCED);
+ } while (!try_cmpxchg(folio_flags(folio, 0), &old_flags, new_flags));
+}
+
static __always_inline
void lruvec_add_folio(struct lruvec *lruvec, struct folio *folio)
{
diff --git a/mm/gup.c b/mm/gup.c
index 8e9ef5ee7498..87a61c67ce2d 100644
--- a/mm/gup.c
+++ b/mm/gup.c
@@ -2895,7 +2895,7 @@ static unsigned long gup_fast_pte_range(pmd_t pmd, pmd_t *pmdp,
gup_put_folio(folio, 1, flags);
goto pte_unmap;
}
- folio_set_referenced(folio);
+ folio_inc_lru_refs_fast(folio);
pages[nr_pages] = page;
nr_pages++;
} while (ptep++, addr += PAGE_SIZE, addr != end);
@@ -2964,7 +2964,7 @@ static unsigned long gup_fast_pmd_leaf(pmd_t orig, pmd_t *pmdp,
for (i = 0; i < nr_pages; i++)
*(pages++) = page++;
- folio_set_referenced(folio);
+ folio_inc_lru_refs_fast(folio);
return nr_pages;
}
@@ -3006,7 +3006,7 @@ static unsigned long gup_fast_pud_leaf(pud_t orig, pud_t *pudp,
for (i = 0; i < nr_pages; i++)
*(pages++) = page++;
- folio_set_referenced(folio);
+ folio_inc_lru_refs_fast(folio);
return nr_pages;
}
--
2.55.0
^ permalink raw reply [flat|nested] 21+ messages in thread
* [PATCH RFC v3 08/17] mm/smap: convert to LRU refs based operations
2026-10-03 12:55 [PATCH RFC v3 00/17] mm/mglru: frequency guided promotion (MGLRU-FG) and flag cleanup Kairui Song via B4 Relay
` (6 preceding siblings ...)
2026-10-03 12:55 ` [PATCH RFC v3 07/17] mm/mglru, gup: mark folios referenced via a fast helper Kairui Song via B4 Relay
@ 2026-10-03 12:55 ` Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 09/17] mm/madvise: adapt for LRU refs based operations in MGLRU Kairui Song via B4 Relay
` (8 subsequent siblings)
16 siblings, 0 replies; 21+ messages in thread
From: Kairui Song via B4 Relay @ 2026-10-03 12:55 UTC (permalink / raw)
To: linux-mm
Cc: Andrew Morton, Johannes Weiner, Muchun Song, Qi Zheng,
Ying Huang, Chris Li, Baoquan He, Nico Pache, Usama Arif,
Michal Hocko, Roman Gushchin, Shakeel Butt, David Hildenbrand,
Lorenzo Stoakes, Barry Song, Axel Rasmussen, Yuanchu Xie, Wei Xu,
Vlastimil Babka, Suren Baghdasaryan, Kemeng Shi, Nhat Pham,
Youngjun Park, Zi Yan, Gregory Price, Matthew Wilcox (Oracle),
Baolin Wang, Ryan Roberts, Dev Jain, Lance Yang, Hugh Dickins,
SeongJae Park, David Rientjes, Yu Zhao, Vernon Yang,
Zicheng Wang, Chen Ridong, Tal Zussman, Kairui Song, Xiang Liu,
Ehab Ababneh, zhaozhengzhuo, Lian Wang, Kunwu Chan, linux-kernel,
cgroups, Kairui Song
From: Kairui Song <kasong@tencent.com>
For MGLRU, switch smap to use the folio refs count API so smap will
report all folio with referenced count >= 1 as "Referenced". Current
smap checking PG_referenced is causing folios to flick between
referenced and not-reference status, because for both MGLRU and
active/inactive LRU, PG_referenced may got cleared on second access.
(Increase of LRU referenced times count for MGLRU, and movig to active
list active/inactive all clears that bit).
After this, we will have a more reliable and useful reading for MGLRU
after the FG change. The behavior is basically identical to what we had
before the series. And there is no behavior change for classical LRU.
Signed-off-by: Kairui Song <kasong@tencent.com>
---
fs/proc/task_mmu.c | 22 +++++++++++++++++++---
1 file changed, 19 insertions(+), 3 deletions(-)
diff --git a/fs/proc/task_mmu.c b/fs/proc/task_mmu.c
index 44147ba0b899..79daaf757412 100644
--- a/fs/proc/task_mmu.c
+++ b/fs/proc/task_mmu.c
@@ -857,6 +857,22 @@ static void smaps_page_accumulate(struct mem_size_stats *mss,
}
}
+static bool smap_check_folio_referenced(struct folio *folio)
+{
+ if (lru_gen_enabled())
+ return folio_lru_refs(folio);
+ else
+ return folio_test_referenced(folio);
+}
+
+static void smap_clear_folio_referenced(struct folio *folio)
+{
+ if (lru_gen_enabled())
+ folio_set_lru_refs(folio, 0);
+ else
+ folio_clear_referenced(folio);
+}
+
static void smaps_account(struct mem_size_stats *mss, struct page *page,
bool compound, bool young, bool dirty, bool locked,
bool present)
@@ -883,7 +899,7 @@ static void smaps_account(struct mem_size_stats *mss, struct page *page,
mss->resident += size;
/* Accumulate the size in pages that have been accessed. */
- if (young || folio_test_young(folio) || folio_test_referenced(folio))
+ if (young || folio_test_young(folio) || smap_check_folio_referenced(folio))
mss->referenced += size;
/*
@@ -1680,7 +1696,7 @@ static int clear_refs_pte_range(pmd_t *pmd, unsigned long addr,
/* Clear accessed and referenced bits. */
pmdp_test_and_clear_young(vma, addr, pmd);
folio_test_clear_young(folio);
- folio_clear_referenced(folio);
+ smap_clear_folio_referenced(folio);
out:
spin_unlock(ptl);
return 0;
@@ -1709,7 +1725,7 @@ static int clear_refs_pte_range(pmd_t *pmd, unsigned long addr,
/* Clear accessed and referenced bits. */
ptep_test_and_clear_young(vma, addr, pte);
folio_test_clear_young(folio);
- folio_clear_referenced(folio);
+ smap_clear_folio_referenced(folio);
}
pte_unmap_unlock(pte - 1, ptl);
cond_resched();
--
2.55.0
^ permalink raw reply [flat|nested] 21+ messages in thread
* [PATCH RFC v3 09/17] mm/madvise: adapt for LRU refs based operations in MGLRU
2026-10-03 12:55 [PATCH RFC v3 00/17] mm/mglru: frequency guided promotion (MGLRU-FG) and flag cleanup Kairui Song via B4 Relay
` (7 preceding siblings ...)
2026-10-03 12:55 ` [PATCH RFC v3 08/17] mm/smap: convert to LRU refs based operations Kairui Song via B4 Relay
@ 2026-10-03 12:55 ` Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 10/17] mm/damon: convert to LRU refs based operations Kairui Song via B4 Relay
` (7 subsequent siblings)
16 siblings, 0 replies; 21+ messages in thread
From: Kairui Song via B4 Relay @ 2026-10-03 12:55 UTC (permalink / raw)
To: linux-mm
Cc: Andrew Morton, Johannes Weiner, Muchun Song, Qi Zheng,
Ying Huang, Chris Li, Baoquan He, Nico Pache, Usama Arif,
Michal Hocko, Roman Gushchin, Shakeel Butt, David Hildenbrand,
Lorenzo Stoakes, Barry Song, Axel Rasmussen, Yuanchu Xie, Wei Xu,
Vlastimil Babka, Suren Baghdasaryan, Kemeng Shi, Nhat Pham,
Youngjun Park, Zi Yan, Gregory Price, Matthew Wilcox (Oracle),
Baolin Wang, Ryan Roberts, Dev Jain, Lance Yang, Hugh Dickins,
SeongJae Park, David Rientjes, Yu Zhao, Vernon Yang,
Zicheng Wang, Chen Ridong, Tal Zussman, Kairui Song, Xiang Liu,
Ehab Ababneh, zhaozhengzhuo, Lian Wang, Kunwu Chan, linux-kernel,
cgroups, Kairui Song
From: Kairui Song <kasong@tencent.com>
For the active/inactive LRU, madvise clears PG_referenced so that one more
access is not enough to reactivate a folio, and keeps PG_workingset on a
folio demoted out of the active list so its refault is still accounted as
a workingset refault (PSI).
MGLRU keeps that history in the folio's refs count instead, which now
lives in PG_referenced, PG_workingset and LRU_REFS_MASK, so writing to
those bits directly corrupts the count. The two hints want different
things:
- MADV_COLD resets the count in folio_deactivate(), which also moves the
folio to the oldest generation.
- MADV_PAGEOUT caps the count at LRU_REFS_WORKINGSET before handing the
folio to reclaim_pages(). The hint marks the folio as cold, so the
eviction shadow records the capped count and does not inflate the
eviction/refault counters of the hot tiers that drive aging, while
folios at or above the workingset threshold keep their PSI
classification.
PG_young is cleared for both LRUs as before, so the referenced walk does
not count stale history as a recent access.
Signed-off-by: Kairui Song <kasong@tencent.com>
---
mm/madvise.c | 56 ++++++++++++++++++++++++++++++++++++++++++--------------
1 file changed, 42 insertions(+), 14 deletions(-)
diff --git a/mm/madvise.c b/mm/madvise.c
index d83ce6abf8c3..c221234fa199 100644
--- a/mm/madvise.c
+++ b/mm/madvise.c
@@ -361,6 +361,44 @@ static inline int madvise_folio_pte_batch(unsigned long addr, unsigned long end,
FPB_MERGE_YOUNG_DIRTY);
}
+static void madvise_cold_prep_folio(struct folio *folio)
+{
+ /*
+ * VM couldn't reclaim the folio unless we clear PG_young.
+ * As a side effect, it makes confuse idle-page tracking
+ * because they will miss recent referenced history.
+ */
+ folio_test_clear_young(folio);
+
+ /*
+ * For the active/inactive LRU, a folio demoted out of the active
+ * list should have PG_workingset so its refault is still accounted
+ * as a workingset refault, and the referenced bit always needs to
+ * be cleared even for inactive folios. Nothing to do for MGLRU as
+ * folio_deactivate() always calls lru_gen_clear_refs().
+ */
+ if (!lru_gen_enabled()) {
+ folio_clear_referenced(folio);
+ if (folio_test_active(folio))
+ folio_set_workingset(folio);
+ }
+}
+
+static void madvise_pageout_prep_folio(struct folio *folio)
+{
+ int refs = folio_lru_refs(folio);
+
+ /*
+ * For MGLRU, drop the count below LRU_REFS_WORKINGSET and cap
+ * the rest at it: MADV is
+ * suggesting the folio is actually cold, so the eviction shadow
+ * should not record its hotness in the high tier counters, but
+ * the workingset classification is preserved for PSI.
+ */
+ if (lru_gen_enabled())
+ folio_set_lru_refs(folio, refs >= LRU_REFS_WORKINGSET ? LRU_REFS_WORKINGSET : 0);
+}
+
static int madvise_cold_or_pageout_pte_range(pmd_t *pmd,
unsigned long addr, unsigned long end,
struct mm_walk *walk)
@@ -438,12 +476,10 @@ static int madvise_cold_or_pageout_pte_range(pmd_t *pmd,
tlb_remove_pmd_tlb_entry(tlb, pmd, addr);
}
- folio_clear_referenced(folio);
- folio_test_clear_young(folio);
- if (folio_test_active(folio))
- folio_set_workingset(folio);
+ madvise_cold_prep_folio(folio);
if (pageout) {
if (folio_isolate_lru(folio)) {
+ madvise_pageout_prep_folio(folio);
if (folio_test_unevictable(folio))
folio_putback_lru(folio);
else
@@ -547,18 +583,10 @@ static int madvise_cold_or_pageout_pte_range(pmd_t *pmd,
tlb_remove_tlb_entries(tlb, pte, nr, addr);
}
- /*
- * We are deactivating a folio for accelerating reclaiming.
- * VM couldn't reclaim the folio unless we clear PG_young.
- * As a side effect, it makes confuse idle-page tracking
- * because they will miss recent referenced history.
- */
- folio_clear_referenced(folio);
- folio_test_clear_young(folio);
- if (folio_test_active(folio))
- folio_set_workingset(folio);
+ madvise_cold_prep_folio(folio);
if (pageout) {
if (folio_isolate_lru(folio)) {
+ madvise_pageout_prep_folio(folio);
if (folio_test_unevictable(folio))
folio_putback_lru(folio);
else
--
2.55.0
^ permalink raw reply [flat|nested] 21+ messages in thread
* [PATCH RFC v3 10/17] mm/damon: convert to LRU refs based operations
2026-10-03 12:55 [PATCH RFC v3 00/17] mm/mglru: frequency guided promotion (MGLRU-FG) and flag cleanup Kairui Song via B4 Relay
` (8 preceding siblings ...)
2026-10-03 12:55 ` [PATCH RFC v3 09/17] mm/madvise: adapt for LRU refs based operations in MGLRU Kairui Song via B4 Relay
@ 2026-10-03 12:55 ` Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 11/17] mm/huge_memory: mark file folio as accessed more accurately on split Kairui Song via B4 Relay
` (6 subsequent siblings)
16 siblings, 0 replies; 21+ messages in thread
From: Kairui Song via B4 Relay @ 2026-10-03 12:55 UTC (permalink / raw)
To: linux-mm
Cc: Andrew Morton, Johannes Weiner, Muchun Song, Qi Zheng,
Ying Huang, Chris Li, Baoquan He, Nico Pache, Usama Arif,
Michal Hocko, Roman Gushchin, Shakeel Butt, David Hildenbrand,
Lorenzo Stoakes, Barry Song, Axel Rasmussen, Yuanchu Xie, Wei Xu,
Vlastimil Babka, Suren Baghdasaryan, Kemeng Shi, Nhat Pham,
Youngjun Park, Zi Yan, Gregory Price, Matthew Wilcox (Oracle),
Baolin Wang, Ryan Roberts, Dev Jain, Lance Yang, Hugh Dickins,
SeongJae Park, David Rientjes, Yu Zhao, Vernon Yang,
Zicheng Wang, Chen Ridong, Tal Zussman, Kairui Song, Xiang Liu,
Ehab Ababneh, zhaozhengzhuo, Lian Wang, Kunwu Chan, linux-kernel,
cgroups, Kairui Song
From: Kairui Song <kasong@tencent.com>
damon_pa_pageout() cleared PG_referenced directly, which under MGLRU
may corrupts the folio's refs count in later commits. That flag is now
bit 0 of the count (LRU_REFS_FLAGS).
Reset the count instead under MGLRU.
Signed-off-by: Kairui Song <kasong@tencent.com>
---
mm/damon/paddr.c | 9 ++++++++-
1 file changed, 8 insertions(+), 1 deletion(-)
diff --git a/mm/damon/paddr.c b/mm/damon/paddr.c
index 2cfdc356b415..4db141b3fce3 100644
--- a/mm/damon/paddr.c
+++ b/mm/damon/paddr.c
@@ -279,7 +279,14 @@ static unsigned long damon_pa_pageout(struct damon_region *r,
else
*sz_filter_passed += folio_size(folio) / addr_unit;
- folio_clear_referenced(folio);
+ /*
+ * DAMON only gets here for regions it measured as cold,
+ * so the hotness can be, and better be dropped.
+ */
+ if (lru_gen_enabled())
+ folio_set_lru_refs(folio, 0);
+ else
+ folio_clear_referenced(folio);
folio_test_clear_young(folio);
if (!folio_isolate_lru(folio))
goto put_folio;
--
2.55.0
^ permalink raw reply [flat|nested] 21+ messages in thread
* [PATCH RFC v3 11/17] mm/huge_memory: mark file folio as accessed more accurately on split
2026-10-03 12:55 [PATCH RFC v3 00/17] mm/mglru: frequency guided promotion (MGLRU-FG) and flag cleanup Kairui Song via B4 Relay
` (9 preceding siblings ...)
2026-10-03 12:55 ` [PATCH RFC v3 10/17] mm/damon: convert to LRU refs based operations Kairui Song via B4 Relay
@ 2026-10-03 12:55 ` Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 12/17] mm/mglru: folio LRU refs based active/inactive number accounting Kairui Song via B4 Relay
` (5 subsequent siblings)
16 siblings, 0 replies; 21+ messages in thread
From: Kairui Song via B4 Relay @ 2026-10-03 12:55 UTC (permalink / raw)
To: linux-mm
Cc: Andrew Morton, Johannes Weiner, Muchun Song, Qi Zheng,
Ying Huang, Chris Li, Baoquan He, Nico Pache, Usama Arif,
Michal Hocko, Roman Gushchin, Shakeel Butt, David Hildenbrand,
Lorenzo Stoakes, Barry Song, Axel Rasmussen, Yuanchu Xie, Wei Xu,
Vlastimil Babka, Suren Baghdasaryan, Kemeng Shi, Nhat Pham,
Youngjun Park, Zi Yan, Gregory Price, Matthew Wilcox (Oracle),
Baolin Wang, Ryan Roberts, Dev Jain, Lance Yang, Hugh Dickins,
SeongJae Park, David Rientjes, Yu Zhao, Vernon Yang,
Zicheng Wang, Chen Ridong, Tal Zussman, Kairui Song, Xiang Liu,
Ehab Ababneh, zhaozhengzhuo, Lian Wang, Kunwu Chan, linux-kernel,
cgroups, Kairui Song
From: Kairui Song <kasong@tencent.com>
The behavior of updating the folio's access info isn't consistent for
huge mapping splitting or ordinary unmapping. The page table's young
flag has to be translated into folio's access info.
Right now it only check and set folio's referenced flag, which isn't
enough since folio flags update on access have its rules. Ordinary
unmapping (zapping) calls folio_mark_accessed(), and it also checks
if the VMA has recency to avoid false updates.
So first just use the right helper here to be more consistent.
Signed-off-by: Kairui Song <kasong@tencent.com>
---
mm/huge_memory.c | 8 ++++----
1 file changed, 4 insertions(+), 4 deletions(-)
diff --git a/mm/huge_memory.c b/mm/huge_memory.c
index ddc631a388b9..f3fb6b97f3ce 100644
--- a/mm/huge_memory.c
+++ b/mm/huge_memory.c
@@ -3094,8 +3094,8 @@ static void __split_huge_pud_locked(struct vm_area_struct *vma, pud_t *pud,
if (!folio_test_dirty(folio) && pud_dirty(old_pud))
folio_mark_dirty(folio);
- if (!folio_test_referenced(folio) && pud_young(old_pud))
- folio_set_referenced(folio);
+ if (pud_young(old_pud) && vma_has_recency(vma))
+ folio_mark_accessed(folio);
folio_remove_rmap_pud(folio, page, vma);
add_mm_counter(vma->vm_mm, mm_counter_file(folio),
-HPAGE_PUD_NR);
@@ -3217,8 +3217,8 @@ static void __split_huge_pmd_locked(struct vm_area_struct *vma, pmd_t *pmd,
folio = page_folio(page);
if (!folio_test_dirty(folio) && pmd_dirty(old_pmd))
folio_mark_dirty(folio);
- if (!folio_test_referenced(folio) && pmd_young(old_pmd))
- folio_set_referenced(folio);
+ if (pmd_young(old_pmd) && vma_has_recency(vma))
+ folio_mark_accessed(folio);
folio_remove_rmap_pmd(folio, page, vma);
add_mm_counter(mm, mm_counter_file(folio), -HPAGE_PMD_NR);
folio_put(folio);
--
2.55.0
^ permalink raw reply [flat|nested] 21+ messages in thread
* [PATCH RFC v3 12/17] mm/mglru: folio LRU refs based active/inactive number accounting
2026-10-03 12:55 [PATCH RFC v3 00/17] mm/mglru: frequency guided promotion (MGLRU-FG) and flag cleanup Kairui Song via B4 Relay
` (10 preceding siblings ...)
2026-10-03 12:55 ` [PATCH RFC v3 11/17] mm/huge_memory: mark file folio as accessed more accurately on split Kairui Song via B4 Relay
@ 2026-10-03 12:55 ` Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 13/17] mm/mglru: make folio_inc_lru_refs lruvec lockless Kairui Song via B4 Relay
` (4 subsequent siblings)
16 siblings, 0 replies; 21+ messages in thread
From: Kairui Song via B4 Relay @ 2026-10-03 12:55 UTC (permalink / raw)
To: linux-mm
Cc: Andrew Morton, Johannes Weiner, Muchun Song, Qi Zheng,
Ying Huang, Chris Li, Baoquan He, Nico Pache, Usama Arif,
Michal Hocko, Roman Gushchin, Shakeel Butt, David Hildenbrand,
Lorenzo Stoakes, Barry Song, Axel Rasmussen, Yuanchu Xie, Wei Xu,
Vlastimil Babka, Suren Baghdasaryan, Kemeng Shi, Nhat Pham,
Youngjun Park, Zi Yan, Gregory Price, Matthew Wilcox (Oracle),
Baolin Wang, Ryan Roberts, Dev Jain, Lance Yang, Hugh Dickins,
SeongJae Park, David Rientjes, Yu Zhao, Vernon Yang,
Zicheng Wang, Chen Ridong, Tal Zussman, Kairui Song, Xiang Liu,
Ehab Ababneh, zhaozhengzhuo, Lian Wang, Kunwu Chan, linux-kernel,
cgroups, Kairui Song
From: Kairui Song <kasong@tencent.com>
Currently the active/inactive LRU sizes exported through /proc/vmstat
and memory.stat are derived from the generation window: folios in the
two newest generations are accounted as active, the rest as inactive.
This has always caused many problems:
- Unmapped file folios stick to the oldest gens, so active files are
always under-reported, unlike classical LRU.
- The reading is jumpy: on aging, the second-newest gen, previously
active, becomes inactive all at once, causing a huge drop of active
folios for no reason from time to time.
- The reading is reversed on swapless machines, and the active anon
number is an oscillating random number. When swappiness is 0 (no
swap or swap is full), MGLRU ages anon by moving the gens directly
without walk or touching any folio, to reduce the overhead. So all
inactive anon folios will suddenly count as active, or vice versa,
which breaks many metric reading components, and the reading itself
doesn't make any sense in any way.
With frequency guided promotion, hotness is already tracked per folio by
its referenced count, so we can do folio level accounting instead: a
folio is active once its refs reach LRU_REFS_ACTIVATED (currently
LRU_REFS_PROTECTED).
On access, folios are accounted to the active part. On forced aging or
PID protecting, folios are reset to LRU_REFS_WORKINGSET, dropping them
to the inactive part. The reading is now smooth and consistent, and
much closer to classical LRU. Note that on swapless machines anon refs
are never reset by aging (the shortcut skips the per-folio walk), so
active anon grows monotonically, but the reading is stable and still
much more meaningful than the previous oscillating random number.
The accounting is now per folio and lruvec independent, so no size
adjustment is needed when reparenting a dying memcg: the hierarchical
stats keep the counts in the parent automatically.
Signed-off-by: Kairui Song <kasong@tencent.com>
---
fs/proc/task_mmu.c | 2 +-
include/linux/mm_inline.h | 79 +++++++++--------
include/linux/mmzone.h | 23 +++--
mm/damon/paddr.c | 2 +-
mm/folio.c | 35 +-------
mm/madvise.c | 5 +-
mm/vmscan.c | 210 +++++++++++++++++++++++++++++-----------------
mm/workingset.c | 2 +-
8 files changed, 201 insertions(+), 157 deletions(-)
diff --git a/fs/proc/task_mmu.c b/fs/proc/task_mmu.c
index 79daaf757412..5083d7e0d078 100644
--- a/fs/proc/task_mmu.c
+++ b/fs/proc/task_mmu.c
@@ -868,7 +868,7 @@ static bool smap_check_folio_referenced(struct folio *folio)
static void smap_clear_folio_referenced(struct folio *folio)
{
if (lru_gen_enabled())
- folio_set_lru_refs(folio, 0);
+ folio_reset_lru_refs(folio);
else
folio_clear_referenced(folio);
}
diff --git a/include/linux/mm_inline.h b/include/linux/mm_inline.h
index cb54913aa184..ee7fcbe2366b 100644
--- a/include/linux/mm_inline.h
+++ b/include/linux/mm_inline.h
@@ -183,24 +183,39 @@ static inline int folio_lru_refs(const struct folio *folio)
return lru_get_refs_flags(READ_ONCE(*const_folio_flags(folio, 0)));
}
-static inline void folio_set_lru_refs(struct folio *folio, unsigned int refs)
+/**
+ * __folio_set_lru_refs - Set a folio's LRU refs.
+ * @folio: the folio
+ * @refs: the new referenced count (0 .. LRU_REFS_MAX)
+ *
+ * Set the folio's LRU refs. The folio must be off the LRU list (e.g.,
+ * isolated), or use folio_inc_lru_refs or folio_reset_lru_refs instead.
+ */
+static inline void __folio_set_lru_refs(struct folio *folio, unsigned int refs)
{
unsigned long new_flags, old_flags = READ_ONCE(*folio_flags(folio, 0));
do {
new_flags = old_flags;
+ VM_WARN_ON_ONCE(lru_get_gen_flags(old_flags) != -1);
lru_set_refs_flags(&new_flags, refs);
} while (!try_cmpxchg(folio_flags(folio, 0), &old_flags, new_flags));
}
#ifdef CONFIG_LRU_GEN
void folio_inc_lru_refs(struct folio *folio, unsigned int flags);
+bool folio_reset_lru_refs(struct folio *folio);
#else
static inline void folio_inc_lru_refs(struct folio *folio, unsigned int flags)
{
/* Should not be called with !CONFIG_LRU_GEN */
WARN_ON_ONCE(1);
}
+
+static inline bool folio_reset_lru_refs(struct folio *folio)
+{
+ return false;
+}
#endif
/**
@@ -214,7 +229,9 @@ static inline void folio_inc_lru_refs(struct folio *folio, unsigned int flags)
*/
static inline void folio_migrate_lru_refs(struct folio *new, const struct folio *old)
{
- folio_set_lru_refs(new, folio_lru_refs(old));
+ VM_WARN_ON_ONCE_FOLIO(folio_test_lru(old), old);
+ VM_WARN_ON_ONCE_FOLIO(folio_test_lru(new), new);
+ __folio_set_lru_refs(new, folio_lru_refs(old));
}
#ifdef CONFIG_LRU_GEN
@@ -265,19 +282,15 @@ static inline int lru_tier_from_refs(unsigned int refs)
return fls(refs - 1);
}
-static inline int folio_lru_gen(const struct folio *folio)
+static inline bool lru_refs_is_active(unsigned int refs)
{
- return lru_get_gen_flags(READ_ONCE(*const_folio_flags(folio, 0)));
+ VM_WARN_ON_ONCE(refs > LRU_REFS_MAX);
+ return refs >= LRU_REFS_ACTIVATED;
}
-static inline bool lru_gen_is_active(const struct lruvec *lruvec, int gen)
+static inline int folio_lru_gen(const struct folio *folio)
{
- unsigned long max_seq = READ_ONCE(lruvec->lrugen.max_seq);
-
- VM_WARN_ON_ONCE(gen >= MAX_NR_GENS);
-
- /* see the comment on MIN_NR_GENS */
- return gen == lru_gen_from_seq(max_seq) || gen == lru_gen_from_seq(max_seq - 1);
+ return lru_get_gen_flags(READ_ONCE(*const_folio_flags(folio, 0)));
}
static inline void lru_gen_update_size(struct lruvec *lruvec, struct folio *folio,
@@ -286,7 +299,6 @@ static inline void lru_gen_update_size(struct lruvec *lruvec, struct folio *foli
int type = folio_is_file_lru(folio);
int zone = folio_zonenum(folio);
int delta = folio_nr_pages(folio);
- enum lru_list lru = type * LRU_INACTIVE_FILE;
struct lru_gen_folio *lrugen = &lruvec->lrugen;
VM_WARN_ON_ONCE(old_gen != -1 && old_gen >= MAX_NR_GENS);
@@ -297,19 +309,6 @@ static inline void lru_gen_update_size(struct lruvec *lruvec, struct folio *foli
atomic_long_sub(delta, &lrugen->nr_pages[old_gen][type][zone]);
if (new_gen >= 0)
atomic_long_add(delta, &lrugen->nr_pages[new_gen][type][zone]);
-
- /* return now if not a promotion */
- if (old_gen < 0 || new_gen < 0)
- return;
-
- /* promotion */
- if (!lru_gen_is_active(lruvec, old_gen) && lru_gen_is_active(lruvec, new_gen)) {
- __update_lru_size(lruvec, lru, zone, -delta);
- __update_lru_size(lruvec, lru + LRU_ACTIVE, zone, delta);
- }
-
- /* demotion requires isolation, e.g., lru_deactivate_fn() */
- VM_WARN_ON_ONCE(lru_gen_is_active(lruvec, old_gen) && !lru_gen_is_active(lruvec, new_gen));
}
static inline unsigned long lru_gen_folio_seq(const struct lruvec *lruvec,
@@ -351,7 +350,7 @@ static inline bool lru_gen_add_folio(struct lruvec *lruvec, struct folio *folio,
{
unsigned long seq;
unsigned long flags;
- int gen = folio_lru_gen(folio);
+ int gen, refs;
int type = folio_is_file_lru(folio);
int zone = folio_zonenum(folio);
int delta = folio_nr_pages(folio);
@@ -359,7 +358,7 @@ static inline bool lru_gen_add_folio(struct lruvec *lruvec, struct folio *folio,
struct lru_gen_folio *lrugen = &lruvec->lrugen;
BUILD_BUG_ON(BIT(LRU_GEN_WIDTH - 1) != MAX_NR_GENS);
- VM_WARN_ON_ONCE_FOLIO(gen != -1, folio);
+ VM_WARN_ON_ONCE_FOLIO(folio_lru_gen(folio) != -1, folio);
if (folio_test_unevictable(folio) || !lrugen->enabled)
return false;
@@ -368,10 +367,12 @@ static inline bool lru_gen_add_folio(struct lruvec *lruvec, struct folio *folio,
gen = lru_gen_from_seq(seq);
flags = (gen + 1UL) << LRU_GEN_PGOFF;
/* see the comment on MIN_NR_GENS about PG_active */
- set_mask_bits(folio_flags(folio, 0), LRU_GEN_MASK | BIT(PG_active), flags);
+ flags = set_mask_bits(folio_flags(folio, 0), LRU_GEN_MASK | BIT(PG_active), flags);
+ /* use the refs from the atomic snapshot to avoid raced update */
+ refs = lru_get_refs_flags(flags);
lru_gen_update_size(lruvec, folio, -1, gen);
- if (lru_gen_is_active(lruvec, gen))
+ if (lru_refs_is_active(refs))
lru += LRU_ACTIVE;
__update_lru_size(lruvec, lru, zone, delta);
@@ -387,24 +388,31 @@ static inline bool lru_gen_add_folio(struct lruvec *lruvec, struct folio *folio,
static inline bool lru_gen_del_folio(struct lruvec *lruvec, struct folio *folio, bool reclaiming)
{
unsigned long flags;
- int gen = folio_lru_gen(folio);
+ int gen, refs;
int zone = folio_zonenum(folio);
int delta = folio_nr_pages(folio);
enum lru_list lru = folio_is_file_lru(folio) * LRU_INACTIVE_FILE;
+ unsigned long max_seq = READ_ONCE(lruvec->lrugen.max_seq);
+ flags = set_mask_bits(folio_flags(folio, 0), LRU_GEN_MASK, 0);
+ gen = lru_get_gen_flags(flags);
+ refs = lru_get_refs_flags(flags);
if (gen < 0)
return false;
VM_WARN_ON_ONCE_FOLIO(folio_test_active(folio), folio);
VM_WARN_ON_ONCE_FOLIO(folio_test_unevictable(folio), folio);
- /* for folio_migrate_flags() */
- flags = !reclaiming && lru_gen_is_active(lruvec, gen) ? BIT(PG_active) : 0;
- flags = set_mask_bits(folio_flags(folio, 0), LRU_GEN_MASK, flags);
- gen = ((flags & LRU_GEN_MASK) >> LRU_GEN_PGOFF) - 1;
+ /*
+ * See the comment in lru_gen_folio_seq. For migration, compaction,
+ * or any other isolation of a hot folio, try best to retain its gen
+ * info. Ideally we would keep the full gen info.
+ */
+ if (!reclaiming && ((max_seq - gen) % MAX_NR_GENS) < MIN_NR_GENS)
+ folio_set_active(folio);
lru_gen_update_size(lruvec, folio, gen, -1);
- if (lru_gen_is_active(lruvec, gen))
+ if (lru_refs_is_active(refs))
lru += LRU_ACTIVE;
__update_lru_size(lruvec, lru, zone, -delta);
list_del(&folio->lru);
@@ -455,6 +463,7 @@ static __always_inline void folio_inc_lru_refs_fast(struct folio *folio)
return;
}
+ BUILD_BUG_ON(LRU_REFS_REFERENCED >= LRU_REFS_ACTIVATED);
old_flags = READ_ONCE(*folio_flags(folio, 0));
do {
if (lru_get_refs_flags(old_flags))
diff --git a/include/linux/mmzone.h b/include/linux/mmzone.h
index e3a438bc2719..e096fd0a5047 100644
--- a/include/linux/mmzone.h
+++ b/include/linux/mmzone.h
@@ -434,15 +434,15 @@ enum lruvec_flags {
* least twice before handing this folio over to the eviction. The first check
* clears the accessed bit from the initial fault; the second check makes sure
* this folio hasn't been used since then. This process, AKA second chance,
- * requires a minimum of two generations, hence MIN_NR_GENS. And to maintain ABI
- * compatibility with the active/inactive LRU, e.g., /proc/vmstat, these two
- * generations are considered active; the rest of generations, if they exist,
- * are considered inactive. See lru_gen_is_active().
+ * requires a minimum of two generations, hence MIN_NR_GENS.
+ *
+ * Active/inactive is per-folio based on the referenced count;
+ * see the comment above MAX_NR_TIERS.
*
* PG_active is always cleared while a folio is on one of lrugen->folios[] so
- * that the sliding window needs not to worry about it. And it's set again when
- * a folio considered active is isolated for non-reclaiming purposes, e.g.,
- * migration. See lru_gen_add_folio() and lru_gen_del_folio().
+ * that the sliding window needs not to worry about it. It is set on a
+ * fault-driven refault (see lru_gen_refault), or isolation (see
+ * lru_gen_del_folio).
*
* MAX_NR_GENS is set to 4 so that the multi-gen LRU can support twice the
* number of categories of the active/inactive LRU when keeping track of
@@ -542,11 +542,15 @@ enum lruvec_flags {
*
* A folio's referenced count never goes backwards except upon gen
* increase as described above, or when explicitly reset by
- * lru_gen_clear_refs(). Refault of a reclaimed folio restores
+ * folio_reset_lru_refs(). Refault of a reclaimed folio restores
* its referenced count, capped at LRU_REFS_PROTECTED, which aligns with
* promotion. Page table refaults of previous workingset folios send
* them to the latest gen, driving aging faster.
*
+ * For the active/inactive LRU ABI (/proc/vmstat), a folio is active
+ * if its referenced count reaches LRU_REFS_ACTIVATED. The threshold
+ * currently set to LRU_REFS_PROTECTED, which is tier 2.
+ *
* MAX_NR_TIERS is set to 4 so that the multi-gen LRU can support twice
* the number of categories of the active/inactive LRU.
*/
@@ -561,6 +565,7 @@ enum lruvec_flags {
#define LRU_REFS_REFERENCED 0x1
#define LRU_REFS_WORKINGSET 0x2
#define LRU_REFS_PROTECTED 0x3
+#define LRU_REFS_ACTIVATED LRU_REFS_PROTECTED
#ifndef __GENERATING_BOUNDS_H
@@ -669,6 +674,8 @@ struct lru_gen_mm_walk {
unsigned long seq;
/* the next address within an mm to scan */
unsigned long next_addr;
+ /* to batch activated pages */
+ int nr_activated[ANON_AND_FILE][MAX_NR_ZONES];
/* to batch promoted pages */
int nr_pages[MAX_NR_GENS][ANON_AND_FILE][MAX_NR_ZONES];
/* to batch the mm stats */
diff --git a/mm/damon/paddr.c b/mm/damon/paddr.c
index 4db141b3fce3..b9408ad29a97 100644
--- a/mm/damon/paddr.c
+++ b/mm/damon/paddr.c
@@ -284,7 +284,7 @@ static unsigned long damon_pa_pageout(struct damon_region *r,
* so the hotness can be, and better be dropped.
*/
if (lru_gen_enabled())
- folio_set_lru_refs(folio, 0);
+ folio_reset_lru_refs(folio);
else
folio_clear_referenced(folio);
folio_test_clear_young(folio);
diff --git a/mm/folio.c b/mm/folio.c
index e200efab814f..9b9a5fe568c9 100644
--- a/mm/folio.c
+++ b/mm/folio.c
@@ -349,35 +349,6 @@ static void __lru_cache_activate_folio(struct folio *folio)
local_unlock(&cpu_fbatches.lock);
}
-#ifdef CONFIG_LRU_GEN
-
-static bool lru_gen_clear_refs(struct folio *folio)
-{
- int gen = folio_lru_gen(folio);
- int type = folio_is_file_lru(folio);
- unsigned long seq;
-
- if (gen < 0)
- return true;
-
- folio_set_lru_refs(folio, 0);
-
- rcu_read_lock();
- seq = READ_ONCE(folio_lruvec(folio)->lrugen.min_seq[type]);
- rcu_read_unlock();
- /* whether can do without shuffling under the LRU lock */
- return gen == lru_gen_from_seq(seq);
-}
-
-#else /* !CONFIG_LRU_GEN */
-
-static bool lru_gen_clear_refs(struct folio *folio)
-{
- return false;
-}
-
-#endif /* CONFIG_LRU_GEN */
-
/**
* folio_mark_accessed - Mark a folio as having seen activity.
* @folio: The folio to mark.
@@ -554,7 +525,7 @@ static void lru_lazyfree(struct lruvec *lruvec, struct folio *folio)
lruvec_del_folio(lruvec, folio);
folio_clear_active(folio);
if (lru_gen_enabled())
- lru_gen_clear_refs(folio);
+ __folio_set_lru_refs(folio, 0);
else
folio_clear_referenced(folio);
/*
@@ -627,7 +598,7 @@ void deactivate_file_folio(struct folio *folio)
if (folio_test_unevictable(folio) || !folio_test_lru(folio))
return;
- if (lru_gen_enabled() && lru_gen_clear_refs(folio))
+ if (lru_gen_enabled() && !folio_reset_lru_refs(folio))
return;
folio_batch_add_and_move(folio, lru_deactivate_file);
@@ -646,7 +617,7 @@ void folio_deactivate(struct folio *folio)
if (folio_test_unevictable(folio) || !folio_test_lru(folio))
return;
- if (lru_gen_enabled() ? lru_gen_clear_refs(folio) : !folio_test_active(folio))
+ if (lru_gen_enabled() ? !folio_reset_lru_refs(folio) : !folio_test_active(folio))
return;
folio_batch_add_and_move(folio, lru_deactivate);
diff --git a/mm/madvise.c b/mm/madvise.c
index c221234fa199..5a2ca6cd78f2 100644
--- a/mm/madvise.c
+++ b/mm/madvise.c
@@ -375,7 +375,8 @@ static void madvise_cold_prep_folio(struct folio *folio)
* list should have PG_workingset so its refault is still accounted
* as a workingset refault, and the referenced bit always needs to
* be cleared even for inactive folios. Nothing to do for MGLRU as
- * folio_deactivate() always calls lru_gen_clear_refs().
+ * folio_deactivate() clears the refs count through
+ * folio_reset_lru_refs().
*/
if (!lru_gen_enabled()) {
folio_clear_referenced(folio);
@@ -396,7 +397,7 @@ static void madvise_pageout_prep_folio(struct folio *folio)
* the workingset classification is preserved for PSI.
*/
if (lru_gen_enabled())
- folio_set_lru_refs(folio, refs >= LRU_REFS_WORKINGSET ? LRU_REFS_WORKINGSET : 0);
+ __folio_set_lru_refs(folio, refs >= LRU_REFS_WORKINGSET ? LRU_REFS_WORKINGSET : 0);
}
static int madvise_cold_or_pageout_pte_range(pmd_t *pmd,
diff --git a/mm/vmscan.c b/mm/vmscan.c
index a8c49262d5b3..ae4a0522a8d5 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -977,17 +977,19 @@ enum folio_references {
void folio_inc_lru_refs(struct folio *folio, unsigned int flags)
{
int max_gen, min_gen;
- int type, refs, old_gen, gen;
+ int file, old_refs, refs, old_gen, gen;
unsigned long new_flags, old_flags, max_seq;
+ long nr_pages = folio_nr_pages(folio);
struct lru_gen_folio *lrugen;
struct lruvec *lruvec = NULL;
- type = folio_is_file_lru(folio);
old_flags = READ_ONCE(*folio_flags(folio, 0));
do {
new_flags = old_flags;
old_gen = lru_get_gen_flags(old_flags);
- refs = lru_get_refs_flags(old_flags) + 1;
+ old_refs = lru_get_refs_flags(old_flags);
+ file = folio_flags_is_file_lru(&old_flags);
+ refs = old_refs + 1;
gen = old_gen;
if (old_gen < 0)
goto out;
@@ -1005,7 +1007,7 @@ void folio_inc_lru_refs(struct folio *folio, unsigned int flags)
}
max_seq = READ_ONCE(lrugen->max_seq);
max_gen = lru_gen_from_seq(max_seq);
- min_gen = lru_gen_from_seq(READ_ONCE(lrugen->min_seq[type]));
+ min_gen = lru_gen_from_seq(READ_ONCE(lrugen->min_seq[file]));
if (old_gen == max_gen)
goto out;
@@ -1037,55 +1039,108 @@ void folio_inc_lru_refs(struct folio *folio, unsigned int flags)
if (gen != old_gen)
lru_gen_update_size(lruvec, folio, old_gen, gen);
+ if (lru_refs_is_active(old_refs) != lru_refs_is_active(refs) && old_gen >= 0) {
+ enum lru_list lru = file * LRU_INACTIVE_FILE;
+
+ __update_lru_size(lruvec, lru + lru_refs_is_active(old_refs),
+ folio_zonenum(folio), -nr_pages);
+ __update_lru_size(lruvec, lru + lru_refs_is_active(refs),
+ folio_zonenum(folio), nr_pages);
+ }
if (lruvec)
lruvec_unlock_irq(lruvec);
}
+/*
+ * Reset the folio's lru refs indicator. The caller doesn't need to hold
+ * the folio lock, isolate the folio, or hold the lruvec lock.
+ */
+bool folio_reset_lru_refs(struct folio *folio)
+{
+ int type, gen, refs;
+ unsigned long seq, new_flags, old_flags;
+ long nr_pages = folio_nr_pages(folio);
+ struct lruvec *lruvec;
+ bool reset = false;
+
+ old_flags = READ_ONCE(*folio_flags(folio, 0));
+ do {
+ new_flags = old_flags;
+ refs = lru_get_refs_flags(old_flags);
+ type = folio_flags_is_file_lru(&old_flags);
+ gen = lru_get_gen_flags(old_flags);
+ if (!refs)
+ break;
+ lru_set_refs_flags(&new_flags, 0);
+ } while (!try_cmpxchg(folio_flags(folio, 0), &old_flags, new_flags));
+
+ if (gen >= 0) {
+ lruvec = folio_lruvec_live_get(folio);
+
+ /* If the folio is on list, caller might also want to demote it */
+ seq = READ_ONCE(lruvec->lrugen.min_seq[type]);
+ reset = (gen != lru_gen_from_seq(seq));
+ if (lru_refs_is_active(refs)) {
+ __update_lru_size(lruvec, type * LRU_FILE + LRU_ACTIVE,
+ folio_zonenum(folio), -nr_pages);
+ __update_lru_size(lruvec, type * LRU_FILE,
+ folio_zonenum(folio), nr_pages);
+ }
+
+ folio_lruvec_live_put(lruvec);
+ }
+
+ return reset;
+}
+
/*
* Update the folio's lru refs indicator during a page table walk or the
* look-around. max_seq can be stale as neither holds the LRU lock.
*
- * Returns the old generation and stores the new generation in @new_gen if
- * the folio is on the LRU and not in the newest generation, or -1 otherwise.
+ * Returns the old generation, or -1 if the folio is off the LRU list.
+ * Old LRU refs, and updated gen and LRU refs info by the successful
+ * cmpxchg are all stored by returning arguments.
*/
static int folio_inc_lru_refs_walk(struct folio *folio, struct lruvec *lruvec,
- const vma_flags_t *vma_flags,
- int *new_gen, int *type)
+ const vma_flags_t *vma_flags, int *new_gen,
+ int *old_refs, int *new_refs, int *type)
{
unsigned long new_flags, old_flags = READ_ONCE(*folio_flags(folio, 0));
unsigned long max_seq = READ_ONCE(lruvec->lrugen.max_seq);
- int refs, gen, min_gen, max_gen, ret;
+ int refs, old_gen, min_gen, max_gen;
max_gen = lru_gen_from_seq(max_seq);
-
do {
- gen = lru_get_gen_flags(old_flags);
+ old_gen = lru_get_gen_flags(old_flags);
refs = lru_get_refs_flags(old_flags) + 1;
*type = folio_flags_is_file_lru(&old_flags);
min_gen = lru_gen_from_seq(READ_ONCE(lruvec->lrugen.min_seq[*type]));
new_flags = old_flags;
- if (gen >= 0 && gen != max_gen) {
- ret = gen;
+ if (old_gen >= 0 && old_gen != max_gen) {
+ *new_refs = min(refs, LRU_REFS_PROTECTED);
/* Promote second page table access or executable */
if (refs > LRU_REFS_REFERENCED || is_exec_file_folio(folio, vma_flags))
*new_gen = max_gen;
/* First access only defers eviction from the oldest gen */
- else if (gen == min_gen)
- *new_gen = (gen + 1) % MAX_NR_GENS;
+ else if (old_gen == min_gen)
+ *new_gen = (old_gen + 1) % MAX_NR_GENS;
else
- *new_gen = gen;
+ *new_gen = old_gen;
lru_set_gen_flags(&new_flags, *new_gen);
- lru_set_refs_flags(&new_flags, min(refs, LRU_REFS_PROTECTED));
+ lru_set_refs_flags(&new_flags, *new_refs);
} else {
- ret = -1;
- lru_set_refs_flags(&new_flags, min(refs, LRU_REFS_MAX));
+ *new_gen = old_gen;
+ *new_refs = min(refs, LRU_REFS_MAX);
+ lru_set_refs_flags(&new_flags, *new_refs);
}
if (new_flags == old_flags)
break;
} while (!try_cmpxchg(folio_flags(folio, 0), &old_flags, new_flags));
- return ret;
+ *old_refs = lru_get_refs_flags(old_flags);
+
+ return old_gen;
}
/*
@@ -3539,10 +3594,11 @@ static bool positive_ctrl_err(struct ctrl_pos *sp, struct ctrl_pos *pv)
* the aging
******************************************************************************/
-static int __folio_inc_gen(struct folio *folio, int old_gen, bool *increased)
+static int __folio_inc_gen(struct lruvec *lruvec, struct folio *folio,
+ int old_gen, bool *increased)
{
unsigned long new_flags, old_flags = READ_ONCE(*folio_flags(folio, 0));
- int refs, new_gen;
+ int file, refs, old_refs, new_gen;
do {
new_gen = lru_get_gen_flags(old_flags);
@@ -3556,11 +3612,24 @@ static int __folio_inc_gen(struct folio *folio, int old_gen, bool *increased)
new_flags = old_flags;
new_gen = (old_gen + 1) % MAX_NR_GENS;
- refs = lru_get_refs_flags(old_flags);
+ old_refs = lru_get_refs_flags(old_flags);
+ file = folio_flags_is_file_lru(&old_flags);
+ refs = min(old_refs, LRU_REFS_WORKINGSET);
lru_set_gen_flags(&new_flags, new_gen);
- lru_set_refs_flags(&new_flags, min(refs, LRU_REFS_WORKINGSET));
+ lru_set_refs_flags(&new_flags, refs);
} while (!try_cmpxchg(folio_flags(folio, 0), &old_flags, new_flags));
+ /* Refs capped below PROTECTED: demote if previously active. */
+ if (lru_refs_is_active(old_refs) != lru_refs_is_active(refs)) {
+ enum lru_list lru = file * LRU_INACTIVE_FILE;
+ int nr_pages = folio_nr_pages(folio);
+
+ __update_lru_size(lruvec, lru + lru_refs_is_active(old_refs),
+ folio_zonenum(folio), -nr_pages);
+ __update_lru_size(lruvec, lru + lru_refs_is_active(refs),
+ folio_zonenum(folio), nr_pages);
+ }
+
if (increased)
*increased = true;
return new_gen;
@@ -3577,7 +3646,7 @@ static int folio_inc_gen(struct lruvec *lruvec, struct folio *folio)
struct lru_gen_folio *lrugen = &lruvec->lrugen;
int new_gen, old_gen = lru_gen_from_seq(lrugen->min_seq[type]);
- new_gen = __folio_inc_gen(folio, old_gen, &gen_increased);
+ new_gen = __folio_inc_gen(lruvec, folio, old_gen, &gen_increased);
if (gen_increased)
lru_gen_update_size(lruvec, folio, old_gen, new_gen);
@@ -3585,18 +3654,26 @@ static int folio_inc_gen(struct lruvec *lruvec, struct folio *folio)
}
static void update_batch_size(struct lru_gen_mm_walk *walk, struct folio *folio,
- int old_gen, int new_gen, int type)
+ int old_gen, int new_gen, int old_refs, int new_refs, int type)
{
int zone = folio_zonenum(folio);
int delta = folio_nr_pages(folio);
- VM_WARN_ON_ONCE(old_gen >= MAX_NR_GENS);
- VM_WARN_ON_ONCE(new_gen >= MAX_NR_GENS);
-
- walk->batched++;
+ /* gen counter update */
+ if (old_gen != new_gen) {
+ VM_WARN_ON_ONCE(old_gen >= MAX_NR_GENS);
+ VM_WARN_ON_ONCE(new_gen >= MAX_NR_GENS);
+ walk->batched++;
+ walk->nr_pages[old_gen][type][zone] -= delta;
+ walk->nr_pages[new_gen][type][zone] += delta;
+ }
- walk->nr_pages[old_gen][type][zone] -= delta;
- walk->nr_pages[new_gen][type][zone] += delta;
+ /* active/inactive counter update */
+ if (lru_refs_is_active(old_refs) != lru_refs_is_active(new_refs)) {
+ walk->nr_activated[type][zone] +=
+ lru_refs_is_active(new_refs) ? delta : -delta;
+ walk->batched++;
+ }
}
static void reset_batch_size(struct lru_gen_mm_walk *walk)
@@ -3608,7 +3685,6 @@ static void reset_batch_size(struct lru_gen_mm_walk *walk)
walk->batched = 0;
for_each_gen_type_zone(gen, type, zone) {
- enum lru_list lru = type * LRU_INACTIVE_FILE;
int delta = walk->nr_pages[gen][type][zone];
if (!delta)
@@ -3616,10 +3692,21 @@ static void reset_batch_size(struct lru_gen_mm_walk *walk)
walk->nr_pages[gen][type][zone] = 0;
atomic_long_add(delta, &lrugen->nr_pages[gen][type][zone]);
+ }
- if (lru_gen_is_active(lruvec, gen))
- lru += LRU_ACTIVE;
- __update_lru_size(lruvec, lru, zone, delta);
+ /* apply batched active/inactive updates */
+ for (type = 0; type < ANON_AND_FILE; type++) {
+ for (zone = 0; zone < MAX_NR_ZONES; zone++) {
+ enum lru_list lru = type * LRU_INACTIVE_FILE;
+ int delta = walk->nr_activated[type][zone];
+
+ if (delta) {
+ __update_lru_size(lruvec, lru + LRU_ACTIVE,
+ zone, delta);
+ __update_lru_size(lruvec, lru, zone, -delta);
+ walk->nr_activated[type][zone] = 0;
+ }
+ }
}
lruvec_unlock_irq(lruvec);
@@ -3775,7 +3862,7 @@ static bool suitable_to_scan(int total, int young)
static void walk_update_folio(struct lru_gen_mm_walk *walk, struct vm_area_struct *vma,
struct lruvec *lruvec, struct folio *folio, bool dirty)
{
- int new_gen, old_gen, type;
+ int new_gen, old_gen, old_refs, new_refs, type;
unsigned int flags = LRU_REF_MAPPED;
if (!folio)
@@ -3788,9 +3875,10 @@ static void walk_update_folio(struct lru_gen_mm_walk *walk, struct vm_area_struc
if (walk) {
old_gen = folio_inc_lru_refs_walk(folio, lruvec, &vma->flags,
- &new_gen, &type);
- if (old_gen >= 0 && old_gen != new_gen)
- update_batch_size(walk, folio, old_gen, new_gen, type);
+ &new_gen, &old_refs, &new_refs, &type);
+ if (old_gen >= 0)
+ update_batch_size(walk, folio, old_gen, new_gen,
+ old_refs, new_refs, type);
} else {
if (is_exec_file_folio(folio, &vma->flags))
flags |= LRU_REF_EXEC;
@@ -4149,6 +4237,7 @@ static void clear_mm_walk(void)
VM_WARN_ON_ONCE(walk && memchr_inv(walk->nr_pages, 0, sizeof(walk->nr_pages)));
VM_WARN_ON_ONCE(walk && memchr_inv(walk->mm_stats, 0, sizeof(walk->mm_stats)));
+ VM_WARN_ON_ONCE(walk && memchr_inv(walk->nr_activated, 0, sizeof(walk->nr_activated)));
current->reclaim_state->mm_walk = NULL;
@@ -4187,8 +4276,6 @@ static bool inc_min_seq(struct lruvec *lruvec, int type, int swappiness)
goto done;
VM_WARN_ON_ONCE(get_nr_gens(lruvec, type) != MAX_NR_GENS);
- VM_WARN_ON_ONCE(lru_gen_is_active(lruvec, old_gen) !=
- lru_gen_is_active(lruvec, target_gen));
/* prevent cold/hot inversion if the type is evictable */
for (zone = 0; zone < MAX_NR_ZONES; zone++) {
struct list_head *target_list = &lrugen->folios[target_gen][type][zone];
@@ -4211,7 +4298,7 @@ static bool inc_min_seq(struct lruvec *lruvec, int type, int swappiness)
prefetchw_next_lru_folio(folio, head);
pos = pos->next;
- new_gen = __folio_inc_gen(folio, old_gen, &gen_increased);
+ new_gen = __folio_inc_gen(lruvec, folio, old_gen, &gen_increased);
/*
* If gen_increased is false, this is a promotion. Put folios
* at the head of the promoted gen. Otherwise, put them at
@@ -4298,8 +4385,7 @@ static void try_to_inc_min_seq(struct lruvec *lruvec, int swappiness)
static bool inc_max_seq(struct lruvec *lruvec, unsigned long seq, int swappiness)
{
bool success;
- int prev, next;
- int type, zone;
+ int type, next;
struct lru_gen_folio *lrugen = &lruvec->lrugen;
restart:
if (seq < READ_ONCE(lrugen->max_seq))
@@ -4325,32 +4411,10 @@ static bool inc_max_seq(struct lruvec *lruvec, unsigned long seq, int swappiness
goto restart;
}
- /*
- * Update the active/inactive LRU sizes for compatibility. Both sides of
- * the current max_seq need to be covered, since max_seq+1 can overlap
- * with min_seq[LRU_GEN_ANON] if swapping is constrained. And if they do
- * overlap, cold/hot inversion happens.
- */
- prev = lru_gen_from_seq(lrugen->max_seq - 1);
- next = lru_gen_from_seq(lrugen->max_seq + 1);
-
- for (type = 0; type < ANON_AND_FILE; type++) {
- for (zone = 0; zone < MAX_NR_ZONES; zone++) {
- enum lru_list lru = type * LRU_INACTIVE_FILE;
- long delta = atomic_long_read(&lrugen->nr_pages[prev][type][zone]) -
- atomic_long_read(&lrugen->nr_pages[next][type][zone]);
-
- if (!delta)
- continue;
-
- __update_lru_size(lruvec, lru, zone, delta);
- __update_lru_size(lruvec, lru + LRU_ACTIVE, zone, -delta);
- }
- }
-
for (type = 0; type < ANON_AND_FILE; type++)
reset_ctrl_pos(lruvec, type, false);
+ next = lru_gen_from_seq(lrugen->max_seq + 1);
WRITE_ONCE(lrugen->timestamps[next], jiffies);
/* make sure preceding modifications appear */
smp_store_release(&lrugen->max_seq, lrugen->max_seq + 1);
@@ -4869,7 +4933,6 @@ static void __lru_gen_reparent_memcg(struct lruvec *child_lruvec, struct lruvec
int zone, int type)
{
struct lru_gen_folio *child_lrugen, *parent_lrugen;
- enum lru_list lru = type * LRU_INACTIVE_FILE;
int i;
child_lrugen = &child_lruvec->lrugen;
@@ -4878,8 +4941,6 @@ static void __lru_gen_reparent_memcg(struct lruvec *child_lruvec, struct lruvec
for (i = 0; i < get_nr_gens(child_lruvec, type); i++) {
int gen = lru_gen_from_seq(child_lrugen->max_seq - i);
long nr_pages = atomic_long_read(&child_lrugen->nr_pages[gen][type][zone]);
- int child_lru_active = lru_gen_is_active(child_lruvec, gen) ? LRU_ACTIVE : 0;
- int parent_lru_active = lru_gen_is_active(parent_lruvec, gen) ? LRU_ACTIVE : 0;
/* Assuming that child pages are colder than parent pages */
list_splice_tail_init(&child_lrugen->folios[gen][type][zone],
@@ -4887,11 +4948,6 @@ static void __lru_gen_reparent_memcg(struct lruvec *child_lruvec, struct lruvec
atomic_long_set(&child_lrugen->nr_pages[gen][type][zone], 0);
atomic_long_add(nr_pages, &parent_lrugen->nr_pages[gen][type][zone]);
-
- if (lru_gen_is_active(child_lruvec, gen) != lru_gen_is_active(parent_lruvec, gen)) {
- __update_lru_size(child_lruvec, lru + child_lru_active, zone, -nr_pages);
- __update_lru_size(parent_lruvec, lru + parent_lru_active, zone, nr_pages);
- }
}
}
@@ -5228,7 +5284,7 @@ static int evict_folios(unsigned long nr_to_scan, struct lruvec *lruvec,
* its first page table access lands in the second newest one.
* See "Referenced count feedback" above.
*/
- folio_set_lru_refs(folio, min(folio_lru_refs(folio), LRU_REFS_WORKINGSET));
+ __folio_set_lru_refs(folio, min(folio_lru_refs(folio), LRU_REFS_WORKINGSET));
if (lru_gen_folio_seq(lruvec, folio, false) == min_seq[type])
folio_set_active(folio);
}
diff --git a/mm/workingset.c b/mm/workingset.c
index 16cdd88e275e..7681af535ca7 100644
--- a/mm/workingset.c
+++ b/mm/workingset.c
@@ -334,7 +334,7 @@ static void lru_gen_refault(struct folio *folio, void *shadow)
mod_lruvec_state(lruvec, WORKINGSET_ACTIVATE_BASE + type, delta);
}
/* Refault is also promotion, cap the refs like folio_inc_lru_refs */
- folio_set_lru_refs(folio, min(refs, LRU_REFS_PROTECTED));
+ __folio_set_lru_refs(folio, min(refs, LRU_REFS_PROTECTED));
}
if (refs >= LRU_REFS_WORKINGSET)
--
2.55.0
^ permalink raw reply [flat|nested] 21+ messages in thread
* [PATCH RFC v3 13/17] mm/mglru: make folio_inc_lru_refs lruvec lockless
2026-10-03 12:55 [PATCH RFC v3 00/17] mm/mglru: frequency guided promotion (MGLRU-FG) and flag cleanup Kairui Song via B4 Relay
` (11 preceding siblings ...)
2026-10-03 12:55 ` [PATCH RFC v3 12/17] mm/mglru: folio LRU refs based active/inactive number accounting Kairui Song via B4 Relay
@ 2026-10-03 12:55 ` Kairui Song via B4 Relay
2026-10-04 17:29 ` KunWu Chan
2026-10-03 12:55 ` [PATCH RFC v3 14/17] mm/mglru: reparent folios from all generations Kairui Song via B4 Relay
` (3 subsequent siblings)
16 siblings, 1 reply; 21+ messages in thread
From: Kairui Song via B4 Relay @ 2026-10-03 12:55 UTC (permalink / raw)
To: linux-mm
Cc: Andrew Morton, Johannes Weiner, Muchun Song, Qi Zheng,
Ying Huang, Chris Li, Baoquan He, Nico Pache, Usama Arif,
Michal Hocko, Roman Gushchin, Shakeel Butt, David Hildenbrand,
Lorenzo Stoakes, Barry Song, Axel Rasmussen, Yuanchu Xie, Wei Xu,
Vlastimil Babka, Suren Baghdasaryan, Kemeng Shi, Nhat Pham,
Youngjun Park, Zi Yan, Gregory Price, Matthew Wilcox (Oracle),
Baolin Wang, Ryan Roberts, Dev Jain, Lance Yang, Hugh Dickins,
SeongJae Park, David Rientjes, Yu Zhao, Vernon Yang,
Zicheng Wang, Chen Ridong, Tal Zussman, Kairui Song, Xiang Liu,
Ehab Ababneh, zhaozhengzhuo, Lian Wang, Kunwu Chan, linux-kernel,
cgroups, Kairui Song
From: Kairui Song <kasong@tencent.com>
The lruvec spinlock in folio_inc_lru_refs was only taken to keep
concurrent aging from corrupting the size counters. This is no
longer needed: the per-generation counters are atomic, and the
active/inactive ABI counters are updated per folio and linearized by
the CAS on the folio flags, so concurrent aging can no longer corrupt
them. The same argument already allowed folio_reset_lru_refs() to
drop the lock.
Use folio_lruvec_live_get()/folio_lruvec_live_put() to hold the RCU
read lock around the lruvec lookup, and drop the spinlock entirely.
Without the lock, aging can advance max_seq while a promotion is in
flight. Generations are indexed by seq % MAX_NR_GENS, so the folio
flags can return to the value read earlier and the cmpxchg succeeds
with a target gen computed from a stale max_seq. Recheck max_seq after
a gen-changing cmpxchg: if it moved, account the committed update and
retry with LRU_REF_FORCE, which promotes the folio to the newest gen.
The retry may count the access twice, which is acceptable and avoids
folio_activate(), whose LRU lock causes latency jitter. smp_rmb()
orders the flags read, including the one returned by a failed
cmpxchg(), before the max_seq read.
Also simplify the post-CAS accounting guards: gen is only assigned
non-negative values after the old_gen < 0 early exit, so gen != old_gen
implies gen >= 0, and lru_gen_update_size() needs no explicit guard.
The active/inactive ABI update keeps its gen >= 0 check because an
off-LRU folio has no lruvec to update.
Signed-off-by: Kairui Song <kasong@tencent.com>
---
include/linux/mm_inline.h | 10 +++++-----
include/linux/mmzone.h | 3 ++-
mm/vmscan.c | 47 ++++++++++++++++++++++++++++-------------------
3 files changed, 35 insertions(+), 25 deletions(-)
diff --git a/include/linux/mm_inline.h b/include/linux/mm_inline.h
index ee7fcbe2366b..b68d68101248 100644
--- a/include/linux/mm_inline.h
+++ b/include/linux/mm_inline.h
@@ -293,10 +293,9 @@ static inline int folio_lru_gen(const struct folio *folio)
return lru_get_gen_flags(READ_ONCE(*const_folio_flags(folio, 0)));
}
-static inline void lru_gen_update_size(struct lruvec *lruvec, struct folio *folio,
- int old_gen, int new_gen)
+static inline void lru_gen_update_size(struct lruvec *lruvec, int type,
+ struct folio *folio, int old_gen, int new_gen)
{
- int type = folio_is_file_lru(folio);
int zone = folio_zonenum(folio);
int delta = folio_nr_pages(folio);
struct lru_gen_folio *lrugen = &lruvec->lrugen;
@@ -371,7 +370,7 @@ static inline bool lru_gen_add_folio(struct lruvec *lruvec, struct folio *folio,
/* use the refs from the atomic snapshot to avoid raced update */
refs = lru_get_refs_flags(flags);
- lru_gen_update_size(lruvec, folio, -1, gen);
+ lru_gen_update_size(lruvec, type, folio, -1, gen);
if (lru_refs_is_active(refs))
lru += LRU_ACTIVE;
__update_lru_size(lruvec, lru, zone, delta);
@@ -389,6 +388,7 @@ static inline bool lru_gen_del_folio(struct lruvec *lruvec, struct folio *folio,
{
unsigned long flags;
int gen, refs;
+ int type = folio_is_file_lru(folio);
int zone = folio_zonenum(folio);
int delta = folio_nr_pages(folio);
enum lru_list lru = folio_is_file_lru(folio) * LRU_INACTIVE_FILE;
@@ -411,7 +411,7 @@ static inline bool lru_gen_del_folio(struct lruvec *lruvec, struct folio *folio,
if (!reclaiming && ((max_seq - gen) % MAX_NR_GENS) < MIN_NR_GENS)
folio_set_active(folio);
- lru_gen_update_size(lruvec, folio, gen, -1);
+ lru_gen_update_size(lruvec, type, folio, gen, -1);
if (lru_refs_is_active(refs))
lru += LRU_ACTIVE;
__update_lru_size(lruvec, lru, zone, -delta);
diff --git a/include/linux/mmzone.h b/include/linux/mmzone.h
index e096fd0a5047..d78e3f97423e 100644
--- a/include/linux/mmzone.h
+++ b/include/linux/mmzone.h
@@ -558,9 +558,10 @@ enum lruvec_flags {
#define LRU_TIER_MIN 0U
#define LRU_TIER_MAX (MAX_NR_TIERS - 1)
-/* Access source flags for folio_inc_lru_refs() */
+/* Flags for folio_inc_lru_refs() */
#define LRU_REF_MAPPED 0x1U
#define LRU_REF_EXEC 0x2U
+#define LRU_REF_FORCE 0x4U
#define LRU_REFS_REFERENCED 0x1
#define LRU_REFS_WORKINGSET 0x2
diff --git a/mm/vmscan.c b/mm/vmscan.c
index ae4a0522a8d5..9cc06a95d931 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -982,38 +982,33 @@ void folio_inc_lru_refs(struct folio *folio, unsigned int flags)
long nr_pages = folio_nr_pages(folio);
struct lru_gen_folio *lrugen;
struct lruvec *lruvec = NULL;
+ bool aged = false;
+retry:
old_flags = READ_ONCE(*folio_flags(folio, 0));
do {
new_flags = old_flags;
old_gen = lru_get_gen_flags(old_flags);
old_refs = lru_get_refs_flags(old_flags);
- file = folio_flags_is_file_lru(&old_flags);
refs = old_refs + 1;
gen = old_gen;
+ file = folio_flags_is_file_lru(&old_flags);
if (old_gen < 0)
goto out;
- /*
- * Lock the lruvec if the folio is on-list. We are already
- * doing lazy promotion so in theory we don't need this,
- * but for now, concurrent aging would still corrupt the
- * size counters. This is a temporary limitation and
- * will be lifted very soon, so the lock here is not a
- * performance concern.
- */
if (!lruvec) {
- lruvec = lruvec_live_lock_irq(folio_lruvec(folio));
+ lruvec = folio_lruvec_live_get(folio);
lrugen = &lruvec->lrugen;
}
+ /* Failed cmpxchg() does not order old_flags before max_seq. */
+ smp_rmb();
max_seq = READ_ONCE(lrugen->max_seq);
max_gen = lru_gen_from_seq(max_seq);
min_gen = lru_gen_from_seq(READ_ONCE(lrugen->min_seq[file]));
if (old_gen == max_gen)
goto out;
-
- if (flags & (LRU_REF_MAPPED | LRU_REF_EXEC)) {
- /* Promote second page table access or executable */
- if (refs > LRU_REFS_REFERENCED || flags & LRU_REF_EXEC)
+ if (flags & (LRU_REF_MAPPED | LRU_REF_EXEC | LRU_REF_FORCE)) {
+ if (refs > LRU_REFS_REFERENCED ||
+ flags & (LRU_REF_EXEC | LRU_REF_FORCE))
gen = max_gen;
/* First access only defers eviction from the oldest gen */
else if (old_gen == min_gen)
@@ -1037,9 +1032,19 @@ void folio_inc_lru_refs(struct folio *folio, unsigned int flags)
break;
} while (!try_cmpxchg(folio_flags(folio, 0), &old_flags, new_flags));
- if (gen != old_gen)
- lru_gen_update_size(lruvec, folio, old_gen, gen);
- if (lru_refs_is_active(old_refs) != lru_refs_is_active(refs) && old_gen >= 0) {
+ if (gen != old_gen) {
+ /*
+ * Gen-index reuse can fool cmpxchg(). Retry locklessly, accepting
+ * an extra reference to avoid folio_activate() latency.
+ */
+ if (unlikely(READ_ONCE(lrugen->max_seq) != max_seq)) {
+ flags = LRU_REF_FORCE;
+ aged = true;
+ }
+ lru_gen_update_size(lruvec, file, folio, old_gen, gen);
+ }
+
+ if (lru_refs_is_active(old_refs) != lru_refs_is_active(refs) && gen >= 0) {
enum lru_list lru = file * LRU_INACTIVE_FILE;
__update_lru_size(lruvec, lru + lru_refs_is_active(old_refs),
@@ -1047,8 +1052,12 @@ void folio_inc_lru_refs(struct folio *folio, unsigned int flags)
__update_lru_size(lruvec, lru + lru_refs_is_active(refs),
folio_zonenum(folio), nr_pages);
}
+ if (aged) {
+ aged = false;
+ goto retry;
+ }
if (lruvec)
- lruvec_unlock_irq(lruvec);
+ folio_lruvec_live_put(lruvec);
}
/*
@@ -3648,7 +3657,7 @@ static int folio_inc_gen(struct lruvec *lruvec, struct folio *folio)
new_gen = __folio_inc_gen(lruvec, folio, old_gen, &gen_increased);
if (gen_increased)
- lru_gen_update_size(lruvec, folio, old_gen, new_gen);
+ lru_gen_update_size(lruvec, type, folio, old_gen, new_gen);
return new_gen;
}
--
2.55.0
^ permalink raw reply [flat|nested] 21+ messages in thread
* [PATCH RFC v3 14/17] mm/mglru: reparent folios from all generations
2026-10-03 12:55 [PATCH RFC v3 00/17] mm/mglru: frequency guided promotion (MGLRU-FG) and flag cleanup Kairui Song via B4 Relay
` (12 preceding siblings ...)
2026-10-03 12:55 ` [PATCH RFC v3 13/17] mm/mglru: make folio_inc_lru_refs lruvec lockless Kairui Song via B4 Relay
@ 2026-10-03 12:55 ` Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 15/17] mm/khugepaged: check folio referenced state via LRU refs under MGLRU Kairui Song via B4 Relay
` (2 subsequent siblings)
16 siblings, 0 replies; 21+ messages in thread
From: Kairui Song via B4 Relay @ 2026-10-03 12:55 UTC (permalink / raw)
To: linux-mm
Cc: Andrew Morton, Johannes Weiner, Muchun Song, Qi Zheng,
Ying Huang, Chris Li, Baoquan He, Nico Pache, Usama Arif,
Michal Hocko, Roman Gushchin, Shakeel Butt, David Hildenbrand,
Lorenzo Stoakes, Barry Song, Axel Rasmussen, Yuanchu Xie, Wei Xu,
Vlastimil Babka, Suren Baghdasaryan, Kemeng Shi, Nhat Pham,
Youngjun Park, Zi Yan, Gregory Price, Matthew Wilcox (Oracle),
Baolin Wang, Ryan Roberts, Dev Jain, Lance Yang, Hugh Dickins,
SeongJae Park, David Rientjes, Yu Zhao, Vernon Yang,
Zicheng Wang, Chen Ridong, Tal Zussman, Kairui Song, Xiang Liu,
Ehab Ababneh, zhaozhengzhuo, Lian Wang, Kunwu Chan, linux-kernel,
cgroups, Kairui Song
From: Kairui Song <kasong@tencent.com>
A lockless promotion racing with aging can reuse a generation index,
see folio_inc_lru_refs(). The retry repairs the folio flags, but
sort_folio() may have already moved the folio to the list of a
generation outside the [min_seq, max_seq] window.
__lru_gen_reparent_memcg() only splices the lists inside the child's
window, so such a folio would be left on a list of the dying memcg.
Splice the lists and move the counters of all MAX_NR_GENS generations
instead. Those outside the window are empty in the common case.
Assisted-by: LLM
Signed-off-by: Kairui Song <kasong@tencent.com>
---
mm/vmscan.c | 3 ++-
1 file changed, 2 insertions(+), 1 deletion(-)
diff --git a/mm/vmscan.c b/mm/vmscan.c
index 9cc06a95d931..1bc8e5056265 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -4947,7 +4947,8 @@ static void __lru_gen_reparent_memcg(struct lruvec *child_lruvec, struct lruvec
child_lrugen = &child_lruvec->lrugen;
parent_lrugen = &parent_lruvec->lrugen;
- for (i = 0; i < get_nr_gens(child_lruvec, type); i++) {
+ /* Lockless promotion may leave folios outside the child's window */
+ for (i = 0; i < MAX_NR_GENS; i++) {
int gen = lru_gen_from_seq(child_lrugen->max_seq - i);
long nr_pages = atomic_long_read(&child_lrugen->nr_pages[gen][type][zone]);
--
2.55.0
^ permalink raw reply [flat|nested] 21+ messages in thread
* [PATCH RFC v3 15/17] mm/khugepaged: check folio referenced state via LRU refs under MGLRU
2026-10-03 12:55 [PATCH RFC v3 00/17] mm/mglru: frequency guided promotion (MGLRU-FG) and flag cleanup Kairui Song via B4 Relay
` (13 preceding siblings ...)
2026-10-03 12:55 ` [PATCH RFC v3 14/17] mm/mglru: reparent folios from all generations Kairui Song via B4 Relay
@ 2026-10-03 12:55 ` Kairui Song via B4 Relay
2026-10-04 2:11 ` Zi Yan
2026-10-03 12:55 ` [PATCH RFC v3 16/17] mm/mglru: make folio_test_workingset() work based on folio LRU refs Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 17/17] Documentation/mm: multi-gen LRU: update for frequency guided promotion Kairui Song via B4 Relay
16 siblings, 1 reply; 21+ messages in thread
From: Kairui Song via B4 Relay @ 2026-10-03 12:55 UTC (permalink / raw)
To: linux-mm
Cc: Andrew Morton, Johannes Weiner, Muchun Song, Qi Zheng,
Ying Huang, Chris Li, Baoquan He, Nico Pache, Usama Arif,
Michal Hocko, Roman Gushchin, Shakeel Butt, David Hildenbrand,
Lorenzo Stoakes, Barry Song, Axel Rasmussen, Yuanchu Xie, Wei Xu,
Vlastimil Babka, Suren Baghdasaryan, Kemeng Shi, Nhat Pham,
Youngjun Park, Zi Yan, Gregory Price, Matthew Wilcox (Oracle),
Baolin Wang, Ryan Roberts, Dev Jain, Lance Yang, Hugh Dickins,
SeongJae Park, David Rientjes, Yu Zhao, Vernon Yang,
Zicheng Wang, Chen Ridong, Tal Zussman, Kairui Song, Xiang Liu,
Ehab Ababneh, zhaozhengzhuo, Lian Wang, Kunwu Chan, linux-kernel,
cgroups, Kairui Song
From: Kairui Song <kasong@tencent.com>
For MGLRU, PG_referenced is only the lowest bit of the folio LRU refs
count, so folios with a higher refs count no longer have the bit set.
folio_pte_referenced() testing the raw bit misses hot folios at
refs >= 2, under-counting the referenced folios of a candidate range
and aborting otherwise good collapses with SCAN_LACK_REFERENCED_PAGE.
Test the refs count directly when MGLRU is enabled, mirroring the
smaps conversion. The classical LRU keeps the plain bit test.
Signed-off-by: Kairui Song <kasong@tencent.com>
---
mm/khugepaged.c | 4 +++-
1 file changed, 3 insertions(+), 1 deletion(-)
diff --git a/mm/khugepaged.c b/mm/khugepaged.c
index 913086eaf17b..be18381c6efd 100644
--- a/mm/khugepaged.c
+++ b/mm/khugepaged.c
@@ -586,7 +586,9 @@ static bool folio_pte_referenced(struct folio *folio,
struct vm_area_struct *vma, unsigned long addr, pte_t pteval)
{
/* The folio was referenced previously ... */
- if (folio_test_young(folio) || folio_test_referenced(folio))
+ if (folio_test_young(folio))
+ return true;
+ if (lru_gen_enabled() ? folio_lru_refs(folio) : folio_test_referenced(folio))
return true;
/* ... or the PTE mapping was recently used */
return pte_young(pteval) || mmu_notifier_test_young(vma->vm_mm, addr);
--
2.55.0
^ permalink raw reply [flat|nested] 21+ messages in thread
* [PATCH RFC v3 16/17] mm/mglru: make folio_test_workingset() work based on folio LRU refs
2026-10-03 12:55 [PATCH RFC v3 00/17] mm/mglru: frequency guided promotion (MGLRU-FG) and flag cleanup Kairui Song via B4 Relay
` (14 preceding siblings ...)
2026-10-03 12:55 ` [PATCH RFC v3 15/17] mm/khugepaged: check folio referenced state via LRU refs under MGLRU Kairui Song via B4 Relay
@ 2026-10-03 12:55 ` Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 17/17] Documentation/mm: multi-gen LRU: update for frequency guided promotion Kairui Song via B4 Relay
16 siblings, 0 replies; 21+ messages in thread
From: Kairui Song via B4 Relay @ 2026-10-03 12:55 UTC (permalink / raw)
To: linux-mm
Cc: Andrew Morton, Johannes Weiner, Muchun Song, Qi Zheng,
Ying Huang, Chris Li, Baoquan He, Nico Pache, Usama Arif,
Michal Hocko, Roman Gushchin, Shakeel Butt, David Hildenbrand,
Lorenzo Stoakes, Barry Song, Axel Rasmussen, Yuanchu Xie, Wei Xu,
Vlastimil Babka, Suren Baghdasaryan, Kemeng Shi, Nhat Pham,
Youngjun Park, Zi Yan, Gregory Price, Matthew Wilcox (Oracle),
Baolin Wang, Ryan Roberts, Dev Jain, Lance Yang, Hugh Dickins,
SeongJae Park, David Rientjes, Yu Zhao, Vernon Yang,
Zicheng Wang, Chen Ridong, Tal Zussman, Kairui Song, Xiang Liu,
Ehab Ababneh, zhaozhengzhuo, Lian Wang, Kunwu Chan, linux-kernel,
cgroups, Kairui Song
From: Kairui Song <kasong@tencent.com>
folio_test_workingset() currently tests the PG_workingset bit, which
is only the second lowest access bit of the folio LRU refs count now.
Folios with a higher refs count carry it in the LRU_REFS_MASK field
and no longer have PG_workingset set, so the bit test misses them.
folio_set_workingset() currently sets the PG_workingset bit with a
plain set_bit(), which also over-promotes folios at a higher refs
count: e.g. a folio at refs 4 is bumped to refs 6, advancing it
toward LRU_REFS_MAX and an unintended generation promotion on the
next access.
Move the test into mm_inline.h as an inline helper based on
folio_lru_refs(), checking refs >= LRU_REFS_WORKINGSET. The set
stays a plain PG_workingset bit operation reserved for the
classical LRU: under MGLRU the refs count is maintained by
folio_inc_lru_refs(), and now triggers a debug WARN when called
while MGLRU is fully on (the switching window is exempt, as the
classical paths legitimately run alongside MGLRU then). The one
PageWorkingset() user in erofs is converted to the folio helper.
Under the classical LRU the refs count is not maintained, so the
test falls back to the PG_workingset bit test, which is the old
behavior.
Signed-off-by: Kairui Song <kasong@tencent.com>
---
fs/btrfs/compression.c | 1 +
fs/erofs/zdata.c | 3 ++-
include/linux/mm_inline.h | 35 +++++++++++++++++++++++++++++++++++
include/linux/page-flags.h | 2 --
mm/filemap.c | 1 +
mm/page_io.c | 1 +
6 files changed, 40 insertions(+), 3 deletions(-)
diff --git a/fs/btrfs/compression.c b/fs/btrfs/compression.c
index c62b5148d5ac..57d24413265e 100644
--- a/fs/btrfs/compression.c
+++ b/fs/btrfs/compression.c
@@ -8,6 +8,7 @@
#include <linux/file.h>
#include <linux/fs.h>
#include <linux/pagemap.h>
+#include <linux/mm_inline.h>
#include <linux/folio_batch.h>
#include <linux/highmem.h>
#include <linux/kthread.h>
diff --git a/fs/erofs/zdata.c b/fs/erofs/zdata.c
index e981e371d6c2..34ee41e1ccdc 100644
--- a/fs/erofs/zdata.c
+++ b/fs/erofs/zdata.c
@@ -6,6 +6,7 @@
*/
#include "compress.h"
#include <linux/psi.h>
+#include <linux/mm_inline.h>
#include <linux/cpuhotplug.h>
#include <trace/events/erofs.h>
@@ -1713,7 +1714,7 @@ static void z_erofs_submit_queue(struct z_erofs_frontend *f,
DBG_BUGON(bvec.bv_len < sb->s_blocksize);
}
- if (unlikely(PageWorkingset(bvec.bv_page)) &&
+ if (unlikely(folio_test_workingset(page_folio(bvec.bv_page))) &&
!memstall) {
psi_memstall_enter(&pflags);
memstall = 1;
diff --git a/include/linux/mm_inline.h b/include/linux/mm_inline.h
index b68d68101248..e651f60cb914 100644
--- a/include/linux/mm_inline.h
+++ b/include/linux/mm_inline.h
@@ -258,6 +258,24 @@ static inline bool lru_gen_enabled(void)
}
#endif
+/**
+ * folio_test_workingset - Test if a folio is in the workingset.
+ * @folio: the folio
+ *
+ * A folio is workingset when its LRU refs count reaches
+ * LRU_REFS_WORKINGSET. Under the classical LRU the refs count never
+ * goes above it, so this is just testing the PG_workingset bit.
+ * NOTE: folio_set_workingset() must not be used under MGLRU, as the
+ * folio refs are tracked by folio_inc_lru_refs(), it triggers a debug
+ * WARN instead for MGLRU.
+ *
+ * Return: true if the folio is workingset.
+ */
+static __always_inline bool folio_test_workingset(const struct folio *folio)
+{
+ return folio_lru_refs(folio) >= LRU_REFS_WORKINGSET;
+}
+
static inline bool lru_gen_in_fault(void)
{
return current->in_lru_fault;
@@ -445,6 +463,11 @@ static inline bool lru_gen_del_folio(struct lruvec *lruvec, struct folio *folio,
{
return false;
}
+
+static inline bool folio_test_workingset(const struct folio *folio)
+{
+ return test_bit(PG_workingset, const_folio_flags(folio, FOLIO_HEAD_PAGE));
+}
#endif /* CONFIG_LRU_GEN */
/**
@@ -473,6 +496,18 @@ static __always_inline void folio_inc_lru_refs_fast(struct folio *folio)
} while (!try_cmpxchg(folio_flags(folio, 0), &old_flags, new_flags));
}
+/*
+ * For the classical LRU only: under MGLRU the PG_workingset bit is
+ * part of the folio refs count maintained by folio_inc_lru_refs(),
+ * and a raw set would corrupt it. The switching window is exempt
+ * because the classical paths legitimately run alongside MGLRU then.
+ */
+static __always_inline void folio_set_workingset(struct folio *folio)
+{
+ VM_WARN_ON_ONCE(lru_gen_enabled() && !lru_gen_switching());
+ set_bit(PG_workingset, folio_flags(folio, FOLIO_HEAD_PAGE));
+}
+
static __always_inline
void lruvec_add_folio(struct lruvec *lruvec, struct folio *folio)
{
diff --git a/include/linux/page-flags.h b/include/linux/page-flags.h
index b0ddc652e76c..f6625dc042da 100644
--- a/include/linux/page-flags.h
+++ b/include/linux/page-flags.h
@@ -549,8 +549,6 @@ PAGEFLAG(LRU, lru, PF_HEAD) __CLEARPAGEFLAG(LRU, lru, PF_HEAD)
FOLIO_FLAG(active, FOLIO_HEAD_PAGE)
__FOLIO_CLEAR_FLAG(active, FOLIO_HEAD_PAGE)
FOLIO_TEST_CLEAR_FLAG(active, FOLIO_HEAD_PAGE)
-PAGEFLAG(Workingset, workingset, PF_HEAD)
- TESTCLEARFLAG(Workingset, workingset, PF_HEAD)
PAGEFLAG(Checked, checked, PF_NO_COMPOUND) /* Used by some filesystems */
/* Xen */
diff --git a/mm/filemap.c b/mm/filemap.c
index b74bc1e5015c..49e827058e14 100644
--- a/mm/filemap.c
+++ b/mm/filemap.c
@@ -21,6 +21,7 @@
#include <linux/gfp.h>
#include <linux/mm.h>
#include <linux/swap.h>
+#include <linux/mm_inline.h>
#include <linux/leafops.h>
#include <linux/syscalls.h>
#include <linux/mman.h>
diff --git a/mm/page_io.c b/mm/page_io.c
index c6824fcd483e..540ae8caa932 100644
--- a/mm/page_io.c
+++ b/mm/page_io.c
@@ -16,6 +16,7 @@
#include <linux/gfp.h>
#include <linux/pagemap.h>
#include <linux/swap.h>
+#include <linux/mm_inline.h>
#include <linux/bio.h>
#include <linux/swapops.h>
#include <linux/writeback.h>
--
2.55.0
^ permalink raw reply [flat|nested] 21+ messages in thread
* [PATCH RFC v3 17/17] Documentation/mm: multi-gen LRU: update for frequency guided promotion
2026-10-03 12:55 [PATCH RFC v3 00/17] mm/mglru: frequency guided promotion (MGLRU-FG) and flag cleanup Kairui Song via B4 Relay
` (15 preceding siblings ...)
2026-10-03 12:55 ` [PATCH RFC v3 16/17] mm/mglru: make folio_test_workingset() work based on folio LRU refs Kairui Song via B4 Relay
@ 2026-10-03 12:55 ` Kairui Song via B4 Relay
16 siblings, 0 replies; 21+ messages in thread
From: Kairui Song via B4 Relay @ 2026-10-03 12:55 UTC (permalink / raw)
To: linux-mm
Cc: Andrew Morton, Johannes Weiner, Muchun Song, Qi Zheng,
Ying Huang, Chris Li, Baoquan He, Nico Pache, Usama Arif,
Michal Hocko, Roman Gushchin, Shakeel Butt, David Hildenbrand,
Lorenzo Stoakes, Barry Song, Axel Rasmussen, Yuanchu Xie, Wei Xu,
Vlastimil Babka, Suren Baghdasaryan, Kemeng Shi, Nhat Pham,
Youngjun Park, Zi Yan, Gregory Price, Matthew Wilcox (Oracle),
Baolin Wang, Ryan Roberts, Dev Jain, Lance Yang, Hugh Dickins,
SeongJae Park, David Rientjes, Yu Zhao, Vernon Yang,
Zicheng Wang, Chen Ridong, Tal Zussman, Kairui Song, Xiang Liu,
Ehab Ababneh, zhaozhengzhuo, Lian Wang, Kunwu Chan, linux-kernel,
cgroups, Kairui Song
From: Kairui Song <kasong@tencent.com>
The design doc still describes the old tier and promotion model:
- Tiers are no longer order_base_2() of the file descriptor accesses.
A folio's tier is derived from its referenced count, which counts
accesses from both page tables and file descriptors.
- A young PTE no longer always promotes a folio to the newest
generation. The first access only moves a folio out of the oldest
generation, a second access promotes it, and executable file folios
are promoted on their first access.
- The eviction no longer uses the first tier as the only baseline.
The type is selected by comparing all tiers of each type, and a tier
is protected by comparing it against the tiers below it.
Update the aging and eviction sections accordingly.
Assisted-by: LLM
Signed-off-by: Kairui Song <kasong@tencent.com>
---
Documentation/mm/multigen_lru.rst | 56 ++++++++++++++++++++++-----------------
1 file changed, 32 insertions(+), 24 deletions(-)
diff --git a/Documentation/mm/multigen_lru.rst b/Documentation/mm/multigen_lru.rst
index 52ed5092022f..e3dab858887d 100644
--- a/Documentation/mm/multigen_lru.rst
+++ b/Documentation/mm/multigen_lru.rst
@@ -95,16 +95,19 @@ at most ``MAX_NR_GENS`` generations. The gen counter stores a value
within ``[1, MAX_NR_GENS]`` while a page is on one of
``lrugen->folios[]``; otherwise it stores zero.
-Each generation is divided into multiple tiers. A page accessed ``N``
-times through file descriptors is in tier ``order_base_2(N)``. Unlike
-generations, tiers do not have dedicated ``lrugen->folios[]``. In
-contrast to moving across generations, which requires the LRU lock,
-moving across tiers only involves atomic operations on
-``folio->flags`` and therefore has a negligible cost. A feedback loop
-modeled after the PID controller monitors refaults over all the tiers
-from anon and file types and decides which tiers from which types to
-evict or protect. The desired effect is to balance refault percentages
-between anon and file types proportional to the swappiness level.
+Each generation is divided into multiple tiers. A page's tier is
+derived from its referenced count, which counts the accesses collected
+from both channels above and is capped at ``LRU_REFS_MAX``: pages
+accessed at most once are in the lowest tier, and the tiers above hold
+pages accessed at least twice. Unlike generations, tiers do not have
+dedicated ``lrugen->folios[]``. In contrast to moving across
+generations, which requires the LRU lock, updating the referenced
+count only involves atomic operations on ``folio->flags`` and
+therefore has a negligible cost. A feedback loop modeled after the PID
+controller monitors refaults over all the tiers from anon and file
+types and decides which tiers from which types to evict or protect.
+The desired effect is to balance refault percentages between anon and
+file types proportional to the swappiness level.
There are two conceptually independent procedures: the aging and the
eviction. They form a closed-loop system, i.e., the page reclaim.
@@ -122,25 +125,30 @@ and calls ``walk_page_range()`` with each ``mm_struct`` on this list
to scan PTEs, and after each iteration, it increments ``max_seq``. For
the latter, when the eviction walks the rmap and finds a young PTE,
the aging scans the adjacent PTEs. For both, on finding a young PTE,
-the aging clears the accessed bit and updates the gen counter of the
-page mapped by this PTE to ``(max_seq%MAX_NR_GENS)+1``.
+the aging clears the accessed bit and bumps the referenced count of
+the page mapped by this PTE. A second access promotes the page by
+updating its gen counter to ``(max_seq%MAX_NR_GENS)+1``; the first
+access only moves a page out of the oldest generation, so that the
+next time it becomes the oldest one, its accessed bit reflects
+whether it has been used again. Executable file pages are promoted on
+their first access, since reclaiming them causes typical I/O
+thrashing.
Eviction
--------
The eviction consumes old generations. Given an ``lruvec``, it
increments ``min_seq`` when ``lrugen->folios[]`` indexed by
-``min_seq%MAX_NR_GENS`` becomes empty. To select a type and a tier to
-evict from, it first compares ``min_seq[]`` to select the older type.
-If both types are equally old, it selects the one whose first tier has
-a lower refault percentage. The first tier contains single-use
-unmapped clean pages, which are the best bet. The eviction sorts a
-page according to its gen counter if the aging has found this page
-accessed through page tables and updated its gen counter. It also
-moves a page to the next generation, i.e., ``min_seq+1``, if this page
-was accessed multiple times through file descriptors and the feedback
-loop has detected outlying refaults from the tier this page is in. To
-this end, the feedback loop uses the first tier as the baseline, for
-the reason stated earlier.
+``min_seq%MAX_NR_GENS`` becomes empty. To select a type to evict from,
+it first compares ``min_seq[]`` to select the older type. If both
+types are equally old, it compares the refault percentages of all the
+tiers of each type proportional to the swappiness level. To select a
+tier to protect, it compares the refault percentage of each tier
+against that of the tiers below it. The lowest tier contains pages
+accessed at most once, which are the best bet. The eviction sorts a
+page according to its gen counter if the aging or an access has
+updated its gen counter. It also moves a page to the next generation,
+i.e., ``min_seq+1``, if the feedback loop has detected outlying
+refaults from the tier this page is in.
Working set protection
----------------------
--
2.55.0
^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [PATCH RFC v3 15/17] mm/khugepaged: check folio referenced state via LRU refs under MGLRU
2026-10-03 12:55 ` [PATCH RFC v3 15/17] mm/khugepaged: check folio referenced state via LRU refs under MGLRU Kairui Song via B4 Relay
@ 2026-10-04 2:11 ` Zi Yan
2026-10-04 3:30 ` Kairui Song
0 siblings, 1 reply; 21+ messages in thread
From: Zi Yan @ 2026-10-04 2:11 UTC (permalink / raw)
To: kasong, linux-mm
Cc: Andrew Morton, Johannes Weiner, Muchun Song, Qi Zheng,
Ying Huang, Chris Li, Baoquan He, Nico Pache, Usama Arif,
Michal Hocko, Roman Gushchin, Shakeel Butt, David Hildenbrand,
Lorenzo Stoakes, Barry Song, Axel Rasmussen, Yuanchu Xie, Wei Xu,
Vlastimil Babka, Suren Baghdasaryan, Kemeng Shi, Nhat Pham,
Youngjun Park, Gregory Price, Matthew Wilcox (Oracle),
Baolin Wang, Ryan Roberts, Dev Jain, Lance Yang, Hugh Dickins,
SeongJae Park, David Rientjes, Yu Zhao, Vernon Yang,
Zicheng Wang, Chen Ridong, Tal Zussman, Kairui Song, Xiang Liu,
Ehab Ababneh, zhaozhengzhuo, Lian Wang, Kunwu Chan, linux-kernel,
cgroups
On Sat Oct 3, 2026 at 8:55 AM EDT, Kairui Song via B4 Relay wrote:
> From: Kairui Song <kasong@tencent.com>
>
> For MGLRU, PG_referenced is only the lowest bit of the folio LRU refs
> count, so folios with a higher refs count no longer have the bit set.
> folio_pte_referenced() testing the raw bit misses hot folios at
> refs >= 2, under-counting the referenced folios of a candidate range
> and aborting otherwise good collapses with SCAN_LACK_REFERENCED_PAGE.
>
> Test the refs count directly when MGLRU is enabled, mirroring the
> smaps conversion. The classical LRU keeps the plain bit test.
Why not reuse the smaps helper function? You can rename the helper to
check_folio_referenced() for general use.
>
> Signed-off-by: Kairui Song <kasong@tencent.com>
> ---
> mm/khugepaged.c | 4 +++-
> 1 file changed, 3 insertions(+), 1 deletion(-)
>
> diff --git a/mm/khugepaged.c b/mm/khugepaged.c
> index 913086eaf17b..be18381c6efd 100644
> --- a/mm/khugepaged.c
> +++ b/mm/khugepaged.c
> @@ -586,7 +586,9 @@ static bool folio_pte_referenced(struct folio *folio,
> struct vm_area_struct *vma, unsigned long addr, pte_t pteval)
> {
> /* The folio was referenced previously ... */
> - if (folio_test_young(folio) || folio_test_referenced(folio))
> + if (folio_test_young(folio))
> + return true;
> + if (lru_gen_enabled() ? folio_lru_refs(folio) : folio_test_referenced(folio))
> return true;
> /* ... or the PTE mapping was recently used */
> return pte_young(pteval) || mmu_notifier_test_young(vma->vm_mm, addr);
--
Best Regards,
Yan, Zi
^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [PATCH RFC v3 15/17] mm/khugepaged: check folio referenced state via LRU refs under MGLRU
2026-10-04 2:11 ` Zi Yan
@ 2026-10-04 3:30 ` Kairui Song
0 siblings, 0 replies; 21+ messages in thread
From: Kairui Song @ 2026-10-04 3:30 UTC (permalink / raw)
To: Zi Yan
Cc: linux-mm, Andrew Morton, Johannes Weiner, Muchun Song, Qi Zheng,
Ying Huang, Chris Li, Baoquan He, Nico Pache, Usama Arif,
Michal Hocko, Roman Gushchin, Shakeel Butt, David Hildenbrand,
Lorenzo Stoakes, Barry Song, Axel Rasmussen, Yuanchu Xie, Wei Xu,
Vlastimil Babka, Suren Baghdasaryan, Kemeng Shi, Nhat Pham,
Youngjun Park, Gregory Price, Matthew Wilcox (Oracle),
Baolin Wang, Ryan Roberts, Dev Jain, Lance Yang, Hugh Dickins,
SeongJae Park, David Rientjes, Yu Zhao, Vernon Yang,
Zicheng Wang, Chen Ridong, Tal Zussman, Xiang Liu, Ehab Ababneh,
zhaozhengzhuo, Lian Wang, Kunwu Chan, linux-kernel, cgroups
On Sun, Oct 4, 2026 at 6:11 AM Zi Yan <ziy@nvidia.com> wrote:
>
> On Sat Oct 3, 2026 at 8:55 AM EDT, Kairui Song via B4 Relay wrote:
> > From: Kairui Song <kasong@tencent.com>
> >
> > For MGLRU, PG_referenced is only the lowest bit of the folio LRU refs
> > count, so folios with a higher refs count no longer have the bit set.
> > folio_pte_referenced() testing the raw bit misses hot folios at
> > refs >= 2, under-counting the referenced folios of a candidate range
> > and aborting otherwise good collapses with SCAN_LACK_REFERENCED_PAGE.
> >
> > Test the refs count directly when MGLRU is enabled, mirroring the
> > smaps conversion. The classical LRU keeps the plain bit test.
>
> Why not reuse the smaps helper function? You can rename the helper to
> check_folio_referenced() for general use.
>
That's a very nice idea, will use that.
> --
> Best Regards,
> Yan, Zi
>
^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [PATCH RFC v3 13/17] mm/mglru: make folio_inc_lru_refs lruvec lockless
2026-10-03 12:55 ` [PATCH RFC v3 13/17] mm/mglru: make folio_inc_lru_refs lruvec lockless Kairui Song via B4 Relay
@ 2026-10-04 17:29 ` KunWu Chan
0 siblings, 0 replies; 21+ messages in thread
From: KunWu Chan @ 2026-10-04 17:29 UTC (permalink / raw)
To: kasong
Cc: linux-mm, Andrew Morton, Johannes Weiner, Muchun Song, Qi Zheng,
Ying Huang, Chris Li, Baoquan He, Nico Pache, Usama Arif,
Michal Hocko, Roman Gushchin, Shakeel Butt, David Hildenbrand,
Lorenzo Stoakes, Barry Song, Axel Rasmussen, Yuanchu Xie, Wei Xu,
Vlastimil Babka, Suren Baghdasaryan, Kemeng Shi, Nhat Pham,
Youngjun Park, Zi Yan, Gregory Price, Matthew Wilcox (Oracle),
Baolin Wang, Ryan Roberts, Dev Jain, Lance Yang, Hugh Dickins,
SeongJae Park, David Rientjes, Yu Zhao, Vernon Yang,
Zicheng Wang, Chen Ridong, Tal Zussman, Kairui Song, Xiang Liu,
Ehab Ababneh, zhaozhengzhuo, Lian Wang, linux-kernel, cgroups
On Sat, Oct 3, 2026 at 8:55 PM Kairui Song via B4 Relay
<devnull+kasong.tencent.com@kernel.org> wrote:
>
> From: Kairui Song <kasong@tencent.com>
>
> The lruvec spinlock in folio_inc_lru_refs was only taken to keep
> concurrent aging from corrupting the size counters. This is no
> longer needed: the per-generation counters are atomic, and the
> active/inactive ABI counters are updated per folio and linearized by
> the CAS on the folio flags, so concurrent aging can no longer corrupt
> them. The same argument already allowed folio_reset_lru_refs() to
> drop the lock.
>
> Use folio_lruvec_live_get()/folio_lruvec_live_put() to hold the RCU
> read lock around the lruvec lookup, and drop the spinlock entirely.
>
> Without the lock, aging can advance max_seq while a promotion is in
> flight. Generations are indexed by seq % MAX_NR_GENS, so the folio
> flags can return to the value read earlier and the cmpxchg succeeds
> with a target gen computed from a stale max_seq. Recheck max_seq after
> a gen-changing cmpxchg: if it moved, account the committed update and
> retry with LRU_REF_FORCE, which promotes the folio to the newest gen.
> The retry may count the access twice, which is acceptable and avoids
> folio_activate(), whose LRU lock causes latency jitter. smp_rmb()
> orders the flags read, including the one returned by a failed
> cmpxchg(), before the max_seq read.
>
> Also simplify the post-CAS accounting guards: gen is only assigned
> non-negative values after the old_gen < 0 early exit, so gen != old_gen
> implies gen >= 0, and lru_gen_update_size() needs no explicit guard.
> The active/inactive ABI update keeps its gen >= 0 check because an
> off-LRU folio has no lruvec to update.
>
> Signed-off-by: Kairui Song <kasong@tencent.com>
> ---
> include/linux/mm_inline.h | 10 +++++-----
> include/linux/mmzone.h | 3 ++-
> mm/vmscan.c | 47 ++++++++++++++++++++++++++++-------------------
> 3 files changed, 35 insertions(+), 25 deletions(-)
>
> diff --git a/include/linux/mm_inline.h b/include/linux/mm_inline.h
> index ee7fcbe2366b..b68d68101248 100644
> --- a/include/linux/mm_inline.h
> +++ b/include/linux/mm_inline.h
> @@ -293,10 +293,9 @@ static inline int folio_lru_gen(const struct folio *folio)
> return lru_get_gen_flags(READ_ONCE(*const_folio_flags(folio, 0)));
> }
>
> -static inline void lru_gen_update_size(struct lruvec *lruvec, struct folio *folio,
> - int old_gen, int new_gen)
> +static inline void lru_gen_update_size(struct lruvec *lruvec, int type,
> + struct folio *folio, int old_gen, int new_gen)
> {
> - int type = folio_is_file_lru(folio);
> int zone = folio_zonenum(folio);
> int delta = folio_nr_pages(folio);
> struct lru_gen_folio *lrugen = &lruvec->lrugen;
> @@ -371,7 +370,7 @@ static inline bool lru_gen_add_folio(struct lruvec *lruvec, struct folio *folio,
> /* use the refs from the atomic snapshot to avoid raced update */
> refs = lru_get_refs_flags(flags);
>
> - lru_gen_update_size(lruvec, folio, -1, gen);
> + lru_gen_update_size(lruvec, type, folio, -1, gen);
> if (lru_refs_is_active(refs))
> lru += LRU_ACTIVE;
> __update_lru_size(lruvec, lru, zone, delta);
> @@ -389,6 +388,7 @@ static inline bool lru_gen_del_folio(struct lruvec *lruvec, struct folio *folio,
> {
> unsigned long flags;
> int gen, refs;
> + int type = folio_is_file_lru(folio);
> int zone = folio_zonenum(folio);
> int delta = folio_nr_pages(folio);
> enum lru_list lru = folio_is_file_lru(folio) * LRU_INACTIVE_FILE;
> @@ -411,7 +411,7 @@ static inline bool lru_gen_del_folio(struct lruvec *lruvec, struct folio *folio,
> if (!reclaiming && ((max_seq - gen) % MAX_NR_GENS) < MIN_NR_GENS)
> folio_set_active(folio);
>
> - lru_gen_update_size(lruvec, folio, gen, -1);
> + lru_gen_update_size(lruvec, type, folio, gen, -1);
> if (lru_refs_is_active(refs))
> lru += LRU_ACTIVE;
> __update_lru_size(lruvec, lru, zone, -delta);
> diff --git a/include/linux/mmzone.h b/include/linux/mmzone.h
> index e096fd0a5047..d78e3f97423e 100644
> --- a/include/linux/mmzone.h
> +++ b/include/linux/mmzone.h
> @@ -558,9 +558,10 @@ enum lruvec_flags {
> #define LRU_TIER_MIN 0U
> #define LRU_TIER_MAX (MAX_NR_TIERS - 1)
>
> -/* Access source flags for folio_inc_lru_refs() */
> +/* Flags for folio_inc_lru_refs() */
> #define LRU_REF_MAPPED 0x1U
> #define LRU_REF_EXEC 0x2U
> +#define LRU_REF_FORCE 0x4U
>
> #define LRU_REFS_REFERENCED 0x1
> #define LRU_REFS_WORKINGSET 0x2
> diff --git a/mm/vmscan.c b/mm/vmscan.c
> index ae4a0522a8d5..9cc06a95d931 100644
> --- a/mm/vmscan.c
> +++ b/mm/vmscan.c
> @@ -982,38 +982,33 @@ void folio_inc_lru_refs(struct folio *folio, unsigned int flags)
> long nr_pages = folio_nr_pages(folio);
> struct lru_gen_folio *lrugen;
> struct lruvec *lruvec = NULL;
> + bool aged = false;
>
> +retry:
> old_flags = READ_ONCE(*folio_flags(folio, 0));
> do {
> new_flags = old_flags;
> old_gen = lru_get_gen_flags(old_flags);
> old_refs = lru_get_refs_flags(old_flags);
> - file = folio_flags_is_file_lru(&old_flags);
> refs = old_refs + 1;
> gen = old_gen;
> + file = folio_flags_is_file_lru(&old_flags);
> if (old_gen < 0)
> goto out;
> - /*
> - * Lock the lruvec if the folio is on-list. We are already
> - * doing lazy promotion so in theory we don't need this,
> - * but for now, concurrent aging would still corrupt the
> - * size counters. This is a temporary limitation and
> - * will be lifted very soon, so the lock here is not a
> - * performance concern.
> - */
> if (!lruvec) {
> - lruvec = lruvec_live_lock_irq(folio_lruvec(folio));
> + lruvec = folio_lruvec_live_get(folio);
> lrugen = &lruvec->lrugen;
> }
> + /* Failed cmpxchg() does not order old_flags before max_seq. */
> + smp_rmb();
> max_seq = READ_ONCE(lrugen->max_seq);
> max_gen = lru_gen_from_seq(max_seq);
> min_gen = lru_gen_from_seq(READ_ONCE(lrugen->min_seq[file]));
> if (old_gen == max_gen)
> goto out;
> -
> - if (flags & (LRU_REF_MAPPED | LRU_REF_EXEC)) {
> - /* Promote second page table access or executable */
> - if (refs > LRU_REFS_REFERENCED || flags & LRU_REF_EXEC)
> + if (flags & (LRU_REF_MAPPED | LRU_REF_EXEC | LRU_REF_FORCE)) {
> + if (refs > LRU_REFS_REFERENCED ||
> + flags & (LRU_REF_EXEC | LRU_REF_FORCE))
> gen = max_gen;
> /* First access only defers eviction from the oldest gen */
> else if (old_gen == min_gen)
> @@ -1037,9 +1032,19 @@ void folio_inc_lru_refs(struct folio *folio, unsigned int flags)
> break;
> } while (!try_cmpxchg(folio_flags(folio, 0), &old_flags, new_flags));
>
> - if (gen != old_gen)
> - lru_gen_update_size(lruvec, folio, old_gen, gen);
> - if (lru_refs_is_active(old_refs) != lru_refs_is_active(refs) && old_gen >= 0) {
> + if (gen != old_gen) {
> + /*
> + * Gen-index reuse can fool cmpxchg(). Retry locklessly, accepting
> + * an extra reference to avoid folio_activate() latency.
> + */
> + if (unlikely(READ_ONCE(lrugen->max_seq) != max_seq)) {
> + flags = LRU_REF_FORCE;
> + aged = true;
> + }
> + lru_gen_update_size(lruvec, file, folio, old_gen, gen);
> + }
> +
> + if (lru_refs_is_active(old_refs) != lru_refs_is_active(refs) && gen >= 0) {
> enum lru_list lru = file * LRU_INACTIVE_FILE;
>
Hi Kairui,
One question about the retry path.
When `max_seq` changes after a gen-changing cmpxchg, the retry uses
`LRU_REF_FORCE` and promotes the folio directly to `max_gen`. In the
normal first `LRU_REF_MAPPED` access path, the same folio is only moved
to the next generation when it is in the oldest generation.
Is the stronger promotion on the retry intentional? In other words,
does detecting that `max_seq` advanced during the update justify
promoting the folio all the way to the newest generation, rather than
re-evaluating the original promotion policy with the new `max_seq`?
Thanks,
Kunwu
> __update_lru_size(lruvec, lru + lru_refs_is_active(old_refs),
> @@ -1047,8 +1052,12 @@ void folio_inc_lru_refs(struct folio *folio, unsigned int flags)
> __update_lru_size(lruvec, lru + lru_refs_is_active(refs),
> folio_zonenum(folio), nr_pages);
> }
> + if (aged) {
> + aged = false;
> + goto retry;
> + }
> if (lruvec)
> - lruvec_unlock_irq(lruvec);
> + folio_lruvec_live_put(lruvec);
> }
>
> /*
> @@ -3648,7 +3657,7 @@ static int folio_inc_gen(struct lruvec *lruvec, struct folio *folio)
>
> new_gen = __folio_inc_gen(lruvec, folio, old_gen, &gen_increased);
> if (gen_increased)
> - lru_gen_update_size(lruvec, folio, old_gen, new_gen);
> + lru_gen_update_size(lruvec, type, folio, old_gen, new_gen);
>
> return new_gen;
> }
>
> --
> 2.55.0
>
>
^ permalink raw reply [flat|nested] 21+ messages in thread
end of thread, other threads:[~2026-10-04 17:29 UTC | newest]
Thread overview: 21+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-03 12:55 [PATCH RFC v3 00/17] mm/mglru: frequency guided promotion (MGLRU-FG) and flag cleanup Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 01/17] mm/memcontrol: allow update of LRU statistic without holding LRU lock Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 02/17] mm/mglru: make generation page counters atomic Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 03/17] mm/memcg: add folio-based lruvec live helper Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 04/17] mm/mglru: frequency guided workingset promotion (MGLRU-FG) Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 05/17] mm/mglru: make folio lru referenced times count a generic API Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 06/17] mm/mglru: move add/del LRU size accounting out of lru_gen_update_size() Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 07/17] mm/mglru, gup: mark folios referenced via a fast helper Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 08/17] mm/smap: convert to LRU refs based operations Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 09/17] mm/madvise: adapt for LRU refs based operations in MGLRU Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 10/17] mm/damon: convert to LRU refs based operations Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 11/17] mm/huge_memory: mark file folio as accessed more accurately on split Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 12/17] mm/mglru: folio LRU refs based active/inactive number accounting Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 13/17] mm/mglru: make folio_inc_lru_refs lruvec lockless Kairui Song via B4 Relay
2026-10-04 17:29 ` KunWu Chan
2026-10-03 12:55 ` [PATCH RFC v3 14/17] mm/mglru: reparent folios from all generations Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 15/17] mm/khugepaged: check folio referenced state via LRU refs under MGLRU Kairui Song via B4 Relay
2026-10-04 2:11 ` Zi Yan
2026-10-04 3:30 ` Kairui Song
2026-10-03 12:55 ` [PATCH RFC v3 16/17] mm/mglru: make folio_test_workingset() work based on folio LRU refs Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 17/17] Documentation/mm: multi-gen LRU: update for frequency guided promotion Kairui Song via B4 Relay
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®