mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH RFC v3 00/17] mm/mglru: frequency guided promotion (MGLRU-FG) and flag cleanup
@ 2026-10-03 12:55 Kairui Song via B4 Relay
  2026-10-03 12:55 ` [PATCH RFC v3 01/17] mm/memcontrol: allow update of LRU statistic without holding LRU lock Kairui Song via B4 Relay
                   ` (16 more replies)
  0 siblings, 17 replies; 21+ messages in thread
From: Kairui Song via B4 Relay @ 2026-10-03 12:55 UTC (permalink / raw)
  To: linux-mm
  Cc: Andrew Morton, Johannes Weiner, Muchun Song, Qi Zheng,
	Ying Huang, Chris Li, Baoquan He, Nico Pache, Usama Arif,
	Michal Hocko, Roman Gushchin, Shakeel Butt, David Hildenbrand,
	Lorenzo Stoakes, Barry Song, Axel Rasmussen, Yuanchu Xie, Wei Xu,
	Vlastimil Babka, Suren Baghdasaryan, Kemeng Shi, Nhat Pham,
	Youngjun Park, Zi Yan, Gregory Price, Matthew Wilcox (Oracle),
	Baolin Wang, Ryan Roberts, Dev Jain, Lance Yang, Hugh Dickins,
	SeongJae Park, David Rientjes, Yu Zhao, Vernon Yang,
	Zicheng Wang, Chen Ridong, Tal Zussman, Kairui Song, Xiang Liu,
	Ehab Ababneh, zhaozhengzhuo, Lian Wang, Kunwu Chan, linux-kernel,
	cgroups, Kairui Song

Hi all,

This is the updated RFC following the idea proposed at LSF/MM/BPF [1],
based on current mm-new. Slightly adjusted the refs increasement part,
so Android test and build kernel test are given even better result.
Retested most other result and they pretty much just match V2's
improvement. Also, dropped the usage of folio_activate in a cold path
to avoid a potential jitter.

In summary, we can see a 10% - 40% higher performance or lower refault in
various different tests, certain workload gets a dramatically reduce of
runtime, while reducing the page flags usage by 1 bit. The gain here is
mostly from real improvement of LRU's ability to distinguish the hotter
workingset. Tested across multiple servers of different archs, desktops,
and Android, all shows very promising results. Compared to V2, V3
is more effected by anon over-reclaim, but provides over-all better
results, that is a problem that should be fixed later or seperately.

It's already very usable, stable, and performing well, but I'll keep
it RFC as this is a major change to LRU, including changing the
Active/Inactive reading, in a good way I think.

It also fixes several long-standing issues including under-accounted PSI
and poor workingset tracking (especially for page cache).

Test results (CLRU means classical LRU):

Build kernel test:
==================
Running make -j48 in a 3G memcg, using ZRAM as swap (256G, lzo-rle) and
holding the kernel and build output on a NVMe drive, 3 runs of 5 swappiness
configurations (15 builds per kernel) [2]; the patched version is better
than mainline at every swappiness value, measuring the total average:

         real     sys    pgpgin  pswpin  pswpout refault_file refault_anon
CLRU    1m37s  13m54s    19.21M   2.43M   10.65M        2.21M        2.49M
Before  1m32s   9m12s    12.44M   2.05M    8.86M         472k        2.03M
After   1m31s   8m30s    11.21M   1.90M    8.99M         391k        1.88M
delta     -1s    -42s     -9.9%   -7.2%    +1.5%       -17.1%        -7.1%

Same test with all 96 threads busy (-j96), which quadruples the reclaim
pressure (pgfault 109M vs 2M per build), 3-4 runs each:

          real         sys    pgpgin  refault_file refault_anon
Before   1m54s      72m12s    111.3M          488k        18.7M
After    1m45s      57m10s    100.0M          362k        17.8M
delta    -9.5s  -15m(-21%)    -10.1%        -25.8%        -4.9%

Same -j48 test, disk swap instead of ZRAM (SSD-backed, the kernel and build
output stay on NVMe). There is a slight regression vs mainline, mostly
from increased swap-out (+17.5% pswpout): while the file working set is
better protected, the 3G memcg pushes more anon out. It still beats CLRU
by a lot, and file refault is lower. Could be related to recent
upstream changes or over-reclaim of anon.

        real      sys   pgpgin  pswpin  pswpout  refault_file  refault_anon
CLRU   4m20s   21m39s   37.54M   3.00M   10.51M         6.22M         3.05M
Before 2m52s   10m11s   10.08M   1.58M    5.04M          422k         1.06M
After  3m04s   11m06s   10.19M   1.68M    5.92M          380k         1.11M
delta   +12s      +9%    +1.1%   +6.5%   +17.5%        -10.1%         +4.3%

For reference, test result from V2:

        real     sys  pgpgin  pswpin  pswpout  refault_file  refault_anon
CLRU   6m06s  31m01s   50.3M   3.20M    13.8M         10.3M         3.35M
Before 2m58s  10m58s   10.30M  1.60M    5.25M          434k         1.06M
After  2m50s  10m38s    8.79M  1.34M    4.82M          377k          844k
delta    -8s    -20s    -15%    -16%     -8%           -13%          -20%

MongoDB YCSB workloadb [3]
==========================
With recordcount:20000000 operationcount:6000000, threads:48,
in a 16G memcg, 3 runs:

CLRU:          98389.94 ops/s
MGLRU Before:  83700.34 ops/s
MGLRU After:   94951.21 ops/s (+13.4%)

There is still a little gap to CLRU, and this is the only test behind
CLRU, which I believe is related to writeback threshold (64 vs 32) which
we can tune later. Test from community didn't show such gap [8].

Chromium & Node.js test [4]
===========================
Using ZRAM as swap, on a 48c96t machine with 128G memory, 64 workers, run
for 1 hour:
                 Total requests:
CLRU:                     63822
MGLRU Before:            132763
MGLRU After:             233774 (+76.0%)

(NOTE: It seems some recent change broken MGLRU's fairness guarteen and
also made this test dramatically faster than a few months ago, which isn't
related to this series and reading are even better now, but I'll take a
deeper look later.)

FIO cached with zipf
====================
Using an NVMe disk, in a 16G cgroup, total file size 40G, this measures
the LRU's theoretical ability to distinguish the hotter portion, 3 test
run each config:

fio --name=fg --numjobs=16 --nrfiles=1 \
    --filename_format="$testdir/rnvmedk.\$jobnum.img" \
    --size=${FILE_MIB}M \
    --buffered=1 --ioengine=sync --rw=randread \
    --random_distribution=zipf:$ZIPF --bs=4k --time_based \
    --ramp_time=45s --runtime=600s --group_reporting

Avg IOPS (higher is better):
+----------+----------+----------+----------+----------+-----------+
| Config   |      0.8 |      0.9 |      1.1 |      1.2 | avg delta |
+----------+----------+----------+----------+----------+-----------+
| CLRU     |  400,000 |  601,667 | 2,074,333| 4,354,333|     +1.2% |
| Before   |  382,667 |  604,667 | 2,071,000| 4,334,667|       --  |
| After    |  434,667 |  669,000 | 2,301,333| 4,800,000|    +11.5% |
+----------+----------+----------+----------+----------+-----------+
Delta per zipf (higher is better): +13.6 / +10.6 / +11.1 / +10.7 %

Throughput-normalized file miss (refault/read, lower is better):
+----------+----------+----------+----------+----------+-----------+
| Config   |      0.8 |      0.9 |      1.1 |      1.2 | avg delta |
+----------+----------+----------+----------+----------+-----------+
| CLRU     |  0.17868 |  0.11840 |  0.03216 |  0.01274 |     -2.8% |
| Before   |  0.18237 |  0.12232 |  0.03296 |  0.01321 |       --  |
| After    |  0.16680 |  0.10994 |  0.02907 |  0.01142 |    -11.0% |
+----------+----------+----------+----------+----------+-----------+
Delta per zipf (lower is better): -8.5 / -10.1 / -11.8 / -13.6 %
(MB/s ~ IOPS x 4 KiB; e.g. 4.80M IOPS ~ 18.75 GB/s.)

On the throughput-normalized file miss rate (refaults per read, the
metric that reflects LRU workingset-detection accuracy), unpatched MGLRU
is ~2–3% worse then CLRU across every zipfian access pattern on
this page-cache read workload which the standard model of real
cache-locality skew. This matches the long complained MGLRU cache issue
from community. However the comparable IOPS largely reflects MGLRU's lower
internal LRU/bookkeeping overhead masking the higher miss rate.

And, the patched MGLRU-FG, lowers the miss rate ~8–14% versus default
MGLRU and ~6–10% versus CLRU (best of all three) while raising
IOPS ~10-11% versus both. It detects the workingset more accurately
than all others while retaining MGLRU's lower overhead than CLRU.

So in summary: MGLFU-FG provides a ~10% gain on zipf access on real
high performance disks compared to CLRU, while unpatched MGLRU is ~1-3%
worse than CLRU.

LevelDB Scan/Get
================

I also retested the LevelDB benchmark from the cache_ext paper [5].
Interestingly, mainline MGLRU already beats CLRU on this one after a
recent change in lru_gen_folio_seq that bumps new folios with refs == 1
to the second-oldest generation. That change accidentally gave random
reads a higher hotness level while making sequential reads much colder:
sequential reads involve many readahead hits, and readahead folios start
with refs == 0, so they're already in the oldest generation, and
folio_mark_accessed() on a readahead hit has almost no effect on generations
in mainline MGLRU. Meanwhile, all direct-hit (random read) folios start
with refs == 1 in the second-oldest generation. As a result, the scan-get
test natively protects the "get" part and sacrifices the "scan" part.
That's not the best solution though. It's unreliable because it depends
on LRU drain timing, and it hurts workloads where the sequential part is
actually hotter (any workload involving a hotter large file and many
small cold files will be affected).

This series improves on that base: it covers ordinary workloads without
hurting the scan-get workload and without relying on that initial bump.

LevelDB Scan / Get, Throughput Total:
CLRU:         4668.8 ops/s
MGLRU:        5026.9 ops/s (faster than CLRU, but hurts other workloads)
MGLRU After:  5029.7 ops/s (fastest in all cases, and no regression)

The hot-sequential and cold-random workload can be easily reproduced with
SQLite and grep. SQLite continuously scans and looks up a small hot
portion of a DB file, while grep iterates over a set of small files much
larger than RAM [6]:

         SQLite scan & lookup time:       Grep iterate time:
CLRU:                       14.51ms               13281.37ms
MGLRU mainline:            567.05ms               13694.47ms
MGLRU After this series:    10.58ms               12930.43ms

The grep cold portion is larger than RAM and accessed only once per
iteration, so there's no promotion of any of it. CLRU handles
this reasonably; mainline MGLRU has a clear regression; MGLRU-FG now
not only recovers but is able to catch some hot parts from the cold grep
workload. This test is somewhat subjective, but the signal is clear.

Android (For reference)
=======================

Testing on Android is quite difficult: lack of mainline support, and it is
very noisy to get a stable result. I did an informal backport of the latest
MGLRU-FG patches onto the 6.1 GKI tree of a Pixel 9 Pro (16 GB RAM, 8 GB
ZRAM, Android 17), preserving the frozen kABI layouts so the implementation
is limited. The "before" and "after" kernels share the same base and differ
only by the FG series: the "before" tree has MGLRU aligned to the current
upstream code, so that unrelated backport deltas cancel out. (And BTW that
alignment also improved the performance by a lot, which matches the report
from previous MGLRU reclaim optimization series [12]).

The workload is a churn loop driven over adb: 34 common apps and 20 Chrome
tabs are launched and cycled, with a brief scroll per app during the cold
build of each iteration, 16 iterations per run, about one hour per run. The
loop exhausts memory in every iteration (8 GB ZRAM full, MemFree down to
~80 MB). The two kernels are run in an interleaved rotation and compared on
matched per-iteration samples (same period, same iteration index), which
cancels most of the run-to-run drift. 6 runs per kernel, roughly more
than 18 hours of device time in total.

The memory-management stack in Android is built around the non-FG MGLRU
behaviour, and the 6.1 base lacks some of the upstream infrastructure. Even
so, FG holds its own, it lowers reclaim traffic on both the file and the anon
side:

                                     Before     After
------------------------------------------------------
workingset_refault_anon               2.69M     2.38M   (-11%)
workingset_refault_file               3.65M     3.16M   (-14%)
pgscan_anon                           8.91M     7.50M   (-16%)
pgsteal_anon                          3.42M     2.97M   (-13%)
pgscan_file                          13.01M    11.07M   (-15%)
pgsteal_file                         11.20M     9.38M   (-16%)
pswpout                               3.82M     3.35M   (-12%)
pgpgin                               40.78M    33.58M   (-18%)

("Before" is the aligned tree without the FG series, "After" adds FG. The
numbers are trimmed means of the 6 interleaved runs per kernel: for each
counter the highest and lowest run are dropped and the middle four averaged;
each counter is the vmstat delta of one full run.)

On the 96 matched per-iteration samples the reductions are significant for
every counter in the table, and also for pgmajfault (-12%); only pgpgout is
neutral. An earlier revision of the series was statistically neutral on the
anon side and reduced the file side only, while this version improves both.

I also ran the Android Jank test from Zicheng [7] on this device (chrome
scroll, FrameTimeline), plus a 34-app keepalive run. No obvious difference
was observed: the run-mean chrome scroll fps of every kernel sits in a
100.6 - 103.6 band on a 120 Hz panel, and the jank ratios are all below 8%
per scroll round (n=3-8 gated jank frames per run, so Poisson noise dominates)
with no consistent ordering across kernels. The keepalive test does not
discriminate on this build: 31-32 of 34 apps stay alive with every kernel.
On the previous build a run kept slightly more apps alive with FG (average
4.8 vs 4.5, peak 23 vs 20), that could be noise or a slight improvement.
Both kernels saturate all 8 cores with no obvious CPU usage difference
(99.8-99.9% utilization, ~518s busy per 65s window).

I also did a test on another Android phone with a 5.15 kernel (Xperia 1 V),
which has all apps in one global memcg (this Pixel 9 Pro has each app in a
separate memcg). The result looks much better there, either due to the memcg
layout or the non-reclaiming anon shadow (the 6.1 kernel reclaims anon
shadow). But in either case, the performance is a positive reading.

So in summary: with the final revision we get lower file refaults, lower
pgpgin, and lower anon reclaim and swap traffic, with no measured
user-visible cost.

Others
======

Additionally, PSI, smaps, and readahead should all benefit from better
accuracy since this series unifies the flag usage between classical
LRU and MGLRU.

Other tests such as MySQL are looking fine, with no regressions. There is
also community test report [8].

Refault distance is not included yet, so MGLRU may respond more slowly to
workingset shifts. That can be added later, as previously demonstrated
[9], [10].

Extra note about future development: this series is actually highly
compatible with ideas like workingset reporting [11]. The "gen climbing
folio" design may appear to conflict with workingset reporting's idea,
but it doesn't. The solution is simple and straightforward: once we can
extend the generation number to a larger value (e.g. 64 or 128), the
refs-driven promotion can stop at a lower gen (e.g. oldest_gen + 16),
leaving the remaining newer generations as perfectly time-gap-separated
bins.

The tier count is not fixed either; we'll need to find a way to tune it
if tiers go beyond 4 though, but that shouldn't be hard, not a blocker
either.

More details are in the individual commit messages. LLM is used to help
improve the comments and tests as I'm really not good at that :)

Link: https://lore.kernel.org/linux-mm/CAMgjq7BoekNjg-Ra3C8M7=8=75su38w=HD782T5E_cxyeCeH_g@mail.gmail.com/ [1]
Link: https://lore.kernel.org/linux-mm/CAGsJ_4xre-x0e+qNVm=KLFnO1dbPkPX5RuecqwvTZu-vS+o8yQ@mail.gmail.com/ [2]
Link: https://github.com/brianfrankcooper/YCSB/blob/master/workloads/workloadb [3]
Link: https://lore.kernel.org/all/20221220214923.1229538-1-yuzhao@google.com/ [4]
Link: https://dl.acm.org/doi/10.1145/3731569.3764820 [5]
Link: https://github.com/ryncsn/emm-test-project/tree/master/sqlite-grep [6]
Link: https://github.com/purplewall1206/android-perf-bench [7]
Link: https://lore.kernel.org/linux-mm/92DCEFFD13221261+20260828102344.1537874-1-zhaozhengzhuo@uniontech.com/ [8]
Link: https://lwn.net/Articles/945266/ [9]
Link: https://lore.kernel.org/linux-mm/20260502-mglru-fg-v1-0-913619b014d9@tencent.com/ [10]
Link: https://lwn.net/Articles/976985/ [11]
Link: https://lore.kernel.org/linux-mm/20260417025123.2971253-1-wxy2009nrrr@163.com/ [12]

Assisted-by: LLM
Signed-off-by: Kairui Song <kasong@tencent.com>
Tested-by: zhaozhengzhuo <zhaozhengzhuo@uniontech.com>
---
Changes in v3:
- Micro-optimization for page flags operations: if the flags are
  unchanged after calculation, skip the cmpxchg.
- Avoid touching folio_activate even in the cold path. I tried multiple
  ways for that lockless promotion, PG_lru in the earlier RFC, and
  previous folio_activate, the new speculative barrier + retry seem the
  best solution.
- Apply tier cap for MADV_PAGEOUT.
- Only promote folio during page table walk or rmap if the folio
  is in the min gen. May worth trying to enlarge the range to 
  (max_seq - MIN_NR_GENS) in next version.
- Fix a potential folio leak caused by reparenting.
- Retest shows great result especially for Android case.
- Link to v2: https://patch.msgid.link/20260911-mglru-fg-v2-0-f26e5cb26da7@tencent.com

Changes in v2:
- Rebased; dropped v1 03/05/06 (already upstream), folded v1 01 and 07
  into patches 01 and 04; new patches 06, 07, 10, 12, 13.
- Make folio_test_workingset() itself arbitrate, dropping the parallel
  folio_is_* helpers; convert the last raw PageWorkingset() user
  (erofs). (Johannes)
- Use LRU_REF_MAPPED/LRU_REF_EXEC flags instead of is_fault/is_exec
  booleans. (Barry)
- Fix syzbot "WARNING in folio_inc_lru_refs" on off-LRU folios.
- Account active/inactive per folio from refs, not the gen window:
  /proc/vmstat and memory.stat no longer jump on aging or reverse on
  swapless machines.
- Make folio_inc_lru_refs() lockless; add folio_inc_lru_refs_fast()
  for the gup fast paths.
- Convert DAMON and khugepaged to the refs-based operations.
- Drop the lru_size WARN_ON_ONCE() and lockdep_assert_held(): the
  counter is lockless now, so transient negatives are expected.
- Link to v1: https://patch.msgid.link/20260804-mglru-fg-v1-0-4d8dad39dad6@tencent.com

---
Kairui Song (17):
      mm/memcontrol: allow update of LRU statistic without holding LRU lock
      mm/mglru: make generation page counters atomic
      mm/memcg: add folio-based lruvec live helper
      mm/mglru: frequency guided workingset promotion (MGLRU-FG)
      mm/mglru: make folio lru referenced times count a generic API
      mm/mglru: move add/del LRU size accounting out of lru_gen_update_size()
      mm/mglru, gup: mark folios referenced via a fast helper
      mm/smap: convert to LRU refs based operations
      mm/madvise: adapt for LRU refs based operations in MGLRU
      mm/damon: convert to LRU refs based operations
      mm/huge_memory: mark file folio as accessed more accurately on split
      mm/mglru: folio LRU refs based active/inactive number accounting
      mm/mglru: make folio_inc_lru_refs lruvec lockless
      mm/mglru: reparent folios from all generations
      mm/khugepaged: check folio referenced state via LRU refs under MGLRU
      mm/mglru: make folio_test_workingset() work based on folio LRU refs
      Documentation/mm: multi-gen LRU: update for frequency guided promotion

 Documentation/mm/multigen_lru.rst |  56 +++--
 fs/btrfs/compression.c            |   1 +
 fs/erofs/zdata.c                  |   3 +-
 fs/proc/task_mmu.c                |  22 +-
 include/linux/memcontrol.h        |  45 +++-
 include/linux/mm_inline.h         | 370 ++++++++++++++++-----------
 include/linux/mmzone.h            | 168 +++++++++----
 include/linux/page-flags.h        |   2 -
 kernel/bounds.c                   |   2 +-
 mm/damon/paddr.c                  |   9 +-
 mm/filemap.c                      |   1 +
 mm/folio.c                        |  85 +------
 mm/gup.c                          |   6 +-
 mm/huge_memory.c                  |   8 +-
 mm/khugepaged.c                   |   4 +-
 mm/madvise.c                      |  57 +++--
 mm/memcontrol.c                   |   6 +-
 mm/migrate.c                      |   2 -
 mm/page_io.c                      |   1 +
 mm/vmscan.c                       | 515 +++++++++++++++++++++++++-------------
 mm/workingset.c                   |  45 ++--
 21 files changed, 882 insertions(+), 526 deletions(-)
---
base-commit: 763ad0211c7b587344f03bc4d1299810aeb736f4
change-id: 20260722-mglru-fg-3a2c8574725b

Best regards,
--  
Kairui Song <kasong@tencent.com>



^ permalink raw reply	[flat|nested] 21+ messages in thread

end of thread, other threads:[~2026-10-04 17:29 UTC | newest]

Thread overview: 21+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-03 12:55 [PATCH RFC v3 00/17] mm/mglru: frequency guided promotion (MGLRU-FG) and flag cleanup Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 01/17] mm/memcontrol: allow update of LRU statistic without holding LRU lock Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 02/17] mm/mglru: make generation page counters atomic Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 03/17] mm/memcg: add folio-based lruvec live helper Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 04/17] mm/mglru: frequency guided workingset promotion (MGLRU-FG) Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 05/17] mm/mglru: make folio lru referenced times count a generic API Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 06/17] mm/mglru: move add/del LRU size accounting out of lru_gen_update_size() Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 07/17] mm/mglru, gup: mark folios referenced via a fast helper Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 08/17] mm/smap: convert to LRU refs based operations Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 09/17] mm/madvise: adapt for LRU refs based operations in MGLRU Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 10/17] mm/damon: convert to LRU refs based operations Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 11/17] mm/huge_memory: mark file folio as accessed more accurately on split Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 12/17] mm/mglru: folio LRU refs based active/inactive number accounting Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 13/17] mm/mglru: make folio_inc_lru_refs lruvec lockless Kairui Song via B4 Relay
2026-10-04 17:29   ` KunWu Chan
2026-10-03 12:55 ` [PATCH RFC v3 14/17] mm/mglru: reparent folios from all generations Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 15/17] mm/khugepaged: check folio referenced state via LRU refs under MGLRU Kairui Song via B4 Relay
2026-10-04  2:11   ` Zi Yan
2026-10-04  3:30     ` Kairui Song
2026-10-03 12:55 ` [PATCH RFC v3 16/17] mm/mglru: make folio_test_workingset() work based on folio LRU refs Kairui Song via B4 Relay
2026-10-03 12:55 ` [PATCH RFC v3 17/17] Documentation/mm: multi-gen LRU: update for frequency guided promotion Kairui Song via B4 Relay

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®