mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v3 0/3] mm: fix workingset refaults in the zswap writeback path
@ 2026-08-25 17:24 Alexandre Ghiti
  2026-08-25 17:24 ` [PATCH v3 1/3] mm/swap: move LRU insertion out of the swap cache allocator Alexandre Ghiti
                   ` (2 more replies)
  0 siblings, 3 replies; 8+ messages in thread
From: Alexandre Ghiti @ 2026-08-25 17:24 UTC (permalink / raw)
  To: Johannes Weiner, Yosry Ahmed, Nhat Pham, Chengming Zhou,
	Andrew Morton, David Hildenbrand, Lorenzo Stoakes,
	Liam R. Howlett, Vlastimil Babka, Mike Rapoport,
	Suren Baghdasaryan, Michal Hocko, Hugh Dickins, Baolin Wang,
	Chris Li, Kairui Song, Kemeng Shi, Baoquan He, Barry Song,
	Youngjun Park, Qi Zheng, Shakeel Butt, Axel Rasmussen,
	Yuanchu Xie, Wei Xu, Joonsoo Kim
  Cc: linux-mm, linux-kernel, Alexandre Ghiti

Note: patch 1 also appears as patch 1 of the zswap dropbehind series [1].
It is the same patch, byte for byte. Both series need it and both are
meant to apply on their own, so it is posted in each; whichever lands
first, the other should drop it.

When an anonymous folio is reclaimed, workingset_eviction() stores a
"shadow" (the eviction cookie) in the swap slot so that a later swap-in
can be recognised as a refault and, if the refault distance is short
enough, the page can be re-activated. This is how anon workingset/refault
detection has worked since commit aae466b0052e ("mm/swap: implement
workingset detection for anonymous LRU").

zswap writeback breaks this in two independent ways:

  - Over-count at writeback: the shrinker allocates a buffer folio in the
    swap cache, and the allocation path counts that folio as a refault.

  - Lost eviction cookie at reclaim: adding the buffer to the swap cache
    overwrites the slot's shadow, so the original cookie is lost; when the
    buffer folio is finally reclaimed a fresh, inaccurate cookie is minted
    in its place.

This series preserves the shadow within zswap itself, without any new
swap-table or swap-slot state. At writeback, instead of erasing the freed
zswap entry, the captured shadow is parked in the zswap tree in its place,
so it outlives the writeback buffer folio.

Finally, the refault evaluation is moved out of the swap-cache allocator
into the swap-in callers, so allocating the writeback buffer is no longer
miscounted as a refault.

Results
-------

Measured with a sysbench OLTP (MariaDB) workload in a memory cgroup sized
so the dataset and InnoDB buffer pool both overcommit it, with the zswap
shrinker on so entries are continuously written back to an NVMe swap
device (classic LRU; MGLRU off). Anon workingset counters over the
measured window, baseline vs this series, mean +/- stddev over 10 runs:

  workingset_refault_anon    383,242 +/- 59,345  ->  199,355 +/- 27,042   -48%
  workingset_activate_anon    54,890 +/- 11,742  ->   22,607 +/-  3,118   -59%
  workingset_restore_anon     16,212 +/-  4,590  ->    7,599 +/-  1,190   -53%

Writeback volume is comparable (zswpwb 183k +/- 16k -> 178k +/- 15k), so
the reduction is not from doing less work. Normalised per transaction the
reduction holds (-53%/-49%/-39%) while the swap work per transaction is
unchanged. A kernel build under the same pressure moves all three
counters in the same direction.

The run-to-run variance of these counters drops as well.

Throughput is unaffected: over the same 10 runs, transactions/s is
18.50 +/- 1.08 -> 18.65 +/- 0.94, i.e. +0.8% with a 95% confidence
interval of +/- 7.7%.

Changes in v3:
- David pointed out that the boolean added to workingset_refault() in v2
  makes the calling code hard to read. The boolean is now internal to
  mm/workingset.c and the callers use workingset_refault() or
  workingset_refault_lru_managed(), the same way remove_mapping() and
  remove_mapping_reclaim() do. The three existing callers are left
  untouched.
- zswap_writeback_entry() now restores the shadow into the slot on the
  error paths taken after the buffer was allocated: the allocation had
  already overwritten it and nothing was parked yet, so it was lost.
- Dropped the "if (!shadow) shadow = ZSWAP_WRITEBACK_NO_SHADOW" fallback,
  which can never be taken: the swap table already marks a swapped out
  slot with xa_mk_value(0), so swap_cache_get_shadow() never returns NULL
  for one.
- Rebased on mm-new.

Changes in v2:
- Sashiko pointed out that v1 evaluated the refault after folio_add_lru(),
  which picks the MGLRU generation before PG_workingset is set. Patch 1
  now moves the LRU insertion out of the swap cache allocator so the
  refault is evaluated before it, as it was originally.
- Sashiko also pointed out that a failed writeback redirties the buffer
  and leaves it in the swap cache with its shadow still parked, so a
  later zswap_store() on it would hand the parked value to
  zswap_entry_free(). zswap_store() now bails out for such a folio:
  writeback already decided that data belongs on disk, so it is written
  there instead of being compressed again, which also keeps the parked
  shadow intact until the folio leaves the swap cache.
- The buffer folio a swap-in consumes is already on the LRU, so the
  refault there cannot activate it by setting PG_active: that leaves the
  flag disagreeing with the list the folio is on, which shows up as an
  mm/memcontrol.c lru_size underflow when it is freed. Such a folio is now
  activated with folio_activate() instead. One still sitting in a per-CPU
  batch cannot be moved safely and just misses the activation; a counter
  on that path measured 0.0007% of the activations over a 10 run test.
  This is only needed because zswap writeback still puts its buffer on the
  LRU: once it stops doing so, workingset_refault_lru_managed() has no
  caller left and can go away.

[1] https://lore.kernel.org/linux-mm/20260818163221.589352-2-alex@ghiti.fr/

v1: https://lore.kernel.org/all/20260817144622.137133-1-alex@ghiti.fr/
v2: https://lore.kernel.org/linux-mm/20260821093606.2231216-1-alex@ghiti.fr/

Alexandre Ghiti (3):
  mm/swap: move LRU insertion out of the swap cache allocator
  mm/swap: refault on swap-in, not in the swap cache allocator
  mm/zswap: preserve the workingset shadow across writeback

 include/linux/zswap.h |  12 +++++
 mm/internal.h         |   1 +
 mm/memory.c           |   5 ++
 mm/shmem.c            |   5 ++
 mm/swap.h             |   6 +--
 mm/swap_state.c       |  50 +++++++++++++++----
 mm/vmscan.c           |   4 +-
 mm/workingset.c       |  60 ++++++++++++++++++-----
 mm/zswap.c            | 110 ++++++++++++++++++++++++++++++++++++++++--
 9 files changed, 224 insertions(+), 29 deletions(-)


base-commit: 1a46b1e97bde62afa7d925bb0dcd9f9748a1d7c3
-- 
2.53.0-Meta


^ permalink raw reply	[flat|nested] 8+ messages in thread

end of thread, other threads:[~2026-08-26 15:23 UTC | newest]

Thread overview: 8+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-08-25 17:24 [PATCH v3 0/3] mm: fix workingset refaults in the zswap writeback path Alexandre Ghiti
2026-08-25 17:24 ` [PATCH v3 1/3] mm/swap: move LRU insertion out of the swap cache allocator Alexandre Ghiti
2026-08-26 13:21   ` Usama Arif
2026-08-26 15:23     ` Alexandre Ghiti
2026-08-25 17:24 ` [PATCH v3 2/3] mm/swap: refault on swap-in, not in " Alexandre Ghiti
2026-08-26 13:31   ` Usama Arif
2026-08-26 15:22     ` Alexandre Ghiti
2026-08-25 17:24 ` [PATCH v3 3/3] mm/zswap: preserve the workingset shadow across writeback Alexandre Ghiti

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®