* [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload
@ 2026-09-26 4:55 Matthias Goergens
2026-09-26 4:55 ` [RFC PATCH v2 1/4] selftests: zram: track owned devices and report cleanup failures Matthias Goergens
` (8 more replies)
0 siblings, 9 replies; 16+ messages in thread
From: Matthias Goergens @ 2026-09-26 4:55 UTC (permalink / raw)
To: Andrew Morton
Cc: Matthias Goergens, Chris Li, Kairui Song, Johannes Weiner,
David Hildenbrand, Michal Hocko, Shakeel Butt, Kemeng Shi,
Nhat Pham, Yosry Ahmed, Youngjun Park, Baoquan He, Barry Song,
linux-mm, linux-kernel, cgroups, linux-api, Alejandro Colomar,
linux-man, Karel Zak, util-linux, Jani Nikula, Joonas Lahtinen,
Rodrigo Vivi, Tvrtko Ursulin, David Airlie, Simona Vetter,
intel-gfx, dri-devel, Rafael J . Wysocki, Pavel Machek,
Catalin Marinas, Will Deacon, linux-pm, linux-arm-kernel,
Shuah Khan, linux-kselftest
This RFC adds offload-only swap areas for deliberate cold-page offload to
backends unsuitable for pressure reclaim. ZFS zvol swap has documented
deadlocks under memory pressure [3]; compressed swap is another target
because writes may need memory despite free logical slots.
Explicit proactive reclaim can use these areas alongside conventional
swap, ordered by priority. Ordinary reclaim cannot initiate non-zero
writes to them, including through retained swap entries. Conventional
capacity is not reserved for emergencies. Operators choose the policy;
the kernel does not measure headroom or make an allocating backend safe.
TMO/Senpai [4] and DAMON_RECLAIM [5] are related cold-page reclaim
approaches. This RFC admits offload-only swap through memory.reclaim and
per-node reclaim, but not DAMON. Other related work includes per-cgroup
zswap writeback control [6], Virtual Swap Space v4 [1] and swap tiers v10
[2].
The util-linux companion [8] proposes swapon --offload-only and an fstab
option. Older swapon silently ignores the fstab token, so persistent
activation needs discussion. Apply the separate i915 fix [9] first to avoid
a pre-existing folio-lock leak when shmem writeback is skipped.
The series fixes zram selftest device tracking and error reporting, adds
the swap policy with DRM eligibility checks, and tests routing, workingset
activation and retained-entry write refusal. It is based on mm-new at
995829088503.
Feedback is particularly welcome on:
- Activation-time eligibility versus backend placement or migration:
is retained-entry refusal, with possible reclaim churn or OOM,
acceptable without migration?
- Eligible-capacity accounting, including overcommit limits and OOM
scoring; cache recovery, cluster invalidation and workingset activation.
- Persistent activation and visibility: current swap listings do not
expose the policy.
Validation includes x86 builds, affected arm64 MTE objects and builds
without swap or memory cgroups. QEMU tests covered routing, retained-entry
recovery, data integrity and cleanup, with negative controls for write
refusal and teardown failure.
I have also been using this policy on my own machine with experimental
bcachefs swap support, with no problems observed so far. Deliberate stress
testing is confined to VMs; I do not deliberately stress-test this machine.
Changes since v1, incorporating Sashiko's public review [7]:
- Improve selftest isolation, retained-swap measurement and memlock skips.
- Track allocated zram devices, wait for udev probes and report cleanup
errors without deleting data through a mount that could not be released.
- Clarify activation, capacity reporting and conventional fallback limits.
[1] https://patchew.org/linux/20260825153238.2695446-1-nphamcs@gmail.com/
[2] https://patchew.org/linux/20260713025644.170839-1-youngjun.park@lge.com/
[3] https://openzfs.github.io/openzfs-docs/Project%20and%20Community/FAQ.html#using-a-zvol-for-a-swap-device-on-linux
[4] https://engineering.fb.com/2022/06/20/data-infrastructure/transparent-memory-offloading-more-memory-at-a-fraction-of-the-cost-and-power/
[5] https://www.kernel.org/doc/html/latest/admin-guide/mm/damon/reclaim.html
[6] https://www.kernel.org/doc/html/latest/admin-guide/mm/zswap.html
[7] https://sashiko.dev/#/patchset/20260914104501.3960616-1-matthias.goergens@gmail.com
[8] https://github.com/util-linux/util-linux/pull/4633
[9] https://lore.kernel.org/all/20260914105606.3997649-1-matthias.goergens@gmail.com/
v1: https://lore.kernel.org/all/20260914104501.3960616-1-matthias.goergens@gmail.com/
Matthias Goergens (4):
selftests: zram: track owned devices and report cleanup failures
mm: restrict offload-only swap to proactive reclaim
selftests: zram: cover offload-only swap policy
selftests: zram: cover retained offload-only entries
Documentation/mm/swap.rst | 105 ++++
drivers/gpu/drm/i915/gem/i915_gem_shrinker.c | 2 +-
.../gpu/drm/i915/gem/selftests/huge_pages.c | 6 +-
drivers/gpu/drm/msm/msm_gem_shrinker.c | 2 +-
drivers/gpu/drm/panthor/panthor_gem.c | 2 +-
drivers/gpu/drm/ttm/ttm_backup.c | 2 +-
drivers/gpu/drm/xe/tests/xe_bo.c | 2 +-
include/linux/swap.h | 29 +-
include/linux/vm_event_item.h | 1 +
mm/memcontrol.c | 20 +-
mm/page_io.c | 24 +-
mm/swapfile.c | 123 ++++-
mm/vmscan.c | 31 +-
mm/vmstat.c | 1 +
tools/testing/selftests/zram/.gitignore | 2 +
tools/testing/selftests/zram/Makefile | 4 +-
tools/testing/selftests/zram/README | 20 +-
tools/testing/selftests/zram/config | 10 +-
tools/testing/selftests/zram/settings | 1 +
tools/testing/selftests/zram/swap_offload.c | 425 +++++++++++++++++
.../selftests/zram/workingset_offload.c | 207 ++++++++
tools/testing/selftests/zram/zram.sh | 9 +
tools/testing/selftests/zram/zram01.sh | 9 +-
tools/testing/selftests/zram/zram02.sh | 9 +-
tools/testing/selftests/zram/zram03.sh | 174 +++++++
tools/testing/selftests/zram/zram04.sh | 164 +++++++
tools/testing/selftests/zram/zram05.sh | 447 ++++++++++++++++++
tools/testing/selftests/zram/zram_lib.sh | 189 ++++++--
28 files changed, 1926 insertions(+), 94 deletions(-)
create mode 100644 tools/testing/selftests/zram/settings
create mode 100644 tools/testing/selftests/zram/swap_offload.c
create mode 100644 tools/testing/selftests/zram/workingset_offload.c
create mode 100755 tools/testing/selftests/zram/zram03.sh
create mode 100755 tools/testing/selftests/zram/zram04.sh
create mode 100755 tools/testing/selftests/zram/zram05.sh
--
2.55.0
^ permalink raw reply [flat|nested] 16+ messages in thread
* [RFC PATCH v2 1/4] selftests: zram: track owned devices and report cleanup failures
2026-09-26 4:55 [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload Matthias Goergens
@ 2026-09-26 4:55 ` Matthias Goergens
2026-09-26 4:55 ` [RFC PATCH v2 2/4] mm: restrict offload-only swap to proactive reclaim Matthias Goergens
` (7 subsequent siblings)
8 siblings, 0 replies; 16+ messages in thread
From: Matthias Goergens @ 2026-09-26 4:55 UTC (permalink / raw)
To: Shuah Khan
Cc: Matthias Goergens, Minchan Kim, Sergey Senozhatsky, linux-kernel,
linux-kselftest, linux-mm, Andrew Morton
hot_add returns the lowest available device ID, which need not equal the
number of existing devices. With devices 0 and 2 present, the tests
allocate device 1 but configure and clean up device 2. Record the returned
IDs and successful swap and mount activations instead of assuming ranges.
Stop zram01 before filling after mount failure. Remove only directories
created for successful mounts, and preserve devices when swapoff or
unmount fails, avoiding deletion through a mount cleanup could not release.
Wait for udev probes before reset and removal: a worker holding the device
open can cause EBUSY. Bound these waits and treat timeouts as warnings,
since unrelated events can delay the global queue. Propagate actual
teardown errors while continuing cleanup of other devices. Preserve both
test results in the registered runner so zram02 cannot hide zram01 failure.
Propagate swap activation and swapoff errors to the test result.
Signed-off-by: Matthias Goergens <matthias.goergens@gmail.com>
---
tools/testing/selftests/zram/zram.sh | 9 ++
tools/testing/selftests/zram/zram01.sh | 9 +-
tools/testing/selftests/zram/zram02.sh | 9 +-
tools/testing/selftests/zram/zram_lib.sh | 161 ++++++++++++++++-------
4 files changed, 136 insertions(+), 52 deletions(-)
diff --git a/tools/testing/selftests/zram/zram.sh b/tools/testing/selftests/zram/zram.sh
index b0b91d9b0dc2..b597cf186835 100755
--- a/tools/testing/selftests/zram/zram.sh
+++ b/tools/testing/selftests/zram/zram.sh
@@ -5,12 +5,21 @@ TCID="zram.sh"
. ./zram_lib.sh
run_zram () {
+local ret status
+
echo "--------------------"
echo "running zram tests"
echo "--------------------"
./zram01.sh
+ret=$?
echo ""
./zram02.sh
+status=$?
+if [ "$status" -ne 0 ] &&
+ { [ "$ret" -eq 0 ] || [ "$ret" -eq "$ksft_skip" ]; }; then
+ ret=$status
+fi
+return "$ret"
}
check_prereqs
diff --git a/tools/testing/selftests/zram/zram01.sh b/tools/testing/selftests/zram/zram01.sh
index 8f4affe34f3e..185f68471b7b 100755
--- a/tools/testing/selftests/zram/zram01.sh
+++ b/tools/testing/selftests/zram/zram01.sh
@@ -33,7 +33,7 @@ zram_algs="lzo"
zram_fill_fs()
{
- for i in $(seq $dev_start $dev_end); do
+ for i in $dev_ids; do
echo "fill zram$i..."
local b=0
while [ true ]; do
@@ -57,19 +57,20 @@ zram_fill_fs()
}
check_prereqs
-zram_load
+zram_load || { zram_cleanup; exit 1; }
zram_max_streams
zram_compress_alg
zram_set_disksizes
zram_set_memlimit
zram_makefs
-zram_mount
+zram_mount || { zram_cleanup; exit 1; }
zram_fill_fs
-zram_cleanup
+zram_cleanup || ERR_CODE=1
if [ $ERR_CODE -ne 0 ]; then
echo "$TCID : [FAIL]"
+ exit 1
else
echo "$TCID : [PASS]"
fi
diff --git a/tools/testing/selftests/zram/zram02.sh b/tools/testing/selftests/zram/zram02.sh
index 2418b0c4ed13..2420dda987d4 100755
--- a/tools/testing/selftests/zram/zram02.sh
+++ b/tools/testing/selftests/zram/zram02.sh
@@ -29,16 +29,17 @@ zram_sizes="1048576" # 1M
zram_mem_limits="1M"
check_prereqs
-zram_load
+zram_load || { zram_cleanup; exit 1; }
zram_max_streams
zram_set_disksizes
zram_set_memlimit
-zram_makeswap
-zram_swapoff
-zram_cleanup
+zram_makeswap || ERR_CODE=1
+zram_swapoff || ERR_CODE=1
+zram_cleanup || ERR_CODE=1
if [ $ERR_CODE -ne 0 ]; then
echo "$TCID : [FAIL]"
+ exit 1
else
echo "$TCID : [PASS]"
fi
diff --git a/tools/testing/selftests/zram/zram_lib.sh b/tools/testing/selftests/zram/zram_lib.sh
index 0d44d83888f9..4b134f70726d 100755
--- a/tools/testing/selftests/zram/zram_lib.sh
+++ b/tools/testing/selftests/zram/zram_lib.sh
@@ -5,10 +5,10 @@
# Author: Alexey Kodanev <alexey.kodanev@oracle.com>
# Modified: Naresh Kamboju <naresh.kamboju@linaro.org>
-dev_makeswap=-1
-dev_mounted=-1
-dev_start=0
-dev_end=-1
+# IDs returned by hot_add, in allocation order; old kernels use 0..dev_num-1.
+dev_ids=""
+dev_swap_ids=""
+dev_mount_ids=""
module_load=-1
sys_control=-1
# Kselftest framework requirement - SKIP code is 4.
@@ -44,32 +44,72 @@ kernel_gte()
return 1
}
+zram_wait_for_udev()
+{
+ # Probing triggered by device changes can still hold the device open.
+ # The queue is global; only the subsequent teardown can establish failure.
+ if command -v udevadm >/dev/null 2>&1; then
+ udevadm settle --timeout=5 ||
+ echo "udev queue did not settle; attempting cleanup" >&2
+ fi
+ return 0
+}
+
zram_cleanup()
{
echo "zram cleanup"
local i=
- for i in $(seq $dev_start $dev_makeswap); do
- swapoff /dev/zram$i
+ local ret=0
+ local busy_ids=""
+ for i in $dev_ids; do
+ case " $dev_swap_ids " in
+ *" $i "*) ;;
+ *)
+ # A signal can arrive after a helper activates swap but
+ # before its caller records the ID.
+ grep -q "^/dev/zram${i}[[:space:]]" /proc/swaps ||
+ continue
+ ;;
+ esac
+ if ! swapoff /dev/zram$i; then
+ ret=1
+ busy_ids="$busy_ids $i"
+ fi
done
- for i in $(seq $dev_start $dev_mounted); do
- umount /dev/zram$i
+ for i in $dev_mount_ids; do
+ if ! umount /dev/zram$i; then
+ ret=1
+ busy_ids="$busy_ids $i"
+ fi
done
- for i in $(seq $dev_start $dev_end); do
- echo 1 > /sys/block/zram${i}/reset
- rm -rf zram$i
+ zram_wait_for_udev
+ for i in $dev_ids; do
+ case " $busy_ids " in
+ *" $i "*) continue ;;
+ esac
+ echo 1 > /sys/block/zram${i}/reset || ret=1
+ case " $dev_mount_ids " in
+ *" $i "*) rmdir "zram$i" || ret=1 ;;
+ esac
done
+ # Reset emits another device-change event before removal.
+ zram_wait_for_udev
if [ $sys_control -eq 1 ]; then
- for i in $(seq $dev_start $dev_end); do
- echo $i > /sys/class/zram-control/hot_remove
+ for i in $dev_ids; do
+ case " $busy_ids " in
+ *" $i "*) continue ;;
+ esac
+ echo $i > /sys/class/zram-control/hot_remove || ret=1
done
fi
if [ $module_load -eq 1 ]; then
- rmmod zram > /dev/null 2>&1
+ rmmod zram || ret=1
fi
+ return "$ret"
}
zram_load()
@@ -80,15 +120,23 @@ zram_load()
if [ -d "/sys/class/zram-control" ]; then
echo "zram modules already loaded, kernel supports" \
"zram-control interface"
- dev_start=$(ls /dev/zram* | wc -w)
- dev_end=$(($dev_start + $dev_num - 1))
sys_control=1
- for i in $(seq $dev_start $dev_end); do
- cat /sys/class/zram-control/hot_add > /dev/null
+ for i in $(seq 1 $dev_num); do
+ if ! id=$(cat /sys/class/zram-control/hot_add); then
+ echo "FAIL zram hot_add failed" >&2
+ return 1
+ fi
+ case "$id" in
+ ''|*[!0-9]*)
+ echo "FAIL invalid zram hot_add ID: $id" >&2
+ return 1
+ ;;
+ esac
+ dev_ids="$dev_ids $id"
done
- echo "all zram devices (/dev/zram$dev_start~$dev_end" \
+ echo "all zram devices ($dev_ids)" \
"successfully created"
return 0
fi
@@ -112,8 +160,11 @@ zram_load()
fi
module_load=1
- dev_end=$(($dev_num - 1))
- echo "all zram devices (/dev/zram0~$dev_end) successfully created"
+ local last=$(($dev_num - 1))
+ for i in $(seq 0 $last); do
+ dev_ids="$dev_ids $i"
+ done
+ echo "all zram devices (/dev/zram0~$last) successfully created"
}
zram_max_streams()
@@ -127,8 +178,10 @@ zram_max_streams()
return 0
fi
- local i=$dev_start
+ set -- $dev_ids
for max_s in $zram_max_streams; do
+ local i=$1
+ shift
local sys_path="/sys/block/zram${i}/max_comp_streams"
echo $max_s > $sys_path || \
echo "FAIL failed to set '$max_s' to $sys_path"
@@ -138,7 +191,6 @@ zram_max_streams()
[ "$max_s" -ne "$max_streams" ] && \
echo "FAIL can't set max_streams '$max_s', get $max_stream"
- i=$(($i + 1))
echo "$sys_path = '$max_streams'"
done
@@ -149,15 +201,17 @@ zram_compress_alg()
{
echo "test that we can set compression algorithm"
- local i=$dev_start
+ set -- $dev_ids
+ local i=$1
local algs=$(cat /sys/block/zram${i}/comp_algorithm)
echo "supported algs: $algs"
for alg in $zram_algs; do
+ local i=$1
+ shift
local sys_path="/sys/block/zram${i}/comp_algorithm"
echo "$alg" > $sys_path || \
echo "FAIL can't set '$alg' to $sys_path"
- i=$(($i + 1))
echo "$sys_path = '$alg'"
done
@@ -167,13 +221,14 @@ zram_compress_alg()
zram_set_disksizes()
{
echo "set disk size to zram device(s)"
- local i=$dev_start
+ set -- $dev_ids
for ds in $zram_sizes; do
+ local i=$1
+ shift
local sys_path="/sys/block/zram${i}/disksize"
echo "$ds" > $sys_path || \
echo "FAIL can't set '$ds' to $sys_path"
- i=$(($i + 1))
echo "$sys_path = '$ds'"
done
@@ -184,13 +239,14 @@ zram_set_memlimit()
{
echo "set memory limit to zram device(s)"
- local i=$dev_start
+ set -- $dev_ids
for ds in $zram_mem_limits; do
+ local i=$1
+ shift
local sys_path="/sys/block/zram${i}/mem_limit"
echo "$ds" > $sys_path || \
echo "FAIL can't set '$ds' to $sys_path"
- i=$(($i + 1))
echo "$sys_path = '$ds'"
done
@@ -200,46 +256,59 @@ zram_set_memlimit()
zram_makeswap()
{
echo "make swap with zram device(s)"
- local i=$dev_start
- for i in $(seq $dev_start $dev_end); do
+ local i
+ local ret=0
+ for i in $dev_ids; do
mkswap /dev/zram$i > err.log 2>&1
if [ $? -ne 0 ]; then
cat err.log
- echo "FAIL mkswap /dev/zram$1 failed"
+ echo "FAIL mkswap /dev/zram$i failed"
+ ret=1
+ continue
fi
swapon /dev/zram$i > err.log 2>&1
if [ $? -ne 0 ]; then
cat err.log
- echo "FAIL swapon /dev/zram$1 failed"
+ echo "FAIL swapon /dev/zram$i failed"
+ ret=1
+ continue
fi
echo "done with /dev/zram$i"
- dev_makeswap=$i
+ dev_swap_ids="$dev_swap_ids $i"
done
- echo "zram making zram mkswap and swapon: OK"
+ [ "$ret" -eq 0 ] && echo "zram making zram mkswap and swapon: OK"
+ return "$ret"
}
zram_swapoff()
{
local i=
- for i in $(seq $dev_start $dev_end); do
+ local failed_ids=""
+ local ret=0
+ for i in $dev_swap_ids; do
swapoff /dev/zram$i > err.log 2>&1
if [ $? -ne 0 ]; then
cat err.log
echo "FAIL swapoff /dev/zram$i failed"
+ ret=1
+ failed_ids="$failed_ids $i"
fi
done
- dev_makeswap=-1
+ dev_swap_ids=$failed_ids
- echo "zram swapoff: OK"
+ [ "$ret" -eq 0 ] && echo "zram swapoff: OK"
+ return "$ret"
}
zram_makefs()
{
- local i=$dev_start
+ set -- $dev_ids
for fs in $zram_filesystems; do
+ local i=$1
+ shift
# if requested fs not supported default it to ext2
which mkfs.$fs > /dev/null 2>&1 || fs=ext2
@@ -249,7 +318,6 @@ zram_makefs()
cat err.log
echo "FAIL failed to make $fs on /dev/zram$i"
fi
- i=$(($i + 1))
echo "zram mkfs.$fs: OK"
done
}
@@ -257,13 +325,18 @@ zram_makefs()
zram_mount()
{
local i=0
- for i in $(seq $dev_start $dev_end); do
+ for i in $dev_ids; do
echo "mount /dev/zram$i"
- mkdir zram$i
- mount /dev/zram$i zram$i > /dev/null || \
+ mkdir "zram$i" || return 1
+ if mount /dev/zram$i "zram$i" > /dev/null; then
+ dev_mount_ids="$dev_mount_ids $i"
+ else
echo "FAIL mount /dev/zram$i failed"
- dev_mounted=$i
+ rmdir "zram$i" || return 1
+ return 1
+ fi
done
echo "zram mount of zram device(s): OK"
+ return 0
}
--
2.55.0
^ permalink raw reply [flat|nested] 16+ messages in thread
* [RFC PATCH v2 2/4] mm: restrict offload-only swap to proactive reclaim
2026-09-26 4:55 [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload Matthias Goergens
2026-09-26 4:55 ` [RFC PATCH v2 1/4] selftests: zram: track owned devices and report cleanup failures Matthias Goergens
@ 2026-09-26 4:55 ` Matthias Goergens
2026-09-26 4:55 ` [RFC PATCH v2 3/4] selftests: zram: cover offload-only swap policy Matthias Goergens
` (6 subsequent siblings)
8 siblings, 0 replies; 16+ messages in thread
From: Matthias Goergens @ 2026-09-26 4:55 UTC (permalink / raw)
To: Andrew Morton
Cc: Matthias Goergens, David Hildenbrand, Lorenzo Stoakes,
Liam R. Howlett, Vlastimil Babka, Mike Rapoport,
Suren Baghdasaryan, Michal Hocko, Jonathan Corbet, Shuah Khan,
Randy Dunlap, Jani Nikula, Joonas Lahtinen, Rodrigo Vivi,
Tvrtko Ursulin, David Airlie, Simona Vetter, Rob Clark,
Dmitry Baryshkov, Abhinav Kumar, Jessica Zhang, Sean Paul,
Marijn Suijten, Boris Brezillon, Steven Price, Liviu Dudau,
Maarten Lankhorst, Maxime Ripard, Thomas Zimmermann,
Christian Koenig, Huang Rui, Matthew Auld, Matthew Brost,
Thomas Hellström, Chris Li, Kairui Song, Kemeng Shi,
Nhat Pham, Baoquan He, Barry Song, Youngjun Park,
Johannes Weiner, Roman Gushchin, Shakeel Butt, Muchun Song,
Qi Zheng, Axel Rasmussen, Yuanchu Xie, Wei Xu, Baolin Wang,
linux-mm, linux-doc, linux-kernel, intel-gfx, dri-devel,
linux-arm-msm, freedreno, intel-xe, cgroups, linux-api,
linux-man, Rafael J . Wysocki, Pavel Machek, linux-pm,
Catalin Marinas, Will Deacon, linux-arm-kernel, Yosry Ahmed,
Chengming Zhou, Ryan Roberts, Alejandro Colomar, Karel Zak,
util-linux
Some swap backends need memory to accept writes, making them useful for
deliberate cold-page offload but unsuitable dependencies for reclaim under
acute memory pressure.
Add SWAP_FLAG_OFFLOAD_ONLY to reserve a swap area for explicitly admitted
reclaim. Cgroup v2 memory.reclaim, per-node reclaim and manual MGLRU
eviction establish admission; ordinary reclaim cannot initiate new,
non-zero backend writes to the area.
Track admission separately from scan_control.proactive and propagate it
through reclaim_state. The MGLRU debugfs interface is not stable ABI.
Filter swap allocation and reclaim capacity by eligibility, including
cached clusters and recovery of unused conventional swap-cache entries.
Preserve physical free-space reporting and workingset accounting.
Allocation filtering alone is insufficient: a partially swapped-in large
folio can retain its swap entry and reach writeout without allocating a
new slot. Refuse such ordinary-reclaim writes by redirtying and activating
the folio. Preserve architecture metadata before refusal, since sibling
faults can restore swap-indexed tags into the resident folio.
Bypass zswap and reject asynchronous page-cluster discard so later writes
cannot escape the admitting context. Exclude marked areas from hibernation
selection. Reads, swapoff and queued or in-flight I/O remain unaffected.
Refusing retained-entry writes can cause repeated reclaim or OOM; moving
these entries to conventional swap would require swap-entry migration.
Signed-off-by: Matthias Goergens <matthias.goergens@gmail.com>
---
Documentation/mm/swap.rst | 105 +++++++++++++++
drivers/gpu/drm/i915/gem/i915_gem_shrinker.c | 2 +-
.../gpu/drm/i915/gem/selftests/huge_pages.c | 6 +-
drivers/gpu/drm/msm/msm_gem_shrinker.c | 2 +-
drivers/gpu/drm/panthor/panthor_gem.c | 2 +-
drivers/gpu/drm/ttm/ttm_backup.c | 2 +-
drivers/gpu/drm/xe/tests/xe_bo.c | 2 +-
include/linux/swap.h | 29 ++++-
include/linux/vm_event_item.h | 1 +
mm/memcontrol.c | 20 ++-
mm/page_io.c | 24 +++-
mm/swapfile.c | 123 ++++++++++++++++--
mm/vmscan.c | 31 +++--
mm/vmstat.c | 1 +
14 files changed, 312 insertions(+), 38 deletions(-)
diff --git a/Documentation/mm/swap.rst b/Documentation/mm/swap.rst
index 78819bd4d745..9acad2c5834f 100644
--- a/Documentation/mm/swap.rst
+++ b/Documentation/mm/swap.rst
@@ -3,3 +3,108 @@
====
Swap
====
+
+Offload-only swap areas
+-----------------------
+
+Passing ``SWAP_FLAG_OFFLOAD_ONLY`` in the flags argument to ``swapon()``
+marks a swap area as eligible for explicit userspace proactive reclaim.
+The stable qualifying interfaces are cgroup v2
+``memory.reclaim`` and ``/sys/devices/system/node/nodeX/reclaim``. The swap
+allocator excludes such an area from kswapd, direct reclaim, and other
+pressure-driven swap allocation. Normal swap priority ordering still applies
+among the areas eligible for the current reclaim context. Proactive reclaim
+can use both conventional and offload-only areas; the flag does not force it
+to choose an offload-only area or reserve conventional capacity exclusively
+for pressure reclaim.
+
+The MGLRU debugfs eviction interface currently establishes the same internal
+proactive-reclaim provenance and can therefore use an offload-only area.
+Debugfs is not a stable userspace ABI, however, so that behaviour is not part
+of this interface's permanent contract.
+
+This permits a system to combine a small conventional swap area, which is
+engineered for forward progress in emergency reclaim, with a larger or more
+complex area used for ordinary cold-page offload. For example, the latter may
+be RAM-compressed or may use a filesystem with compression, checksums, or
+redundancy. Making every such write path safe in direct reclaim can require
+backend-specific reserves, preallocation, non-blocking allocation, and
+recursion rules. The flag restricts when writes may be initiated; backends
+still need to handle reads and complete previously admitted writes under
+memory pressure.
+
+For a RAM-compressed area such as zram, unused logical slots also do not imply
+that enough physical memory remains to store their future contents. Static
+swap priority cannot express that distinction or provide late fallback after a
+selected area's write fails.
+
+The policy is attached to an activated swap area, not to its underlying
+physical storage. A raw swap partition and a filesystem swapfile on the same
+device are separate areas and may use different policies. The kernel does not
+infer this policy from the block driver, filesystem, or swap priority.
+Changing an active area's policy requires swapoff followed by reactivation.
+
+``/proc/swaps`` does not expose the policy. Reported swap totals and free
+space include offload-only areas, so free swap space does not necessarily
+mean that pressure reclaim can allocate from it.
+
+This is a reclaim-provenance policy, not a measurement of current memory
+headroom. Userspace should only request proactive offload while its own
+watermark or PSI policy considers memory pressure low. Since
+``memory.reclaim`` can be delegated, that policy must also account for
+requests from delegated cgroups.
+
+DAMON reclaim and ``MADV_PAGEOUT`` do not currently establish the proactive
+reclaim context, so they cannot allocate slots from an offload-only area.
+Offload-only areas are also ineligible for hibernation image allocation.
+
+Offload-only areas bypass zswap stores. Zswap writeback may run after the
+proactive context which selected the slot has ended, so admitting the folio to
+zswap would otherwise defer the backend write beyond that context. The
+hierarchical cgroup v2 ``memory.zswap.writeback=0`` policy remains
+authoritative: when zswap is enabled, it also refuses direct proactive writes
+to an offload-only area. Marking an area offload-only does not override a
+cgroup policy which disables all swapping attempts to devices.
+
+The flag controls both allocation of new swap slots and newly initiated
+non-zero backend writes. A folio can retain its swap entry after swapin. If
+ordinary reclaim later tries to rewrite such an offload-only entry, the VM
+redirties and activates the folio instead; proactive reclaim may retry the
+write. Zero-filled folios may still update the in-memory swap zeromap without
+backend I/O.
+
+The flag does not prevent reads, swapoff, or writes which are already queued or
+in flight. It therefore does not by itself provide a forward progress
+guarantee for an I/O path which allocates memory: earlier writes must still be
+able to complete, and reads must remain reclaim-safe. Repeatedly refusing
+retained-entry writes can also reduce reclaim efficiency and lead to OOM while
+the dirty folios remain resident.
+
+Page-cluster discard is incompatible with an offload-only area because its
+work item can run after the context which freed the entries has ended. Swapon
+therefore rejects a resolved page-cluster discard policy combined with
+``SWAP_FLAG_OFFLOAD_ONLY``. Swapon-time discard is permitted because it
+completes synchronously during activation. The existing discard precedence
+still applies: requesting both discard-once and discard-pages selects
+discard-once. The bare ``SWAP_FLAG_DISCARD`` request enables page-cluster
+discard and is therefore rejected on a discard-capable offload-only area;
+add ``SWAP_FLAG_DISCARD_ONCE`` to ``SWAP_FLAG_DISCARD`` for synchronous
+activation-time discard.
+Discard requests which the swap area does not support remain ignored.
+
+Architecture-specific swap metadata preparation still runs before a retained
+write is refused, so that metadata remains coherent with the dirty resident
+folio. This policy controls swap-backend I/O; it does not promise that core VM
+or architecture preparation performs no allocation.
+
+With ``CONFIG_VM_EVENT_COUNTERS``, ``/proc/vmstat`` reports
+``swpout_offload_refused`` in base pages. The counter advances when ordinary
+reclaim refuses a newly initiated write through a retained offload-only
+entry. Repeated refusals of the same folio are counted again. It is not a
+count of skipped areas during new-slot allocation.
+
+An offload-only area should therefore be configured with a reclaim-safe swap
+area as fallback for new swap allocations. This does not migrate retained
+offload-only entries or retry a failed backend write on another area. If no
+eligible swap space remains, swap allocation fails and the existing reclaim
+and OOM policy applies.
diff --git a/drivers/gpu/drm/i915/gem/i915_gem_shrinker.c b/drivers/gpu/drm/i915/gem/i915_gem_shrinker.c
index e0d1f369a163..b58e61f15ab1 100644
--- a/drivers/gpu/drm/i915/gem/i915_gem_shrinker.c
+++ b/drivers/gpu/drm/i915/gem/i915_gem_shrinker.c
@@ -21,7 +21,7 @@
static bool swap_available(void)
{
- return get_nr_swap_pages() > 0;
+ return get_nr_swap_pages_eligible() > 0;
}
static bool can_release_pages(struct drm_i915_gem_object *obj)
diff --git a/drivers/gpu/drm/i915/gem/selftests/huge_pages.c b/drivers/gpu/drm/i915/gem/selftests/huge_pages.c
index 44718e728291..73b62b065510 100644
--- a/drivers/gpu/drm/i915/gem/selftests/huge_pages.c
+++ b/drivers/gpu/drm/i915/gem/selftests/huge_pages.c
@@ -1895,8 +1895,8 @@ static int igt_shrink_thp(void *arg)
i915_gem_context_unlock_engines(ctx);
/*
* Nuke everything *before* we unpin the pages so we can be reasonably
- * sure that when later checking get_nr_swap_pages() that some random
- * leftover object doesn't steal the remaining swap space.
+ * sure that when later checking get_nr_swap_pages_eligible() that some
+ * random leftover object doesn't steal the remaining swap space.
*/
i915_gem_shrink(NULL, i915, -1UL, NULL,
I915_SHRINK_BOUND |
@@ -1910,7 +1910,7 @@ static int igt_shrink_thp(void *arg)
* Now that the pages are *unpinned* shrinking should invoke
* shmem to truncate our pages, if we have available swap.
*/
- should_swap = get_nr_swap_pages() > 0;
+ should_swap = get_nr_swap_pages_eligible() > 0;
i915_gem_shrink(NULL, i915, -1UL, NULL,
I915_SHRINK_BOUND |
I915_SHRINK_UNBOUND |
diff --git a/drivers/gpu/drm/msm/msm_gem_shrinker.c b/drivers/gpu/drm/msm/msm_gem_shrinker.c
index 9d2788f79ace..da6a66b747b0 100644
--- a/drivers/gpu/drm/msm/msm_gem_shrinker.c
+++ b/drivers/gpu/drm/msm/msm_gem_shrinker.c
@@ -21,7 +21,7 @@ module_param(enable_eviction, bool, 0600);
static bool can_swap(void)
{
- return enable_eviction && get_nr_swap_pages() > 0;
+ return enable_eviction && get_nr_swap_pages_eligible() > 0;
}
static bool can_block(struct shrink_control *sc)
diff --git a/drivers/gpu/drm/panthor/panthor_gem.c b/drivers/gpu/drm/panthor/panthor_gem.c
index 72908be5e144..23488f417cf7 100644
--- a/drivers/gpu/drm/panthor/panthor_gem.c
+++ b/drivers/gpu/drm/panthor/panthor_gem.c
@@ -1371,7 +1371,7 @@ panthor_dummy_bo_create(struct panthor_device *ptdev)
static bool can_swap(void)
{
- return get_nr_swap_pages() > 0;
+ return get_nr_swap_pages_eligible() > 0;
}
static bool can_block(struct shrink_control *sc)
diff --git a/drivers/gpu/drm/ttm/ttm_backup.c b/drivers/gpu/drm/ttm/ttm_backup.c
index 0c2d53a13b2a..bf9d41bdf652 100644
--- a/drivers/gpu/drm/ttm/ttm_backup.c
+++ b/drivers/gpu/drm/ttm/ttm_backup.c
@@ -206,7 +206,7 @@ u64 ttm_backup_bytes_avail(void)
* number also depends on shmem actually swapping out backed-up
* shmem objects without too much buffering.
*/
- return (u64)get_nr_swap_pages() << PAGE_SHIFT;
+ return (u64)get_nr_swap_pages_eligible() << PAGE_SHIFT;
}
EXPORT_SYMBOL_GPL(ttm_backup_bytes_avail);
diff --git a/drivers/gpu/drm/xe/tests/xe_bo.c b/drivers/gpu/drm/xe/tests/xe_bo.c
index 6a17e13d58cf..14d6bb8e41c9 100644
--- a/drivers/gpu/drm/xe/tests/xe_bo.c
+++ b/drivers/gpu/drm/xe/tests/xe_bo.c
@@ -695,7 +695,7 @@ static int shrink_test_run_device(struct xe_device *xe)
}
to_alloc = ram * 2;
- ram_and_swap = ram + get_nr_swap_pages() * PAGE_SIZE;
+ ram_and_swap = ram + get_nr_swap_pages_eligible() * PAGE_SIZE;
if (to_alloc > ram_and_swap)
purgeable = to_alloc - ram_and_swap;
purgeable += div64_u64(purgeable, 5);
diff --git a/include/linux/swap.h b/include/linux/swap.h
index 78974da6810e..75350c450c47 100644
--- a/include/linux/swap.h
+++ b/include/linux/swap.h
@@ -21,10 +21,11 @@
#define SWAP_FLAG_DISCARD 0x10000 /* enable discard for swap */
#define SWAP_FLAG_DISCARD_ONCE 0x20000 /* discard swap area at swapon-time */
#define SWAP_FLAG_DISCARD_PAGES 0x40000 /* discard page-clusters after use */
+#define SWAP_FLAG_OFFLOAD_ONLY 0x80000 /* only use for proactive reclaim */
#define SWAP_FLAGS_VALID (SWAP_FLAG_PRIO_MASK | SWAP_FLAG_PREFER | \
SWAP_FLAG_DISCARD | SWAP_FLAG_DISCARD_ONCE | \
- SWAP_FLAG_DISCARD_PAGES)
+ SWAP_FLAG_DISCARD_PAGES | SWAP_FLAG_OFFLOAD_ONLY)
/*
* MAX_SWAPFILES defines the maximum number of swaptypes: things which can
* be swapped to. The swap type and the offset into that swap type are
@@ -140,12 +141,20 @@ union swap_header {
struct reclaim_state {
/* pages reclaimed outside of LRU-based reclaim */
unsigned long reclaimed;
+ /* this reclaim context may use offload-only swap */
+ bool allow_offload_swap;
#ifdef CONFIG_LRU_GEN
/* per-thread mm walk data */
struct lru_gen_mm_walk *mm_walk;
#endif
};
+static inline bool current_reclaim_allows_offload_swap(void)
+{
+ return current->reclaim_state &&
+ current->reclaim_state->allow_offload_swap;
+}
+
/*
* mm_account_reclaimed_pages(): account reclaimed pages outside of LRU-based
* reclaim
@@ -201,6 +210,7 @@ enum {
SWP_STABLE_WRITES = (1 << 11), /* no overwrite PG_writeback pages */
SWP_SYNCHRONOUS_IO = (1 << 12), /* synchronous IO is efficient */
SWP_HIBERNATION = (1 << 13), /* pinned for hibernation */
+ SWP_OFFLOAD_ONLY = (1 << 14), /* proactive-reclaim swap only */
/* add others here before... */
};
@@ -389,6 +399,9 @@ static inline long get_nr_swap_pages(void)
return atomic_long_read(&nr_swap_pages);
}
+long get_nr_swap_pages_eligible(void);
+bool folio_swap_full(struct folio *folio);
+
extern void si_swapinfo(struct sysinfo *);
extern int pin_hibernation_swap_type(dev_t device, sector_t offset);
extern void unpin_hibernation_swap_type(int type);
@@ -443,10 +456,16 @@ static inline void put_swap_device(struct swap_info_struct *si)
}
#define get_nr_swap_pages() 0L
+#define get_nr_swap_pages_eligible() 0L
#define total_swap_pages 0L
#define total_swapcache_pages() 0UL
#define vm_swap_full() 0
+static inline bool folio_swap_full(struct folio *folio)
+{
+ return false;
+}
+
#define si_swapinfo(val) \
do { (val)->freeswap = (val)->totalswap = 0; } while (0)
#define free_folio_and_swap_cache(folio) \
@@ -531,6 +550,7 @@ static inline void mem_cgroup_uncharge_swap(unsigned short id, unsigned int nr_p
long mem_cgroup_get_folio_swap_margin(struct folio *folio);
extern long mem_cgroup_get_nr_swap_pages(struct mem_cgroup *memcg);
+long mem_cgroup_get_nr_swap_pages_eligible(struct mem_cgroup *memcg);
extern bool mem_cgroup_swap_full(struct folio *folio);
#else
static inline int mem_cgroup_try_charge_swap(struct folio *folio)
@@ -553,9 +573,14 @@ static inline long mem_cgroup_get_nr_swap_pages(struct mem_cgroup *memcg)
return get_nr_swap_pages();
}
+static inline long mem_cgroup_get_nr_swap_pages_eligible(struct mem_cgroup *memcg)
+{
+ return get_nr_swap_pages_eligible();
+}
+
static inline bool mem_cgroup_swap_full(struct folio *folio)
{
- return vm_swap_full();
+ return folio_swap_full(folio);
}
#endif
diff --git a/include/linux/vm_event_item.h b/include/linux/vm_event_item.h
index 2628ccda076a..fc8458b314a5 100644
--- a/include/linux/vm_event_item.h
+++ b/include/linux/vm_event_item.h
@@ -32,6 +32,7 @@
HIGHMEM_ZONE(xx) xx##_MOVABLE, DEVICE_ZONE(xx)
enum vm_event_item { PGPGIN, PGPGOUT, PSWPIN, PSWPOUT,
+ SWPOUT_OFFLOAD_REFUSED,
FOR_ALL_ZONES(PGALLOC)
FOR_ALL_ZONES(ALLOCSTALL)
FOR_ALL_ZONES(PGSCAN_SKIP)
diff --git a/mm/memcontrol.c b/mm/memcontrol.c
index 1460cba53588..1412084d2f43 100644
--- a/mm/memcontrol.c
+++ b/mm/memcontrol.c
@@ -6002,16 +6002,28 @@ void __mem_cgroup_uncharge_swap(unsigned short id, unsigned int nr_pages)
rcu_read_unlock();
}
-long mem_cgroup_get_nr_swap_pages(struct mem_cgroup *memcg)
+static long
+mem_cgroup_get_nr_swap_pages_with_limit(struct mem_cgroup *memcg,
+ long nr_swap_pages)
{
- long nr_swap_pages = get_nr_swap_pages();
-
if (!mem_cgroup_disabled() && !do_memsw_account())
nr_swap_pages = min(nr_swap_pages, page_counter_margin(&memcg->swap));
return nr_swap_pages;
}
+long mem_cgroup_get_nr_swap_pages(struct mem_cgroup *memcg)
+{
+ return mem_cgroup_get_nr_swap_pages_with_limit(memcg,
+ get_nr_swap_pages());
+}
+
+long mem_cgroup_get_nr_swap_pages_eligible(struct mem_cgroup *memcg)
+{
+ return mem_cgroup_get_nr_swap_pages_with_limit(memcg,
+ get_nr_swap_pages_eligible());
+}
+
/**
* mem_cgroup_get_folio_swap_margin - get a folio's memcg swap margin
* @folio: folio whose memcg margin is queried
@@ -6042,7 +6054,7 @@ bool mem_cgroup_swap_full(struct folio *folio)
VM_BUG_ON_FOLIO(!folio_test_locked(folio), folio);
- if (vm_swap_full())
+ if (folio_swap_full(folio))
return true;
if (do_memsw_account() || !folio_memcg_charged(folio))
return ret;
diff --git a/mm/page_io.c b/mm/page_io.c
index 1da4ff484f09..1b281a1d6df0 100644
--- a/mm/page_io.c
+++ b/mm/page_io.c
@@ -203,6 +203,7 @@ static void swap_zeromap_folio_clear(struct folio *folio)
*/
int swap_writeout(struct swap_io_ctx *ctx, struct folio *folio)
{
+ struct swap_info_struct *sis = __swap_entry_to_info(folio->swap);
int ret = 0;
if (folio_free_swap(folio))
@@ -210,7 +211,9 @@ int swap_writeout(struct swap_io_ctx *ctx, struct folio *folio)
/*
* Arch code may have to preserve more data than just the folio
- * contents, e.g. memory tags.
+ * contents, e.g. memory tags. Do this before refusing a retained
+ * offload-only entry below: a later sibling swap-PTE fault can restore
+ * swap-indexed metadata into this resident folio.
*/
ret = arch_prepare_to_swap(folio);
if (ret) {
@@ -228,6 +231,18 @@ int swap_writeout(struct swap_io_ctx *ctx, struct folio *folio)
goto out_unlock;
}
+ /*
+ * A folio can retain an existing swap entry after swapin. Do not let
+ * ordinary reclaim use an offload-only entry through that path.
+ */
+ if ((READ_ONCE(sis->flags) & SWP_OFFLOAD_ONLY) &&
+ !current_reclaim_allows_offload_swap()) {
+ count_vm_events(SWPOUT_OFFLOAD_REFUSED,
+ folio_nr_pages(folio));
+ folio_mark_dirty(folio);
+ return AOP_WRITEPAGE_ACTIVATE;
+ }
+
/*
* Clear bits this folio occupies in the zeromap to prevent zero data
* being read in from any previous zero writes that occupied the same
@@ -235,7 +250,12 @@ int swap_writeout(struct swap_io_ctx *ctx, struct folio *folio)
*/
swap_zeromap_folio_clear(folio);
- if (zswap_store(folio)) {
+ /*
+ * Zswap writeback can happen much later from pressure reclaim or its
+ * shrinker workqueue. Do not let it defer an offload-only backend write
+ * beyond the proactive reclaim context which admitted the swap slot.
+ */
+ if (!(READ_ONCE(sis->flags) & SWP_OFFLOAD_ONLY) && zswap_store(folio)) {
count_mthp_stat(folio_order(folio), MTHP_STAT_ZSWPOUT);
goto out_unlock;
}
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 280dd906eb18..b72d598a9973 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -65,13 +65,17 @@ static void move_cluster(struct swap_info_struct *si,
*/
static DEFINE_SPINLOCK(swap_lock);
static unsigned int nr_swapfiles;
-atomic_long_t nr_swap_pages;
/*
* Some modules use swappable objects and may try to swap them out under
* memory pressure (via the shrinker). Before doing so, they may wish to
* check to see if any swap space is available.
+ *
+ * This remains the raw free-space counter for accounting users. Reclaim
+ * decisions subtract nr_swap_pages_offload_only below when necessary.
*/
+atomic_long_t nr_swap_pages;
EXPORT_SYMBOL_GPL(nr_swap_pages);
+static atomic_long_t nr_swap_pages_offload_only;
/* protected with swap_lock. reading in vm_swap_full() doesn't need lock */
long total_swap_pages;
#define DEF_SWAP_PRIO -1
@@ -120,6 +124,7 @@ atomic_t nr_rotate_swap = ATOMIC_INIT(0);
struct percpu_swap_cluster {
struct swap_info_struct *si[SWAP_NR_ORDERS];
unsigned long offset[SWAP_NR_ORDERS];
+ bool allow_offload_swap[SWAP_NR_ORDERS];
local_lock_t lock;
};
@@ -163,6 +168,31 @@ static long swap_usage_in_pages(struct swap_info_struct *si)
return atomic_long_read(&si->inuse_pages) & SWAP_USAGE_COUNTER_MASK;
}
+static bool swap_area_needs_reclaim(struct swap_info_struct *si)
+{
+ if (vm_swap_full())
+ return true;
+
+ /*
+ * Free offload-only slots must not keep a full conventional area
+ * pinned in swapcache. Recover that area's unused cache entries
+ * even when the raw pool is not full. This is independent of the
+ * current task: proactive reclaim can fill the conventional area too.
+ */
+ return !(READ_ONCE(si->flags) & SWP_OFFLOAD_ONLY) &&
+ atomic_long_read(&nr_swap_pages_offload_only) > 0 &&
+ swap_usage_in_pages(si) == si->pages;
+}
+
+/* The caller must hold the lock on a folio in swapcache. */
+bool folio_swap_full(struct folio *folio)
+{
+ VM_WARN_ON_FOLIO(!folio_test_locked(folio), folio);
+ VM_WARN_ON_FOLIO(!folio_test_swapcache(folio), folio);
+
+ return swap_area_needs_reclaim(__swap_entry_to_info(folio->swap));
+}
+
/* Reclaim the swap entry anyway if possible */
#define TTRS_ANYWAY 0x1
/*
@@ -903,7 +933,7 @@ static bool cluster_scan_range(struct swap_info_struct *si,
if (swp_tb_is_null(swp_tb))
continue;
if (swp_tb_is_folio(swp_tb) && !__swp_tb_get_count(swp_tb)) {
- if (!vm_swap_full())
+ if (!swap_area_needs_reclaim(si))
return false;
*need_reclaim = true;
continue;
@@ -1014,6 +1044,8 @@ static unsigned int alloc_swap_scan_cluster(struct swap_info_struct *si,
if (si->flags & SWP_SOLIDSTATE) {
this_cpu_write(percpu_swap_cluster.offset[order], next);
this_cpu_write(percpu_swap_cluster.si[order], si);
+ this_cpu_write(percpu_swap_cluster.allow_offload_swap[order],
+ current_reclaim_allows_offload_swap());
} else {
si->global_cluster->next[order] = next;
}
@@ -1161,7 +1193,7 @@ static unsigned long cluster_alloc_swap_entry(struct swap_info_struct *si,
}
/* Try reclaim full clusters if free and nonfull lists are drained */
- if (vm_swap_full())
+ if (swap_area_needs_reclaim(si))
swap_reclaim_full_clusters(si, false);
if (order < PMD_ORDER) {
@@ -1315,9 +1347,11 @@ static void swap_range_alloc(struct swap_info_struct *si,
unsigned int nr_entries)
{
if (swap_usage_add(si, nr_entries)) {
- if (vm_swap_full())
+ if (swap_area_needs_reclaim(si))
schedule_work(&si->reclaim_work);
}
+ if (si->flags & SWP_OFFLOAD_ONLY)
+ atomic_long_sub(nr_entries, &nr_swap_pages_offload_only);
atomic_long_sub(nr_entries, &nr_swap_pages);
}
@@ -1346,6 +1380,8 @@ static void swap_range_free(struct swap_info_struct *si, unsigned long offset,
* only after the above cleanups are done.
*/
smp_wmb();
+ if (si->flags & SWP_OFFLOAD_ONLY)
+ atomic_long_add(nr_entries, &nr_swap_pages_offload_only);
atomic_long_add(nr_entries, &nr_swap_pages);
swap_usage_sub(si, nr_entries);
}
@@ -1366,6 +1402,31 @@ static bool get_swap_device_info(struct swap_info_struct *si)
return true;
}
+static bool swap_area_eligible(struct swap_info_struct *si)
+{
+ if (!(READ_ONCE(si->flags) & SWP_OFFLOAD_ONLY))
+ return true;
+
+ return current_reclaim_allows_offload_swap();
+}
+
+long get_nr_swap_pages_eligible(void)
+{
+ long nr_pages;
+
+ if (current_reclaim_allows_offload_swap())
+ return get_nr_swap_pages();
+
+ /*
+ * The reads are intentionally unpaired. This is a capacity hint; the
+ * allocator enforces eligibility. Clamp a transient negative result.
+ */
+ nr_pages = get_nr_swap_pages() -
+ atomic_long_read(&nr_swap_pages_offload_only);
+ return max(nr_pages, 0L);
+}
+EXPORT_SYMBOL_GPL(get_nr_swap_pages_eligible);
+
/*
* Fast path try to get swap entries with specified order from current
* CPU's swap entry pool (a cluster).
@@ -1385,6 +1446,20 @@ static bool swap_alloc_fast(struct folio *folio)
offset = this_cpu_read(percpu_swap_cluster.offset[order]);
if (!si || !offset || !get_swap_device_info(si))
return false;
+ if (!swap_area_eligible(si)) {
+ put_swap_device(si);
+ return false;
+ }
+ /*
+ * Pressure reclaim may cache a lower-priority conventional area while
+ * an offload-only area is ineligible. Drop that cache on a context
+ * change so proactive reclaim returns to the normal priority search.
+ */
+ if (this_cpu_read(percpu_swap_cluster.allow_offload_swap[order]) !=
+ current_reclaim_allows_offload_swap()) {
+ put_swap_device(si);
+ return false;
+ }
ci = swap_cluster_lock(si, offset);
if (cluster_is_usable(ci, order)) {
@@ -1407,6 +1482,9 @@ static void swap_alloc_slow(struct folio *folio)
spin_lock(&swap_avail_lock);
start_over:
plist_for_each_entry_safe(si, next, &swap_avail_head, avail_list) {
+ if (!swap_area_eligible(si))
+ continue;
+
/* Rotate the device and switch to a new cluster */
plist_requeue(&si->avail_list, &swap_avail_head);
spin_unlock(&swap_avail_lock);
@@ -1738,8 +1816,8 @@ static int swap_dup_entries_cluster(struct swap_info_struct *si,
*
* Context: Caller needs to hold the folio lock.
* Return: %0 on success, %-E2BIG if splitting the folio might allow swapout,
- * %-ENOSPC if no global swap space is available, or %-ENOMEM if splitting
- * would not help.
+ * %-ENOSPC if no global swap space is eligible for the caller, or %-ENOMEM
+ * if splitting would not help.
*/
int folio_alloc_swap(struct folio *folio)
{
@@ -1790,7 +1868,7 @@ int folio_alloc_swap(struct folio *folio)
return 0;
failed:
- if (get_nr_swap_pages() <= 0)
+ if (get_nr_swap_pages_eligible() <= 0)
return -ENOSPC;
if (mem_cgroup_get_folio_swap_margin(folio) <= 0)
return -ENOMEM;
@@ -2180,7 +2258,7 @@ swp_entry_t swap_alloc_hibernation_slot(int type)
struct swap_cluster_info *ci;
swp_entry_t entry = {0};
- if (!si)
+ if (!si || (si->flags & SWP_OFFLOAD_ONLY))
goto fail;
/*
@@ -2247,7 +2325,8 @@ static int __find_hibernation_swap_type(dev_t device, sector_t offset)
for (type = 0; type < nr_swapfiles; type++) {
struct swap_info_struct *sis = swap_info[type];
- if (!(sis->flags & SWP_WRITEOK))
+ if (!(sis->flags & SWP_WRITEOK) ||
+ (sis->flags & SWP_OFFLOAD_ONLY))
continue;
if (device == sis->bdev->bd_dev) {
@@ -2434,7 +2513,8 @@ int find_first_swap(dev_t *device)
for (type = 0; type < nr_swapfiles; type++) {
struct swap_info_struct *sis = swap_info[type];
- if (!(sis->flags & SWP_WRITEOK))
+ if (!(sis->flags & SWP_WRITEOK) ||
+ (sis->flags & SWP_OFFLOAD_ONLY))
continue;
*device = sis->bdev->bd_dev;
spin_unlock(&swap_lock);
@@ -2474,7 +2554,8 @@ unsigned int count_swap_pages(int type, int free)
struct swap_info_struct *sis = swap_info[type];
spin_lock(&sis->lock);
- if (sis->flags & SWP_WRITEOK) {
+ if ((sis->flags & SWP_WRITEOK) &&
+ !(sis->flags & SWP_OFFLOAD_ONLY)) {
n = sis->pages;
if (free)
n -= swap_usage_in_pages(sis);
@@ -3083,6 +3164,8 @@ static int setup_swap_extents(struct swap_info_struct *sis,
static void _enable_swap_info(struct swap_info_struct *si)
{
+ if (si->flags & SWP_OFFLOAD_ONLY)
+ atomic_long_add(si->pages, &nr_swap_pages_offload_only);
atomic_long_add(si->pages, &nr_swap_pages);
total_swap_pages += si->pages;
@@ -3231,6 +3314,8 @@ SYSCALL_DEFINE1(swapoff, const char __user *, specialfile)
spin_lock(&p->lock);
del_from_avail_list(p, true);
plist_del(&p->list, &swap_active_head);
+ if (p->flags & SWP_OFFLOAD_ONLY)
+ atomic_long_sub(p->pages, &nr_swap_pages_offload_only);
atomic_long_sub(p->pages, &nr_swap_pages);
total_swap_pages -= p->pages;
spin_unlock(&p->lock);
@@ -3727,7 +3812,6 @@ SYSCALL_DEFINE2(swapon, const char __user *, specialfile, int, swap_flags)
if (swap_flags & ~SWAP_FLAGS_VALID)
return -EINVAL;
-
if (!capable(CAP_SYS_ADMIN))
return -EPERM;
@@ -3841,6 +3925,9 @@ SYSCALL_DEFINE2(swapon, const char __user *, specialfile, int, swap_flags)
if (error)
goto bad_swap_unlock_inode;
+ if (swap_flags & SWAP_FLAG_OFFLOAD_ONLY)
+ si->flags |= SWP_OFFLOAD_ONLY;
+
if ((swap_flags & SWAP_FLAG_DISCARD) &&
si->bdev && bdev_max_discard_sectors(si->bdev)) {
/*
@@ -3863,6 +3950,18 @@ SYSCALL_DEFINE2(swapon, const char __user *, specialfile, int, swap_flags)
else if (swap_flags & SWAP_FLAG_DISCARD_PAGES)
si->flags &= ~SWP_AREA_DISCARD;
+ /*
+ * Cluster discard can run later from discard_work, after the
+ * context which freed the entries has ended. Swapon-time discard
+ * is explicit and synchronous, but page discard cannot honour
+ * offload provenance.
+ */
+ if ((si->flags & SWP_OFFLOAD_ONLY) &&
+ (si->flags & SWP_PAGE_DISCARD)) {
+ error = -EINVAL;
+ goto bad_swap_unlock_inode;
+ }
+
/* issue a swapon-time discard if it's still required */
if (si->flags & SWP_AREA_DISCARD) {
int err = discard_swap(si);
diff --git a/mm/vmscan.c b/mm/vmscan.c
index aaceed4759ee..0633feb5d88b 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -123,6 +123,9 @@ struct scan_control {
/* Proactive reclaim invoked by userspace */
unsigned int proactive:1;
+ /* This reclaim context may use offload-only swap */
+ unsigned int allow_offload_swap:1;
+
/*
* Cgroup memory below memory.low is protected as long as we
* don't threaten to OOM. If any cgroup is reclaimed at
@@ -291,14 +294,18 @@ static inline bool is_exec_file_folio(const struct folio *folio,
}
static void set_task_reclaim_state(struct task_struct *task,
- struct reclaim_state *rs)
+ struct scan_control *sc)
{
+ struct reclaim_state *rs = sc ? &sc->reclaim_state : NULL;
+
/* Check for an overwrite */
WARN_ON_ONCE(rs && task->reclaim_state);
/* Check for the nulling of an already-nulled member */
WARN_ON_ONCE(!rs && !task->reclaim_state);
+ if (rs)
+ rs->allow_offload_swap = sc->allow_offload_swap;
task->reclaim_state = rs;
}
@@ -418,7 +425,7 @@ static inline bool can_reclaim_anon_pages(struct mem_cgroup *memcg,
* And under GFP_NOIO, is there enough swapcached anon to make
* scanning anon worthwhile?
*/
- if (get_nr_swap_pages() > 0 &&
+ if (get_nr_swap_pages_eligible() > 0 &&
!reclaimable_anon_is_low(memcg, nid, sc))
return true;
} else {
@@ -426,7 +433,7 @@ static inline bool can_reclaim_anon_pages(struct mem_cgroup *memcg,
* Is the memcg below its swap limit, and under GFP_NOIO does
* it have enough swapcached anon to make scanning worthwhile?
*/
- if (mem_cgroup_get_nr_swap_pages(memcg) > 0 &&
+ if (mem_cgroup_get_nr_swap_pages_eligible(memcg) > 0 &&
!reclaimable_anon_is_low(memcg, nid, sc))
return true;
}
@@ -2853,7 +2860,7 @@ static int get_swappiness(struct lruvec *lruvec, struct scan_control *sc)
return 0;
if (!can_demote(pgdat->node_id, sc, memcg) &&
- mem_cgroup_get_nr_swap_pages(memcg) < MIN_LRU_BATCH)
+ mem_cgroup_get_nr_swap_pages_eligible(memcg) < MIN_LRU_BATCH)
return 0;
return swappiness;
@@ -5926,6 +5933,7 @@ static ssize_t lru_gen_seq_write(struct file *file, const char __user *src,
.reclaim_idx = MAX_NR_ZONES - 1,
.gfp_mask = GFP_KERNEL,
.proactive = true,
+ .allow_offload_swap = true,
};
buf = kvmalloc(len + 1, GFP_KERNEL);
@@ -5937,7 +5945,7 @@ static ssize_t lru_gen_seq_write(struct file *file, const char __user *src,
return -EFAULT;
}
- set_task_reclaim_state(current, &sc.reclaim_state);
+ set_task_reclaim_state(current, &sc);
flags = memalloc_noreclaim_save();
blk_start_plug(&plug);
if (!set_mm_walk(NULL, true)) {
@@ -6941,7 +6949,7 @@ unsigned long try_to_free_pages(struct zonelist *zonelist, int order,
if (throttle_direct_reclaim(sc.gfp_mask, zonelist, nodemask))
return 1;
- set_task_reclaim_state(current, &sc.reclaim_state);
+ set_task_reclaim_state(current, &sc);
trace_mm_vmscan_direct_reclaim_begin(sc.gfp_mask, order, NULL);
nr_reclaimed = do_try_to_free_pages(zonelist, &sc);
@@ -6974,6 +6982,8 @@ unsigned long try_to_free_mem_cgroup_pages(struct mem_cgroup *memcg,
.may_unmap = 1,
.may_swap = !!(reclaim_options & MEMCG_RECLAIM_MAY_SWAP),
.proactive = !!(reclaim_options & MEMCG_RECLAIM_PROACTIVE),
+ .allow_offload_swap =
+ !!(reclaim_options & MEMCG_RECLAIM_PROACTIVE),
};
/*
* Traverse the ZONELIST_FALLBACK zonelist of the current node to put
@@ -6982,7 +6992,7 @@ unsigned long try_to_free_mem_cgroup_pages(struct mem_cgroup *memcg,
*/
struct zonelist *zonelist = node_zonelist(numa_node_id(), sc.gfp_mask);
- set_task_reclaim_state(current, &sc.reclaim_state);
+ set_task_reclaim_state(current, &sc);
trace_mm_vmscan_memcg_reclaim_begin(sc.gfp_mask, 0, memcg);
noreclaim_flag = memalloc_noreclaim_save();
@@ -7272,7 +7282,7 @@ static int balance_pgdat(pg_data_t *pgdat, int order, int highest_zoneidx)
trace_mm_vmscan_balance_pgdat_begin(pgdat->node_id, order,
highest_zoneidx);
- set_task_reclaim_state(current, &sc.reclaim_state);
+ set_task_reclaim_state(current, &sc);
psi_memstall_enter(&pflags);
__fs_reclaim_acquire(_THIS_IP_);
@@ -7769,7 +7779,7 @@ unsigned long shrink_all_memory(unsigned long nr_to_reclaim)
fs_reclaim_acquire(sc.gfp_mask);
noreclaim_flag = memalloc_noreclaim_save();
- set_task_reclaim_state(current, &sc.reclaim_state);
+ set_task_reclaim_state(current, &sc);
nr_reclaimed = do_try_to_free_pages(zonelist, &sc);
@@ -7950,7 +7960,7 @@ static unsigned long __node_reclaim(struct pglist_data *pgdat,
* We need to be able to allocate from the reserves for RECLAIM_UNMAP
*/
noreclaim_flag = memalloc_noreclaim_save();
- set_task_reclaim_state(p, &sc->reclaim_state);
+ set_task_reclaim_state(p, sc);
do {
shrink_node(pgdat, sc);
@@ -8136,6 +8146,7 @@ int user_proactive_reclaim(char *buf,
.may_unmap = 1,
.may_swap = 1,
.proactive = 1,
+ .allow_offload_swap = 1,
};
if (test_and_set_bit_lock(PGDAT_RECLAIM_LOCKED,
diff --git a/mm/vmstat.c b/mm/vmstat.c
index a3e809c57f29..b924c715886e 100644
--- a/mm/vmstat.c
+++ b/mm/vmstat.c
@@ -1331,6 +1331,7 @@ const char * const vmstat_text[] = {
[I(PGPGOUT)] = "pgpgout",
[I(PSWPIN)] = "pswpin",
[I(PSWPOUT)] = "pswpout",
+ [I(SWPOUT_OFFLOAD_REFUSED)] = "swpout_offload_refused",
#define OFF (NR_VM_ZONE_STAT_ITEMS + NR_VM_NUMA_EVENT_ITEMS + \
NR_VM_NODE_STAT_ITEMS + NR_VM_STAT_ITEMS)
--
2.55.0
^ permalink raw reply [flat|nested] 16+ messages in thread
* [RFC PATCH v2 3/4] selftests: zram: cover offload-only swap policy
2026-09-26 4:55 [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload Matthias Goergens
2026-09-26 4:55 ` [RFC PATCH v2 1/4] selftests: zram: track owned devices and report cleanup failures Matthias Goergens
2026-09-26 4:55 ` [RFC PATCH v2 2/4] mm: restrict offload-only swap to proactive reclaim Matthias Goergens
@ 2026-09-26 4:55 ` Matthias Goergens
2026-09-26 4:55 ` [RFC PATCH v2 4/4] selftests: zram: cover retained offload-only entries Matthias Goergens
` (5 subsequent siblings)
8 siblings, 0 replies; 16+ messages in thread
From: Matthias Goergens @ 2026-09-26 4:55 UTC (permalink / raw)
To: Shuah Khan
Cc: Matthias Goergens, Minchan Kim, Sergey Senozhatsky, linux-kernel,
linux-kselftest, Chris Li, Kairui Song, Johannes Weiner,
Yosry Ahmed, Nhat Pham, linux-mm, Andrew Morton
Add zram tests for proactive offload, fallback to conventional swap under
ordinary pressure, zswap bypass, discard policy and data recovery after
swapoff. Pin reclaim transitions to one CPU to exercise cached-cluster
eligibility changes, and use conventional swap as a positive zswap control.
Check that discard-once takes precedence when both discard modes are set.
Check that offload-only capacity preserves file workingset activation
with a calibrated refault workload. Wait for memory.stat to reflect the
initial footprint before reclaim, and distinguish insufficient calibration
from helper failures.
Skip unsupported kernels and cgroup setups. Restore swap devices,
controller delegation and modified zswap/MGLRU settings on exit. These
tests require an exclusive environment because swap priorities and some
of the settings are global.
Signed-off-by: Matthias Goergens <matthias.goergens@gmail.com>
---
tools/testing/selftests/zram/.gitignore | 2 +
tools/testing/selftests/zram/Makefile | 4 +-
tools/testing/selftests/zram/README | 14 +-
tools/testing/selftests/zram/config | 6 +-
tools/testing/selftests/zram/swap_offload.c | 222 ++++++++++++++++++
.../selftests/zram/workingset_offload.c | 207 ++++++++++++++++
tools/testing/selftests/zram/zram03.sh | 174 ++++++++++++++
tools/testing/selftests/zram/zram04.sh | 164 +++++++++++++
tools/testing/selftests/zram/zram_lib.sh | 28 +++
9 files changed, 817 insertions(+), 4 deletions(-)
create mode 100644 tools/testing/selftests/zram/swap_offload.c
create mode 100644 tools/testing/selftests/zram/workingset_offload.c
create mode 100755 tools/testing/selftests/zram/zram03.sh
create mode 100755 tools/testing/selftests/zram/zram04.sh
diff --git a/tools/testing/selftests/zram/.gitignore b/tools/testing/selftests/zram/.gitignore
index 088cd9bad87a..74b0217c60f0 100644
--- a/tools/testing/selftests/zram/.gitignore
+++ b/tools/testing/selftests/zram/.gitignore
@@ -1,2 +1,4 @@
# SPDX-License-Identifier: GPL-2.0-only
err.log
+swap_offload
+workingset_offload
diff --git a/tools/testing/selftests/zram/Makefile b/tools/testing/selftests/zram/Makefile
index 7f78eb1b59cb..781d10a20e56 100644
--- a/tools/testing/selftests/zram/Makefile
+++ b/tools/testing/selftests/zram/Makefile
@@ -1,9 +1,9 @@
# SPDX-License-Identifier: GPL-2.0
all:
-TEST_PROGS := zram.sh
+TEST_GEN_FILES := swap_offload workingset_offload
+TEST_PROGS := zram.sh zram03.sh zram04.sh
TEST_FILES := zram01.sh zram02.sh zram_lib.sh
EXTRA_CLEAN := err.log
include ../lib.mk
-
diff --git a/tools/testing/selftests/zram/README b/tools/testing/selftests/zram/README
index 82921c75681c..cd7f389ed3c6 100644
--- a/tools/testing/selftests/zram/README
+++ b/tools/testing/selftests/zram/README
@@ -21,10 +21,22 @@ ZRAM Testcases
zram_lib.sh: create library with initialization/cleanup functions
zram.sh: For sanity check of CONFIG_ZRAM and to run zram01 and zram02
-Two functional tests: zram01 and zram02:
+The legacy runner executes two functional tests:
zram01.sh: creates general purpose ram disks with ext4 filesystems
zram02.sh: creates block device for swap
+Offload-only swap tests, registered separately:
+zram03.sh: checks proactive routing, pressure fallback and zswap bypass
+zram04.sh: checks file workingset activation with offload-only swap
+
+Run these tests as root in an exclusive disposable VM with the offload-only
+swap policy and the options listed in config. They change global swap,
+zswap and reclaim settings. Use the initial cgroup namespace with an
+unrestricted cgroup v2 memory hierarchy mounted at /sys/fs/cgroup. A child's
+memory.zswap.writeback value does not reveal restrictions in its ancestors.
+
+zram04 needs a disk-backed TMPDIR.
+
Commands required for testing:
- bc
- dd
diff --git a/tools/testing/selftests/zram/config b/tools/testing/selftests/zram/config
index e0cc47e2c7e2..c59b8c3806a5 100644
--- a/tools/testing/selftests/zram/config
+++ b/tools/testing/selftests/zram/config
@@ -1,2 +1,6 @@
+CONFIG_CGROUPS=y
+CONFIG_MEMCG=y
+CONFIG_SWAP=y
CONFIG_ZSMALLOC=y
-CONFIG_ZRAM=m
+CONFIG_ZRAM=y
+CONFIG_ZSWAP=y
diff --git a/tools/testing/selftests/zram/swap_offload.c b/tools/testing/selftests/zram/swap_offload.c
new file mode 100644
index 000000000000..b2b94cd6ee3a
--- /dev/null
+++ b/tools/testing/selftests/zram/swap_offload.c
@@ -0,0 +1,222 @@
+// SPDX-License-Identifier: GPL-2.0
+#define _GNU_SOURCE
+
+#include <errno.h>
+#include <fcntl.h>
+#include <sched.h>
+#include <signal.h>
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <sys/mman.h>
+#include <sys/syscall.h>
+#include <unistd.h>
+
+#define SWAP_FLAG_PREFER 0x8000
+#define SWAP_FLAG_DISCARD 0x10000
+#define SWAP_FLAG_DISCARD_ONCE 0x20000
+#define SWAP_FLAG_DISCARD_PAGES 0x40000
+#define SWAP_FLAG_OFFLOAD_ONLY 0x80000
+
+static int activate(const char *path, int priority, int discard_flags)
+{
+ int flags = SWAP_FLAG_PREFER | SWAP_FLAG_OFFLOAD_ONLY |
+ priority | discard_flags;
+
+ if (syscall(SYS_swapon, path, flags)) {
+ perror("swapon");
+ return 1;
+ }
+
+ return 0;
+}
+
+static int reject_page_discard(const char *path, int priority)
+{
+ int flags = SWAP_FLAG_PREFER | SWAP_FLAG_OFFLOAD_ONLY |
+ SWAP_FLAG_DISCARD | SWAP_FLAG_DISCARD_PAGES | priority;
+ int ret;
+
+ errno = 0;
+ ret = syscall(SYS_swapon, path, flags);
+ if (ret == -1 && errno == EINVAL)
+ return 0;
+ if (!ret) {
+ syscall(SYS_swapoff, path);
+ fprintf(stderr, "offload-only page discard was accepted\n");
+ } else {
+ fprintf(stderr, "swapon returned unexpected error: %s\n",
+ strerror(errno));
+ }
+ return 1;
+}
+
+static int accept_discard_once_pages(const char *path, int priority)
+{
+ int flags = SWAP_FLAG_PREFER | SWAP_FLAG_OFFLOAD_ONLY |
+ SWAP_FLAG_DISCARD | SWAP_FLAG_DISCARD_ONCE |
+ SWAP_FLAG_DISCARD_PAGES | priority;
+
+ if (syscall(SYS_swapon, path, flags)) {
+ perror("swapon discard-once+discard-pages");
+ return 1;
+ }
+ if (syscall(SYS_swapoff, path)) {
+ perror("swapoff discard-once+discard-pages");
+ return 1;
+ }
+ return 0;
+}
+
+static int join_cgroup(const char *procs)
+{
+ char pid[32];
+ int fd, len;
+
+ fd = open(procs, O_WRONLY);
+ if (fd < 0) {
+ perror("open cgroup.procs");
+ return 1;
+ }
+
+ len = snprintf(pid, sizeof(pid), "%d\n", getpid());
+ if (write(fd, pid, len) != len) {
+ perror("write cgroup.procs");
+ close(fd);
+ return 1;
+ }
+ close(fd);
+ return 0;
+}
+
+static int pin_to_one_cpu(const char *pid_arg)
+{
+ cpu_set_t allowed, selected;
+ char *end;
+ long pid;
+ int cpu;
+
+ errno = 0;
+ pid = strtol(pid_arg, &end, 10);
+ if (errno || *end || pid <= 0) {
+ fprintf(stderr, "invalid pid: %s\n", pid_arg);
+ return 1;
+ }
+
+ if (sched_getaffinity(pid, sizeof(allowed), &allowed)) {
+ perror("sched_getaffinity");
+ return 1;
+ }
+ for (cpu = 0; cpu < CPU_SETSIZE; cpu++)
+ if (CPU_ISSET(cpu, &allowed))
+ break;
+ if (cpu == CPU_SETSIZE) {
+ fprintf(stderr, "pid %ld has no allowed CPU\n", pid);
+ return 1;
+ }
+
+ CPU_ZERO(&selected);
+ CPU_SET(cpu, &selected);
+ if (sched_setaffinity(pid, sizeof(selected), &selected)) {
+ perror("sched_setaffinity");
+ return 1;
+ }
+ return 0;
+}
+
+static int allocate(const char *size_arg, const char *procs,
+ const char *ready, const char *verified)
+{
+ unsigned char *memory;
+ unsigned long size;
+ unsigned long page_size;
+ char *end;
+ sigset_t signals;
+ int signal;
+ int fd;
+
+ errno = 0;
+ size = strtoul(size_arg, &end, 0);
+ if (errno || *end || !size) {
+ fprintf(stderr, "invalid allocation size: %s\n", size_arg);
+ return 1;
+ }
+ if (join_cgroup(procs))
+ return 1;
+
+ memory = mmap(NULL, size, PROT_READ | PROT_WRITE,
+ MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
+ if (memory == MAP_FAILED) {
+ perror("mmap");
+ return 1;
+ }
+
+ page_size = getpagesize();
+ for (unsigned long i = 0; i < size; i += page_size)
+ memset(memory + i, i / page_size % 251 + 1,
+ page_size < size - i ? page_size : size - i);
+ if (mprotect(memory, size, PROT_READ)) {
+ perror("mprotect");
+ return 1;
+ }
+
+ sigemptyset(&signals);
+ sigaddset(&signals, SIGUSR1);
+ if (sigprocmask(SIG_BLOCK, &signals, NULL)) {
+ perror("sigprocmask");
+ return 1;
+ }
+
+ fd = open(ready, O_WRONLY | O_CREAT | O_EXCL, 0600);
+ if (fd < 0) {
+ perror("create ready file");
+ return 1;
+ }
+ close(fd);
+
+ errno = sigwait(&signals, &signal);
+ if (errno) {
+ perror("sigwait");
+ return 1;
+ }
+ for (unsigned long i = 0; i < size; i++) {
+ unsigned char expected = i / page_size % 251 + 1;
+
+ if (memory[i] != expected) {
+ fprintf(stderr,
+ "data mismatch at %lu: got %u, expected %u\n",
+ i, memory[i], expected);
+ return 1;
+ }
+ }
+
+ fd = open(verified, O_WRONLY | O_CREAT | O_EXCL, 0600);
+ if (fd < 0) {
+ perror("create verified file");
+ return 1;
+ }
+ close(fd);
+
+ for (;;)
+ pause();
+}
+
+int main(int argc, char **argv)
+{
+ if (argc == 4 && !strcmp(argv[1], "activate"))
+ return activate(argv[2], atoi(argv[3]),
+ SWAP_FLAG_DISCARD | SWAP_FLAG_DISCARD_ONCE);
+ if (argc == 4 && !strcmp(argv[1], "reject-page-discard"))
+ return reject_page_discard(argv[2], atoi(argv[3]));
+ if (argc == 4 && !strcmp(argv[1], "accept-discard-once-pages"))
+ return accept_discard_once_pages(argv[2], atoi(argv[3]));
+ if (argc == 3 && !strcmp(argv[1], "pin"))
+ return pin_to_one_cpu(argv[2]);
+ if (argc == 6 && !strcmp(argv[1], "allocate"))
+ return allocate(argv[2], argv[3], argv[4], argv[5]);
+
+ fprintf(stderr,
+ "usage: %s activate DEVICE PRIORITY | reject-page-discard DEVICE PRIORITY | accept-discard-once-pages DEVICE PRIORITY | pin PID | allocate BYTES CGROUP.PROCS READY VERIFIED\n",
+ argv[0]);
+ return 1;
+}
diff --git a/tools/testing/selftests/zram/workingset_offload.c b/tools/testing/selftests/zram/workingset_offload.c
new file mode 100644
index 000000000000..7b82797489c2
--- /dev/null
+++ b/tools/testing/selftests/zram/workingset_offload.c
@@ -0,0 +1,207 @@
+// SPDX-License-Identifier: GPL-2.0
+#define _GNU_SOURCE
+
+#include <errno.h>
+#include <fcntl.h>
+#include <signal.h>
+#include <stdbool.h>
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <sys/mman.h>
+#include <sys/stat.h>
+#include <unistd.h>
+
+#define ANON_SIZE (64UL << 20)
+#define TARGET_SIZE (16UL << 20)
+#define FILLER_SIZE (32UL << 20)
+
+static int join_cgroup(const char *path)
+{
+ char pid[32];
+ int fd, len;
+
+ fd = open(path, O_WRONLY);
+ if (fd < 0)
+ return -1;
+ len = snprintf(pid, sizeof(pid), "%d\n", getpid());
+ if (write(fd, pid, len) != len) {
+ close(fd);
+ return -1;
+ }
+ return close(fd);
+}
+
+static int create_file(const char *path, size_t size)
+{
+ unsigned char page[4096];
+ size_t offset;
+ int fd;
+
+ fd = open(path, O_CREAT | O_TRUNC | O_RDWR, 0600);
+ if (fd < 0)
+ return -1;
+ memset(page, 0xa5, sizeof(page));
+ for (offset = 0; offset < size; offset += sizeof(page)) {
+ if (pwrite(fd, page, sizeof(page), offset) != sizeof(page)) {
+ close(fd);
+ return -1;
+ }
+ }
+ if (fsync(fd) || posix_fadvise(fd, 0, size, POSIX_FADV_DONTNEED)) {
+ close(fd);
+ return -1;
+ }
+ return fd;
+}
+
+static int cache_file(int fd, size_t size)
+{
+ unsigned char byte;
+ size_t offset;
+
+ if (posix_fadvise(fd, 0, size, POSIX_FADV_RANDOM))
+ return -1;
+ for (offset = 0; offset < size; offset += getpagesize())
+ if (pread(fd, &byte, 1, offset) != 1)
+ return -1;
+ return 0;
+}
+
+static unsigned long memory_stat(const char *path, const char *key)
+{
+ unsigned long value;
+ char name[64];
+ FILE *file;
+
+ file = fopen(path, "r");
+ if (!file)
+ return 0;
+ while (fscanf(file, "%63s %lu", name, &value) == 2) {
+ if (!strcmp(name, key)) {
+ fclose(file);
+ return value;
+ }
+ }
+ fclose(file);
+ return 0;
+}
+
+static int touch_ready(const char *path)
+{
+ int fd = open(path, O_WRONLY | O_CREAT | O_EXCL, 0600);
+
+ if (fd < 0)
+ return -1;
+ return close(fd);
+}
+
+static bool calibration_sufficient(size_t pages, unsigned long selected,
+ unsigned long refaulted,
+ unsigned long anon)
+{
+ return selected >= pages / 2 && refaulted * 100 >= selected * 80 &&
+ anon >= TARGET_SIZE;
+}
+
+int main(int argc, char **argv)
+{
+ unsigned long refault_before, refault_after;
+ unsigned long activate_before, activate_after;
+ unsigned long anon, selected = 0;
+ unsigned char *resident, *anon_memory;
+ unsigned char byte, checksum = 0;
+ char stat_path[4096];
+ sigset_t signals;
+ size_t pages, i;
+ void *mapping;
+ int target_fd, filler_fd, signal;
+
+ if (argc != 5) {
+ fprintf(stderr, "usage: %s CGROUP.PROCS TARGET FILLER READY\n",
+ argv[0]);
+ return 1;
+ }
+ if (join_cgroup(argv[1])) {
+ perror("join cgroup");
+ return 1;
+ }
+
+ anon_memory = mmap(NULL, ANON_SIZE, PROT_READ | PROT_WRITE,
+ MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
+ if (anon_memory == MAP_FAILED) {
+ perror("mmap anonymous");
+ return 1;
+ }
+ for (i = 0; i < ANON_SIZE; i += getpagesize())
+ anon_memory[i] = i / getpagesize() % 251 + 1;
+
+ target_fd = create_file(argv[2], TARGET_SIZE);
+ filler_fd = create_file(argv[3], FILLER_SIZE);
+ if (target_fd < 0 || filler_fd < 0) {
+ perror("create files");
+ return 1;
+ }
+ if (cache_file(target_fd, TARGET_SIZE) ||
+ cache_file(filler_fd, FILLER_SIZE)) {
+ perror("populate page cache");
+ return 1;
+ }
+
+ sigemptyset(&signals);
+ sigaddset(&signals, SIGUSR1);
+ if (sigprocmask(SIG_BLOCK, &signals, NULL) || touch_ready(argv[4])) {
+ perror("prepare signal");
+ return 1;
+ }
+ errno = sigwait(&signals, &signal);
+ if (errno) {
+ perror("sigwait");
+ return 1;
+ }
+
+ mapping = mmap(NULL, TARGET_SIZE, PROT_READ, MAP_SHARED, target_fd, 0);
+ if (mapping == MAP_FAILED) {
+ perror("mmap target");
+ return 1;
+ }
+ pages = TARGET_SIZE / getpagesize();
+ resident = calloc(pages, 1);
+ if (!resident || mincore(mapping, TARGET_SIZE, resident)) {
+ perror("mincore");
+ return 1;
+ }
+ munmap(mapping, TARGET_SIZE);
+
+ snprintf(stat_path, sizeof(stat_path), "%.*s/memory.stat",
+ (int)(strlen(argv[1]) - strlen("/cgroup.procs")), argv[1]);
+ refault_before = memory_stat(stat_path, "workingset_refault_file");
+ activate_before = memory_stat(stat_path, "workingset_activate_file");
+ for (i = 0; i < pages; i++) {
+ if (resident[i] & 1)
+ continue;
+ if (pread(target_fd, &byte, 1, i * getpagesize()) != 1) {
+ perror("refault target");
+ return 1;
+ }
+ selected++;
+ }
+ refault_after = memory_stat(stat_path, "workingset_refault_file");
+ activate_after = memory_stat(stat_path, "workingset_activate_file");
+ anon = memory_stat(stat_path, "active_anon") +
+ memory_stat(stat_path, "inactive_anon");
+
+ for (i = 0; i < ANON_SIZE; i += getpagesize())
+ checksum ^= anon_memory[i];
+ printf("selected=%lu refault=%lu activate=%lu anon=%lu checksum=%u\n",
+ selected, refault_after - refault_before,
+ activate_after - activate_before, anon, checksum);
+
+ free(resident);
+ close(filler_fd);
+ close(target_fd);
+ if (!calibration_sufficient(pages, selected,
+ refault_after - refault_before, anon))
+ return 2;
+ return (activate_after - activate_before) * 100 < selected * 80;
+}
diff --git a/tools/testing/selftests/zram/zram03.sh b/tools/testing/selftests/zram/zram03.sh
new file mode 100755
index 000000000000..951f70c1d878
--- /dev/null
+++ b/tools/testing/selftests/zram/zram03.sh
@@ -0,0 +1,174 @@
+#!/bin/sh
+# SPDX-License-Identifier: GPL-2.0
+# Test proactive-only swap allocation and pressure fallback.
+
+set -eu
+
+# shellcheck source=zram_lib.sh
+. ./zram_lib.sh
+
+TCID="zram03"
+cg="/sys/fs/cgroup/zram-offload-$$"
+cgroup_root="/sys/fs/cgroup"
+tmp=""
+ready=""
+verified=""
+allocator_pid=""
+offload=""
+zswap_enabled=""
+
+fail()
+{
+ echo "$TCID: [FAIL] $*" >&2
+ exit 1
+}
+
+skip()
+{
+ echo "$TCID: [SKIP] $*" >&2
+ exit "$ksft_skip"
+}
+
+cleanup()
+{
+ local status=$?
+ set +e
+ if [ -n "$allocator_pid" ]; then
+ kill "$allocator_pid"
+ wait "$allocator_pid"
+ fi
+ if [ -n "$tmp" ]; then
+ rm -rf "$tmp" || status=1
+ fi
+ if [ -n "$cg" ] && [ -d "$cg" ]; then
+ rmdir "$cg" || status=1
+ fi
+ cgroup_disable_memory_controller "$cgroup_root" || status=1
+ if [ -n "$dev_ids" ]; then
+ zram_cleanup || status=1
+ fi
+ if [ -n "$zswap_enabled" ]; then
+ echo "$zswap_enabled" > /sys/module/zswap/parameters/enabled || status=1
+ fi
+ exit "$status"
+}
+
+swap_used_kb()
+{
+ awk -v device="$1" '$1 == device { print $4 }' /proc/swaps
+}
+
+zram_orig_data_size()
+{
+ awk '{ print $1 }' "/sys/block/${1##*/}/mm_stat"
+}
+
+wait_file()
+{
+ for _ in $(seq 1 400); do
+ [ -e "$1" ] && return 0
+ sleep 0.05
+ done
+ return 1
+}
+
+check_prereqs
+# The feature marker also requires CONFIG_VM_EVENT_COUNTERS.
+grep -q '^swpout_offload_refused ' /proc/vmstat ||
+ skip "offload refusal counters are unavailable"
+[ -x ./swap_offload ] || skip "swap_offload helper is unavailable"
+[ -e /sys/fs/cgroup/cgroup.controllers ] || skip "cgroup v2 controllers are unavailable"
+grep -qw memory /sys/fs/cgroup/cgroup.controllers || skip "memory controller is unavailable"
+[ -e /sys/module/zswap/parameters/enabled ] || skip "zswap is unavailable"
+zswap_enabled=$(cat /sys/module/zswap/parameters/enabled)
+
+# Keep every reclaim transition on one CPU. This makes the test exercise the
+# cached-cluster provenance mismatch rather than relying on scheduler placement.
+./swap_offload pin "$$" || skip "cannot pin test to one allowed CPU"
+
+# Swap priorities are global. Put both test areas ahead of any existing swap
+# so distro-managed zram or another high-priority area cannot absorb the test
+# allocations and make the routing assertions fail spuriously.
+max_prio=$(awk 'BEGIN { max = -1 } NR > 1 && $5 > max { max = $5 } END { print max }' /proc/swaps)
+[ "$max_prio" -le 32765 ] || skip "cannot outrank existing swap priority $max_prio"
+safe_prio=$((max_prio + 1))
+offload_prio=$((max_prio + 2))
+
+dev_num=2
+zram_sizes="67108864 67108864"
+trap cleanup EXIT
+trap 'exit 129' HUP
+trap 'exit 130' INT
+trap 'exit 143' TERM
+zram_load
+echo Y > /sys/module/zswap/parameters/enabled || skip "cannot enable zswap"
+zram_set_disksizes
+
+set -- $dev_ids
+safe="/dev/zram${1}"
+offload="/dev/zram${2}"
+safe_id=$1
+offload_id=$2
+mkswap "$safe" >/dev/null
+mkswap "$offload" >/dev/null
+swapon -p "$safe_prio" "$safe"
+dev_swap_ids=" $safe_id"
+./swap_offload reject-page-discard "$offload" "$offload_prio" ||
+ fail "offload-only page discard was accepted"
+./swap_offload accept-discard-once-pages "$offload" "$offload_prio" ||
+ fail "resolved offload-only discard-once was rejected"
+./swap_offload activate "$offload" "$offload_prio"
+dev_swap_ids="$dev_swap_ids $offload_id"
+
+cgroup_enable_memory_controller "$cgroup_root" ||
+ skip "cannot enable the cgroup v2 memory controller"
+mkdir "$cg" || skip "cannot create test cgroup"
+[ -e "$cg/memory.max" ] || skip "cgroup v2 memory controller is unavailable"
+echo max > "$cg/memory.swap.max"
+echo 1 > "$cg/memory.zswap.writeback" ||
+ skip "cannot enable zswap writeback for test cgroup"
+
+tmp_dir=$(mktemp -d "${TMPDIR:-/tmp}/zram-offload.XXXXXX") ||
+ fail "cannot create readiness directory"
+tmp=$tmp_dir
+ready="$tmp/ready"
+verified="$tmp/verified"
+./swap_offload allocate 67108864 "$cg/cgroup.procs" "$ready" "$verified" &
+allocator_pid=$!
+wait_file "$ready" || fail "allocator did not become ready"
+
+echo "16M swappiness=max" > "$cg/memory.reclaim" ||
+ fail "proactive reclaim failed"
+offload_before=$(swap_used_kb "$offload")
+safe_before=$(swap_used_kb "$safe")
+offload_orig=$(zram_orig_data_size "$offload")
+[ "${offload_before:-0}" -gt 0 ] || fail "proactive reclaim missed offload area"
+[ "${safe_before:-0}" -eq 0 ] || fail "proactive reclaim used lower-priority safe area"
+[ "$((offload_orig + 1048576))" -ge "$((offload_before * 1024))" ] ||
+ fail "zswap deferred the offload-only backend write"
+
+echo 24M > "$cg/memory.max"
+sleep 1
+offload_pressure=$(swap_used_kb "$offload")
+safe_pressure=$(swap_used_kb "$safe")
+safe_orig=$(zram_orig_data_size "$safe")
+[ "${safe_pressure:-0}" -gt 0 ] || fail "pressure reclaim missed safe fallback"
+[ "$offload_pressure" -le "$offload_before" ] ||
+ fail "pressure reclaim allocated offload-only slots"
+[ "$safe_orig" -lt "$((safe_pressure * 1024))" ] ||
+ fail "zswap did not retain any unmarked safe-area pages"
+
+echo max > "$cg/memory.max"
+echo "4M swappiness=max" > "$cg/memory.reclaim" ||
+ fail "second proactive reclaim failed"
+offload_after=$(swap_used_kb "$offload")
+[ "$offload_after" -gt "$offload_pressure" ] ||
+ fail "proactive priority was not restored after pressure reclaim"
+
+swapoff "$offload" || fail "swapoff could not recover offloaded pages"
+offload=""
+dev_swap_ids=" $safe_id"
+kill -USR1 "$allocator_pid" || fail "cannot request data verification"
+wait_file "$verified" || fail "allocator did not verify recovered data"
+kill -0 "$allocator_pid" || fail "allocator died during swapoff"
+echo "$TCID: [PASS]"
diff --git a/tools/testing/selftests/zram/zram04.sh b/tools/testing/selftests/zram/zram04.sh
new file mode 100755
index 000000000000..a351d08906f2
--- /dev/null
+++ b/tools/testing/selftests/zram/zram04.sh
@@ -0,0 +1,164 @@
+#!/bin/sh
+# SPDX-License-Identifier: GPL-2.0
+# Reproduce offload-only swap leaking into the workingset refault heuristic.
+
+set -eu
+
+# shellcheck source=zram_lib.sh
+. ./zram_lib.sh
+
+TCID="zram04"
+cg="/sys/fs/cgroup/zram-workingset-$$"
+cg_created=0
+cgroup_root="/sys/fs/cgroup"
+tmp=""
+ready=""
+worker=""
+tmp_fs=""
+mglru=""
+zswap_enabled=""
+
+fail()
+{
+ echo "$TCID: [FAIL] $*" >&2
+ exit 1
+}
+
+skip()
+{
+ echo "$TCID: [SKIP] $*" >&2
+ exit "$ksft_skip"
+}
+
+cleanup()
+{
+ local status=$?
+ set +e
+ [ -n "$worker" ] && kill "$worker"
+ [ -n "$worker" ] && wait "$worker"
+ if [ -n "$tmp" ]; then
+ rm -rf "$tmp" || status=1
+ fi
+ if [ "$cg_created" -eq 1 ]; then
+ rmdir "$cg" || status=1
+ fi
+ cgroup_disable_memory_controller "$cgroup_root" || status=1
+ if [ -n "$dev_ids" ]; then
+ zram_cleanup || status=1
+ fi
+ if [ -n "$mglru" ]; then
+ echo "$mglru" > /sys/kernel/mm/lru_gen/enabled || status=1
+ fi
+ if [ -n "$zswap_enabled" ]; then
+ echo "$zswap_enabled" > /sys/module/zswap/parameters/enabled || status=1
+ fi
+ exit "$status"
+}
+
+wait_helper_ready()
+{
+ for _ in $(seq 1 400); do
+ [ -e "$ready" ] && return 0
+ worker_state=$(awk '{ print $3 }' "/proc/$worker/stat" \
+ 2>/dev/null || :)
+ if ! kill -0 "$worker" 2>/dev/null ||
+ [ "$worker_state" = Z ]; then
+ if wait "$worker"; then
+ status=0
+ else
+ status=$?
+ fi
+ worker=""
+ [ "$status" -eq 2 ] &&
+ skip "workingset calibration was insufficient"
+ fail "workingset helper exited with status $status before readiness"
+ fi
+ sleep 0.05
+ done
+ fail "workingset helper timed out before readiness"
+}
+
+check_prereqs
+# The feature marker also requires CONFIG_VM_EVENT_COUNTERS.
+grep -q '^swpout_offload_refused ' /proc/vmstat ||
+ skip "offload refusal counters are unavailable"
+[ -x ./swap_offload ] || skip "swap_offload helper is unavailable"
+[ -x ./workingset_offload ] || skip "workingset helper is unavailable"
+[ -e /sys/fs/cgroup/cgroup.controllers ] || skip "cgroup v2 is unavailable"
+grep -qw memory /sys/fs/cgroup/cgroup.controllers ||
+ skip "cgroup v2 memory controller is unavailable"
+[ "$(awk 'END { print NR }' /proc/swaps)" -eq 1 ] ||
+ skip "test requires no pre-existing swap"
+tmp_fs=$(stat -f -c %T "${TMPDIR:-/var/tmp}") ||
+ skip "cannot identify the test filesystem"
+[ "$tmp_fs" != tmpfs ] ||
+ skip "test files require a disk-backed filesystem"
+
+[ -e /sys/kernel/mm/lru_gen/enabled ] &&
+ mglru=$(cat /sys/kernel/mm/lru_gen/enabled)
+[ -e /sys/module/zswap/parameters/enabled ] &&
+ zswap_enabled=$(cat /sys/module/zswap/parameters/enabled)
+trap cleanup EXIT
+trap 'exit 129' HUP
+trap 'exit 130' INT
+trap 'exit 143' TERM
+[ -n "$mglru" ] && echo 0 > /sys/kernel/mm/lru_gen/enabled
+[ -n "$zswap_enabled" ] && echo N > /sys/module/zswap/parameters/enabled
+
+dev_num=1
+zram_sizes="134217728"
+zram_load
+zram_set_disksizes
+set -- $dev_ids
+offload="/dev/zram${1}"
+mkswap "$offload" >/dev/null
+./swap_offload activate "$offload" 1
+dev_swap_ids=" $1"
+
+cgroup_enable_memory_controller "$cgroup_root" ||
+ skip "cannot enable the cgroup v2 memory controller"
+mkdir "$cg" || skip "cannot create test cgroup"
+cg_created=1
+[ -e "$cg/memory.max" ] || skip "cgroup v2 memory controller is unavailable"
+echo max > "$cg/memory.max"
+echo max > "$cg/memory.swap.max"
+tmp=$(mktemp -d "${TMPDIR:-/var/tmp}/zram-workingset.XXXXXX")
+ready="$tmp/ready"
+
+./workingset_offload "$cg/cgroup.procs" "$tmp/target" "$tmp/filler" \
+ "$ready" &
+worker=$!
+wait_helper_ready
+
+# Shadow retention uses local LRU/slab statistics updated by memcg flushing.
+# Check that the 64 MiB anonymous and 48 MiB file setup is visible before
+# evicting file pages.
+stats_ready=0
+for _ in $(seq 1 50); do
+ if awk '
+ $1 == "active_anon" || $1 == "inactive_anon" { anon += $2 }
+ $1 == "file" { file = $2 }
+ END { exit !(anon >= 67108864 && file >= 50331648) }
+ ' "$cg/memory.stat"; then
+ stats_ready=1
+ break
+ fi
+ sleep 0.1
+done
+[ "$stats_ready" -eq 1 ] ||
+ skip "initial working-set statistics did not become visible"
+
+echo "48M swappiness=0" > "$cg/memory.reclaim" ||
+ skip "file-only proactive reclaim failed"
+kill -USR1 "$worker"
+if wait "$worker"; then
+ worker=""
+ echo "$TCID: [PASS]"
+ exit 0
+else
+ status=$?
+fi
+worker=""
+[ "$status" -eq 2 ] && skip "workingset calibration was insufficient"
+echo "$TCID: [FAIL] refaulted file pages were not activated" >&2
+exit 1
diff --git a/tools/testing/selftests/zram/zram_lib.sh b/tools/testing/selftests/zram/zram_lib.sh
index 4b134f70726d..cf55bcd1b09d 100755
--- a/tools/testing/selftests/zram/zram_lib.sh
+++ b/tools/testing/selftests/zram/zram_lib.sh
@@ -17,6 +17,10 @@ kernel_version=`uname -r | cut -d'.' -f1,2`
kernel_major=${kernel_version%.*}
kernel_minor=${kernel_version#*.}
+# Whether this test enabled the memory controller on its cgroup parent. Tests
+# must leave a delegation which was already present alone.
+cgroup_memory_controller_enabled=0
+
trap INT
check_prereqs()
@@ -30,6 +34,30 @@ check_prereqs()
fi
}
+cgroup_enable_memory_controller()
+{
+ local cgroup_root=$1
+
+ if grep -qw memory "$cgroup_root/cgroup.subtree_control"; then
+ return 0
+ fi
+
+ if ! echo +memory > "$cgroup_root/cgroup.subtree_control"; then
+ return 1
+ fi
+
+ cgroup_memory_controller_enabled=1
+}
+
+cgroup_disable_memory_controller()
+{
+ local cgroup_root=$1
+
+ [ "$cgroup_memory_controller_enabled" -eq 1 ] || return 0
+ echo -memory > "$cgroup_root/cgroup.subtree_control" || return 1
+ cgroup_memory_controller_enabled=0
+}
+
kernel_gte()
{
major=${1%.*}
--
2.55.0
^ permalink raw reply [flat|nested] 16+ messages in thread
* [RFC PATCH v2 4/4] selftests: zram: cover retained offload-only entries
2026-09-26 4:55 [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload Matthias Goergens
` (2 preceding siblings ...)
2026-09-26 4:55 ` [RFC PATCH v2 3/4] selftests: zram: cover offload-only swap policy Matthias Goergens
@ 2026-09-26 4:55 ` Matthias Goergens
2026-09-26 10:16 ` [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload Kairui Song
` (4 subsequent siblings)
8 siblings, 0 replies; 16+ messages in thread
From: Matthias Goergens @ 2026-09-26 4:55 UTC (permalink / raw)
To: Shuah Khan
Cc: Matthias Goergens, Minchan Kim, Sergey Senozhatsky, linux-kernel,
linux-kselftest, Chris Li, Kairui Song, Johannes Weiner,
Yosry Ahmed, Nhat Pham, linux-mm, Andrew Morton
Faulting and dirtying one page of a swapped-out large folio can leave
sibling PTEs referring to its existing offload-only swap allocation.
Exercise this path with zram behind dm-delay: ordinary reclaim must not
rewrite the allocation, while proactive reclaim must still be able to.
Check backing-device writes, refusal counters and data integrity. Derive
the folio and expected I/O sizes from the architecture's PMD huge-page
size. Keep this test separate from basic routing coverage because it
also requires transparent huge pages, MADV_COLLAPSE and dm-delay.
Measure retained swap in the target VMA and lock ancillary mappings to
keep them out of backing-device write counts.
Signed-off-by: Matthias Goergens <matthias.goergens@gmail.com>
---
tools/testing/selftests/zram/Makefile | 2 +-
tools/testing/selftests/zram/README | 8 +-
tools/testing/selftests/zram/config | 4 +
tools/testing/selftests/zram/settings | 1 +
tools/testing/selftests/zram/swap_offload.c | 205 ++++++++-
tools/testing/selftests/zram/zram05.sh | 447 ++++++++++++++++++++
6 files changed, 664 insertions(+), 3 deletions(-)
create mode 100644 tools/testing/selftests/zram/settings
create mode 100755 tools/testing/selftests/zram/zram05.sh
diff --git a/tools/testing/selftests/zram/Makefile b/tools/testing/selftests/zram/Makefile
index 781d10a20e56..b83e6a1563b2 100644
--- a/tools/testing/selftests/zram/Makefile
+++ b/tools/testing/selftests/zram/Makefile
@@ -2,7 +2,7 @@
all:
TEST_GEN_FILES := swap_offload workingset_offload
-TEST_PROGS := zram.sh zram03.sh zram04.sh
+TEST_PROGS := zram.sh zram03.sh zram04.sh zram05.sh
TEST_FILES := zram01.sh zram02.sh zram_lib.sh
EXTRA_CLEAN := err.log
diff --git a/tools/testing/selftests/zram/README b/tools/testing/selftests/zram/README
index cd7f389ed3c6..8c11d70d2a1f 100644
--- a/tools/testing/selftests/zram/README
+++ b/tools/testing/selftests/zram/README
@@ -28,6 +28,7 @@ zram02.sh: creates block device for swap
Offload-only swap tests, registered separately:
zram03.sh: checks proactive routing, pressure fallback and zswap bypass
zram04.sh: checks file workingset activation with offload-only swap
+zram05.sh: checks retained-entry write refusal, recovery and data integrity
Run these tests as root in an exclusive disposable VM with the offload-only
swap policy and the options listed in config. They change global swap,
@@ -35,7 +36,9 @@ zswap and reclaim settings. Use the initial cgroup namespace with an
unrestricted cgroup v2 memory hierarchy mounted at /sys/fs/cgroup. A child's
memory.zswap.writeback value does not reveal restrictions in its ancestors.
-zram04 needs a disk-backed TMPDIR.
+zram04 needs a disk-backed TMPDIR. zram05 additionally needs dm-delay,
+transparent huge pages and enough locked-memory allowance for its fixture;
+it skips when those prerequisites cannot be established.
Commands required for testing:
- bc
@@ -46,6 +49,9 @@ Commands required for testing:
- swapon
- swapoff
- mkfs/ mkfs.ext4
+ - dmsetup (zram05)
+ - blockdev (zram05)
+ - stat (zram04 and zram05)
For more information please refer:
kernel-source-tree/Documentation/admin-guide/blockdev/zram.rst
diff --git a/tools/testing/selftests/zram/config b/tools/testing/selftests/zram/config
index c59b8c3806a5..018429d8f5ca 100644
--- a/tools/testing/selftests/zram/config
+++ b/tools/testing/selftests/zram/config
@@ -1,6 +1,10 @@
CONFIG_CGROUPS=y
+CONFIG_BLK_DEV_DM=y
+CONFIG_DM_DELAY=y
CONFIG_MEMCG=y
CONFIG_SWAP=y
+CONFIG_TRANSPARENT_HUGEPAGE=y
+CONFIG_VM_EVENT_COUNTERS=y
CONFIG_ZSMALLOC=y
CONFIG_ZRAM=y
CONFIG_ZSWAP=y
diff --git a/tools/testing/selftests/zram/settings b/tools/testing/selftests/zram/settings
new file mode 100644
index 000000000000..6091b45d226b
--- /dev/null
+++ b/tools/testing/selftests/zram/settings
@@ -0,0 +1 @@
+timeout=120
diff --git a/tools/testing/selftests/zram/swap_offload.c b/tools/testing/selftests/zram/swap_offload.c
index b2b94cd6ee3a..0c01c97d8476 100644
--- a/tools/testing/selftests/zram/swap_offload.c
+++ b/tools/testing/selftests/zram/swap_offload.c
@@ -3,6 +3,7 @@
#include <errno.h>
#include <fcntl.h>
+#include <limits.h>
#include <sched.h>
#include <signal.h>
#include <stdio.h>
@@ -17,6 +18,7 @@
#define SWAP_FLAG_DISCARD_ONCE 0x20000
#define SWAP_FLAG_DISCARD_PAGES 0x40000
#define SWAP_FLAG_OFFLOAD_ONLY 0x80000
+#define KSFT_SKIP 4
static int activate(const char *path, int priority, int discard_flags)
{
@@ -201,11 +203,210 @@ static int allocate(const char *size_arg, const char *procs,
pause();
}
+static int create_marker(const char *path, unsigned char *memory,
+ unsigned long size)
+{
+ char *temporary;
+ int fd, ret = 1;
+
+ if (asprintf(&temporary, "%s.XXXXXX", path) < 0) {
+ perror("asprintf marker path");
+ return 1;
+ }
+ fd = mkstemp(temporary);
+ if (fd < 0) {
+ perror(path);
+ goto out_free;
+ }
+ if (memory && dprintf(fd, "%d %08lx-%08lx\n", getpid(),
+ (unsigned long)memory,
+ (unsigned long)(memory + size)) < 0) {
+ perror("write marker");
+ close(fd);
+ goto out_unlink;
+ }
+ if (close(fd)) {
+ perror("close marker");
+ goto out_unlink;
+ }
+ /* Publish only after the payload is complete for the shell reader. */
+ if (rename(temporary, path)) {
+ perror("publish marker");
+ goto out_unlink;
+ }
+ ret = 0;
+out_unlink:
+ unlink(temporary);
+out_free:
+ free(temporary);
+ return ret;
+}
+
+static unsigned char retained_byte(unsigned long offset, long page_size,
+ int touched)
+{
+ unsigned char value = offset / page_size % 251 + 1;
+
+ if (touched && !offset)
+ value ^= 0x5a;
+ return value;
+}
+
+static int verify_retained(unsigned char *memory, unsigned long size,
+ long page_size, int full)
+{
+ unsigned long limit = full ? size : 1;
+
+ for (unsigned long offset = 0; offset < limit; offset++) {
+ unsigned char expected = retained_byte(offset, page_size, 1);
+
+ if (memory[offset] != expected) {
+ fprintf(stderr,
+ "retained data mismatch at %lu: got %u, expected %u\n",
+ offset, memory[offset], expected);
+ return 1;
+ }
+ }
+ return 0;
+}
+
+static int retained(const char *size_arg, const char *procs,
+ const char *ready, const char *touched,
+ const char *verified)
+{
+ unsigned long allocation_size, mapping_start, size;
+ unsigned char *mapping, *memory;
+ char *end;
+ sigset_t signals;
+ long page_size;
+ int signal;
+
+ errno = 0;
+ size = strtoul(size_arg, &end, 0);
+ if (errno || *end || !size || (size & (size - 1))) {
+ fprintf(stderr, "allocation size must be a power of two\n");
+ return 1;
+ }
+ /*
+ * Keep existing runtime mappings out of the backing-write measurement.
+ * Enable future locking only after creating the unlocked target mapping.
+ */
+ if (mlockall(MCL_CURRENT)) {
+ int error = errno;
+
+ perror("mlockall ancillary mappings");
+ if (error == EPERM || error == ENOMEM)
+ return KSFT_SKIP;
+ return 1;
+ }
+ if (join_cgroup(procs))
+ return 1;
+
+ page_size = sysconf(_SC_PAGESIZE);
+ if (page_size <= 0) {
+ perror("sysconf _SC_PAGESIZE");
+ return 1;
+ }
+ if (size < (unsigned long)page_size || size % page_size) {
+ fprintf(stderr, "allocation size must contain whole pages\n");
+ return 1;
+ }
+ if (size > ULONG_MAX / 2) {
+ fprintf(stderr, "allocation size is too large\n");
+ return 1;
+ }
+ allocation_size = size * 2;
+ mapping = mmap(NULL, allocation_size, PROT_READ | PROT_WRITE,
+ MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
+ if (mapping == MAP_FAILED) {
+ int error = errno;
+
+ perror("mmap");
+ if (error == EAGAIN || error == ENOMEM)
+ return KSFT_SKIP;
+ return 1;
+ }
+ mapping_start = (unsigned long)mapping;
+ memory = (unsigned char *)((mapping_start + size - 1) & ~(size - 1));
+ if ((unsigned long)memory != mapping_start)
+ munmap((void *)mapping_start,
+ (unsigned long)memory - mapping_start);
+ munmap(memory + size, allocation_size - size -
+ ((unsigned long)memory - mapping_start));
+ if (mlockall(MCL_FUTURE)) {
+ perror("mlockall future mappings");
+ return 1;
+ }
+ if (madvise(memory, size, MADV_HUGEPAGE)) {
+ perror("madvise MADV_HUGEPAGE");
+ return 1;
+ }
+ for (unsigned long offset = 0; offset < size; offset += page_size)
+ memset(memory + offset, retained_byte(offset, page_size, 0),
+ page_size);
+#ifdef MADV_COLLAPSE
+ if (madvise(memory, size, MADV_COLLAPSE)) {
+ int error = errno;
+
+ perror("madvise MADV_COLLAPSE");
+ if (error == EAGAIN || error == EINVAL || error == ENOMEM)
+ return KSFT_SKIP;
+ return 1;
+ }
+#else
+ fprintf(stderr, "MADV_COLLAPSE is unavailable\n");
+ return KSFT_SKIP;
+#endif
+
+ sigemptyset(&signals);
+ sigaddset(&signals, SIGUSR1);
+ sigaddset(&signals, SIGUSR2);
+ sigaddset(&signals, SIGALRM);
+ if (sigprocmask(SIG_BLOCK, &signals, NULL)) {
+ perror("sigprocmask");
+ return 1;
+ }
+ if (create_marker(ready, memory, size))
+ return 1;
+
+ for (;;) {
+ errno = sigwait(&signals, &signal);
+ if (errno) {
+ perror("sigwait");
+ return 1;
+ }
+ if (signal == SIGUSR1) {
+ unsigned char value;
+
+ if (mprotect(memory, page_size, PROT_READ)) {
+ perror("mprotect read");
+ return 1;
+ }
+ value = memory[0];
+ if (mprotect(memory, page_size, PROT_READ | PROT_WRITE)) {
+ perror("mprotect write");
+ return 1;
+ }
+ memory[0] = value ^ 0x5a;
+ if (create_marker(touched, NULL, 0))
+ return 1;
+ } else {
+ if (verify_retained(memory, size, page_size,
+ signal == SIGALRM))
+ return 1;
+ if (create_marker(verified, NULL, 0))
+ return 1;
+ }
+ }
+}
+
int main(int argc, char **argv)
{
if (argc == 4 && !strcmp(argv[1], "activate"))
return activate(argv[2], atoi(argv[3]),
SWAP_FLAG_DISCARD | SWAP_FLAG_DISCARD_ONCE);
+ if (argc == 4 && !strcmp(argv[1], "activate-no-discard"))
+ return activate(argv[2], atoi(argv[3]), 0);
if (argc == 4 && !strcmp(argv[1], "reject-page-discard"))
return reject_page_discard(argv[2], atoi(argv[3]));
if (argc == 4 && !strcmp(argv[1], "accept-discard-once-pages"))
@@ -214,9 +415,11 @@ int main(int argc, char **argv)
return pin_to_one_cpu(argv[2]);
if (argc == 6 && !strcmp(argv[1], "allocate"))
return allocate(argv[2], argv[3], argv[4], argv[5]);
+ if (argc == 7 && !strcmp(argv[1], "retained"))
+ return retained(argv[2], argv[3], argv[4], argv[5], argv[6]);
fprintf(stderr,
- "usage: %s activate DEVICE PRIORITY | reject-page-discard DEVICE PRIORITY | accept-discard-once-pages DEVICE PRIORITY | pin PID | allocate BYTES CGROUP.PROCS READY VERIFIED\n",
+ "usage: %s activate|activate-no-discard DEVICE PRIORITY | reject-page-discard DEVICE PRIORITY | accept-discard-once-pages DEVICE PRIORITY | pin PID | allocate BYTES CGROUP.PROCS READY VERIFIED | retained BYTES CGROUP.PROCS READY TOUCHED VERIFIED\n",
argv[0]);
return 1;
}
diff --git a/tools/testing/selftests/zram/zram05.sh b/tools/testing/selftests/zram/zram05.sh
new file mode 100755
index 000000000000..c37f670e2824
--- /dev/null
+++ b/tools/testing/selftests/zram/zram05.sh
@@ -0,0 +1,447 @@
+#!/bin/sh
+# SPDX-License-Identifier: GPL-2.0
+# Test retained offload-only entries under ordinary and proactive reclaim.
+
+set -eu
+
+# shellcheck source=zram_lib.sh
+. ./zram_lib.sh
+
+TCID="zram05"
+cg="/sys/fs/cgroup/zram-retained-$$"
+cg_created=0
+cgroup_root="/sys/fs/cgroup"
+tmp=""
+ready=""
+touched=""
+verified=""
+holder_pid=""
+safe=""
+safe_active=0
+offload_backing=""
+offload=""
+offload_active=0
+dm_name="zram-retained-$$"
+dm_active=0
+dm_node_created=0
+zswap_enabled=""
+thp_size=0
+thp_kib=0
+thp_sectors=0
+page_kib=0
+
+fail()
+{
+ echo "$TCID: [FAIL] $*" >&2
+ exit 1
+}
+
+skip()
+{
+ echo "$TCID: [SKIP] $*" >&2
+ exit "$ksft_skip"
+}
+
+cleanup()
+{
+ local status=$?
+ set +e
+ if [ -n "$holder_pid" ]; then
+ kill "$holder_pid"
+ wait "$holder_pid"
+ fi
+ if [ "$offload_active" -eq 1 ] ||
+ { [ -n "$offload" ] &&
+ awk -v device="$offload" '$1 == device { found = 1 }
+ END { exit !found }' /proc/swaps; }; then
+ if swapoff "$offload"; then
+ offload_active=0
+ else
+ status=1
+ fi
+ fi
+ if [ "$safe_active" -eq 1 ]; then
+ if swapoff "$safe"; then
+ safe_active=0
+ dev_swap_ids=""
+ else
+ status=1
+ fi
+ fi
+ if [ "$dm_active" -eq 1 ]; then
+ zram_wait_for_udev
+ if dmsetup --noudevsync --noudevrules remove "$dm_name"; then
+ dm_active=0
+ else
+ status=1
+ fi
+ fi
+ if [ "$dm_node_created" -eq 1 ] && [ "$dm_active" -eq 0 ]; then
+ rm -f "$offload" || status=1
+ fi
+ if [ -n "$tmp" ]; then
+ rm -rf "$tmp" || status=1
+ fi
+ if [ "$cg_created" -eq 1 ]; then
+ rmdir "$cg" || status=1
+ fi
+ cgroup_disable_memory_controller "$cgroup_root" || status=1
+ if [ -n "$dev_ids" ]; then
+ # A live dm mapping still needs its zram backing device.
+ [ "$dm_active" -eq 0 ] || dev_ids=" $safe_id"
+ zram_cleanup || status=1
+ fi
+ if [ -n "$zswap_enabled" ]; then
+ echo "$zswap_enabled" > /sys/module/zswap/parameters/enabled || status=1
+ fi
+ exit "$status"
+}
+
+wait_file()
+{
+ for _ in $(seq 1 400); do
+ [ -e "$1" ] && return 0
+ if [ -n "$holder_pid" ]; then
+ holder_state=$(awk '{ print $3 }' \
+ "/proc/$holder_pid/stat" 2>/dev/null || :)
+ else
+ holder_state=""
+ fi
+ if [ -n "$holder_pid" ] &&
+ { ! kill -0 "$holder_pid" 2>/dev/null ||
+ [ "$holder_state" = Z ]; }; then
+ if wait "$holder_pid"; then
+ helper_status=0
+ else
+ helper_status=$?
+ fi
+ holder_pid=""
+ return 2
+ fi
+ sleep 0.05
+ done
+ return 1
+}
+
+require_helper_file()
+{
+ if wait_file "$1"; then
+ return 0
+ else
+ status=$?
+ fi
+ if [ "$status" -eq 2 ]; then
+ [ "$helper_status" -eq "$ksft_skip" ] && skip "$2 is unavailable"
+ fail "retained helper exited with status $helper_status while $2"
+ fi
+ fail "retained helper timed out while $2"
+}
+
+written_sectors()
+{
+ awk '{ print $7 }' "/sys/block/${offload_backing##*/}/stat"
+}
+
+vmstat_value()
+{
+ awk -v name="$1" '$1 == name { print $2 }' /proc/vmstat
+}
+
+memcg_counter()
+{
+ awk -v name="$1" '$1 == name { print $2; found = 1; exit }
+ END { if (!found) exit 1 }' "$cg/memory.stat"
+}
+
+measure_thp_vma_swap()
+{
+ # Measure only the helper's target; process-wide VmSwap includes other
+ # mappings. Before reclaim, require the entire unlocked target to be huge.
+ awk -v want="$thp_vma_range" -v expected="$thp_kib" -v initial="$1" '
+ function emit() {
+ if (range == want) {
+ matches++
+ if (size != expected || swap < 0 || locked != 0 ||
+ (initial && (anon != expected || swap != 0)))
+ invalid = 1
+ matched_swap = swap
+ }
+ }
+ /^[[:xdigit:]]+-[[:xdigit:]]+[[:space:]]/ {
+ emit()
+ range = $1
+ size = anon = locked = swap = -1
+ next
+ }
+ $1 == "Size:" { size = $2; next }
+ $1 == "AnonHugePages:" { anon = $2; next }
+ $1 == "Locked:" { locked = $2; next }
+ $1 == "Swap:" { swap = $2; next }
+ END {
+ emit()
+ if (matches != 1 || invalid)
+ exit 1
+ print matched_swap
+ }' "/proc/$holder_pid/smaps" > "$tmp/thp-vma-swap"
+}
+
+wait_offload_quiet()
+{
+ # A bio queued inside dm-delay is not yet visible in either the backing
+ # device statistics or the mapped device's inflight counters.
+ sleep 4
+ previous=-1
+ stable=0
+ for _ in $(seq 1 240); do
+ current=$(written_sectors)
+ read -r reads writes < "$dm_inflight"
+ if [ "$current" -eq "$previous" ] && \
+ [ "$reads" -eq 0 ] && [ "$writes" -eq 0 ]; then
+ stable=$((stable + 1))
+ [ "$stable" -ge 20 ] && return 0
+ else
+ stable=0
+ fi
+ previous=$current
+ sleep 0.05
+ done
+ return 1
+}
+
+check_prereqs
+[ -x ./swap_offload ] || skip "swap_offload helper is unavailable"
+command -v dmsetup >/dev/null 2>&1 || skip "dmsetup is unavailable"
+command -v blockdev >/dev/null 2>&1 || skip "blockdev is unavailable"
+command -v stat >/dev/null 2>&1 || skip "stat is unavailable"
+[ -e /sys/fs/cgroup/cgroup.controllers ] ||
+ skip "cgroup v2 controllers are unavailable"
+grep -qw memory /sys/fs/cgroup/cgroup.controllers ||
+ skip "memory controller is unavailable"
+[ -d /sys/kernel/mm/transparent_hugepage ] ||
+ skip "transparent huge pages are unavailable"
+[ -r /sys/kernel/mm/transparent_hugepage/hpage_pmd_size ] ||
+ skip "PMD huge-page size is unavailable"
+grep -q '^swpout_offload_refused ' /proc/vmstat ||
+ skip "offload refusal counters are unavailable"
+
+thp_size=$(cat /sys/kernel/mm/transparent_hugepage/hpage_pmd_size)
+case "$thp_size" in
+ ''|*[!0-9]*) skip "invalid PMD huge-page size: $thp_size" ;;
+esac
+[ "$thp_size" -gt 0 ] || skip "PMD huge-page size is zero"
+page_kib=$(awk '/KernelPageSize:/ { print $2; exit }' /proc/self/smaps)
+[ "${page_kib:-0}" -gt 0 ] || skip "cannot determine the base page size"
+page_size=$((page_kib * 1024))
+[ "$page_size" -gt 0 ] || skip "base page size is zero"
+[ $((thp_size % page_size)) -eq 0 ] ||
+ skip "PMD huge-page size is not page aligned"
+thp_kib=$((thp_size / 1024))
+thp_sectors=$((thp_size / 512))
+expected_retained_kib=$((thp_kib - page_kib))
+expected_refused=$((thp_size / page_size))
+
+tmp_dir=$(mktemp -d "${TMPDIR:-/var/tmp}/zram-retained.XXXXXX") ||
+ skip "cannot create temporary directory"
+tmp=$tmp_dir
+ready="$tmp/ready"
+touched="$tmp/touched"
+verified="$tmp/verified"
+trap cleanup EXIT
+trap 'exit 129' HUP
+trap 'exit 130' INT
+trap 'exit 143' TERM
+if [ -e /sys/module/zswap/parameters/enabled ]; then
+ zswap_enabled=$(cat /sys/module/zswap/parameters/enabled)
+ echo N > /sys/module/zswap/parameters/enabled ||
+ skip "cannot disable zswap"
+fi
+
+# Swap priorities are global. The ordinary zram device must be preferred to
+# any pre-existing swap, while the delayed offload device remains first for
+# eligible proactive reclaim.
+max_prio=$(awk 'BEGIN { max = -1 } NR > 1 && $5 > max { max = $5 } END { print max }' /proc/swaps)
+[ "$max_prio" -le 32765 ] ||
+ skip "cannot outrank existing swap priority $max_prio"
+safe_prio=$((max_prio + 1))
+offload_prio=$((max_prio + 2))
+
+dev_num=2
+zram_size=$((thp_size * 4))
+[ "$zram_size" -ge 67108864 ] || zram_size=67108864
+zram_sizes="$zram_size $zram_size"
+zram_load
+zram_set_disksizes
+set -- $dev_ids
+safe="/dev/zram${1}"
+offload_backing="/dev/zram${2}"
+safe_id=$1
+
+sectors=$(blockdev --getsz "$offload_backing") ||
+ skip "cannot read offload backing size"
+[ "$sectors" -gt 0 ] || skip "offload backing has zero size"
+dm_table="0 $sectors delay $offload_backing 0 0 $offload_backing 0 3000"
+dmsetup --noudevsync --noudevrules create "$dm_name" --table "$dm_table" ||
+ skip "cannot create delayed offload device"
+dm_active=1
+dmsetup --noudevsync --noudevrules info --columns --noheadings \
+ --separator ' ' -o major,minor "$dm_name" > "$tmp/dm-devno" ||
+ skip "cannot identify delayed offload device"
+read -r dm_major dm_minor < "$tmp/dm-devno"
+offload="/dev/mapper/$dm_name"
+if [ ! -e "$offload" ]; then
+ mkdir -p /dev/mapper
+ if mknod "$offload" b "$dm_major" "$dm_minor"; then
+ dm_node_created=1
+ fi
+fi
+[ -b "$offload" ] &&
+ [ "$(stat -L -c '%t:%T' "$offload")" = \
+ "$(printf '%x:%x' "$dm_major" "$dm_minor")" ] ||
+ skip "delayed offload device node is unavailable or has the wrong number"
+dm_inflight="/sys/dev/block/$dm_major:$dm_minor/inflight"
+[ -r "$dm_inflight" ] || skip "cannot observe delayed offload I/O"
+
+mkswap "$safe" >/dev/null || fail "cannot initialise safe swap"
+mkswap "$offload" >/dev/null || fail "cannot initialise offload swap"
+swapon -p "$safe_prio" "$safe" || fail "cannot activate safe swap"
+dev_swap_ids=" $safe_id"
+safe_active=1
+./swap_offload activate-no-discard "$offload" "$offload_prio" ||
+ fail "cannot activate offload-only swap"
+offload_active=1
+
+cgroup_enable_memory_controller "$cgroup_root" ||
+ skip "cannot enable the cgroup v2 memory controller"
+mkdir "$cg" || skip "cannot create test cgroup"
+cg_created=1
+[ -e "$cg/memory.max" ] ||
+ skip "cgroup v2 memory controller is unavailable"
+echo max > "$cg/memory.swap.max"
+
+./swap_offload retained "$thp_size" "$cg/cgroup.procs" "$ready" \
+ "$touched" "$verified" &
+holder_pid=$!
+require_helper_file "$ready" "creating a PMD-sized anonymous huge folio"
+read -r reported_pid thp_vma_range < "$ready"
+[ "$reported_pid" -eq "$holder_pid" ] ||
+ fail "retained helper reported the wrong pid"
+# On tmpfs, the helper's marker data is reclaimable and charged to its cgroup.
+rm "$ready" || fail "cannot remove consumed readiness marker"
+measure_thp_vma_swap 1 ||
+ fail "target is not an unlocked PMD-sized huge VMA before reclaim"
+echo "$TCID: thp_vma=$thp_vma_range"
+ancillary_locked_kib=$(awk '/VmLck:/ { print $2 }' "/proc/$holder_pid/status")
+[ "${ancillary_locked_kib:-0}" -gt 0 ] ||
+ fail "ancillary helper mappings are not locked"
+echo "$TCID: ancillary_locked_kib=$ancillary_locked_kib target_locked_kib=0"
+
+# This is the only reclaim before the retained entry is made dirty. It puts
+# the whole folio on offload-only swap so the later single-page fault leaves
+# sibling swap PTEs referring to the existing slot.
+setup_thp_before=$(memcg_counter thp_swpout) ||
+ skip "large-folio swapout counter is unavailable"
+setup_fallback_before=$(memcg_counter thp_swpout_fallback) ||
+ skip "large-folio fallback counter is unavailable"
+echo "$thp_size swappiness=max" > "$cg/memory.reclaim" ||
+ echo "$TCID: setup reclaim was incomplete; checking resulting state" >&2
+wait_offload_quiet || fail "offload device did not quiesce after setup"
+setup_thp_after=$(memcg_counter thp_swpout) ||
+ fail "large-folio swapout counter disappeared"
+setup_fallback_after=$(memcg_counter thp_swpout_fallback) ||
+ fail "large-folio fallback counter disappeared"
+[ "$setup_thp_after" -ge "$setup_thp_before" ] &&
+ [ "$setup_fallback_after" -ge "$setup_fallback_before" ] ||
+ fail "large-folio counters went backwards"
+if [ "$setup_thp_after" -eq "$setup_thp_before" ]; then
+ [ "$setup_fallback_after" -gt "$setup_fallback_before" ] &&
+ skip "PMD-sized folio fell back to base-page swapout"
+ fail "setup did not swap out a PMD-sized folio"
+fi
+kill -0 "$holder_pid" || fail "retained helper died during setup reclaim"
+kill -USR1 "$holder_pid" || fail "cannot request the dirty-page transition"
+require_helper_file "$touched" "faulting and dirtying the retained entry"
+measure_thp_vma_swap 0 ||
+ fail "PMD-sized huge VMA was missing, split, merged, or changed size"
+read -r retained_kib < "$tmp/thp-vma-swap" ||
+ fail "cannot read PMD-sized huge VMA swap usage"
+[ "${retained_kib:-0}" -eq "$expected_retained_kib" ] ||
+ fail "retained $retained_kib KiB, expected $expected_retained_kib KiB"
+echo "$TCID: retained_kib=$retained_kib"
+
+wait_offload_quiet || fail "offload device did not quiesce before pressure"
+ordinary_before=$(written_sectors)
+refused_before=$(vmstat_value swpout_offload_refused)
+echo $((thp_size / 2)) > "$cg/memory.high" ||
+ fail "ordinary pressure reclaim failed"
+wait_offload_quiet || fail "offload device did not quiesce after pressure"
+ordinary_after=$(written_sectors)
+refused_after=$(vmstat_value swpout_offload_refused)
+ordinary_writes=$((ordinary_after - ordinary_before))
+refused_pages=$((refused_after - refused_before))
+echo "$TCID: ordinary_sectors=$ordinary_writes refused_pages=$refused_pages"
+[ "$ordinary_writes" -eq 0 ] ||
+ fail "ordinary reclaim wrote $ordinary_writes offload sectors"
+[ "$refused_pages" -ge "$expected_refused" ] ||
+ fail "ordinary reclaim refused $refused_pages pages, expected at least $expected_refused"
+kill -0 "$holder_pid" || fail "retained helper died after refused write"
+rm -f "$verified"
+kill -USR2 "$holder_pid" || fail "cannot request dirty-byte verification"
+require_helper_file "$verified" "verifying the dirty retained byte"
+
+# Proactive reclaim must still be able to rewrite the retained PMD-sized slot.
+# Repeated requests make the test insensitive to a short-lived writeback
+# collision; the backing-sector delta still requires exactly one folio write.
+wait_offload_quiet || fail "offload device did not quiesce before recovery"
+recovery_before=$(written_sectors)
+echo max > "$cg/memory.high"
+for _ in $(seq 1 8); do
+ echo "$thp_size swappiness=max" > "$cg/memory.reclaim" 2>/dev/null || :
+done
+wait_offload_quiet || fail "offload device did not quiesce after recovery"
+recovery_after=$(written_sectors)
+recovery_writes=$((recovery_after - recovery_before))
+echo "$TCID: recovery_sectors=$recovery_writes"
+[ "$recovery_writes" -eq "$thp_sectors" ] ||
+ fail "proactive recovery wrote $recovery_writes sectors, expected $thp_sectors"
+kill -0 "$holder_pid" || fail "retained helper died during recovery"
+rm -f "$verified"
+kill -ALRM "$holder_pid" || fail "cannot request full data verification"
+require_helper_file "$verified" "verifying all retained data"
+kill -0 "$holder_pid" || fail "retained helper failed full data verification"
+
+kill "$holder_pid" || fail "cannot stop retained helper"
+if wait "$holder_pid"; then
+ :
+else
+ status=$?
+ [ "$status" -eq 143 ] || fail "retained helper exited with status $status"
+fi
+holder_pid=""
+swapoff "$offload" || fail "cannot deactivate offload-only swap"
+offload_active=0
+swapoff "$safe" || fail "cannot deactivate safe swap"
+safe_active=0
+dev_swap_ids=""
+dmsetup --noudevsync --noudevrules remove "$dm_name" ||
+ fail "cannot remove delayed offload device"
+dm_active=0
+if [ "$dm_node_created" -eq 1 ]; then
+ rm -f "$offload" || fail "cannot remove delayed offload device node"
+ dm_node_created=0
+fi
+rmdir "$cg" || fail "cannot remove test cgroup"
+cg=""
+cg_created=0
+cgroup_disable_memory_controller "$cgroup_root" ||
+ fail "cannot restore the cgroup memory controller"
+zram_cleanup || fail "cannot clean up zram devices"
+dev_ids=""
+if [ -n "$zswap_enabled" ]; then
+ echo "$zswap_enabled" > /sys/module/zswap/parameters/enabled ||
+ fail "cannot restore zswap"
+ zswap_enabled=""
+fi
+rm -rf "$tmp" || fail "cannot remove temporary files"
+tmp=""
+
+echo "$TCID: [PASS]"
--
2.55.0
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload
2026-09-26 4:55 [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload Matthias Goergens
` (3 preceding siblings ...)
2026-09-26 4:55 ` [RFC PATCH v2 4/4] selftests: zram: cover retained offload-only entries Matthias Goergens
@ 2026-09-26 10:16 ` Kairui Song
2026-09-28 15:19 ` Matthias Goergens
2026-09-26 23:32 ` Chris Li
` (3 subsequent siblings)
8 siblings, 1 reply; 16+ messages in thread
From: Kairui Song @ 2026-09-26 10:16 UTC (permalink / raw)
To: Matthias Goergens
Cc: Andrew Morton, Chris Li, Youngjun Park, Baoquan He,
Johannes Weiner, David Hildenbrand, Michal Hocko, Shakeel Butt,
Kemeng Shi, Nhat Pham, Yosry Ahmed, Barry Song, linux-mm,
linux-kernel, cgroups, linux-api, Alejandro Colomar, linux-man,
Karel Zak, util-linux, Jani Nikula, Joonas Lahtinen,
Rodrigo Vivi, Tvrtko Ursulin, David Airlie, Simona Vetter,
intel-gfx, dri-devel, Rafael J . Wysocki, Pavel Machek,
Catalin Marinas, Will Deacon, linux-pm, linux-arm-kernel,
Shuah Khan, linux-kselftest
On Sat, Sep 26, 2026 at 12:55 PM Matthias Goergens
<matthias.goergens@gmail.com> wrote:
>
> This RFC adds offload-only swap areas for deliberate cold-page offload to
> backends unsuitable for pressure reclaim. ZFS zvol swap has documented
> deadlocks under memory pressure [3]; compressed swap is another target
> because writes may need memory despite free logical slots.
>
> Explicit proactive reclaim can use these areas alongside conventional
> swap, ordered by priority. Ordinary reclaim cannot initiate non-zero
> writes to them, including through retained swap entries. Conventional
> capacity is not reserved for emergencies. Operators choose the policy;
> the kernel does not measure headroom or make an allocating backend safe.
>
> TMO/Senpai [4] and DAMON_RECLAIM [5] are related cold-page reclaim
> approaches. This RFC admits offload-only swap through memory.reclaim and
> per-node reclaim, but not DAMON. Other related work includes per-cgroup
> zswap writeback control [6], Virtual Swap Space v4 [1] and swap tiers v10
> [2].
>
> The util-linux companion [8] proposes swapon --offload-only and an fstab
> option. Older swapon silently ignores the fstab token, so persistent
> activation needs discussion. Apply the separate i915 fix [9] first to avoid
> a pre-existing folio-lock leak when shmem writeback is skipped.
>
> The series fixes zram selftest device tracking and error reporting, adds
> the swap policy with DRM eligibility checks, and tests routing, workingset
> activation and retained-entry write refusal. It is based on mm-new at
> 995829088503.
>
> Feedback is particularly welcome on:
>
> - Activation-time eligibility versus backend placement or migration:
> is retained-entry refusal, with possible reclaim churn or OOM,
> acceptable without migration?
> - Eligible-capacity accounting, including overcommit limits and OOM
> scoring; cache recovery, cluster invalidation and workingset activation.
> - Persistent activation and visibility: current swap listings do not
> expose the policy.
>
> Validation includes x86 builds, affected arm64 MTE objects and builds
> without swap or memory cgroups. QEMU tests covered routing, retained-entry
> recovery, data integrity and cleanup, with negative controls for write
> refusal and teardown failure.
>
> I have also been using this policy on my own machine with experimental
> bcachefs swap support, with no problems observed so far. Deliberate stress
> testing is confined to VMs; I do not deliberately stress-test this machine.
>
> Changes since v1, incorporating Sashiko's public review [7]:
>
> - Improve selftest isolation, retained-swap measurement and memlock skips.
> - Track allocated zram devices, wait for udev probes and report cleanup
> errors without deleting data through a mount that could not be released.
> - Clarify activation, capacity reporting and conventional fallback limits.
>
> [1] https://patchew.org/linux/20260825153238.2695446-1-nphamcs@gmail.com/
> [2] https://patchew.org/linux/20260713025644.170839-1-youngjun.park@lge.com/
> [3] https://openzfs.github.io/openzfs-docs/Project%20and%20Community/FAQ.html#using-a-zvol-for-a-swap-device-on-linux
> [4] https://engineering.fb.com/2022/06/20/data-infrastructure/transparent-memory-offloading-more-memory-at-a-fraction-of-the-cost-and-power/
> [5] https://www.kernel.org/doc/html/latest/admin-guide/mm/damon/reclaim.html
> [6] https://www.kernel.org/doc/html/latest/admin-guide/mm/zswap.html
> [7] https://sashiko.dev/#/patchset/20260914104501.3960616-1-matthias.goergens@gmail.com
> [8] https://github.com/util-linux/util-linux/pull/4633
> [9] https://lore.kernel.org/all/20260914105606.3997649-1-matthias.goergens@gmail.com/
>
> v1: https://lore.kernel.org/all/20260914104501.3960616-1-matthias.goergens@gmail.com/
>
> Matthias Goergens (4):
> selftests: zram: track owned devices and report cleanup failures
> mm: restrict offload-only swap to proactive reclaim
> selftests: zram: cover offload-only swap policy
> selftests: zram: cover retained offload-only entries
>
> Documentation/mm/swap.rst | 105 ++++
> drivers/gpu/drm/i915/gem/i915_gem_shrinker.c | 2 +-
> .../gpu/drm/i915/gem/selftests/huge_pages.c | 6 +-
> drivers/gpu/drm/msm/msm_gem_shrinker.c | 2 +-
> drivers/gpu/drm/panthor/panthor_gem.c | 2 +-
> drivers/gpu/drm/ttm/ttm_backup.c | 2 +-
> drivers/gpu/drm/xe/tests/xe_bo.c | 2 +-
> include/linux/swap.h | 29 +-
> include/linux/vm_event_item.h | 1 +
> mm/memcontrol.c | 20 +-
> mm/page_io.c | 24 +-
> mm/swapfile.c | 123 ++++-
> mm/vmscan.c | 31 +-
> mm/vmstat.c | 1 +
> tools/testing/selftests/zram/.gitignore | 2 +
> tools/testing/selftests/zram/Makefile | 4 +-
> tools/testing/selftests/zram/README | 20 +-
> tools/testing/selftests/zram/config | 10 +-
> tools/testing/selftests/zram/settings | 1 +
> tools/testing/selftests/zram/swap_offload.c | 425 +++++++++++++++++
> .../selftests/zram/workingset_offload.c | 207 ++++++++
> tools/testing/selftests/zram/zram.sh | 9 +
> tools/testing/selftests/zram/zram01.sh | 9 +-
> tools/testing/selftests/zram/zram02.sh | 9 +-
> tools/testing/selftests/zram/zram03.sh | 174 +++++++
> tools/testing/selftests/zram/zram04.sh | 164 +++++++
> tools/testing/selftests/zram/zram05.sh | 447 ++++++++++++++++++
> tools/testing/selftests/zram/zram_lib.sh | 189 ++++++--
> 28 files changed, 1926 insertions(+), 94 deletions(-)
Hi Matthias
That's a lot of changes for a rather limited usage. IIUC what it want
to archive is, for different kind of reclaim, swap operations should
only go to certain devices?
This really sounds like another usage of the swap tiering design: A
default per-cgroup tier setup, and a one time tier limit during
proactive reclaim.
Just like we already have a "swappiness=" parameter in the reclaim
interface, perhaps adding a "swap.tier=" would be better and much
cleaner to achieve the same goal? Based on Youngjun's work here:
https://lore.kernel.org/all/20260916183437.2946306-1-youngjun.park@lge.com/
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload
2026-09-26 4:55 [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload Matthias Goergens
` (4 preceding siblings ...)
2026-09-26 10:16 ` [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload Kairui Song
@ 2026-09-26 23:32 ` Chris Li
2026-09-27 0:03 ` Chris Li
2026-09-27 17:33 ` Andy Lutomirski
` (2 subsequent siblings)
8 siblings, 1 reply; 16+ messages in thread
From: Chris Li @ 2026-09-26 23:32 UTC (permalink / raw)
To: Matthias Goergens
Cc: Andrew Morton, Kairui Song, Johannes Weiner, David Hildenbrand,
Michal Hocko, Shakeel Butt, Kemeng Shi, Nhat Pham, Yosry Ahmed,
Youngjun Park, Baoquan He, Barry Song, linux-mm, linux-kernel,
cgroups, linux-api, Alejandro Colomar, linux-man, Karel Zak,
util-linux, Jani Nikula, Joonas Lahtinen, Rodrigo Vivi,
Tvrtko Ursulin, David Airlie, Simona Vetter, intel-gfx,
dri-devel, Rafael J . Wysocki, Pavel Machek, Catalin Marinas,
Will Deacon, linux-pm, linux-arm-kernel, Shuah Khan,
linux-kselftest
On Fri, Sep 25, 2026 at 6:55 PM Matthias Goergens
<matthias.goergens@gmail.com> wrote:
> The series fixes zram selftest device tracking and error reporting, adds
> the swap policy with DRM eligibility checks, and tests routing, workingset
> activation and retained-entry write refusal. It is based on mm-new at
> 995829088503.
Hi Matthias,
I can't find mm-new 995829088503 anymore.
Your patch does not apply cleanly to the current tip of mm-new
mm-unstable, or mm-stable branches.
Do you have a git repo with your patch applied?
mm-new is whatever Andrew can find on the mailing list, regardless of
whether it has been reviewed. That mm-new branch is very unstable,
even more so than mm-unstable, which should at least have some review
already. Maybe consider using mm-stable or mm-unstable as base next
time.
Chris
>
> Feedback is particularly welcome on:
>
> - Activation-time eligibility versus backend placement or migration:
> is retained-entry refusal, with possible reclaim churn or OOM,
> acceptable without migration?
> - Eligible-capacity accounting, including overcommit limits and OOM
> scoring; cache recovery, cluster invalidation and workingset activation.
> - Persistent activation and visibility: current swap listings do not
> expose the policy.
>
> Validation includes x86 builds, affected arm64 MTE objects and builds
> without swap or memory cgroups. QEMU tests covered routing, retained-entry
> recovery, data integrity and cleanup, with negative controls for write
> refusal and teardown failure.
>
> I have also been using this policy on my own machine with experimental
> bcachefs swap support, with no problems observed so far. Deliberate stress
> testing is confined to VMs; I do not deliberately stress-test this machine.
>
> Changes since v1, incorporating Sashiko's public review [7]:
>
> - Improve selftest isolation, retained-swap measurement and memlock skips.
> - Track allocated zram devices, wait for udev probes and report cleanup
> errors without deleting data through a mount that could not be released.
> - Clarify activation, capacity reporting and conventional fallback limits.
>
> [1] https://patchew.org/linux/20260825153238.2695446-1-nphamcs@gmail.com/
> [2] https://patchew.org/linux/20260713025644.170839-1-youngjun.park@lge.com/
> [3] https://openzfs.github.io/openzfs-docs/Project%20and%20Community/FAQ.html#using-a-zvol-for-a-swap-device-on-linux
> [4] https://engineering.fb.com/2022/06/20/data-infrastructure/transparent-memory-offloading-more-memory-at-a-fraction-of-the-cost-and-power/
> [5] https://www.kernel.org/doc/html/latest/admin-guide/mm/damon/reclaim.html
> [6] https://www.kernel.org/doc/html/latest/admin-guide/mm/zswap.html
> [7] https://sashiko.dev/#/patchset/20260914104501.3960616-1-matthias.goergens@gmail.com
> [8] https://github.com/util-linux/util-linux/pull/4633
> [9] https://lore.kernel.org/all/20260914105606.3997649-1-matthias.goergens@gmail.com/
>
> v1: https://lore.kernel.org/all/20260914104501.3960616-1-matthias.goergens@gmail.com/
>
> Matthias Goergens (4):
> selftests: zram: track owned devices and report cleanup failures
> mm: restrict offload-only swap to proactive reclaim
> selftests: zram: cover offload-only swap policy
> selftests: zram: cover retained offload-only entries
>
> Documentation/mm/swap.rst | 105 ++++
> drivers/gpu/drm/i915/gem/i915_gem_shrinker.c | 2 +-
> .../gpu/drm/i915/gem/selftests/huge_pages.c | 6 +-
> drivers/gpu/drm/msm/msm_gem_shrinker.c | 2 +-
> drivers/gpu/drm/panthor/panthor_gem.c | 2 +-
> drivers/gpu/drm/ttm/ttm_backup.c | 2 +-
> drivers/gpu/drm/xe/tests/xe_bo.c | 2 +-
> include/linux/swap.h | 29 +-
> include/linux/vm_event_item.h | 1 +
> mm/memcontrol.c | 20 +-
> mm/page_io.c | 24 +-
> mm/swapfile.c | 123 ++++-
> mm/vmscan.c | 31 +-
> mm/vmstat.c | 1 +
> tools/testing/selftests/zram/.gitignore | 2 +
> tools/testing/selftests/zram/Makefile | 4 +-
> tools/testing/selftests/zram/README | 20 +-
> tools/testing/selftests/zram/config | 10 +-
> tools/testing/selftests/zram/settings | 1 +
> tools/testing/selftests/zram/swap_offload.c | 425 +++++++++++++++++
> .../selftests/zram/workingset_offload.c | 207 ++++++++
> tools/testing/selftests/zram/zram.sh | 9 +
> tools/testing/selftests/zram/zram01.sh | 9 +-
> tools/testing/selftests/zram/zram02.sh | 9 +-
> tools/testing/selftests/zram/zram03.sh | 174 +++++++
> tools/testing/selftests/zram/zram04.sh | 164 +++++++
> tools/testing/selftests/zram/zram05.sh | 447 ++++++++++++++++++
> tools/testing/selftests/zram/zram_lib.sh | 189 ++++++--
> 28 files changed, 1926 insertions(+), 94 deletions(-)
> create mode 100644 tools/testing/selftests/zram/settings
> create mode 100644 tools/testing/selftests/zram/swap_offload.c
> create mode 100644 tools/testing/selftests/zram/workingset_offload.c
> create mode 100755 tools/testing/selftests/zram/zram03.sh
> create mode 100755 tools/testing/selftests/zram/zram04.sh
> create mode 100755 tools/testing/selftests/zram/zram05.sh
>
> --
> 2.55.0
>
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload
2026-09-26 23:32 ` Chris Li
@ 2026-09-27 0:03 ` Chris Li
0 siblings, 0 replies; 16+ messages in thread
From: Chris Li @ 2026-09-27 0:03 UTC (permalink / raw)
To: Matthias Goergens
Cc: Andrew Morton, Kairui Song, Johannes Weiner, David Hildenbrand,
Michal Hocko, Shakeel Butt, Kemeng Shi, Nhat Pham, Yosry Ahmed,
Youngjun Park, Baoquan He, Barry Song, linux-mm, linux-kernel,
cgroups, linux-api, Alejandro Colomar, linux-man, Karel Zak,
util-linux, Jani Nikula, Joonas Lahtinen, Rodrigo Vivi,
Tvrtko Ursulin, David Airlie, Simona Vetter, intel-gfx,
dri-devel, Rafael J . Wysocki, Pavel Machek, Catalin Marinas,
Will Deacon, linux-pm, linux-arm-kernel, Shuah Khan,
linux-kselftest
On Sat, Sep 26, 2026 at 1:32 PM Chris Li <chrisl@kernel.org> wrote:
>
> On Fri, Sep 25, 2026 at 6:55 PM Matthias Goergens
> <matthias.goergens@gmail.com> wrote:
> > The series fixes zram selftest device tracking and error reporting, adds
> > the swap policy with DRM eligibility checks, and tests routing, workingset
> > activation and retained-entry write refusal. It is based on mm-new at
> > 995829088503.
>
> Hi Matthias,
>
> I can't find mm-new 995829088503 anymore.
>
> Your patch does not apply cleanly to the current tip of mm-new
> mm-unstable, or mm-stable branches.
> Do you have a git repo with your patch applied?
Never mind, I have the conflict resolved on mm-unstable.
Sorry I usually need to have the patch appliable to use my own review scripts.
Chris
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload
2026-09-26 4:55 [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload Matthias Goergens
` (5 preceding siblings ...)
2026-09-26 23:32 ` Chris Li
@ 2026-09-27 17:33 ` Andy Lutomirski
2026-09-28 15:24 ` Matthias Goergens
2026-09-28 0:19 ` Chris Li
2026-09-28 5:19 ` Christoph Hellwig
8 siblings, 1 reply; 16+ messages in thread
From: Andy Lutomirski @ 2026-09-27 17:33 UTC (permalink / raw)
To: Matthias Goergens
Cc: Andrew Morton, Chris Li, Kairui Song, Johannes Weiner,
David Hildenbrand, Michal Hocko, Shakeel Butt, Kemeng Shi,
Nhat Pham, Yosry Ahmed, Youngjun Park, Baoquan He, Barry Song,
linux-mm, linux-kernel, cgroups, linux-api, Alejandro Colomar,
linux-man, Karel Zak, util-linux, Jani Nikula, Joonas Lahtinen,
Vivi Rodrigo, Tvrtko Ursulin, David Airlie, Simona Vetter,
intel-gfx, dri-devel, Rafael J Wysocki, Pavel Machek,
Catalin Marinas, Will Deacon, linux-pm, linux-arm-kernel,
Shuah Khan, linux-kselftest, Matthias Goergens
> On Sep 25, 2026, at 9:55 PM, Matthias Goergens <matthias.goergens@gmail.com> wrote:
>
> This RFC adds offload-only swap areas for deliberate cold-page offload to
> backends unsuitable for pressure reclaim. ZFS zvol swap has documented
> deadlocks under memory pressure [3]; compressed swap is another target
> because writes may need memory despite free logical slots.
This seems unnecessarily annoying to configure. It seems to me that you’re creating a bizarrely named administrative flag that means, roughly, “swapping to this location may allocate memory”. Why can’t the kernel figure this out itself?
For that matter, how well does this even work in practice? If I configure two swap devices, one “offload-only” and one conventional, it seems like the relative fullness of the devices will be mostly an accident of what triggers swap as the system is running. If the conventional swap fills up, is there a means to proactively empty it? Under OOM conditions, will anything preferentially kill tasks that reference memory in conventional swap?
Is the locking and recursion structure of the mm code such that we won’t have deadlocks where cgroup triggers swap to an “offload-only” device, which allocates, which causes memory pressure, which then tries to swap more to conventional swap? (Maybe this works fine.)
For that matter, if the system is under overall memory pressure, why do you care what triggered the particular swap operation that is being processed?
The remainder of the writeup is IMO somewhat incoherent. To the extent that AI was used, can you read it and make sure it makes sense?
And the first few hundred lines of patch I skimmed seem like incomprehensible churn. Adding “eligible” to a bunch of calls does nothing to explain what’s going on.
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload
2026-09-26 4:55 [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload Matthias Goergens
` (6 preceding siblings ...)
2026-09-27 17:33 ` Andy Lutomirski
@ 2026-09-28 0:19 ` Chris Li
2026-09-28 15:19 ` Matthias Goergens
2026-09-28 5:19 ` Christoph Hellwig
8 siblings, 1 reply; 16+ messages in thread
From: Chris Li @ 2026-09-28 0:19 UTC (permalink / raw)
To: Matthias Goergens
Cc: Andrew Morton, Kairui Song, Johannes Weiner, David Hildenbrand,
Michal Hocko, Shakeel Butt, Kemeng Shi, Nhat Pham, Yosry Ahmed,
Youngjun Park, Baoquan He, Barry Song, linux-mm, linux-kernel,
cgroups, linux-api, Alejandro Colomar, linux-man, Karel Zak,
util-linux, Jani Nikula, Joonas Lahtinen, Rodrigo Vivi,
Tvrtko Ursulin, David Airlie, Simona Vetter, intel-gfx,
dri-devel, Rafael J . Wysocki, Pavel Machek, Catalin Marinas,
Will Deacon, linux-pm, linux-arm-kernel, Shuah Khan,
linux-kselftest
On Fri, Sep 25, 2026 at 6:55 PM Matthias Goergens
<matthias.goergens@gmail.com> wrote:
>
> This RFC adds offload-only swap areas for deliberate cold-page offload to
> backends unsuitable for pressure reclaim. ZFS zvol swap has documented
> deadlocks under memory pressure [3]; compressed swap is another target
> because writes may need memory despite free logical slots.
What is the high-level user-visible impact of this series?
I am curious: if we never run out of swap file space on the
non-offload swap area, does that mean we don't need this patch series?
The series discusses proactive reclaim and direct reclaim. The
offload-only swap area is for proactive reclaim. However, direct
reclaim can use both types of swap areas. Wouldn't using the
non-offload case cause memory allocation which you want to avoid?
I was wondering if we should actually flip things around. Mark the
type of swap device that avoids memory allocation during swap out.
Then direct reclaim prioritizes using those.That would make the happy
path allocate less memory during direct reclaim. Right now this series
does not seem to guarantee that direct reclaim stays away from the
area that allocates memory.
Chris
>
> Explicit proactive reclaim can use these areas alongside conventional
> swap, ordered by priority. Ordinary reclaim cannot initiate non-zero
> writes to them, including through retained swap entries. Conventional
> capacity is not reserved for emergencies. Operators choose the policy;
> the kernel does not measure headroom or make an allocating backend safe.
>
> TMO/Senpai [4] and DAMON_RECLAIM [5] are related cold-page reclaim
> approaches. This RFC admits offload-only swap through memory.reclaim and
> per-node reclaim, but not DAMON. Other related work includes per-cgroup
> zswap writeback control [6], Virtual Swap Space v4 [1] and swap tiers v10
> [2].
>
> The util-linux companion [8] proposes swapon --offload-only and an fstab
> option. Older swapon silently ignores the fstab token, so persistent
> activation needs discussion. Apply the separate i915 fix [9] first to avoid
> a pre-existing folio-lock leak when shmem writeback is skipped.
>
> The series fixes zram selftest device tracking and error reporting, adds
> the swap policy with DRM eligibility checks, and tests routing, workingset
> activation and retained-entry write refusal. It is based on mm-new at
> 995829088503.
>
> Feedback is particularly welcome on:
>
> - Activation-time eligibility versus backend placement or migration:
> is retained-entry refusal, with possible reclaim churn or OOM,
> acceptable without migration?
> - Eligible-capacity accounting, including overcommit limits and OOM
> scoring; cache recovery, cluster invalidation and workingset activation.
> - Persistent activation and visibility: current swap listings do not
> expose the policy.
>
> Validation includes x86 builds, affected arm64 MTE objects and builds
> without swap or memory cgroups. QEMU tests covered routing, retained-entry
> recovery, data integrity and cleanup, with negative controls for write
> refusal and teardown failure.
>
> I have also been using this policy on my own machine with experimental
> bcachefs swap support, with no problems observed so far. Deliberate stress
> testing is confined to VMs; I do not deliberately stress-test this machine.
>
> Changes since v1, incorporating Sashiko's public review [7]:
>
> - Improve selftest isolation, retained-swap measurement and memlock skips.
> - Track allocated zram devices, wait for udev probes and report cleanup
> errors without deleting data through a mount that could not be released.
> - Clarify activation, capacity reporting and conventional fallback limits.
>
> [1] https://patchew.org/linux/20260825153238.2695446-1-nphamcs@gmail.com/
> [2] https://patchew.org/linux/20260713025644.170839-1-youngjun.park@lge.com/
> [3] https://openzfs.github.io/openzfs-docs/Project%20and%20Community/FAQ.html#using-a-zvol-for-a-swap-device-on-linux
> [4] https://engineering.fb.com/2022/06/20/data-infrastructure/transparent-memory-offloading-more-memory-at-a-fraction-of-the-cost-and-power/
> [5] https://www.kernel.org/doc/html/latest/admin-guide/mm/damon/reclaim.html
> [6] https://www.kernel.org/doc/html/latest/admin-guide/mm/zswap.html
> [7] https://sashiko.dev/#/patchset/20260914104501.3960616-1-matthias.goergens@gmail.com
> [8] https://github.com/util-linux/util-linux/pull/4633
> [9] https://lore.kernel.org/all/20260914105606.3997649-1-matthias.goergens@gmail.com/
>
> v1: https://lore.kernel.org/all/20260914104501.3960616-1-matthias.goergens@gmail.com/
>
> Matthias Goergens (4):
> selftests: zram: track owned devices and report cleanup failures
> mm: restrict offload-only swap to proactive reclaim
> selftests: zram: cover offload-only swap policy
> selftests: zram: cover retained offload-only entries
>
> Documentation/mm/swap.rst | 105 ++++
> drivers/gpu/drm/i915/gem/i915_gem_shrinker.c | 2 +-
> .../gpu/drm/i915/gem/selftests/huge_pages.c | 6 +-
> drivers/gpu/drm/msm/msm_gem_shrinker.c | 2 +-
> drivers/gpu/drm/panthor/panthor_gem.c | 2 +-
> drivers/gpu/drm/ttm/ttm_backup.c | 2 +-
> drivers/gpu/drm/xe/tests/xe_bo.c | 2 +-
> include/linux/swap.h | 29 +-
> include/linux/vm_event_item.h | 1 +
> mm/memcontrol.c | 20 +-
> mm/page_io.c | 24 +-
> mm/swapfile.c | 123 ++++-
> mm/vmscan.c | 31 +-
> mm/vmstat.c | 1 +
> tools/testing/selftests/zram/.gitignore | 2 +
> tools/testing/selftests/zram/Makefile | 4 +-
> tools/testing/selftests/zram/README | 20 +-
> tools/testing/selftests/zram/config | 10 +-
> tools/testing/selftests/zram/settings | 1 +
> tools/testing/selftests/zram/swap_offload.c | 425 +++++++++++++++++
> .../selftests/zram/workingset_offload.c | 207 ++++++++
> tools/testing/selftests/zram/zram.sh | 9 +
> tools/testing/selftests/zram/zram01.sh | 9 +-
> tools/testing/selftests/zram/zram02.sh | 9 +-
> tools/testing/selftests/zram/zram03.sh | 174 +++++++
> tools/testing/selftests/zram/zram04.sh | 164 +++++++
> tools/testing/selftests/zram/zram05.sh | 447 ++++++++++++++++++
> tools/testing/selftests/zram/zram_lib.sh | 189 ++++++--
> 28 files changed, 1926 insertions(+), 94 deletions(-)
> create mode 100644 tools/testing/selftests/zram/settings
> create mode 100644 tools/testing/selftests/zram/swap_offload.c
> create mode 100644 tools/testing/selftests/zram/workingset_offload.c
> create mode 100755 tools/testing/selftests/zram/zram03.sh
> create mode 100755 tools/testing/selftests/zram/zram04.sh
> create mode 100755 tools/testing/selftests/zram/zram05.sh
>
> --
> 2.55.0
>
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload
2026-09-26 4:55 [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload Matthias Goergens
` (7 preceding siblings ...)
2026-09-28 0:19 ` Chris Li
@ 2026-09-28 5:19 ` Christoph Hellwig
2026-09-28 15:19 ` Matthias Goergens
8 siblings, 1 reply; 16+ messages in thread
From: Christoph Hellwig @ 2026-09-28 5:19 UTC (permalink / raw)
To: Matthias Goergens
Cc: Andrew Morton, Chris Li, Kairui Song, Johannes Weiner,
David Hildenbrand, Michal Hocko, Shakeel Butt, Kemeng Shi,
Nhat Pham, Yosry Ahmed, Youngjun Park, Baoquan He, Barry Song,
linux-mm, linux-kernel, cgroups, linux-api, Alejandro Colomar,
linux-man, Karel Zak, util-linux, Jani Nikula, Joonas Lahtinen,
Rodrigo Vivi, Tvrtko Ursulin, David Airlie, Simona Vetter,
intel-gfx, dri-devel, Rafael J . Wysocki, Pavel Machek,
Catalin Marinas, Will Deacon, linux-pm, linux-arm-kernel,
Shuah Khan, linux-kselftest
On Sat, Sep 26, 2026 at 12:55:13PM +0800, Matthias Goergens wrote:
> This RFC adds offload-only swap areas for deliberate cold-page offload to
> backends unsuitable for pressure reclaim. ZFS zvol swap has documented
> deadlocks under memory pressure [3]; compressed swap is another target
> because writes may need memory despite free logical slots.
Out of tree code not matter. Out of tree that violates the linux
license terms even less.
Either way block devices must be safe to be called from paging paths,
not just for swap but also for file system based paging. So this is a
bug in the implementation, and not something worked around in the swap
code.
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload
2026-09-28 0:19 ` Chris Li
@ 2026-09-28 15:19 ` Matthias Goergens
2026-09-29 0:27 ` Chris Li
0 siblings, 1 reply; 16+ messages in thread
From: Matthias Goergens @ 2026-09-28 15:19 UTC (permalink / raw)
To: Chris Li
Cc: Andrew Morton, Kairui Song, Johannes Weiner, David Hildenbrand,
Michal Hocko, Shakeel Butt, Kemeng Shi, Nhat Pham, Yosry Ahmed,
Youngjun Park, Baoquan He, Barry Song, linux-mm, linux-kernel,
cgroups, linux-api, Alejandro Colomar, linux-man, Karel Zak,
util-linux, Jani Nikula, Joonas Lahtinen, Rodrigo Vivi,
Tvrtko Ursulin, David Airlie, Simona Vetter, intel-gfx,
dri-devel, Rafael J . Wysocki, Pavel Machek, Catalin Marinas,
Will Deacon, linux-pm, linux-arm-kernel, Shuah Khan,
linux-kselftest
Hi Chris,
Thanks for resolving the conflicts yourself and reviewing it.
> What is the high-level user-visible impact of this series?
Being able to use more interesting swap backends and logic when the
machine is not under memory pressure. It is much easier to write a
swap backend that may occasionally allocate memory than one that is
guaranteed never to, so today such backends are either unsafe as swap
or ruled out (btrfs, for example, refuses swapfiles that are
copy-on-write, checksummed or compressed).
> I am curious: if we never run out of swap file space on the
> non-offload swap area, does that mean we don't need this patch series?
No: running out of conventional swap isn't the point. Without the
series, any active swap area may be written under pressure, so a
backend that may allocate can't be used as swap at all, however much
conventional swap there is next to it.
> However, direct reclaim can use both types of swap areas.
In v2 it can't: only memory.reclaim, per-node reclaim and MGLRU's
debugfs eviction may write new data to an offload-only area; direct
reclaim, kswapd, MADV_PAGEOUT and DAMON reclaim may not. But I think
your suggestion to flip it round is closer to what I want: mark the
areas whose writes may allocate, and keep pressure reclaim away from
those, rather than tying it to what started the reclaim. I'll work
that into v3, together with Kairui's suggestion to build on the swap
tiers work.
Thanks,
Matthias
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload
2026-09-28 5:19 ` Christoph Hellwig
@ 2026-09-28 15:19 ` Matthias Goergens
0 siblings, 0 replies; 16+ messages in thread
From: Matthias Goergens @ 2026-09-28 15:19 UTC (permalink / raw)
To: Christoph Hellwig
Cc: Andrew Morton, Chris Li, Kairui Song, Johannes Weiner,
David Hildenbrand, Michal Hocko, Shakeel Butt, Kemeng Shi,
Nhat Pham, Yosry Ahmed, Youngjun Park, Baoquan He, Barry Song,
linux-mm, linux-kernel, cgroups, linux-api, Alejandro Colomar,
linux-man, Karel Zak, util-linux, Jani Nikula, Joonas Lahtinen,
Rodrigo Vivi, Tvrtko Ursulin, David Airlie, Simona Vetter,
intel-gfx, dri-devel, Rafael J . Wysocki, Pavel Machek,
Catalin Marinas, Will Deacon, linux-pm, linux-arm-kernel,
Shuah Khan, linux-kselftest
Hi Christoph,
> Either way block devices must be safe to be called from paging paths,
> not just for swap but also for file system based paging. So this is
> a bug in the implementation, and not something worked around in the
> swap code.
Agreed that a backend used for paging has to make progress under
pressure, and the zvol example was a poor lead. The series isn't
meant to excuse a backend that can deadlock; that would still be a bug
to fix in the backend.
What I'm after is different: backends that are correct but need
memory to accept a write, where the better policy is to use them for
cold-page offload when memory isn't tight and keep pressure reclaim on
areas that don't need to allocate. Two in-tree examples: zram
allocates on writes and fails with -ENOMEM, with no fallback, while its
logical size still shows free slots; and filesystem swapfiles must
give up copy-on-write, checksums and compression today, because swap
writes bypass the filesystem. I'll make that the case for v3 instead.
Thanks,
Matthias
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload
2026-09-26 10:16 ` [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload Kairui Song
@ 2026-09-28 15:19 ` Matthias Goergens
0 siblings, 0 replies; 16+ messages in thread
From: Matthias Goergens @ 2026-09-28 15:19 UTC (permalink / raw)
To: Kairui Song
Cc: Andrew Morton, Chris Li, Youngjun Park, Baoquan He,
Johannes Weiner, David Hildenbrand, Michal Hocko, Shakeel Butt,
Kemeng Shi, Nhat Pham, Yosry Ahmed, Barry Song, linux-mm,
linux-kernel, cgroups, linux-api, Alejandro Colomar, linux-man,
Karel Zak, util-linux, Jani Nikula, Joonas Lahtinen,
Rodrigo Vivi, Tvrtko Ursulin, David Airlie, Simona Vetter,
intel-gfx, dri-devel, Rafael J . Wysocki, Pavel Machek,
Catalin Marinas, Will Deacon, linux-pm, linux-arm-kernel,
Shuah Khan, linux-kselftest, Kairui Song
Hi Kairui,
> Just like we already have a "swappiness=" parameter in the reclaim
> interface, perhaps adding a "swap.tier=" would be better and much
> cleaner to achieve the same goal? Based on Youngjun's work here:
> https://lore.kernel.org/all/20260916183437.2946306-1-youngjun.park@lge.com/
Thanks, I think that's the right base. I tried it: on top of
Youngjun's v11, a "swap.tier=" argument to memory.reclaim is about 40
lines. In a VM with zram at a higher priority than a disk swap
device, and a cgroup whose tier mask allowed only the disk, ordinary
reclaim in that cgroup stayed on the disk, while memory.reclaim with
"swap.tier=" pointing at zram went to zram only.
What it doesn't cover yet is reclaim outside any configured cgroup:
under global pressure, a cgroup nobody configured still reached the
zram tier. So for v3 I'd like to add a system-wide limit on which
tiers pressure reclaim may use, on top of the per-cgroup settings.
I'll wait for Youngjun's v12 and build on that.
Thanks,
Matthias
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload
2026-09-27 17:33 ` Andy Lutomirski
@ 2026-09-28 15:24 ` Matthias Goergens
0 siblings, 0 replies; 16+ messages in thread
From: Matthias Goergens @ 2026-09-28 15:24 UTC (permalink / raw)
To: Andy Lutomirski
Cc: Andrew Morton, Chris Li, Kairui Song, Johannes Weiner,
David Hildenbrand, Michal Hocko, Shakeel Butt, Kemeng Shi,
Nhat Pham, Yosry Ahmed, Youngjun Park, Baoquan He, Barry Song,
linux-mm, linux-kernel, cgroups, linux-api, Alejandro Colomar,
linux-man, Karel Zak, util-linux, Jani Nikula, Joonas Lahtinen,
Vivi Rodrigo, Tvrtko Ursulin, David Airlie, Simona Vetter,
intel-gfx, dri-devel, Rafael J Wysocki, Pavel Machek,
Catalin Marinas, Will Deacon, linux-pm, linux-arm-kernel,
Shuah Khan, linux-kselftest
Hi Andy,
Thanks for reading it, and for the direct feedback.
> if the system is under overall memory pressure, why do you care what
> triggered the particular swap operation that is being processed?
You're right, and that question gets at what I actually want better
than the series does. The goal is: when there is no memory pressure,
swap-out may do more work, including allocating memory (compression,
copy-on-write, filesystem-backed swap), because moving cold pages out
to make room for page cache is what swap is for most of the time.
Under real pressure, swap-out must not depend on that.
The RFC was the smallest change I could think of in that direction.
It used "who started the reclaim" as a stand-in for "is there
pressure", and that is wrong both ways: memory.reclaim during global
pressure may allocate, while kswapd with plenty of free memory may
not. Building on Kairui's suggestion in this thread (swap tiers), or
on virtual swap, may well be the better route, and I'm looking at both
for v3.
> Why can't the kernel figure this out itself?
It can: the backend knows whether its writes may allocate (zram knows
it compresses, a filesystem knows at swapon whether the swapfile is
copy-on-write or compressed), so it should declare that, rather than
the administrator setting a flag.
On the practical questions: v2 does nothing to balance fullness
between areas or to empty conventional swap, and does not change OOM
selection. Those are fair gaps and I'll address them, or say
explicitly what is out of scope, in v3.
On the recursion question: within the reclaiming task it can't
happen. memory.reclaim and per-node reclaim both run with PF_MEMALLOC
set (memalloc_noreclaim_save() in try_to_free_mem_cgroup_pages() and
__node_reclaim()), so an allocation the backend makes in that task
never enters direct reclaim; it either fails or, unless it passes
__GFP_NOMEMALLOC, dips into the reserves. Two things do need care. A
backend allocating there can drain the emergency reserves unless it
passes __GFP_NOMEMALLOC or fails fast. And work the backend hands to
another thread, such as a filesystem or zvol worker, runs without
PF_MEMALLOC and can enter reclaim itself; that is only safe if that
reclaim never waits for the write the worker is meant to complete.
> The remainder of the writeup is IMO somewhat incoherent.
Agreed. I cut the v1 cover letter down too far and lost the
definitions it depended on. v3's will start from the goal above and
define its terms.
> Adding "eligible" to a bunch of calls does nothing to explain what's
> going on.
It exists because v2 made "how much swap is free" depend on who asks.
Reclaim itself checks free swap before scanning anonymous pages
(can_reclaim_anon_pages(), MGLRU's get_swappiness()) and before
allocating a slot (folio_alloc_swap()), and callers such as the GPU
shrinkers check it before pushing objects towards swap. With an
offload-only area, each of those would otherwise count space that the
calling reclaim may not use. In v3 I'd keep get_nr_swap_pages()
meaning "usable by ordinary reclaim", so those checks stay as they are
and only the offload path asks for a different count.
Thanks,
Matthias
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload
2026-09-28 15:19 ` Matthias Goergens
@ 2026-09-29 0:27 ` Chris Li
0 siblings, 0 replies; 16+ messages in thread
From: Chris Li @ 2026-09-29 0:27 UTC (permalink / raw)
To: Matthias Goergens
Cc: Andrew Morton, Kairui Song, Johannes Weiner, David Hildenbrand,
Michal Hocko, Shakeel Butt, Kemeng Shi, Nhat Pham, Yosry Ahmed,
Youngjun Park, Baoquan He, Barry Song, linux-mm, linux-kernel,
cgroups, linux-api, Alejandro Colomar, linux-man, Karel Zak,
util-linux, Jani Nikula, Joonas Lahtinen, Rodrigo Vivi,
Tvrtko Ursulin, David Airlie, Simona Vetter, intel-gfx,
dri-devel, Rafael J . Wysocki, Pavel Machek, Catalin Marinas,
Will Deacon, linux-pm, linux-arm-kernel, Shuah Khan,
linux-kselftest
On Mon, Sep 28, 2026 at 5:19 AM Matthias Goergens
<matthias.goergens@gmail.com> wrote:
>
> Hi Chris,
>
> Thanks for resolving the conflicts yourself and reviewing it.
>
> > What is the high-level user-visible impact of this series?
>
> Being able to use more interesting swap backends and logic when the
> machine is not under memory pressure. It is much easier to write a
> swap backend that may occasionally allocate memory than one that is
> guaranteed never to, so today such backends are either unsafe as swap
I actually don't know of a swap backend that absolutely will not
allocate memory on swap out yet. Some backends allocate more than
others. Because the proactive reclaim can write to the non-offloaded
swap backend. That brings me back to my original question, in what way
this series helps.
So the answer seems to be that previously, some backends were unusable
by swap, due to possible memory allocation. With this series, those
back ends are usable for the proactive reclaim now. It is a 0 to 0.5
improvement because direct reclaim can't use it yet.
> or ruled out (btrfs, for example, refuses swapfiles that are
> copy-on-write, checksummed or compressed).
>
> > I am curious: if we never run out of swap file space on the
> > non-offload swap area, does that mean we don't need this patch series?
>
> No: running out of conventional swap isn't the point. Without the
> series, any active swap area may be written under pressure, so a
> backend that may allocate can't be used as swap at all, however much
> conventional swap there is next to it.
But proactive reclaim can also use the non-offload swap area. So if we
have plenty of non-offload swap area, both proactive reclaim and
direct reclaim can use that non-offload area. The value of this series
isn't apparent. If you don't have a traditional swap-capable backend,
then your non-offload area is zero. That is considered a special case
of running out of non-offload swap area: never having a traditional
swap area in the first place.
>
> > However, direct reclaim can use both types of swap areas.
>
> In v2 it can't: only memory.reclaim, per-node reclaim and MGLRU's
Sorry I meant proactive reclaim can use both types, nothing constrains
proactive reclaim.
> debugfs eviction may write new data to an offload-only area; direct
> reclaim, kswapd, MADV_PAGEOUT and DAMON reclaim may not. But I think
> your suggestion to flip it round is closer to what I want: mark the
Flipping it around might provide additional benefit. I know some users
maintain off tree patches to turn off zswap on the direct reclaim path
exactly because zswap might allocate more memory before it can free
some. If the kernel can be smart about it. It provides additional
value to the status quo.
> areas whose writes may allocate, and keep pressure reclaim away from
> those, rather than tying it to what started the reclaim. I'll work
> that into v3, together with Kairui's suggestion to build on the swap
> tiers work.
Yes, I feel that cluster-level marking might not be needed, you can
remember which si has the offload flag. However the per cpu cache
needs to know which context might require using a different cached
cluster. That is the trickiest part. The rest of the patch is just
enough plumbing to preserve whether the swap-out context is proactive
or not.
Chris
^ permalink raw reply [flat|nested] 16+ messages in thread
end of thread, other threads:[~2026-09-29 0:27 UTC | newest]
Thread overview: 16+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-26 4:55 [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload Matthias Goergens
2026-09-26 4:55 ` [RFC PATCH v2 1/4] selftests: zram: track owned devices and report cleanup failures Matthias Goergens
2026-09-26 4:55 ` [RFC PATCH v2 2/4] mm: restrict offload-only swap to proactive reclaim Matthias Goergens
2026-09-26 4:55 ` [RFC PATCH v2 3/4] selftests: zram: cover offload-only swap policy Matthias Goergens
2026-09-26 4:55 ` [RFC PATCH v2 4/4] selftests: zram: cover retained offload-only entries Matthias Goergens
2026-09-26 10:16 ` [RFC PATCH v2 0/4] mm/swap: reserve swap areas for deliberate offload Kairui Song
2026-09-28 15:19 ` Matthias Goergens
2026-09-26 23:32 ` Chris Li
2026-09-27 0:03 ` Chris Li
2026-09-27 17:33 ` Andy Lutomirski
2026-09-28 15:24 ` Matthias Goergens
2026-09-28 0:19 ` Chris Li
2026-09-28 15:19 ` Matthias Goergens
2026-09-29 0:27 ` Chris Li
2026-09-28 5:19 ` Christoph Hellwig
2026-09-28 15:19 ` Matthias Goergens
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®