mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH 0/2] mm: zswap: reduce request contention on loads
@ 2026-10-06  0:22 Usama Arif
  2026-10-06  0:22 ` [PATCH 1/2] mm: zswap: use separate compression and decompression requests Usama Arif
                   ` (2 more replies)
  0 siblings, 3 replies; 4+ messages in thread
From: Usama Arif @ 2026-10-06  0:22 UTC (permalink / raw)
  To: Andrew Morton, chengming.zhou, dsterba, hannes, linux-kernel,
	linux-mm, nphamcs, terrelln, yosry, riel, shakeel.butt, alex,
	senozhatsky, kernel-team
  Cc: Usama Arif

Stores and loads share a per-CPU acomp request and mutex. A low-priority
store can be preempted right after the compressor drops its stream
lock, while it still holds the zswap mutex, and a higher-priority load
on that CPU then waits for the store to run again. This follows the work
from Sergey Senozhatsky's zram series which splits it for the same
reason [1].

Patch 1 gives compression and decompression separate requests, waits
and mutexes, so loads no longer wait for stores, though they can still
wait for each other. Patch 2 decompresses with an on-stack request when
the algorithm is synchronous and needs no request context, which covers
all in-tree software compressors, so those loads take no zswap lock.
Asynchronous algorithms keep the per-CPU request and mutex. For software
compressors the series allocates the same number of requests as before;
each per-CPU context grows by 72 bytes, and the load path is about 270
bytes deeper on x86-64.

The series does not fix two related cases:
- Stores still serialize on the compression mutex, so a high-priority
  task that reclaims (direct reclaim, MADV_PAGEOUT) can still wait for
  a preempted store.
- On PREEMPT_RT the codec stream locks are preemptible, so a load can
  still wait for a preempted store inside the codec.

The numbers below are the slowest read per run, as a median (min-max)
of 5 runs. Each run is 12 seconds in a zstd VM with lazy preemption,
vm.page-cluster=0 and swap on /dev/ram0. With 1 vCPU, four nice +10
workers page memory out and read it back while a nice 0 task spins. A
nice -19 reader pages out its own buffer and measures how long each
read of it takes. With 8 vCPUs there are 16 workers, 8 spinning tasks
and 8 readers.

              Before series (ms)      With series (ms)
  1 vCPU      22.3 (21.6-22.6)        0.97 (0.72-1.4)
  8 vCPUs     314 (97-2542)           7.0 (5.0-98)

Reads over 10 ms fell from 26-35 per run to none with 1 vCPU, and from
3-18 per run to at most one with 8 vCPUs. The benchmark and test programs
were written with the help of an LLM.

[1] https://lore.kernel.org/all/20261005122036.718976-10-senozhatsky@chromium.org/
 
Usama Arif (2):
  mm: zswap: use separate compression and decompression requests
  mm: zswap: use stack requests for synchronous decompression

 mm/zswap.c | 136 ++++++++++++++++++++++++++++++++++-------------------
 1 file changed, 88 insertions(+), 48 deletions(-)

-- 
2.53.0-Meta


^ permalink raw reply	[flat|nested] 4+ messages in thread

end of thread, other threads:[~2026-10-06  9:18 UTC | newest]

Thread overview: 4+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-06  0:22 [PATCH 0/2] mm: zswap: reduce request contention on loads Usama Arif
2026-10-06  0:22 ` [PATCH 1/2] mm: zswap: use separate compression and decompression requests Usama Arif
2026-10-06  0:22 ` [PATCH 2/2] mm: zswap: use stack requests for synchronous decompression Usama Arif
2026-10-06  9:18 ` [PATCH 0/2] mm: zswap: reduce request contention on loads Usama Arif

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®