From: Johannes Weiner <hannes@cmpxchg.org>
To: Baoquan He <baoquan.he@linux.dev>
Cc: Baoquan He <hebaoquan@kylinos.cn>,
linux-mm@kvack.org, akpm@linux-foundation.org, chrisl@kernel.org,
kasong@tencent.com, nphamcs@gmail.com, baohua@kernel.org,
youngjun.park@lge.com, yosry@kernel.org,
shikemeng@huaweicloud.com, chengming.zhou@linux.dev,
david@kernel.org, linux-kernel@vger.kernel.org,
kunwu.chan@gmail.com
Subject: Re: [PATCH v3 00/14] mm, swap: extendable swap devices (xswap)
Date: Thu, 17 Sep 2026 09:17:12 -0400 [thread overview]
Message-ID: <20260917131712.GA1344@cmpxchg.org> (raw)
In-Reply-To: <aquWKJoP5JLP8Ci7@fedora>
On Thu, Sep 17, 2026 at 03:31:23PM +0800, Baoquan He wrote:
> On 09/16/26 at 12:45pm, Johannes Weiner wrote:
> > On Wed, Sep 16, 2026 at 06:19:07PM +0800, Baoquan He wrote:
> > > xswap is a swap device with no backing storage. Swapped-out pages live
> > > in zswap. Its cluster_info[] array lives in a VM_SPARSE vmalloc area,
> > > and the area is grown and shrunk on demand as swap usage changes.
> > >
> > > The problem being solved is the static size of compressed swap. Both
> > > zram and zswap need the size fixed in advance, and neither gives memory
> > > back when the workload shrinks. The solution should be a device whose
> > > size can scale up/down as per usage. xswap does that by mapping the
> > > metadata lazily instead of reserving it for the whole range.
> > >
> > > Design
> > > ------
> > > - si->cluster_info[] stays a plain array. Access is still
> > > &si->cluster_info[offset / SWAPFILE_CLUSTER]: no per-access branch, no
> > > RCU discipline, no tear-down state machine, no NULL return.
> > > - Only an initial chunk is mapped at creation. The rest of the address
> > > space is reserved, not allocated, so an idle device costs nothing.
> > > - Growth is driven by allocation. When no free cluster is left and the
> > > address space has room, the next chunk is mapped and added to the free
> > > list. No userspace involvement.
> > > - Shrink is driven by frees. The free tail is scanned, and whole chunks
> > > are unmapped once the mapped range is at most half in use and several
> > > chunks can go. One chunk is left mapped as slack, so the next
> > > allocation does not map it straight back. A ceiling lowered below the
> > > mapped range skips the half-in-use rule and is enforced at once.
> >
> > If the swap maintainers prefer the VM_SPARSE route, I'm happy to defer
> > to them on that.
> >
> > However, from the cgroup and zswap camp, two stipulations that I
> > reasoned out in the other thread[1]:
>
> >
> > 1. You must not charge compression space as swap space to the cgroup.
>
> Hmm, I don't have a stance on this. However, isn't this an issue
> zswap/zram have been doing? It feels like an independent issue which
> should be done separately?
If you have 3 containers using compression space, and two of them have
writeback enabled to a shared swapfile, the memory.swap.* controls
need to work to manage fair access to that swapfile. They do not work
if compression space itself is conflated in.
Right now zswap entries actually consume physical swapfile space, even
before writeback. Charging the space is correct. But the whole point
is to decouple compression space from physical swap space.
This is not something that can be done later. It would be a dramatic
user-visible change to how the resource is categorized and managed.
> > 2. You must make the compression space large enough to be outside the
> > range where users can hit space limits before hitting memory limits.
>
> We may need a way to define 'large enough' at first.
I've tried to lay this out in the other thread, and highlighted the
usability issues that result from hitting compression space limits
prematurely. It's kind of your call whether you want to seriously
engage with this or not.
But ultimately it's your claim that a static size can be made to work,
so it's on you to make a convincing case.
> > That also means not allowing setups where this is possible.
>
> And the limit is only an optional knob. If the admin does not set it,
> the device grows to the full address space, so there is no space limit
> to hit at all. It already behaves the way you want by default. The knob
> is only for admins who want a ceiling, they can use it or not. I hope
> this would not be a problem for your use case.
No, I've laid this out already as well.
This isn't about "my" usecase. It's about designing a coherent
interface that works well with a large number of usecases, and other
pieces of kernel infrastructure commonly used in conjunction.
The other proposal in the room needs no such interface. The burden of
proof for adding one is on you.
> > [1] https://lore.kernel.org/linux-mm/aqLi6cIjD2wJwk0B@cmpxchg.org/
next prev parent reply other threads:[~2026-09-17 13:17 UTC|newest]
Thread overview: 19+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-16 10:19 Baoquan He
2026-09-16 10:19 ` [PATCH v3 01/14] mm: xswap support for zswap Baoquan He
2026-09-16 10:19 ` [PATCH v3 02/14] mm, swap: add CONFIG_XSWAP and xswap fields to swap_info_struct Baoquan He
2026-09-16 10:19 ` [PATCH v3 03/14] mm, swap: refactor free_swap_cluster_info to take swap_info_struct Baoquan He
2026-09-16 10:19 ` [PATCH v3 04/14] mm, swap: add xswap cluster grow via VM_SPARSE vmalloc Baoquan He
2026-09-16 10:19 ` [PATCH v3 05/14] mm, swap: add sysfs create interface for xswap Baoquan He
2026-09-16 10:19 ` [PATCH v3 06/14] mm, swap: add xswap grow trigger on cluster allocation Baoquan He
2026-09-16 10:19 ` [PATCH v3 07/14] mm, swap: add xswap_try_shrink and shrink trigger on cluster free Baoquan He
2026-09-16 10:19 ` [PATCH v3 08/14] mm, swap: free backing pages in xswap_unmap_clusters Baoquan He
2026-09-16 10:19 ` [PATCH v3 09/14] mm, swap: defer xswap shrink to workqueue to avoid lock recursion Baoquan He
2026-09-16 10:19 ` [PATCH v3 10/14] mm, swap: refactor swapoff and add xswap_destroy Baoquan He
2026-09-16 10:19 ` [PATCH v3 11/14] mm, swap: require zswap for xswap devices Baoquan He
2026-09-16 10:19 ` [PATCH v3 12/14] mm, swap: cap xswap growth at nr_clusters Baoquan He
2026-09-16 10:19 ` [PATCH v3 13/14] mm, swap: add sysfs per-device size limit for xswap Baoquan He
2026-09-16 10:19 ` [PATCH v3 14/14] mm, swap: shrink xswap to the ceiling when it drops Baoquan He
2026-09-16 16:45 ` [PATCH v3 00/14] mm, swap: extendable swap devices (xswap) Johannes Weiner
2026-09-17 7:31 ` Baoquan He
2026-09-17 13:17 ` Johannes Weiner [this message]
2026-09-17 10:04 ` Baoquan He
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260917131712.GA1344@cmpxchg.org \
--to=hannes@cmpxchg.org \
--cc=akpm@linux-foundation.org \
--cc=baohua@kernel.org \
--cc=baoquan.he@linux.dev \
--cc=chengming.zhou@linux.dev \
--cc=chrisl@kernel.org \
--cc=david@kernel.org \
--cc=hebaoquan@kylinos.cn \
--cc=kasong@tencent.com \
--cc=kunwu.chan@gmail.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=nphamcs@gmail.com \
--cc=shikemeng@huaweicloud.com \
--cc=yosry@kernel.org \
--cc=youngjun.park@lge.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®