mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Baoquan He <baoquan.he@linux.dev>
To: Chris Li <chrisl@kernel.org>, Johannes Weiner <hannes@cmpxchg.org>
Cc: Baoquan He <hebaoquan@kylinos.cn>,
	linux-mm@kvack.org, akpm@linux-foundation.org,
	kasong@tencent.com, nphamcs@gmail.com, baohua@kernel.org,
	youngjun.park@lge.com, yosry@kernel.org,
	shikemeng@huaweicloud.com, chengming.zhou@linux.dev,
	david@kernel.org, linux-kernel@vger.kernel.org,
	kunwu.chan@gmail.com
Subject: Re: [PATCH v3 00/14] mm, swap: extendable swap devices (xswap)
Date: Mon, 21 Sep 2026 18:04:46 +0800	[thread overview]
Message-ID: <arEBPsFJvzhQZk9W@fedora> (raw)
In-Reply-To: <CACePvbXQRBnTAivJp7mh-y3PQjjUeO4KxzV8+nnhC5UzMomT0w@mail.gmail.com>

On 09/20/26 at 11:52pm, Chris Li wrote:
> On Thu, Sep 17, 2026 at 3:17 AM Johannes Weiner <hannes@cmpxchg.org> wrote:
> >
> > On Thu, Sep 17, 2026 at 03:31:23PM +0800, Baoquan He wrote:
> > > On 09/16/26 at 12:45pm, Johannes Weiner wrote:
> > > > On Wed, Sep 16, 2026 at 06:19:07PM +0800, Baoquan He wrote:
> > > > > xswap is a swap device with no backing storage. Swapped-out pages live
> > > > > in zswap. Its cluster_info[] array lives in a VM_SPARSE vmalloc area,
> > > > > and the area is grown and shrunk on demand as swap usage changes.
> > > > >
> > > > > The problem being solved is the static size of compressed swap. Both
> > > > > zram and zswap need the size fixed in advance, and neither gives memory
> > > > > back when the workload shrinks. The solution should be a device whose
> > > > > size can scale up/down as per usage. xswap does that by mapping the
> > > > > metadata lazily instead of reserving it for the whole range.
> > > > >
> > > > > Design
> > > > > ------
> > > > > - si->cluster_info[] stays a plain array. Access is still
> > > > >   &si->cluster_info[offset / SWAPFILE_CLUSTER]: no per-access branch, no
> > > > >   RCU discipline, no tear-down state machine, no NULL return.
> > > > > - Only an initial chunk is mapped at creation. The rest of the address
> > > > >   space is reserved, not allocated, so an idle device costs nothing.
> > > > > - Growth is driven by allocation. When no free cluster is left and the
> > > > >   address space has room, the next chunk is mapped and added to the free
> > > > >   list. No userspace involvement.
> > > > > - Shrink is driven by frees. The free tail is scanned, and whole chunks
> > > > >   are unmapped once the mapped range is at most half in use and several
> > > > >   chunks can go. One chunk is left mapped as slack, so the next
> > > > >   allocation does not map it straight back. A ceiling lowered below the
> > > > >   mapped range skips the half-in-use rule and is enforced at once.
> > > >
> > > > If the swap maintainers prefer the VM_SPARSE route, I'm happy to defer
> > > > to them on that.
> > > >
> > > > However, from the cgroup and zswap camp, two stipulations that I
> > > > reasoned out in the other thread[1]:
> > >
> > > >
> > > > 1. You must not charge compression space as swap space to the cgroup.
> > >
> 
> Sorry let me push back on that. That is already existing user-visible
> behavior. Changing that will break our deployment using zswap. I don't
> think we should change that.
> 
> See more in my reply in the other email thread.
> 
> https://lore.kernel.org/linux-mm/CACePvbVaPDnva8X-Xmz84r7j5HjTuih-w5phpw7cerK2uPnK6w@mail.gmail.com/

Thank both for valuable input. I am thinking if we can add an counter
like memory.swap.disk.* or memory.swap.backing.*, then we won't break
the existing behaviour, and also cover the use case Johannes mentioned
where different cgroup have different swapout target setting on xswap.

> 
> > > Hmm, I don't have a stance on this. However, isn't this an issue
> > > zswap/zram have been doing? It feels like an independent issue which
> > > should be done separately?
> >
> > If you have 3 containers using compression space, and two of them have
> > writeback enabled to a shared swapfile, the memory.swap.* controls
> > need to work to manage fair access to that swapfile. They do not work
> > if compression space itself is conflated in.
> >
> > Right now zswap entries actually consume physical swapfile space, even
> > before writeback. Charging the space is correct. But the whole point
> > is to decouple compression space from physical swap space.
> >
> > This is not something that can be done later. It would be a dramatic
> > user-visible change to how the resource is categorized and managed.
> >
> > > > 2. You must make the compression space large enough to be outside the
> > > >    range where users can hit space limits before hitting memory limits.
> > >
> > > We may need a way to define 'large enough' at first.
> >
> > I've tried to lay this out in the other thread, and highlighted the
> > usability issues that result from hitting compression space limits
> > prematurely. It's kind of your call whether you want to seriously
> > engage with this or not.
> >
> > But ultimately it's your claim that a static size can be made to work,
> > so it's on you to make a convincing case.
> >
> > > >    That also means not allowing setups where this is possible.
> > >
> > > And the limit is only an optional knob. If the admin does not set it,
> > > the device grows to the full address space, so there is no space limit
> > > to hit at all. It already behaves the way you want by default. The knob
> > > is only for admins who want a ceiling, they can use it or not. I hope
> > > this would not be a problem for your use case.
> >
> > No, I've laid this out already as well.
> >
> > This isn't about "my" usecase. It's about designing a coherent
> > interface that works well with a large number of usecases, and other
> > pieces of kernel infrastructure commonly used in conjunction.
> >
> > The other proposal in the room needs no such interface. The burden of
> > proof for adding one is on you.
> >
> > > > [1] https://lore.kernel.org/linux-mm/aqLi6cIjD2wJwk0B@cmpxchg.org/

  reply	other threads:[~2026-09-21 10:04 UTC|newest]

Thread overview: 25+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-16 10:19 Baoquan He
2026-09-16 10:19 ` [PATCH v3 01/14] mm: xswap support for zswap Baoquan He
2026-09-16 10:19 ` [PATCH v3 02/14] mm, swap: add CONFIG_XSWAP and xswap fields to swap_info_struct Baoquan He
2026-09-16 10:19 ` [PATCH v3 03/14] mm, swap: refactor free_swap_cluster_info to take swap_info_struct Baoquan He
2026-09-16 10:19 ` [PATCH v3 04/14] mm, swap: add xswap cluster grow via VM_SPARSE vmalloc Baoquan He
2026-09-16 10:19 ` [PATCH v3 05/14] mm, swap: add sysfs create interface for xswap Baoquan He
2026-09-16 10:19 ` [PATCH v3 06/14] mm, swap: add xswap grow trigger on cluster allocation Baoquan He
2026-09-16 10:19 ` [PATCH v3 07/14] mm, swap: add xswap_try_shrink and shrink trigger on cluster free Baoquan He
2026-09-16 10:19 ` [PATCH v3 08/14] mm, swap: free backing pages in xswap_unmap_clusters Baoquan He
2026-09-16 10:19 ` [PATCH v3 09/14] mm, swap: defer xswap shrink to workqueue to avoid lock recursion Baoquan He
2026-09-16 10:19 ` [PATCH v3 10/14] mm, swap: refactor swapoff and add xswap_destroy Baoquan He
2026-09-16 10:19 ` [PATCH v3 11/14] mm, swap: require zswap for xswap devices Baoquan He
2026-09-16 10:19 ` [PATCH v3 12/14] mm, swap: cap xswap growth at nr_clusters Baoquan He
2026-09-16 10:19 ` [PATCH v3 13/14] mm, swap: add sysfs per-device size limit for xswap Baoquan He
2026-09-16 10:19 ` [PATCH v3 14/14] mm, swap: shrink xswap to the ceiling when it drops Baoquan He
2026-09-16 16:45 ` [PATCH v3 00/14] mm, swap: extendable swap devices (xswap) Johannes Weiner
2026-09-17  7:31   ` Baoquan He
2026-09-17 13:17     ` Johannes Weiner
2026-09-21  9:52       ` Chris Li
2026-09-21 10:04         ` Baoquan He [this message]
2026-09-17 10:04 ` Baoquan He
2026-09-18  0:13 ` Nhat Pham
2026-09-18  0:19   ` Nhat Pham
2026-09-18 18:07   ` Nhat Pham
2026-09-21  6:45   ` Baoquan He

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=arEBPsFJvzhQZk9W@fedora \
    --to=baoquan.he@linux.dev \
    --cc=akpm@linux-foundation.org \
    --cc=baohua@kernel.org \
    --cc=chengming.zhou@linux.dev \
    --cc=chrisl@kernel.org \
    --cc=david@kernel.org \
    --cc=hannes@cmpxchg.org \
    --cc=hebaoquan@kylinos.cn \
    --cc=kasong@tencent.com \
    --cc=kunwu.chan@gmail.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=nphamcs@gmail.com \
    --cc=shikemeng@huaweicloud.com \
    --cc=yosry@kernel.org \
    --cc=youngjun.park@lge.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®