mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Baoquan He <baoquan.he@linux.dev>
To: Gregory Price <gourry@gourry.net>
Cc: "Chris Li" <chrisl@kernel.org>, "Kairui Song" <ryncsn@gmail.com>,
	"Johannes Weiner" <hannes@cmpxchg.org>,
	"Nhat Pham" <nphamcs@gmail.com>,
	"Michal Hocko" <mhocko@kernel.org>,
	"Roman Gushchin" <roman.gushchin@linux.dev>,
	"Shakeel Butt" <shakeel.butt@linux.dev>,
	"Yosry Ahmed" <yosry@kernel.org>,
	"David Hildenbrand" <david@kernel.org>,
	"Muchun Song" <muchun.song@linux.dev>,
	"Kemeng Shi" <shikemeng@huaweicloud.com>,
	"Barry Song" <baohua@kernel.org>,
	"YoungJun Park" <youngjun.park@lge.com>,
	"Chengming Zhou" <chengming.zhou@linux.dev>,
	"Lorenzo Stoakes (Oracle)" <ljs@kernel.org>,
	"Liam R. Howlett" <liam@infradead.org>,
	"Vlastimil Babka (SUSE)" <vbabka@kernel.org>,
	"Mike Rapoport" <rppt@kernel.org>,
	"Suren Baghdasaryan" <surenb@google.com>,
	"Qi Zheng" <qi.zheng@linux.dev>,
	"Axel Rasmussen" <axelrasmussen@google.com>,
	"Yuanchu Xie" <yuanchu@google.com>, "Wei Xu" <weixugc@google.com>,
	"Rik van Riel" <riel@surriel.com>,
	"Wenchao Hao" <haowenchao22@gmail.com>,
	"Jonathan Corbet" <corbet@lwn.net>,
	"Hugh Dickins" <hughd@google.com>,
	"Baolin Wang" <baolin.wang@linux.alibaba.com>,
	"Tejun Heo" <tj@kernel.org>, "Michal Koutný" <mkoutny@suse.com>,
	"Shuah Khan" <skhan@linuxfoundation.org>,
	"Kunwu Chan" <kunwu.chan@linux.dev>,
	"Meta kernel team" <kernel-team@meta.com>,
	"Linux Memory Management List" <linux-mm@kvack.org>,
	"Linux Kernel Mailing List" <linux-kernel@vger.kernel.org>,
	linux-doc@vger.kernel.org,
	"open list:CONTROL GROUP - MEMORY RESOURCE CONTROLLER (MEMCG)"
	<cgroups@vger.kernel.org>,
	"Andrew Morton" <akpm@linux-foundation.org>,
	"Joshua Hahn" <joshua.hahnjy@gmail.com>
Subject: Re: Path forward for Virtualized Swap?
Date: Wed, 23 Sep 2026 17:39:44 +0800	[thread overview]
Message-ID: <arOeYHEf9aO7AU5X@MiWiFi-R3L-srv> (raw)
In-Reply-To: <arFe215hUcjLhvmS@gourry-fedora-PF4VCD3F>

On 09/21/26 at 01:02pm, Gregory Price wrote:
> On Mon, Sep 21, 2026 at 06:32:56AM -1000, Chris Li wrote:
> > On Mon, Sep 21, 2026 at 3:27 AM Gregory Price <gourry@gourry.net> wrote:
> > >
> > > This would preserve the existing memory.swap semantics while allowing
> > > both backing resources to be constrained independently.
> > 
> > It sounds like you want memory.tiers have limit enforced.
> > 
> > I suppose it is possible. Again I want to see how people would
> > actually use this feature.
> > 
> 
> Possible, but arguably not needed. the swap and zswap counters already
> work for this existing interaction.
> 
> As I pointed to in my response to Rik, in every reasonable use of
> pswap+zswap the global swap counter is pointless.
> 
> So then pswap=swap and we're left with zswap and swap.
> 
> And I'm not convinced your reading of the swap counter as a limit on the
> *logical* memory allowed to be swapped out is actually accurate.
> 
>   memory.swap.current
>         The total amount of swap currently being used by the cgroup
>         and its descendants.
> 
>   memory.swap.max
>         Swap usage hard limit.  If a cgroup's swap usage reaches this
>         limit, anonymous memory of the cgroup will not be swapped out.
> 
> There is no documentation I can find that has ever documented these
> counters as "the amount of memory requiring a fault".  If you put a
> compression system in front of physical swap - the counters as-described
> would still be accurate, while your reading would be broken.
> 
> "swap" here is highly implied to mean "storage" as opposed to memory,
> which is why "zswap" defines its limits in terms of memory.
> 
>   memory.zswap.current
>         The total amount of memory consumed by the zswap compression
>         backend.
> 
>   memory.zswap.max
>         Zswap usage hard limit. If a cgroup's zswap pool reaches this
>         limit, it will refuse to take any more stores before existing
>         entries fault back in or are written out to disk.
> 
> If you're presently using swap.max to mean the "logical amount of memory
> allowed to be swapped" - then your usage does not meet the definition of
> the knob.  You need to justify that your use case cannot be expressed
> via memory.min/low controls:
> 
>   memory.min
>         Hard memory protection.  If the memory usage of a cgroup
>         is within its effective min boundary, the cgroup's memory
>         won't be reclaimed under any conditions. If there is no
>         unprotected reclaimable memory available, OOM killer
>         is invoked. Above the effective min boundary (or
>         effective low boundary if it is higher), pages are reclaimed
>         proportionally to the overage, reducing reclaim pressure for
>         smaller overages.
> 
> That's an SLO interface.  memory.swap is a provisioning interface.
> 
> As it stands, I'm left viewing zswap's counter inclusion in swap as more
> of a bug than a feature - they account for different things (memory vs
> storage usage).

It's hard to say. When 37e84351198b ("mm: memcontrol: charge swap to
cgroup2") introduced memory.swap.*, it was clearly defined as charging
"the actual number of swap entries used by a cgroup". Please see the
commit log. With that, zswap still reserved a swap slot even when the
data never reached disk. And not to mention zram, it's backend is RAM,
but not physical disk.                                          
                                                                                                      
Now some deployments do use memory.swap.* as an SLO signal, and that is                             
real use cases as Chris and Kairui told. So I don't think this is about
who is right and who is wrong.
                                                                                                      
To keep the existing deployment working and at the same time give the
physical slot its own knob, I think the solution is to add a memory.pswap.*
counter as you suggested. And that is not something we think of from a
brain storm, it comes from real deployments which already depend on the
current memory.swap.* behavior.

Thanks
Baoquan

  reply	other threads:[~2026-09-23  9:39 UTC|newest]

Thread overview: 110+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-04 21:14 Nhat Pham
2026-09-07  5:51 ` Kairui Song
2026-09-08 16:36   ` Nhat Pham
2026-09-11 16:09     ` Kairui Song
2026-09-11 16:57       ` Nhat Pham
2026-09-11 18:14         ` Kairui Song
2026-09-11 19:03           ` Nhat Pham
2026-09-12  8:47             ` Kairui Song
2026-09-14 16:49               ` Nhat Pham
2026-09-08 18:30   ` Johannes Weiner
2026-09-09 16:41     ` Nhat Pham
2026-09-09 17:47       ` Nhat Pham
2026-09-12  9:00       ` Kairui Song
2026-09-12 11:51         ` Johannes Weiner
2026-09-19  7:23           ` Chris Li
2026-09-20  1:06             ` Rik van Riel
2026-09-21  0:18               ` Chris Li
2026-09-21  0:35                 ` Rik van Riel
2026-09-21  0:53                   ` Chris Li
2026-09-14 16:10         ` Nhat Pham
2026-09-10 23:27   ` Nhat Pham
2026-09-07 11:30 ` David Hildenbrand (Arm)
2026-09-08 16:45   ` Nhat Pham
2026-09-10 10:56     ` David Hildenbrand (Arm)
2026-09-10 16:22       ` Nhat Pham
2026-09-10 17:57         ` David Hildenbrand (Arm)
2026-09-11 16:20         ` Kairui Song
2026-09-11 16:56           ` David Hildenbrand (Arm)
2026-09-14 15:00           ` Baoquan He
2026-09-19  6:50             ` Chris Li
2026-09-21 18:55               ` Nhat Pham
2026-09-22 13:38                 ` Chris Li
2026-09-22 14:43                   ` Shakeel Butt
2026-09-22 14:56                     ` Chris Li
2026-09-22 15:07                       ` Johannes Weiner
2026-09-22 15:30                         ` Chris Li
2026-09-22 15:45                           ` Shakeel Butt
2026-09-23  6:44                             ` Chris Li
2026-09-22 15:46                           ` Johannes Weiner
2026-09-22 15:28                       ` Baoquan He
2026-09-10  7:09 ` Baoquan He
2026-09-10 16:39   ` Shakeel Butt
2026-09-11 13:06     ` Baoquan He
2026-09-11 16:45       ` Shakeel Butt
2026-09-15  5:48         ` Baoquan He
2026-09-15  6:44           ` Baoquan He
2026-09-19  6:55             ` Chris Li
2026-09-19 16:15               ` Rik van Riel
2026-09-19 21:21                 ` Chris Li
2026-09-19 22:50                   ` Rik van Riel
2026-09-21  0:13                     ` Chris Li
2026-09-22 17:54                       ` Nhat Pham
2026-09-23  7:44                         ` Chris Li
2026-09-23 16:30                           ` Shakeel Butt
2026-09-19  7:07           ` Chris Li
2026-09-10 17:03   ` Johannes Weiner
2026-09-11 12:27     ` Baoquan He
2026-09-11 16:21       ` Johannes Weiner
2026-09-19  8:45     ` Chris Li
2026-09-19 16:08       ` Gregory Price
2026-09-19 19:02         ` Chris Li
2026-09-20  1:14           ` Rik van Riel
2026-09-20  1:53           ` Gregory Price
2026-09-21  9:43             ` Chris Li
2026-09-21 12:53               ` Gregory Price
2026-09-21 16:11                 ` Chris Li
2026-09-22 17:14                   ` Rik van Riel
2026-09-22 17:21                     ` Nhat Pham
2026-09-23  7:20                       ` Chris Li
2026-09-22 17:31                     ` Johannes Weiner
2026-09-23  7:32                       ` Chris Li
2026-09-23 13:55                         ` Johannes Weiner
2026-09-21 14:06               ` Rik van Riel
2026-09-21 16:15                 ` Chris Li
2026-09-21 15:16               ` Rik van Riel
2026-09-21 16:31                 ` Chris Li
2026-09-21 16:36                   ` Rik van Riel
2026-09-21 18:10                     ` Chris Li
2026-09-21 18:52                       ` Rik van Riel
2026-09-22 13:32                         ` Chris Li
2026-09-22 14:59                           ` Rik van Riel
2026-09-22 15:23                             ` Chris Li
2026-09-22 15:43                               ` Johannes Weiner
2026-09-23  6:10                                 ` Chris Li
2026-09-22 15:48                               ` Rik van Riel
2026-09-23  7:00                                 ` Chris Li
2026-09-23 14:21                                   ` Rik van Riel
2026-09-22 17:32                               ` Nhat Pham
2026-09-23  7:37                                 ` Chris Li
2026-09-23 14:11                                   ` Johannes Weiner
2026-09-23 15:39                                   ` Nhat Pham
2026-09-22 15:20                           ` Gregory Price
2026-09-22 15:43                             ` Chris Li
2026-09-22 15:52                               ` Rik van Riel
2026-09-23  7:10                                 ` Chris Li
2026-09-23 12:52                                   ` Klara Modin
2026-09-23 15:43                                     ` Nhat Pham
2026-09-23 17:26                                       ` Klara Modin
2026-09-21 10:01           ` Kairui Song
2026-09-21 13:27             ` Gregory Price
2026-09-21 15:27               ` Rik van Riel
2026-09-21 15:48                 ` Gregory Price
2026-09-21 16:32               ` Chris Li
2026-09-21 17:02                 ` Gregory Price
2026-09-23  9:39                   ` Baoquan He [this message]
2026-09-23 12:11                     ` Gregory Price
2026-09-23 14:34               ` Kairui Song
2026-09-21 18:11             ` Nhat Pham
2026-09-22 17:16           ` Johannes Weiner
2026-09-10 17:16   ` Nhat Pham

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=arOeYHEf9aO7AU5X@MiWiFi-R3L-srv \
    --to=baoquan.he@linux.dev \
    --cc=akpm@linux-foundation.org \
    --cc=axelrasmussen@google.com \
    --cc=baohua@kernel.org \
    --cc=baolin.wang@linux.alibaba.com \
    --cc=cgroups@vger.kernel.org \
    --cc=chengming.zhou@linux.dev \
    --cc=chrisl@kernel.org \
    --cc=corbet@lwn.net \
    --cc=david@kernel.org \
    --cc=gourry@gourry.net \
    --cc=hannes@cmpxchg.org \
    --cc=haowenchao22@gmail.com \
    --cc=hughd@google.com \
    --cc=joshua.hahnjy@gmail.com \
    --cc=kernel-team@meta.com \
    --cc=kunwu.chan@linux.dev \
    --cc=liam@infradead.org \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=mhocko@kernel.org \
    --cc=mkoutny@suse.com \
    --cc=muchun.song@linux.dev \
    --cc=nphamcs@gmail.com \
    --cc=qi.zheng@linux.dev \
    --cc=riel@surriel.com \
    --cc=roman.gushchin@linux.dev \
    --cc=rppt@kernel.org \
    --cc=ryncsn@gmail.com \
    --cc=shakeel.butt@linux.dev \
    --cc=shikemeng@huaweicloud.com \
    --cc=skhan@linuxfoundation.org \
    --cc=surenb@google.com \
    --cc=tj@kernel.org \
    --cc=vbabka@kernel.org \
    --cc=weixugc@google.com \
    --cc=yosry@kernel.org \
    --cc=youngjun.park@lge.com \
    --cc=yuanchu@google.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®