From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-oi1-f179.google.com (mail-oi1-f179.google.com [209.85.167.179]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E44B5330B29 for ; Mon, 22 Jun 2026 23:19:51 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.167.179 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1782170394; cv=none; b=BlwLvj1+iIviXhP9nWxhGy5MBnhqD2UFpFxa9Qmss0iBfFL3MM17R8ht5qvqFbHV829OsrLHZTpS28naJ+v52HsQgTbT7BPzCMC6Z/f9iDCGMmzc7g/SEXlkPymfddsbwq+sBn8a5JT8+eZtcSIuqFhiy4pSMTA+7Sacgp1s60c= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1782170394; c=relaxed/simple; bh=Xlw43IzFtb/tFwtGy/6ziKdecpraNqKX+Avh7kTsw8U=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=RleGj+eRhCnIybKI9hmo7KR0ZM0brxcgG9HhM5WVDPl2M8jZVQwuOIm8dONunBHlYsgierhSp6+4SiZswQ+jYY6K+L8CLbXX1VJ1KopwSq32lmp3cdxz+G4P/4PXKxpNk2xD9oz1ZccoYm0Gkv95tOKx6xtRf0wLxNBuAhKOCDE= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=n0hNcIjU; arc=none smtp.client-ip=209.85.167.179 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="n0hNcIjU" Received: by mail-oi1-f179.google.com with SMTP id 5614622812f47-4877a7b451dso3557121b6e.0 for ; Mon, 22 Jun 2026 16:19:51 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1782170391; x=1782775191; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to; bh=UoK1Vry3ZA37o4N+PhVs+8t5kaUWOE3Fz+zyZS4k++4=; b=n0hNcIjU1B55QsN2INYUU7h1Y+BbVdsDyRdvYFudarhNRBsY2Gno2bUBL5DCg2pvK8 MaBkfb1iz7C0SrJzb6wzTW/olUT6ME972o7uRzf9jILlPdz41Op327vrpYzRdCTg5VkE 4mUk9v8XzfJwCW6zqo/PNuDIiLChBGQeuxQlBVUURVd1EtzoeWOmRNDVAfSJMW6ufrIx QS2PXQCUod7gjbJd2A2epe+s5fzOk1g9N5fRBL8nvCPNKGe1hrRG986VEg2H8L9mBRqT 6KXpzRVyJfu1rKvODrRZbrrZVjZxreAGXeZI5LILKsEQOAEW6CN/01Qqtbly9rZ4Ls0U qAWg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1782170391; x=1782775191; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to; bh=UoK1Vry3ZA37o4N+PhVs+8t5kaUWOE3Fz+zyZS4k++4=; b=n9tPl7d9pXmVeZZxHLw6cMOkyY29haCvl/NcBmUhUgUkLKY+saYt9UCr+9NYliBWVk i/MwUTViLEkOYEkAgFRp1UAPY+Z9AC48IFjXNqlixRyfB42+wAvxjcgrPbnCynNRsybB Nz/rbcSiNio39Y3/pFTCeME2Vrq3y8QSg7oqZv+BQg4lXTmiyg6/yVwCc4p0Dlo5mGfN BBks6hlP6FLgcv5PksM7X6kOTwoP1tVfp15Hlyd1HZ1RNPTRMTg6PWVkV/8F7T6wCpPV PjovdliB0DV+/GXx778G632KViZb+xUc5G/qJ76lsZrjIeYDmtSoGVyiSBru7cEWL9sW inhw== X-Forwarded-Encrypted: i=1; AFNElJ/uEJyfm8U76nJuqEkry1hg+Zf2KycmSU4jezjh9cq/pNt4BnJ2UUFntXam0rbQ0vZss+APH0vA6FfIPFg=@vger.kernel.org X-Gm-Message-State: AOJu0Yz7lKhjsZ7tHDvwkh9eYqI95pM3MIe1q4M0xuuy0hNvBYance/6 Lp8nHPf/WT3TPBpTRAC+2hsS5m51fhbwXXLAmFCjtnl3V60PYTuiq43H X-Gm-Gg: AfdE7clZueljkTZJYz+1/LV2X3mRdXYBbJDDduX1M4hMo5SE0jpn1AJ4l4M2eEJ4vtS Q+reEcco5oVIjUbIunVuSR4RoI5dlCpmPoRXxBHVmGuhjq2HrSg0TusnodDGASrupDyLYxAjWhv whw5ZlFjH9orRCxdSFChuDYxanyqc3Q/lR/OJdtqfQ99KuQ6oG6nKTJ8hkm/DC9dNY925cqPqxk vC7LcaxdTwRzGug3qSUEQkW+pGwcFBZGW1Vzw4wayhDnKoSwQg5qxrgmDMbRj+zZyk3uQ3fDzO9 FFl6f+4HMylsqQgPR9CAg7rl9VSvSHVRfAAg0JRjws4pW1VS1rm+6JMo2+ZOqB4Udo7no5zeduc slsynpsldVB7o+BbxL6lYpXk7c/PN0gxTjQTJ72k+ZJI1cfeFlXR2THHV4KsJR8aQAzvF4iQtOM Z9FkYxpJWxmSgheeHkclOtn/B+xdp3Mi41MKK+GOfU5w== X-Received: by 2002:a05:6808:191c:b0:487:4c11:7a17 with SMTP id 5614622812f47-4896acb178emr14254437b6e.36.1782170390846; Mon, 22 Jun 2026 16:19:50 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:7::]) by smtp.gmail.com with ESMTPSA id 5614622812f47-48aec0e5e53sm5453907b6e.8.2026.06.22.16.19.50 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 22 Jun 2026 16:19:50 -0700 (PDT) From: Joshua Hahn To: Yosry Ahmed Cc: Youngjun Park , Shakeel Butt , akpm@linux-foundation.org, chrisl@kernel.org, youngjun.park@lge.com, linux-mm@kvack.org, cgroups@vger.kernel.org, linux-kernel@vger.kernel.org, kasong@tencent.com, hannes@cmpxchg.org, mhocko@kernel.org, roman.gushchin@linux.dev, muchun.song@linux.dev, shikemeng@huaweicloud.com, nphamcs@gmail.com, baoquan.he@linux.dev, baohua@kernel.org, gunho.lee@lge.com, taejoon.song@lge.com, hyungjun.cho@lge.com, mkoutny@suse.com, baver.bae@lge.com, matia.kim@lge.com Subject: Re: [PATCH v9 3/6] mm: memcontrol: add interface for swap tier selection Date: Mon, 22 Jun 2026 16:19:47 -0700 Message-ID: <20260622231948.1002174-1-joshua.hahnjy@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit On Mon, 22 Jun 2026 15:26:17 -0700 Yosry Ahmed wrote: > On Mon, Jun 22, 2026 at 3:10 PM Joshua Hahn wrote: > > > > On Mon, 22 Jun 2026 14:21:30 -0700 Yosry Ahmed wrote: > > > > > On Sat, Jun 20, 2026 at 11:17 AM Youngjun Park wrote: > > > > > > > > Introduce memory.swap.tiers.max, a flat-keyed file listing each > > > > tier defined in /sys/kernel/mm/swap/tiers with its state, "max" > > > > (allowed, the default) or "0" (disabled). A tier is one bit in the > > > > cgroup's tier mask, so writing " max" or " 0" sets or > > > > clears that bit. > > > > > > > > Since the current use case lacks amount control, it only supports > > > > "max" (on) and "0" (off). Therefore, it does not track per-tier swap > > > > usage, relying instead on a fast runtime bitmask check. > > > > > > > > We maintain both `mask` and `effective_mask`. The `effective_mask` is > > > > strictly bounded by the parent (e.g., if a parent is "0", the child's > > > > effective state is "0" even if its `mask` is "max"). Maintaining this > > > > separately avoids costly cgroup tree traversals to check ancestors at > > > > runtime. > > > > > > > > Suggested-by: Shakeel Butt > > > > Suggested-by: Yosry Ahmed > > > > Signed-off-by: Youngjun Park > > > > --- > > > > Documentation/admin-guide/cgroup-v2.rst | 20 +++++ > > > > Documentation/mm/swap-tier.rst | 9 +++ > > > > include/linux/memcontrol.h | 5 ++ > > > > mm/memcontrol.c | 67 ++++++++++++++++ > > > > mm/swap_state.c | 5 +- > > > > mm/swap_tier.c | 102 +++++++++++++++++++++++- > > > > mm/swap_tier.h | 57 +++++++++++-- > > > > 7 files changed, 255 insertions(+), 10 deletions(-) > > > > > > > > diff --git a/Documentation/admin-guide/cgroup-v2.rst b/Documentation/admin-guide/cgroup-v2.rst > > > > index 6efd0095ed99..4843ffcfd110 100644 > > > > --- a/Documentation/admin-guide/cgroup-v2.rst > > > > +++ b/Documentation/admin-guide/cgroup-v2.rst > > > > @@ -1850,6 +1850,26 @@ The following nested keys are defined. > > > > Swap usage hard limit. If a cgroup's swap usage reaches this > > > > limit, anonymous memory of the cgroup will not be swapped out. > > > > > > > > + memory.swap.tiers.max > > > > + A read-write flat-keyed file which exists on non-root > > > > + cgroups. The default is "max" for every tier. > > > > Hi Yosry, > > > > Sorry, I feel like I'm joining the party late. Apologies if I'm missing > > some context or repeating a discussion that's already been had. > > Please let me know if that is the case. > > > > One quick tangent: > > I was chatting with Nhat last week about swap tiers and its relation to > > memory tiering. Nhat brought up a good point, which is that while both > > swap tiers and memory tiers provide a clear hierarchy of performance, > > only memory tiering allows for movement between the tiers. > > AFAICT, swap tiering does not allow for direct migration from a higher > > tier swap backend to a lower tier swap backend if the higher tier > > backend runs out of memory. > > > > In that sense, I'm not entirely sure if we need to enforce similar > > semantics across swap tiering and memory tiering; it seems like there > > are some fundamental differences anyways to how we treat these tiers. > > > > > I wonder what should the default behavior be if memory.swap.max is set > > > to a value other than "max". Should the limits in > > > memory.swap.tiers.max auto-scale or remain as "max"? We probably want > > > to keep the behavior consistent with memory tiering. > > > > > > Shakeel/Joshua, WDYT? > > > > I think that the motivation behind these tiers is different for swap > > and memory. Tiered memory limits is motivated by preventing one > > workload from conusming all of a valuable resource, while swap tiers > > seems more to do with excluding certain workloads from using performant > > tiers and ensuring other workloads stay on those performant tiers. > > > > IOW memory tiers exist for fairness, but it seems like swap tiers exist > > for workload performance tiering. But maybe there's a usecase out there > > that would want fairness to apply in the swap tiers as well that I am > > not seeing. > > I am not sure what use cases exist, but I think it's possible we end > up wanting to enforce fairness for swap tiers as well. Maybe not as > aggressively as memory (e.g. to avoid wearing out SSDs), but maybe at > least proactively through userspace? > > At the end of the day, faster swap tiers are also valuable resources > that we probably don't want a few workloads to hog. I also think the > interfaces being consistent makes everyone's lives easier, even if > it's a bit of an overkill for swap tiers. I see, thank you for the explanation. That makes sense to me. > > If that is the case, I think auto-scaling makes sense but can be a bit > > tricky, since there is no universal tiered ratio; each workload will > > have different tiers it can swap to, so they will all have to calculate > > their own ratios. Tiered memory limits escapes this difficulty since we > > assume all memory can be placed on all tiers, so we have a system-wide > > ratio : -) > > Hmm I don't follow. It's also possible (maybe not initially) that a > memcg cannot use specific memory tiers, right? I am not sure what the > difference is. You're right, I was speaking more to the current state of memory tiers. The majority of the feedack I received was that we already have too many memcg knobs, so I just opted to make tiered memcg limits a cgroup mount, with no ability for individual memcgs to tune their limits or opt-in/out. What do you think Yosry? Would it make sense for us to be able to tune these values? Personally I think it makes sense but just wanted to make the basic features merged before I went to push for making those knobs tunable. If we want to make the tuning the same across swap & memory we should probably align on the file names and how we interact with them. Thanks, Joshua