From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-qt1-f176.google.com (mail-qt1-f176.google.com [209.85.160.176]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 70978324B31 for ; Thu, 25 Jun 2026 13:36:29 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.160.176 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1782394591; cv=none; b=nk/1FmNG15m01XvViXZvql7T8HMc2QlCcH7m3FdKIqRdM4E1Fos4pHJJGC4O5ZXIXu0VXh4Ps6qFccxEASjIEFBVdIZp5IpvtklALUCU0moUPAwa+cjr8PsrR95NuZgqb1IHl/tj1yw4GkoSfJkeQaZjc+x+yBiMX4Q5LzDjQpM= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1782394591; c=relaxed/simple; bh=6LuWcAsRkFsr+tincYKQ9pOPkvbasAdSEjAa/E+9WEo=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=Tb7U0oOjDa6ekiVnogCi/hCyyDr2fPLAMaF8Xj53X0ur1+LYVadRpJyZHW7ttk941O3dF4pK+WMOOgEzPK1XL3FhMQ12Lhq+gk4vrErwIEUPyW5WK1MBMDse0vIymPhhyAGBx4tl8h6D+Ynl3EzFRAwUtyYAbXzkRiFeAyFxAoc= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=cmpxchg.org; spf=pass smtp.mailfrom=cmpxchg.org; dkim=pass (2048-bit key) header.d=cmpxchg.org header.i=@cmpxchg.org header.b=BOkIymgy; arc=none smtp.client-ip=209.85.160.176 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=cmpxchg.org Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=cmpxchg.org Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=cmpxchg.org header.i=@cmpxchg.org header.b="BOkIymgy" Received: by mail-qt1-f176.google.com with SMTP id d75a77b69052e-51a1fe8f578so22938501cf.2 for ; Thu, 25 Jun 2026 06:36:29 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=cmpxchg.org; s=google; t=1782394588; x=1782999388; darn=vger.kernel.org; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:from:to:cc:subject:date:message-id:reply-to; bh=ALlzMSFTkD2qwY6cH7Q83YaJoqNCS+acziBP11A7pe0=; b=BOkIymgyGLuqpNPlkE0es657hY3SpLrWbGXWFAA3azYZoF/VgFsM4jnJks8yHSJ9Ge 2106Kv20xXbOkYeatg9QUuUa6S1/Ct4feAf8ZY1wIWYVp8th43YX185tTVgxIe1SDZmr J9SO4rkwjJmkhrnLlgVT9vEsUrXifwiHKOz86Ey+eRI/sYY2IRY/YwHLRGLK0Mdvk0er 7PWaoYS5NkkZrh+VWyxv6tJGrALrulKEgiql17adc3NoqVywSnvAaQGuvqsvFJzcsaMd oKw/JssUSSRXu0quPV3lU16rNn9pNYGHrATg/8xnJkFnbsfyqnaqCzmKvUotVrcZfQZi Ho5w== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1782394588; x=1782999388; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-gg:x-gm-message-state:from:to:cc :subject:date:message-id:reply-to; bh=ALlzMSFTkD2qwY6cH7Q83YaJoqNCS+acziBP11A7pe0=; b=iOrPaFJSauFyNRDVU9QuozR++Y57tGIaG1YkjyB29+DIwACVBXM3eA+ikhahhuids4 FTzlyRklszNcwzOiLK2zG6gHkAzjKzKyblZ5W5YCVccB1WmGmOr6AeeOL0KNeX3kVb0J RSHrPGLBURWOz3hxwBLiA/kmXIrrsgFu4FOYTwcTePjkLxSLiRzTRpQiwYrdsREGCT56 5Yxif8klv6p/+76VcgYjM4eDiGqUxkIwVfeGDRABXqg6WY0UCPvR6si/ArPMcC/Cr/3k Gd05m3RUOAtzdz61IILlzqo1R73po1Q3mjWvjPjRX+QrPK/U7wa2Vz6UCxqXnoRnpQj8 CESw== X-Forwarded-Encrypted: i=1; AFNElJ9FGlwh3ADvwsKVZrwRoFBxdU1gMjAuoZZ+icV49c6+YMcmjul8h5UxUC8dyWjZmXnAnXea4ZTstd+WX88=@vger.kernel.org X-Gm-Message-State: AOJu0Yy4rcmIuMPx6M/bL7GZhBA13Sqdn8N6F1Ad0EIclrqif7HrUXmM WZpW3SdSFEKtjB+oFpAzs9pNJQe1fJ0c4H4gwwAnLep56ZpsFW8sWDnRxajCg4+DbcI= X-Gm-Gg: AfdE7ck836ilao3Uj1nBLZcRdsJuC435fGTF7Q0bW4jv59qODSoBpOwyst6VXZAbdat 8f+v/PDQ3xUWgNDHLUDOewwUrX9eB8KmnYAgn+btRPr3vIY+yL0Hg9ctMOBohYJnSw8DDjUJFvJ VdedPsRjPQh+mS8TE6p8HeFb9u99IqrITa6zr1Wuf2BeQN/R072ch//oXttTFi1OK5aB4kBM+rb fLkC/KleKrFgTHbMsZ0WikONmJUufZWQhmJaUF3Fjnb4ismC5ijKnb8mJXynZI8KZwyA++/zZm6 bdybkVTjUE8fNNaXySDpByIBSi2LoRIxJK3oxDUyKX8+a6XNnqsrXTsJT+gwLjwxaHimkZKTCVR XU4Ap5tUwYw5knKAIZ7VY+I/05J9IRfrGVYqYIXHL0Sf/N+QIWSh95sR125LtJU6cGP2cHOnRDO HJ+qjsgvUTqHo= X-Received: by 2002:a05:622a:180f:b0:519:51a9:fc84 with SMTP id d75a77b69052e-51a727cb368mr33707201cf.41.1782394588132; Thu, 25 Jun 2026 06:36:28 -0700 (PDT) Received: from localhost ([2603:7001:f100:500:365a:60ff:fe62:ff29]) by smtp.gmail.com with ESMTPSA id d75a77b69052e-51a51b12242sm69457131cf.29.2026.06.25.06.36.27 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 25 Jun 2026 06:36:27 -0700 (PDT) Date: Thu, 25 Jun 2026 09:36:23 -0400 From: Johannes Weiner To: "David Hildenbrand (Arm)" Cc: Barry Song , akpm@linux-foundation.org, axelrasmussen@google.com, baolin.wang@linux.alibaba.com, dev.jain@arm.com, kasong@tencent.com, lance.yang@linux.dev, liam@infradead.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org, ljs@kernel.org, npache@redhat.com, qi.zheng@linux.dev, ryan.roberts@arm.com, shakeel.butt@linux.dev, weixugc@google.com, yuanchu@google.com, zhaonanzhe@xiaomi.com, ziy@nvidia.com, Michal Hocko , Roman Gushchin Subject: Re: [RFC PATCH] mm: Avoiding split large folios if swap has no space Message-ID: References: <5790c4a4-d502-4180-82f5-47de5809a4fe@kernel.org> <20260620081017.89085-1-baohua@kernel.org> <4aa8350e-712f-4380-b3bf-2ff06cf2a35d@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: On Thu, Jun 25, 2026 at 09:49:56AM +0200, David Hildenbrand (Arm) wrote: > >> > >> But now I wonder whether we would also want to check "is there any free swap > >> space", not just "is there any swap". > > > > I don't quite understand you. get_nr_swap_pages() returns > > nr_swap_pages, which increases or decreases as swap is allocated or > > freed. I guess it just reflects how many swaps we currently have > > available? > > Indeed, I was confused by the function name it's "free swap pages". So all goof :) > > > > >> > >> > >> Essentially, try returning -E2BIG if there is the chance to swap out after > >> split, and -ENOSPC / -ENOMEM if a split wouldn't help. > >> > >>> } > >>> > >>> again: > >>> @@ -1769,11 +1772,13 @@ int folio_alloc_swap(struct folio *folio) > >>> } > >>> > >>> /* Need to call this even if allocation failed, for MEMCG_SWAP_FAIL. */ > >>> - if (unlikely(mem_cgroup_try_charge_swap(folio))) > >>> + if (unlikely(mem_cgroup_try_charge_swap(folio))) { > >>> swap_cache_del_folio(folio); > >>> + return -ENOMEM; > >> > >> Here we wouldn't have the information whether we could charge after a split. > >> > >> So that would require a rework to signal this more cleanly to the caller. > > > > Yep. The tricky part is that mem_cgroup_try_charge_swap() cannot > > return how much swap quota is available in the memcg. Do you prefer to > > add an output argument to mem_cgroup_try_charge_swap() to expose > > that > That would probably be cleanest, if that is easily possible. We would want to > get memcg maintainer feedback on that. > > @memcg folks: we'd like to know whether splitting a large folio would make > mem_cgroup_try_charge_swap() succeed on a split (smaller) part, to distinguish > "there is no way we can swap out anything, don't split" vs. "we could swap out, > split". It's technically doable, but is this worth the bother? The remaining headroom is less than a large folio. You can split this one, but you cannot even swap out all of its subpages anymore? From the cgroup side, we don't need the limit to be obeyed this rigidly. We overcharge temporarily in other places if it's convenient to do so. A fuzz factor around the limit is acceptable. But if you still want to do it, here is how: The page_counter_try_charge() in __mem_cgroup_try_charge_swap() walks the hierarchy upwards. If it fails, it will store the first level that failed against its limit. You can do the mem_cgroup_margin() math against this counter to determine headroom. An ancestor *could* be more restrictive, so you need to finish the hierarchy walk to the root and use the min() of all the swap.max - page_counter_read(swap). Then return that in a return argument from __mem_cgroup_try_charge_swap().