From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-oi2-f43.google.com (mail-oi2-f43.google.com [74.125.231.235]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CE7C951D536 for ; Fri, 18 Sep 2026 18:03:01 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.231.235 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789754585; cv=none; b=FzYgvngeCN4I43MqJGBrSpQXWiwWia3msqbwifiOhSPechu4Alxo9Skgb5Z1Vd2kfL/8Yra5fUW7T0UNuZJwQbEWzfW06Jm6xNBUvsmh46HGpGd42qihSadvgnzoxgOm0ew9biB6cFm5BrUpQAc1zWwxNVyuXR9SfbpofU5iIGk= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789754585; c=relaxed/simple; bh=248meyirULeKvPgqKHcpHfv3DScKvY8uep/W5NAxx1s=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=aFnsxHELmhiWF+CDTs3ZQSBLqWgvw35KJ+0T+D1jrjDicewJTK/5b6g2xNUSoCvY654YpFXGa0e6by+V22089LAojsGu98CQIRo64f9MWUCLORQTr+bWyIyKGQR/5xawueZGM1osdsF347dwDz7mIvY1Rr/XvdTosMJSFvLIVxs= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=K3zgXMub; arc=none smtp.client-ip=74.125.231.235 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="K3zgXMub" Received: by mail-oi2-f43.google.com with SMTP id 46e09a7af769-7f4f0d1779dso938539a34.2 for ; Fri, 18 Sep 2026 11:03:01 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1789754580; x=1790359380; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=leUxoZNbrIHhJwMkVXew1sBjFqiId1Gr67nxTcdiA8U=; b=K3zgXMubFzyRxcXj6UaURZWLgP8JGNZ9k+S8ptsKVpiH0+quEGRhzVcew29tLaHNGC J961N1sOL1y2qMWgshUIedJzrcgEybXUiUOL/yeQGPkgU6J+RfQ/r2XVFlRwD/I5dYzy Brq7E+TSuirObrPLZ1MVU47LZi9HzeEWFZghRprpW0Y9pbt/6j0cl0mkb3qjjh/OTupb IbMHRsn4QqvFT7jdIc4gSum4hV5EQCMYMjGQE71fPNDgj2m8gDlv0ZXe3ZsFk1Dj//O0 PzhAtsWwZgj97lsHWinTMGD/HwgviefqzlFKnlQn5QHGRDRau69FvItijwTr7yd4MtAg W9qA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789754580; x=1790359380; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=leUxoZNbrIHhJwMkVXew1sBjFqiId1Gr67nxTcdiA8U=; b=FyAds++XJLPazFyn+iOoUVq0gjNLgcntto1iPW86oWgkEcOkGtOaNrELrXJeYAP+ZY hR/Lnj0JZv+jWxn9Rs+Yq5k21pW9BceuoGZH7R3AckM/Abv/QQ3ZhVv5JXucvpgCgXkO ++QBzO+ZZkMr46hITMbYVcdj0DpJ5Q4VWu+yZNOBlfEnO9diizVv9EvJLiBiYfNJ8Hvq P5SoATwluLJuk5GMJybLoxfgMviMmradeIUROYchmu65597jS72M51AnbIHgpWqxdCE+ JZxQPS4r5y9mczi8o428O4RbdgVVGuhoiKN2VVjGXPUMyi5tDpsgpj4kh94AOO3vRiPt /GTg== X-Forwarded-Encrypted: i=1; AKwUvBzr+GJ/YNUqu3PLkBYMFMwuQZqABGYmAABkbaoqVLo2U7q4/nFvRjNhvtM3DbZhSHgbLq5XcymqkyEYks0=@vger.kernel.org X-Gm-Message-State: AFuF++nHDbRE58xGv2Y3oj0fLvM6ElokOsiUG5egcfLryIdnyZ5rgBRn /0GWRNiccVtYAGUR4WlJ26wIRSJgLrVB42MngXzEM1BmdUDdM/GaQdeD X-Gm-Gg: AYBFou3LqxcHdms7p+LNkYv0C8CydNtdsUbnm8K3bZ6X6DvCcOE6eSCw6ARNcEXgEI4 z4HyaZ+92rWfuxDEazVXSbTNKvNlaQT1VA5ebqc1b80fGXNku55vT0EoR/JULfJ8IRbf5NuVned DzaMDzxVgE3AO/XPTQWo5CDkizOBnQnrdmWKrE7x73ITWpWUVGYUiDoEzSNTc88OmfZ2FJjTCEV 0h6zLhbIiphbJ5Y+5sQ6UZPuD75phMViqZK/EcKwec2VhJq3ttwBggZWuK0N1myPRWAPwqEaOuk CqYWw/iPIR11LCREaNow205pKCmxAajHLhqOfv7MgeXV557MRDLqVioWHNUenkv+Nup66UXF9IF YzX6YSc/NPffpVZ5N6WYviiXKoJ0tdl8b8Wo0xzuTGcLU3GDkkv+4Q7rlIT/xlC+qO6sithrXMq X6qOfLBz3cuo41lzb0BGm9a7KfmfjJvHkCNmBPZWCmb035jbthzuzcuBuKB6dGrKXJF7y1MPN9R ethYkMTM9hdRJAY8YaNcfDq2cFXx/aY X-Received: by 2002:a05:6830:650a:b0:804:ca33:4aff with SMTP id 46e09a7af769-80de0e50a9cmr3677995a34.9.1789754580097; Fri, 18 Sep 2026 11:03:00 -0700 (PDT) Received: from localhost ([2a03:2880:10ff:43::]) by smtp.gmail.com with ESMTPSA id 46e09a7af769-8107e894949sm141880a34.17.2026.09.18.11.02.59 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 18 Sep 2026 11:02:59 -0700 (PDT) From: Nhat Pham To: akpm@linux-foundation.org Cc: chrisl@kernel.org, kasong@tencent.com, hannes@cmpxchg.org, mhocko@kernel.org, roman.gushchin@linux.dev, shakeel.butt@linux.dev, yosry@kernel.org, david@kernel.org, muchun.song@linux.dev, shikemeng@huaweicloud.com, baoquan.he@linux.dev, baohua@kernel.org, youngjun.park@lge.com, chengming.zhou@linux.dev, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, qi.zheng@linux.dev, axelrasmussen@google.com, yuanchu@google.com, weixugc@google.com, riel@surriel.com, gourry@gourry.net, haowenchao22@gmail.com, corbet@lwn.net, hughd@google.com, baolin.wang@linux.alibaba.com, tj@kernel.org, mkoutny@suse.com, skhan@linuxfoundation.org, kunwu.chan@linux.dev, kernel-team@meta.com, nphamcs@gmail.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, cgroups@vger.kernel.org Subject: [PATCH v5 10/11] mm, swap: defer memcg_table allocation for physical swap clusters Date: Fri, 18 Sep 2026 11:02:40 -0700 Message-ID: <20260918180241.3424851-11-nphamcs@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260918180241.3424851-1-nphamcs@gmail.com> References: <20260918180241.3424851-1-nphamcs@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit Stop allocating a memcg table for every physical swap cluster that only ever holds vswap backings. The table costs SWAPFILE_CLUSTER * sizeof(unsigned short) per cluster, 1 KB per 2 MB of swap on a 64-bit kernel with 4 KB pages. On a vswap-heavy workload, where zswap writeback is the only consumer of physical swap, that is the common case. Such clusters never have their memcg_table read or written: vswap-layer charging records on the vswap cluster's table, not the physical one. Allocate eagerly only where the table is known to be needed: every vswap cluster, and, when vswap is off, every physical cluster, since none of its slots is then a vswap backing. A physical cluster otherwise allocates on its first direct-use slot, and skips entirely if it only holds vswap backings. Hibernation slots have no folio and record no cgroup, so they do not trigger it. That deferred allocation is on the allocator's fast path and can fail; the allocation it serves then fails too, and the caller falls back to another cluster. Signed-off-by: Nhat Pham --- mm/swapfile.c | 83 +++++++++++++++++++++++++++++++++++++++------------ 1 file changed, 64 insertions(+), 19 deletions(-) diff --git a/mm/swapfile.c b/mm/swapfile.c index cf07d87d2301..39d1840b0d36 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -472,7 +472,8 @@ static void swap_cluster_free_table(struct swap_cluster_info *ci) swap_cluster_free_count_table(table); } -static int swap_cluster_alloc_table(struct swap_cluster_info *ci, gfp_t gfp) +static int swap_cluster_alloc_table(struct swap_info_struct *si, + struct swap_cluster_info *ci, gfp_t gfp) { struct swap_table *table = NULL; struct folio *folio; @@ -493,7 +494,14 @@ static int swap_cluster_alloc_table(struct swap_cluster_info *ci, gfp_t gfp) return -ENOMEM; #ifdef CONFIG_MEMCG - if (!mem_cgroup_disabled()) { + /* + * A physical cluster under vswap may hold only vswap backings, which + * record their memcg on the vswap cluster's table, not this one. Such + * clusters defer memcg_table allocation until they hand out a slot + * that maps directly into the PTEs. + */ + if ((!vswap_is_enabled() || swap_is_vswap(si)) && + !mem_cgroup_disabled()) { VM_WARN_ON_ONCE(ci->memcg_table); ci->memcg_table = kzalloc_obj(*ci->memcg_table, gfp); if (!ci->memcg_table) { @@ -571,8 +579,8 @@ swap_cluster_populate(struct swap_info_struct *si, lockdep_assert_held(&si->global_cluster_lock); lockdep_assert_held(&ci->lock); - if (!swap_cluster_alloc_table(ci, __GFP_HIGH | __GFP_NOMEMALLOC | - __GFP_NOWARN)) + if (!swap_cluster_alloc_table(si, ci, __GFP_HIGH | __GFP_NOMEMALLOC | + __GFP_NOWARN)) return ci; /* @@ -585,8 +593,8 @@ swap_cluster_populate(struct swap_info_struct *si, spin_unlock(&si->global_cluster_lock); local_unlock(&percpu_swap_cluster.lock); - ret = swap_cluster_alloc_table(ci, __GFP_HIGH | __GFP_NOMEMALLOC | - GFP_KERNEL); + ret = swap_cluster_alloc_table(si, ci, __GFP_HIGH | __GFP_NOMEMALLOC | + GFP_KERNEL); /* * Back to atomic context. We might have migrated to a new CPU with a @@ -863,7 +871,7 @@ static int swap_cluster_setup_bad_slot(struct swap_info_struct *si, ci = cluster_info + idx; /* Need to allocate swap table first for initial bad slot marking. */ - if (!ci->count && swap_cluster_alloc_table(ci, GFP_KERNEL)) + if (!ci->count && swap_cluster_alloc_table(si, ci, GFP_KERNEL)) return -ENOMEM; spin_lock(&ci->lock); /* Check for duplicated bad swap slots. */ @@ -1086,7 +1094,9 @@ static bool __swap_cluster_alloc_entries(struct swap_info_struct *si, /* Try use a new cluster for current CPU and allocate from it. */ static unsigned int alloc_swap_scan_cluster(struct swap_info_struct *si, struct swap_cluster_info *ci, - struct folio *folio, unsigned long offset) + struct folio *folio, + unsigned long offset, + bool *nomem) { unsigned int next = SWAP_ENTRY_INVALID, found = SWAP_ENTRY_INVALID; unsigned long start = ALIGN_DOWN(offset, SWAPFILE_CLUSTER); @@ -1115,6 +1125,23 @@ static unsigned int alloc_swap_scan_cluster(struct swap_info_struct *si, if (!ret) continue; } +#ifdef CONFIG_MEMCG + /* + * Lazy-allocate memcg_table on the first direct-use slot of a + * physical cluster. + */ + if (vswap_is_enabled() && folio && + !folio_test_swapcache(folio) && !mem_cgroup_disabled() && + !ci->memcg_table) { + ci->memcg_table = kzalloc_obj(*ci->memcg_table, + GFP_ATOMIC | __GFP_NOWARN); + if (!ci->memcg_table) { + if (nomem) + *nomem = true; + goto out; + } + } +#endif if (!__swap_cluster_alloc_entries(si, ci, folio, offset % SWAPFILE_CLUSTER)) break; found = offset; @@ -1124,7 +1151,15 @@ static unsigned int alloc_swap_scan_cluster(struct swap_info_struct *si, break; } out: - relocate_cluster(si, ci); + /* + * On a discard-capable device, relocating a cluster whose memcg_table + * allocation failed queues a discard for slots that were never used, + * which folio_alloc_phys_swap() reads as progress and retries on. + */ + if (nomem && *nomem && !ci->count) + __free_cluster(si, ci); + else + relocate_cluster(si, ci); swap_cluster_unlock(ci); if (swap_is_vswap(si)) { this_cpu_write(percpu_vswap_cluster.offset[order], next); @@ -1145,7 +1180,13 @@ static unsigned int alloc_swap_scan_list(struct swap_info_struct *si, bool scan_all) { unsigned int found = SWAP_ENTRY_INVALID; + bool nomem = false; + /* + * In rare cases alloc_swap_scan_cluster() can fail due to + * memcg_table allocation failure. Short-circuit to avoid looping + * over the list indefinitely. + */ do { struct swap_cluster_info *ci = isolate_lock_cluster(si, list); unsigned long offset; @@ -1153,10 +1194,10 @@ static unsigned int alloc_swap_scan_list(struct swap_info_struct *si, if (!ci) break; offset = cluster_offset(si, ci); - found = alloc_swap_scan_cluster(si, ci, folio, offset); + found = alloc_swap_scan_cluster(si, ci, folio, offset, &nomem); if (found) break; - } while (scan_all); + } while (scan_all && !nomem); return found; } @@ -1177,7 +1218,8 @@ static unsigned int vswap_alloc_cluster(struct swap_info_struct *si, spin_lock_init(&ci_dyn->ci.lock); INIT_LIST_HEAD(&ci_dyn->ci.list); - if (swap_cluster_alloc_table(&ci_dyn->ci, GFP_ATOMIC | __GFP_NOWARN)) { + if (swap_cluster_alloc_table(si, &ci_dyn->ci, + GFP_ATOMIC | __GFP_NOWARN)) { kfree(ci_dyn); return SWAP_ENTRY_INVALID; } @@ -1203,7 +1245,7 @@ static unsigned int vswap_alloc_cluster(struct swap_info_struct *si, } offset = cluster_offset(si, ci); - return alloc_swap_scan_cluster(si, ci, folio, offset); + return alloc_swap_scan_cluster(si, ci, folio, offset, NULL); } static void swap_reclaim_full_clusters(struct swap_info_struct *si, bool force) @@ -1310,7 +1352,8 @@ static unsigned long cluster_alloc_swap_entry(struct swap_info_struct *si, if (cluster_is_usable(ci, order)) { if (cluster_is_empty(ci)) offset = cluster_offset(si, ci); - found = alloc_swap_scan_cluster(si, ci, folio, offset); + found = alloc_swap_scan_cluster(si, ci, folio, offset, + NULL); } else { swap_cluster_unlock(ci); } @@ -1354,7 +1397,7 @@ static unsigned long cluster_alloc_swap_entry(struct swap_info_struct *si, if (order < PMD_ORDER) { /* * Scan only one fragment cluster is good enough. Order 0 - * allocation will surely success, and large allocation + * allocation rarely fails, and large allocation * failure is not critical. Scanning one cluster still * keeps the list rotated and reclaimed (for clean swap cache). */ @@ -1369,7 +1412,7 @@ static unsigned long cluster_alloc_swap_entry(struct swap_info_struct *si, /* Order 0 stealing from higher order */ for (int o = 1; o < SWAP_NR_ORDERS; o++) { /* - * Clusters here have at least one usable slots and can't fail order 0 + * Clusters here have at least one usable slots and rarely fail order 0 * allocation, but reclaim may drop si->lock and race with another user. */ found = alloc_swap_scan_list(si, &si->frag_clusters[o], folio, true); @@ -1590,7 +1633,7 @@ static swp_entry_t swap_alloc_fast(struct folio *folio) if (ci && cluster_is_usable(ci, order)) { if (cluster_is_empty(ci)) offset = cluster_offset(si, ci); - found = alloc_swap_scan_cluster(si, ci, folio, offset); + found = alloc_swap_scan_cluster(si, ci, folio, offset, NULL); } else if (ci) { swap_cluster_unlock(ci); } @@ -1981,7 +2024,8 @@ static bool vswap_alloc(struct folio *folio) if (ci && cluster_is_usable(ci, order)) { if (cluster_is_empty(ci)) offset = cluster_offset(vswap_si, ci); - alloc_swap_scan_cluster(vswap_si, ci, folio, offset); + alloc_swap_scan_cluster(vswap_si, ci, folio, offset, + NULL); } else if (ci) { swap_cluster_unlock(ci); } @@ -2860,7 +2904,8 @@ swp_entry_t swap_alloc_hibernation_slot(int type) if (pcp_si == si && pcp_offset) { ci = swap_cluster_lock(si, pcp_offset); if (cluster_is_usable(ci, 0)) - offset = alloc_swap_scan_cluster(si, ci, NULL, pcp_offset); + offset = alloc_swap_scan_cluster(si, ci, NULL, + pcp_offset, NULL); else swap_cluster_unlock(ci); } -- 2.53.0-Meta