From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mta1.migadu.com (out-101.mta1.migadu.com [95.215.58.101]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 71CA6224D6 for ; Sat, 22 Aug 2026 02:47:22 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=95.215.58.101 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787366845; cv=none; b=THeDQ3E0Xvw+aqeogAatsH5Gmo74LPP6OoKrghqyhpwxi4diSZa9c+bJp3bkBpTFIk7Tul8fJhJ6+u71BFX/IG8DsAOSJC8McKWUbvJQQzmUIzrp/FsR64d3mOYhQlD6hzJq6JJx9/B6gG2SJO050JkjULVM1+m2qF74AvceZtM= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787366845; c=relaxed/simple; bh=eVVRhhBLN/+FkM0gx4duUnqi0A2acoQjl/ZnvoeEHME=; h=From:To:Cc:Subject:Date:Message-ID:MIME-Version; b=DJ39u6MdAxMCL0y0gzHmL1fFn8VFW4JkXJHByCIDwfm95GaonUnC3AFkl32XnyEynnZ0g+O3tq5GA7RxDWsnNn47SxJvGGOuoXjTTOxHyM4KZNTAivGZUCFzMmh6Uy+ezAulpJc9+S0FgSazh2ZSWo4MVYpXhL1sphm0NHk1fb0= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=gQX8HsSM; arc=none smtp.client-ip=95.215.58.101 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="gQX8HsSM" X-Envelope-To: linux-kernel@vger.kernel.org DKIM-Signature: a=rsa-sha256; bh=eVVRhhBLN/+FkM0gx4duUnqi0A2acoQjl/ZnvoeEHME=; c=simple/simple; d=linux.dev; h=from:to:subject:date:message-id:mime-version:content-type; s=key1; t=1787366841; v=1; x=1787971641; b=gQX8HsSM9DqTFD4agTBDOyYobA6sHis3+7otPremcuTr7orWFzBJ/jPQ/EEC506SVxs1U4ae g8c1NvQlwifvnOWX2Jd+gMT7XUDJjx8Icko6SY+WDmt565PVx+18BaxiHPExdB6OK+HK70lROjw J/ZQER65vq7hxnEyAyy+mmQk= X-Envelope-To: linux-kernel@vger.kernel.org Received: from localhost (2a03:2880:10ff:5c::) by smtp.migadu.com with ESMTPS id 07ae24895d946565; Sat, 22 Aug 2026 02:47:11 +0000 X-Mizu-Trace-ID: 07ae24895d946565 X-Migadu-Flow: FLOW_OUT From: Shakeel Butt To: Andrew Morton Cc: Johannes Weiner , Michal Hocko , Muchun Song , Qi Zheng , Roman Gushchin , Meta kernel team , linux-mm@kvack.org, linux-kernel@vger.kernel.org, stable@vger.kernel.org Subject: [PATCH] memcg: move LRU size accounting on reparenting instead of copying it Date: Fri, 21 Aug 2026 19:47:07 -0700 Message-ID: <20260822024707.77192-1-shakeel.butt@linux.dev> X-Mailer: git-send-email 2.53.0 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit When a memory cgroup is offlined its LRU folios are reparented to the parent. lruvec_reparent_lru() splices the child's lists into the parent's and credits the parent with the child's per-zone lru_zone_size[], but never clears the child's copy, so the size is copied rather than moved. lru_gen_reparent_memcg() does the same for MGLRU. The parent is left correct, credited with exactly the folios it took over. The stale value sits on the child and nothing will correct it: folio->memcg_data now resolves to the parent, so every later update_lru_size() for those folios goes there. Dying cgroups are not freed immediately and mem_cgroup_iter() still walks them, so shrink_lruvec() keeps being called on them. get_scan_count() reads the phantom counter through lruvec_lru_size() and the scan loop then grinds through nr[] in SWAP_CLUSTER_MAX steps against an empty list, for as long as the dead cgroup lives. Under MGLRU the MGLRU scanner runs instead, but count_shadow_nodes() sums all of NR_LRU_LISTS through lruvec_lru_size() and over-budgets the shadow node limit just the same. On one 251 GiB host a sweep of every mz->lru_zone_size[] found 380 counters describing folios on no list at all: 124777314 pages, 476 GiB, 1.89x the machine's RAM, across 57 cgroups. All were on memcgs with CSS_DYING set and CSS_ONLINE clear, and parent/child pairs reported byte-identical sizes. LRU_UNEVICTABLE needs its size moved too. Its list is deliberately not spliced because lruvec_init() poisons the head - the unevictable LRU is imaginary and folios are never threaded on it - but the size is kept by lruvec_add_folio()/lruvec_del_folio() and those folios account to the parent from here on. This depends on commit bf4ade7dbd76 ("memcg: keep folio's objcg same as its node") and must not be backported ahead of it. Without that invariant a folio's objcg can belong to another node, so a folio already spliced onto the parent's list can still resolve to the child's lruvec until the objcg's node is reparented in a later iteration of memcg_reparent_objcgs(); clearing the child's counter early then lets lruvec_del_folio() underflow it and trip the WARN_ONCE()/VM_BUG_ON() in mem_cgroup_update_lru_size(). Fixes: 07a6e9a2c199 ("mm: vmscan: prepare for reparenting traditional LRU folios") Fixes: f304652609ea ("mm: vmscan: prepare for reparenting MGLRU folios") Cc: # After: bf4ade7dbd76: memcg: keep folio's objcg same as its node Signed-off-by: Shakeel Butt --- mm/folio.c | 9 +++++++++ mm/vmscan.c | 5 +++++ 2 files changed, 14 insertions(+) diff --git a/mm/folio.c b/mm/folio.c index 59c477120b9a..c02dcea9c03c 100644 --- a/mm/folio.c +++ b/mm/folio.c @@ -1130,7 +1130,16 @@ static void lruvec_reparent_lru(struct lruvec *child_lruvec, for_each_managed_zone_pgdat(zone, NODE_DATA(nid), zid, MAX_NR_ZONES - 1) { unsigned long size = mem_cgroup_get_zone_lru_size(child_lruvec, lru, zid); + if (!size) + continue; + + /* + * The folios are accounted to the parent from now on, so the + * size has to be moved, not just copied. Leaving it behind + * makes the dying child describe folios it no longer owns. + */ mem_cgroup_update_lru_size(parent_lruvec, lru, zid, size); + mem_cgroup_update_lru_size(child_lruvec, lru, zid, -(long)size); } } diff --git a/mm/vmscan.c b/mm/vmscan.c index fe7f0c52a18c..561eeec5628c 100644 --- a/mm/vmscan.c +++ b/mm/vmscan.c @@ -4635,7 +4635,12 @@ void lru_gen_reparent_memcg(struct mem_cgroup *memcg, struct mem_cgroup *parent, for_each_managed_zone_pgdat(zone, NODE_DATA(nid), zid, MAX_NR_ZONES - 1) { unsigned long size = mem_cgroup_get_zone_lru_size(child_lruvec, lru, zid); + if (!size) + continue; + + /* Move the accounting, do not duplicate it. */ mem_cgroup_update_lru_size(parent_lruvec, lru, zid, size); + mem_cgroup_update_lru_size(child_lruvec, lru, zid, -(long)size); } } } -- 2.53.0-Meta