From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-ed1-f41.google.com (mail-ed1-f41.google.com [209.85.208.41]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E6F444052C2 for ; Mon, 24 Aug 2026 11:47:57 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.208.41 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787572081; cv=none; b=W/Nb/WDxZDbt3/rErt7OUIQmGwvUGUHrCIWvbkRf7dNAYjkw/WMGRZiIhnD1y23e3/eI0ZRGwnolwZ1Qa5Xb33RbkYKsNMyVjYJhC6TRevmTavLWmXw6DPAlIUHvUK34sh5RoFbkFFwQYctEKhd81CbWOz7j1cBX9R7bT7x6ozw= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787572081; c=relaxed/simple; bh=gY6G/UMz+0s4p+36zv5CU7QNwLHvEo+zz24SYFx1VZQ=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=F0MmUjIFmGidhsMhvoyJRO4dmqC3Do+rH0Uv87yCVE/prxU76ROII/Txrfwg8AN/8HejSmwd8ZzX1gYyA6FGy1xqZyuyyKQbHlQwxhzRXUcLKpNZlp4DgCgGq/tn7WltIlhGQzcAv6BGfow9rYdiRgLlt5HJ4qX+IYZlaJQs6RY= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=suse.com; spf=pass smtp.mailfrom=suse.com; dkim=pass (2048-bit key) header.d=suse.com header.i=@suse.com header.b=Y6nUCitI; arc=none smtp.client-ip=209.85.208.41 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=suse.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=suse.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=suse.com header.i=@suse.com header.b="Y6nUCitI" Received: by mail-ed1-f41.google.com with SMTP id 4fb4d7f45d1cf-69c600f76ccso5605810a12.0 for ; Mon, 24 Aug 2026 04:47:57 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=suse.com; s=google; t=1787572076; x=1788176876; darn=vger.kernel.org; h=in-reply-to:content-disposition:content-type:mime-version :references:message-id:subject:cc:to:from:date:from:to:cc:subject :date:message-id:reply-to:content-type; bh=S1OQ+o11GFCuRHAOrkLxBgVaNwncS1WphwbZ+fSPfb4=; b=Y6nUCitIZq+B7E4yMFTf8C/1Xy51lxpzsZiDPJ9uH5xZjBu395VhJqJMQxFOJ8bue9 veWu7g0tLIGevhQPoDTLGY5AuFGhL8LHY1E5exc9NUotcSsSFKh+nkI5QYTLpqMGBDsU OwUnNOB63b8zwi0xc8IWVgaKLrBhlC6syGDaMW7yOgymKQdZ1JXcWNFzaAo+tptsE7x+ iSWX2fl7R18vtG33GStdMAgguAABHFe9w2LNfF+lLLz4mA9aCdMnlVvMPjrIf4jO4Kuc WK9JgUxwFaQqLEpjZrS5RDnYW/L1PSix3WcXEozGf+Ox5G1+uFkxCtakhYKcyA8llqPM MYGw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787572076; x=1788176876; h=in-reply-to:content-disposition:content-type:mime-version :references:message-id:subject:cc:to:from:date:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=S1OQ+o11GFCuRHAOrkLxBgVaNwncS1WphwbZ+fSPfb4=; b=gnCbyP22zNevFsTpfd27odURf78+2tUrKypye14Gr/eMweJV1XRNZYTkzr4j7fiN6U ftDKS5XFwXtK1WXqj01WzwAK3fyKRIYgk1aOIpgFfTXxsDwDKOVBRkOU92ueTM9lQsnc 7oNP6tWap+shDnZ9+C9m/VjOMLkdFpfC597xR09zTEvYQFSztNScmSC2dcfVs9yF3e6r LPEXKrz75nWRcOL+FQOkK7NQv+6S3DxiiCZ9LOrzB7fGfKgVdRD7Ec4Jl7xaHiWY0IM0 iErwFraH49CGeIB+tNMe/KwyU2y4acNcck3y1QLek2oxTaRiZPV771dGHkI0pUrs+LKi 6iGw== X-Forwarded-Encrypted: i=1; AHgh+RqPU7Jdl3Vk1g76+2TiH0eqigZFYs3vIMQNGnNZImQ5g3CxHlHVRNJ5XyQon3sHs0U2ZE/P9tbvHjBSN8I=@vger.kernel.org X-Gm-Message-State: AFuF++nAKFpXaBhmmqSe7GmEhzpkjglTbSlmshffP43pVzuXtMKiwvm0 h918Fg7V1zgvhE1yE5FFOwejb3rU2V5y+sMzRgAbnyglFe4LIT5cyQOObbjKJQvlrfs= X-Gm-Gg: AR+sD11P7Ji/1rURoINiyGRUjktO0UFgb3zpQ/SPRk6sKd+8DHr13r6jzOnO0kIer7w IgXMUQK7OniWyuHRtz/XDgMi/c0kgr4VAHn9Q/C3IohpJP1+ir7IOchEvi5m7GLWsLqJ81w1TAf YBrHAYE7u8iPHt14ziHK/ZFnSSk2mD/UkWE1baEIb9B3ic5I0tgeyRpEntiTPd57BQvbBpEkTRd 1eb/t1Ry4IcDOewtA021COubipNn5knJHcv13f2uPnhU9HCOxVJY9CJVKwK1NvarLYHc4EQMUFT WjMPry/+NqevW9z2OqJNvbZgIS7PCiowJGshSUA0zTFJDvAKW//GbGHCbxUgU3vMDpIfPxtdf3k 4/syBrChN7CM3rOaF2rjC1A3ug+Ohz+DyrLlpwAAweSHASyE/5+yDjZAkIQ5gCl0XV1sEMbjzI1 DnjvAcAfA/O94FD+yxHKZoZSTGh8/czP1/qwmUEpK/QejNvH2iwRjYj1X0Jg/h99XDM6zvHg/M8 xr2Xumagud8 X-Received: by 2002:a05:6402:450f:b0:6a0:a65a:a0c2 with SMTP id 4fb4d7f45d1cf-6a42f1a0226mr29775246a12.7.1787572076095; Mon, 24 Aug 2026 04:47:56 -0700 (PDT) Received: from localhost (109-81-81-112.rct.o2.cz. [109.81.81.112]) by smtp.gmail.com with ESMTPSA id 4fb4d7f45d1cf-6a59e00108asm12355655a12.2.2026.08.24.04.47.55 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 24 Aug 2026 04:47:55 -0700 (PDT) Date: Mon, 24 Aug 2026 13:47:54 +0200 From: Michal Hocko To: Shakeel Butt Cc: Andrew Morton , Johannes Weiner , Muchun Song , Qi Zheng , Roman Gushchin , Meta kernel team , linux-mm@kvack.org, linux-kernel@vger.kernel.org, stable@vger.kernel.org Subject: Re: [PATCH] memcg: move LRU size accounting on reparenting instead of copying it Message-ID: References: <20260822024707.77192-1-shakeel.butt@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <20260822024707.77192-1-shakeel.butt@linux.dev> On Fri 21-08-26 19:47:07, Shakeel Butt wrote: > When a memory cgroup is offlined its LRU folios are reparented to the > parent. lruvec_reparent_lru() splices the child's lists into the > parent's and credits the parent with the child's per-zone > lru_zone_size[], but never clears the child's copy, so the size is > copied rather than moved. lru_gen_reparent_memcg() does the same for > MGLRU. > > The parent is left correct, credited with exactly the folios it took > over. The stale value sits on the child and nothing will correct it: > folio->memcg_data now resolves to the parent, so every later > update_lru_size() for those folios goes there. > > Dying cgroups are not freed immediately and mem_cgroup_iter() still > walks them, so shrink_lruvec() keeps being called on them. > get_scan_count() reads the phantom counter through lruvec_lru_size() and > the scan loop then grinds through nr[] in SWAP_CLUSTER_MAX steps against > an empty list, for as long as the dead cgroup lives. Under MGLRU the > MGLRU scanner runs instead, but count_shadow_nodes() sums all of > NR_LRU_LISTS through lruvec_lru_size() and over-budgets the shadow node > limit just the same. > > On one 251 GiB host a sweep of every mz->lru_zone_size[] found 380 > counters describing folios on no list at all: 124777314 pages, 476 GiB, > 1.89x the machine's RAM, across 57 cgroups. All were on memcgs with > CSS_DYING set and CSS_ONLINE clear, and parent/child pairs reported > byte-identical sizes. > > LRU_UNEVICTABLE needs its size moved too. Its list is deliberately not > spliced because lruvec_init() poisons the head - the unevictable LRU is > imaginary and folios are never threaded on it - but the size is kept by > lruvec_add_folio()/lruvec_del_folio() and those folios account to the > parent from here on. > > This depends on commit bf4ade7dbd76 ("memcg: keep folio's objcg same as > its node") and must not be backported ahead of it. Without that > invariant a folio's objcg can belong to another node, so a folio already > spliced onto the parent's list can still resolve to the child's lruvec > until the objcg's node is reparented in a later iteration of > memcg_reparent_objcgs(); clearing the child's counter early then lets > lruvec_del_folio() underflow it and trip the WARN_ONCE()/VM_BUG_ON() in > mem_cgroup_update_lru_size(). > > Fixes: 07a6e9a2c199 ("mm: vmscan: prepare for reparenting traditional LRU folios") > Fixes: f304652609ea ("mm: vmscan: prepare for reparenting MGLRU folios") > Cc: # After: bf4ade7dbd76: memcg: keep folio's objcg same as its node > Signed-off-by: Shakeel Butt Acked-by: Michal Hocko Thanks! > --- > mm/folio.c | 9 +++++++++ > mm/vmscan.c | 5 +++++ > 2 files changed, 14 insertions(+) > > diff --git a/mm/folio.c b/mm/folio.c > index 59c477120b9a..c02dcea9c03c 100644 > --- a/mm/folio.c > +++ b/mm/folio.c > @@ -1130,7 +1130,16 @@ static void lruvec_reparent_lru(struct lruvec *child_lruvec, > for_each_managed_zone_pgdat(zone, NODE_DATA(nid), zid, MAX_NR_ZONES - 1) { > unsigned long size = mem_cgroup_get_zone_lru_size(child_lruvec, lru, zid); > > + if (!size) > + continue; > + > + /* > + * The folios are accounted to the parent from now on, so the > + * size has to be moved, not just copied. Leaving it behind > + * makes the dying child describe folios it no longer owns. > + */ > mem_cgroup_update_lru_size(parent_lruvec, lru, zid, size); > + mem_cgroup_update_lru_size(child_lruvec, lru, zid, -(long)size); > } > } > > diff --git a/mm/vmscan.c b/mm/vmscan.c > index fe7f0c52a18c..561eeec5628c 100644 > --- a/mm/vmscan.c > +++ b/mm/vmscan.c > @@ -4635,7 +4635,12 @@ void lru_gen_reparent_memcg(struct mem_cgroup *memcg, struct mem_cgroup *parent, > for_each_managed_zone_pgdat(zone, NODE_DATA(nid), zid, MAX_NR_ZONES - 1) { > unsigned long size = mem_cgroup_get_zone_lru_size(child_lruvec, lru, zid); > > + if (!size) > + continue; > + > + /* Move the accounting, do not duplicate it. */ > mem_cgroup_update_lru_size(parent_lruvec, lru, zid, size); > + mem_cgroup_update_lru_size(child_lruvec, lru, zid, -(long)size); > } > } > } > -- > 2.53.0-Meta -- Michal Hocko SUSE Labs