From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-qk1-f181.google.com (mail-qk1-f181.google.com [209.85.222.181]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C4A42347BDB for ; Thu, 12 Mar 2026 14:26:17 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.222.181 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1773325580; cv=none; b=fWZb5lEq0Gcv1JOBbJs6lSuL0FbeUFJ0t4us2XPX6NfirMH2LGIwhBPhzz278rS1Oe9isll50zzDxpg3XvqPPp4EMZsOufPrZnI6vW6XwveA8APwo4M5Dh7WZLop/hpzUl5hiB7ellixDTDE/ALL50iJuRMlEn719xLfUdJN1as= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1773325580; c=relaxed/simple; bh=NJuQ1+viDSkq9xU0DMf3BlMO1JT3ftacNFgriwKWLzY=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=Ymw7O9A5I4WtUy6KS0o5Zcxty+X/E1/EtPZLLWoEQCSsvxcUc5kN/XgcphViFZkNoGPklj1FuHYu2iD/IKOInI2KElynKHA1vNs7NbPYKyFoGWmLMv+K1s1lkb/Bh7AtiJolqVzCeJ5n1rCvMU9x5IW45cBLZLVHCYdlCjzOIMg= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=cmpxchg.org; spf=pass smtp.mailfrom=cmpxchg.org; dkim=pass (2048-bit key) header.d=cmpxchg.org header.i=@cmpxchg.org header.b=RO5xhblh; arc=none smtp.client-ip=209.85.222.181 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=cmpxchg.org Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=cmpxchg.org Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=cmpxchg.org header.i=@cmpxchg.org header.b="RO5xhblh" Received: by mail-qk1-f181.google.com with SMTP id af79cd13be357-8c9f6b78ca4so137119685a.0 for ; Thu, 12 Mar 2026 07:26:17 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=cmpxchg.org; s=google; t=1773325576; x=1773930376; darn=vger.kernel.org; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:from:to:cc:subject:date:message-id:reply-to; bh=dhyHTRhO4qLOZKTEybzXY/CiyPdSzoBcqB+Ti2SR1/g=; b=RO5xhblhPOC6vB3/XRFALTgfaOFKxn4o86+uwnzBxLtxqpwnLoSoLASMT/27153y2n iE3onggyL/TvAqeqc0hCnJPmavRtrPf+PjkeDS1CN4KVHQlupbxfHzUKYe/NEwu8MS35 Vn0fkt/j7N++YfV5P3U4ThOZ6vH0H7/J93GB7pSmzuOn150jEu9R5A/zx3Rz/vWoTkzM j9A3GwC+uVUSmoAJMJIa64S57hWqLA8XsJwwLB0BI4C6M307kmkxUPu2p0XI9wZgvZEO KJWi2vAe4n54E9cpUlY7VOQF+GBYvEDf74W3FjClJMXJ4I8mSb+DEVtai2L5/QcB/Ust dF0Q== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1773325576; x=1773930376; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-gg:x-gm-message-state:from:to:cc :subject:date:message-id:reply-to; bh=dhyHTRhO4qLOZKTEybzXY/CiyPdSzoBcqB+Ti2SR1/g=; b=B1DHTtT2oSmnaaR4UDz6nU930+QWpeNYZERNbnStm7D8j/nypA4OL25r8mzeQBKtSs J+RjybxkZUnmR3w0a0F3sDEx7IivsbbOdF6p0kxchyLDleyncw0arIUyW1VpLvsphauP z9g50diHYVQU4emGXTLEPEFq5LInDPbw5FD1d2i5TBtNZDtZJInAOrthxNv8JKfHb0Jr n5t7JeZ5834x3/euWNRtVudyooLDS5+V5aor19lsqDCDfewvi5vzYqugUbdHa8765YXB YVeKrINEEF56E2waapF864LlylgzBPihGvIqi+zkMfQId1sxlrDDKOi5xkpvDYCy+iNa ywUQ== X-Forwarded-Encrypted: i=1; AJvYcCWFnjxv3Z1tq2E9YLS4wRCAOb9cbx+W2fJrzN2LHpMZiQJkcmUqK8x7Cu2UHEVNmdq/BvveyX+PqKexZEs=@vger.kernel.org X-Gm-Message-State: AOJu0YxKLGBugCpXko3ErzXllPbAJJLp1DUiOjDzuCRXj6mGBA8Ue5G5 K1IYFXIWiu2Xq7vA06nfLlycR40lHixt1Vc1Y/t1OmCrbtaj6lOnSrKRPd2XP7cbtic= X-Gm-Gg: ATEYQzzuADfz7nQGN6ka+WQ4HEgNZYXGuv/mGoz8FwlcSYBmLo6gSPXzATdqboUBvUO fP/4UlJo8f8vi8Ave2waJml4WLLsF2nWhAhmBP/TwoKz4AfXttdpd6Sr/+b82oVomm6U+0Ml9tO /zv6oN6gsiBpuPHgml19Q7RaODgW/HFmYE/FYzZsj7LWrRnO8ve5UfsVN9pv5MYMSAJ2MZECzjm OJe303Lnu7n1JseKlClvFVsOqgmhQcIYreMMJJHL4BVo4VqKji4WYAO5AoMKDDmPKObUwDIVySM nGO3gnRV1/EBuPVUOnHtryVGy8yVCRS01P4e6Y0ETLvRTtiB0deThggeilvdDna3Cvf0ShZ1tiE mHz2fG6CxKMHrfn85sy6qvdwsbL6IVRdorOP9qLRcqc1P4uutu4kJe/TyerCi8kLH8+OdLAeEF7 b/u0XKBw0uTQFOmJTkTqqIQURg5sQu90tq X-Received: by 2002:a05:620a:414c:b0:8cd:9b4c:1470 with SMTP id af79cd13be357-8cda1ab570fmr851624885a.60.1773325576500; Thu, 12 Mar 2026 07:26:16 -0700 (PDT) Received: from localhost ([2603:7000:c00:3a00:365a:60ff:fe62:ff29]) by smtp.gmail.com with ESMTPSA id af79cd13be357-8cda210d972sm391016085a.28.2026.03.12.07.26.15 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 12 Mar 2026 07:26:15 -0700 (PDT) Date: Thu, 12 Mar 2026 10:26:14 -0400 From: Johannes Weiner To: Dave Chinner Cc: Andrew Morton , David Hildenbrand , Zi Yan , "Liam R. Howlett" , Usama Arif , Kiryl Shutsemau , Dave Chinner , Roman Gushchin , linux-mm@kvack.org, linux-kernel@vger.kernel.org Subject: Re: [PATCH] mm: switch deferred split shrinker to list_lru Message-ID: References: <20260311154358.150977-1-hannes@cmpxchg.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: On Thu, Mar 12, 2026 at 09:23:35AM +1100, Dave Chinner wrote: > On Wed, Mar 11, 2026 at 11:43:58AM -0400, Johannes Weiner wrote: > > The deferred split queue handles cgroups in a suboptimal fashion. The > > queue is per-NUMA node or per-cgroup, not the intersection. That means > > on a cgrouped system, a node-restricted allocation entering reclaim > > can end up splitting large pages on other nodes: > > > > alloc/unmap > > deferred_split_folio() > > list_add_tail(memcg->split_queue) > > set_shrinker_bit(memcg, node, deferred_shrinker_id) > > > > for_each_zone_zonelist_nodemask(restricted_nodes) > > mem_cgroup_iter() > > shrink_slab(node, memcg) > > shrink_slab_memcg(node, memcg) > > if test_shrinker_bit(memcg, node, deferred_shrinker_id) > > deferred_split_scan() > > walks memcg->split_queue > > > > The shrinker bit adds an imperfect guard rail. As soon as the cgroup > > has a single large page on the node of interest, all large pages owned > > by that memcg, including those on other nodes, will be split. > > > > list_lru properly sets up per-node, per-cgroup lists. As a bonus, it > > streamlines a lot of the list operations and reclaim walks. It's used > > widely by other major shrinkers already. Convert the deferred split > > queue as well. > > > > The list_lru per-memcg heads are instantiated on demand when the first > > object of interest is allocated for a cgroup, by calling > > memcg_list_lru_alloc(). Add calls to where splittable pages are > > created: anon faults, swapin faults, khugepaged collapse. > > > > These calls create all possible node heads for the cgroup at once, so > > the migration code (between nodes) doesn't need any special care. > > > > The folio_test_partially_mapped() state is currently protected and > > serialized wrt LRU state by the deferred split queue lock. To > > facilitate the transition, add helpers to the list_lru API to allow > > caller-side locking. > > > > Signed-off-by: Johannes Weiner > > --- > > include/linux/huge_mm.h | 6 +- > > include/linux/list_lru.h | 48 ++++++ > > include/linux/memcontrol.h | 4 - > > include/linux/mmzone.h | 12 -- > > mm/huge_memory.c | 326 +++++++++++-------------------------- > > mm/internal.h | 2 +- > > mm/khugepaged.c | 7 + > > mm/list_lru.c | 197 ++++++++++++++-------- > > mm/memcontrol.c | 12 +- > > mm/memory.c | 52 +++--- > > mm/mm_init.c | 14 -- > > 11 files changed, 310 insertions(+), 370 deletions(-) > > Can you please split this up into multiple patches (i.e. one logical > change per patch) to make it easier to review? No problem, I'll do that and send out a v2. The list_lru changes started as only the locking functions, then things kept creeping in... Thanks