From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-qt1-f172.google.com (mail-qt1-f172.google.com [209.85.160.172]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B765435C180 for ; Fri, 3 Apr 2026 19:45:37 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.160.172 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1775245540; cv=none; b=DipKUyh824LRSkC4uOQ88EF3lAHGeEG5oXYDG3HD+nvNSkfeHxJznNA32h/Q9iHEj3B76Logseq/P3A0QuoPs/y4BqCF22yW9oeZXDVlOoJTI7EBxhCOEqx4XLt2/gvmCTo8AXGOWcDTSqtqHFrOAE0V+j8EkqaOWXzerR4xkjs= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1775245540; c=relaxed/simple; bh=Z/+YvZ2tyvOjI0Jh+d23e4YZQo/ntJmf18niT0iJaew=; h=From:To:Cc:Subject:Date:Message-ID:MIME-Version; b=hdVGTzaZzGM9FL9r6rKWjbSDf+2opLFQ/Du5vFFdBqkbaFMpAMcw7Nzq1ihjidpeTGmei+UvVOG6vllRqAR2n7np/z9lqhBXeZQPkxsbJ8NuQTpKEGF3QgY8GtoGgQ0+1JDwnY82r7SMDSEK/dg0EXxZrGDQNT+VP1FFdO1GeKg= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=cmpxchg.org; spf=pass smtp.mailfrom=cmpxchg.org; dkim=pass (2048-bit key) header.d=cmpxchg.org header.i=@cmpxchg.org header.b=Uf1qFrnT; arc=none smtp.client-ip=209.85.160.172 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=cmpxchg.org Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=cmpxchg.org Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=cmpxchg.org header.i=@cmpxchg.org header.b="Uf1qFrnT" Received: by mail-qt1-f172.google.com with SMTP id d75a77b69052e-50b2b289925so20028871cf.2 for ; Fri, 03 Apr 2026 12:45:37 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=cmpxchg.org; s=google; t=1775245536; x=1775850336; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:from:to:cc:subject:date:message-id:reply-to; bh=MkCnZMc51A5ktMZ7NvfpuZ1Tvt5XekG9z+r9y+GS9pg=; b=Uf1qFrnTwyvFUblOe5YU7uzPzJYXPy953J+m3xSVAB+z1tZk//ngwMWLCPI/sVYGHz 8TG8qFLtdzZaLcP30rlX8bOWYA1P8X5+qLNsYixY703pg8sFMYdqGxQUAYStH9muld8G +zAuGRWSHaxPPIQ16FbYkGggrGYIfqPQoERY1TsVVbVUFUmCC4mRNfUN0g4Tw3/a7MJ1 E5RkmOAv5rSBc0fd/cA+E1Z4fG/t6lG4SqKjRNXSL/GPpBswGJ7I3PCG7gps286by3yY cf5OzOxKWxFy4xbpmn116Ihf7O2WWk51U1Aa1lNNEdOVDIrQ76Xxd1ouMSVR4dUCQ38L fK/Q== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1775245536; x=1775850336; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:x-gm-gg:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to; bh=MkCnZMc51A5ktMZ7NvfpuZ1Tvt5XekG9z+r9y+GS9pg=; b=h8pFT6Ma84ncBRazMV5Zpy55P5PB1KNQH5/799RHKzBaYv7XJ6ywm9Fy/VjPFiOXkw 4vxqd23510h8m0ZHWSvnspPZA2PkF1g4wXoYMoP1FtiM48kbc3TbGLZHKwc6EMp4XfvN P3qSXvmM7v8V/66Y+IuvYUEUI0tiB2K8PCCoeDWwdyRYntqjLjQLovozCotyzzmEHFwO MA5wUYC4bO6dDR6YRe08u0HC5pA2PNsOFEKoyPORO9WOJ+Sk4HdoOSBAgmnc8jJEWzeG s7yTsHaTwKTg5swMFfgmxtTKSoIMGQnAQixyGmf2t5MQx4iAlQ5E+6irCMgIk7GGZyJr GHrw== X-Forwarded-Encrypted: i=1; AJvYcCUZE5kjYlMb/9zOlYVJgLrdGHTjV2k9DQqWehASIvwuZzvS+iEET2gIQ26n5vgPc+3Ag/TNv+u3BsbK8mQ=@vger.kernel.org X-Gm-Message-State: AOJu0YzRiOiNT8MUsrYSd0ku776b/8EAgT6hMWTxj6MPvnSDE6uPzMcE 4SAmDwm/kd15pUBakbSKhk1rfp+/uyi0j+Efn5XzAxQtV+LO0O8/ds4pltUojEJVJ/ev9KULyrE 4nttL X-Gm-Gg: ATEYQzxTHhz3dyDw0HRJIXFXQ6abKOYg/CJ9VmZedyMg1P+eeyI2Ng3fKyVxDQIW4o5 PZ21XLrmHPjm2bLs+rBEAlyf9bwFm0ZmR5zgCTAbKcE1Ri4QLMwJUcYrcXfFncxusQW2i7sLTwf /lZu+9pi4PXBlBTsNfTTPvhEj4FLNgE+wyHmbgZWPMkcpOYCSq/bMJUotXwXfYnvqH7paHmR/IW oby4PLE8AtnTevYriAcBPghfzfDgyljEf3LyK5OUWclXdAuXoPlRbe9IHv3TCIaikXx4PLNY/vM uDlZAMZgMQ3Qx/iCiKq/HxAoI/EUQBiG27XOkAjuLocwHjqkiD5/m+BHj1UixATUVYVZsrgKxPn y4N2aaZaCZAxHo9zXuZar4MKbvJzixHNF9UmZIFnOT1wDLrlirLORCZZBI/yk1QYlMDJNA4i8rz UW2mTyS4+eIx2zW8wdSn31gqojzGeA9+yO X-Received: by 2002:ac8:6e8b:0:b0:50b:8689:bd4f with SMTP id d75a77b69052e-50d62b0bf34mr47427901cf.55.1775245536394; Fri, 03 Apr 2026 12:45:36 -0700 (PDT) Received: from localhost ([2603:7000:c00:3a00:365a:60ff:fe62:ff29]) by smtp.gmail.com with ESMTPSA id 6a1803df08f44-8a5976e223dsm65418606d6.45.2026.04.03.12.45.35 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 03 Apr 2026 12:45:35 -0700 (PDT) From: Johannes Weiner To: linux-mm@kvack.org Cc: Vlastimil Babka , Zi Yan , David Hildenbrand , Lorenzo Stoakes , "Liam R. Howlett" , Rik van Riel , linux-kernel@vger.kernel.org Subject: [RFC 0/2] mm: page_alloc: pcp buddy allocator Date: Fri, 3 Apr 2026 15:40:33 -0400 Message-ID: <20260403194526.477775-1-hannes@cmpxchg.org> X-Mailer: git-send-email 2.53.0 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit Hi, this is an RFC for making the page allocator scale better with higher thread counts and larger memory quantities. In Meta production, we're seeing increasing zone->lock contention that was traced back to a few different paths. A prominent one is the userspace allocator, jemalloc. Allocations happen from page faults on all CPUs running the workload. Frees are cached for reuse, but the caches are periodically purged back to the kernel from a handful of purger threads. This breaks affinity between allocations and frees: Both sides use their own PCPs - one side depletes them, the other one overfills them. Both sides routinely hit the zone->locked slowpath. My understanding is that tcmalloc has a similar architecture. Another contributor to contention is process exits, where large numbers of pages are freed at once. The current PCP can only reduce lock time when pages are reused. Reuse is unlikely because it's an avalanche of free pages on a CPU busy walking page tables. Every time the PCP overflows, the drain acquires the zone->lock and frees pages one by one, trying to merge buddies together. The idea proposed here is this: instead of single pages, make the PCP grab entire pageblocks, split them outside the zone->lock. That CPU then takes ownership of the block, and all frees route back to that PCP instead of the freeing CPU's local one. This has several benefits: 1. It's right away coarser/fewer allocations transactions under the zone->lock. 1a. Even if no full free blocks are available (memory pressure or small zone), with splitting available at the PCP level means the PCP can still grab chunks larger than the requested order from the zone->lock freelists, and dole them out on its own time. 2. The pages free back to where the allocations happen, increasing the odds of reuse and reducing the chances of zone->lock slowpaths. 3. The page buddies come back into one place, allowing upfront merging under the local pcp->lock. This makes coarser/fewer freeing transactions under the zone->lock. The big concern is fragmentation. Movable allocations tend to be a mix of short-lived anon and long-lived file cache pages. By the time the PCP needs to drain due to thresholds or pressure, the blocks might not be fully re-assembled yet. To prevent gobbling up and fragmenting ever more blocks, partial blocks are remembered on drain and their pages queued last on the zone freelist. When a PCP refills, it first tries to recover any such fragment blocks. On small or pressured machines, the PCP degrades to its previous behavior. If a whole block doesn't fit the pcp->high limit, or a whole block isn't available, the refill grabs smaller chunks that aren't marked for ownership. The free side will use the local PCP as before. I still need to run broader benchmarks, but I've been consistently seeing a 3-4% reduction in %sys time for simple kernel builds on my 32-way, 32G RAM test machine. A synthetic test on the same machine that allocates on many CPUs and frees on just a few sees a consistent 1% increase in throughput. I would expect those numbers to increase with higher concurrency and larger memory volumes, but verifying that is TBD. Sending an RFC to get an early gauge on direction. Based on 0257f64bdac7fdca30fa3cae0df8b9ecbec7733a. include/linux/mmzone.h | 38 ++- include/linux/page-flags.h | 9 + mm/debug.c | 1 + mm/internal.h | 17 + mm/mm_init.c | 25 +- mm/page_alloc.c | 784 +++++++++++++++++++++++++++++++------------ mm/sparse.c | 3 +- 7 files changed, 622 insertions(+), 255 deletions(-)