From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pl1-f181.google.com (mail-pl1-f181.google.com [209.85.214.181]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C521736AB4B for ; Fri, 12 Jun 2026 18:09:48 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.181 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1781287789; cv=none; b=gswDqQHY/LXlrjMOFYrNhr5AQ9bUgrqTPJaPiNp00rauwWCkdAPoJGmxZGOG5/wRWb5KVlfohxfjcbTsR7qkiuQlgNPk6LNqxtPq1D70p6lLjF+t1pT5fx6VrldIPNEMKY0K4q9PBLf2xFQAJns9X59+mVBp9KKqUYMWchrPPC4= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1781287789; c=relaxed/simple; bh=Q+9pcS/pBPiF1oPivDHV9NFFwG8jSo9lvD/DEY4Gbz8=; h=From:To:Cc:Subject:Date:Message-Id:In-Reply-To:References: MIME-Version; b=rLc0gkCZSvaciEZVjht+TmDZP/R+aX7R7FLPLcIFop3UdfBWzDzFMkN9J7kaBJ3c5slEsIYYbyRziDI8ljoQpxdjnedRIoSyTm0n0WqI3hznlQzSo3AxRoSg+RmfJd5ikEERyMaJqi6vA0oiAnNZQfvrF3cOdEvdTyYumYGLzv8= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=exnvcMTS; arc=none smtp.client-ip=209.85.214.181 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="exnvcMTS" Received: by mail-pl1-f181.google.com with SMTP id d9443c01a7336-2c0c32f6ce1so9559145ad.2 for ; Fri, 12 Jun 2026 11:09:48 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1781287788; x=1781892588; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to; bh=ScGPgP09d6CTH8k9i/xxOIHmej0/0GzUcvmFJcghR1c=; b=exnvcMTSryJCB10wmfVBruyL+VgXwh3SL6wMQMUYGfLFl4mL9nFurdGngCH2cnv5Eo 5hpD0vfAc7hCjroEHc7JjqfbcvUT3Zyq9uh3KvAilUVAotAFEg514vG4ruEekyX/HiyH jlR/g5RaQ0qZSeMVcfO8yG2c/cgWE3Fo766aGf6bDCCpVD3vB6vlXdlf96xSpEjTb4Ao udcs/yLYw+qZkXwirkuAELTjjylMP6DPp9lzp4jf0zf5bUIFk94Lc7MH4n2nO8N5OaDj AH+0mmRckcI/M8C/0iZMU0MoQFwnTOslfTQaL8LbPdfbF6ANB7CSSsiIdgE8hOeM4X0r 4tPg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1781287788; x=1781892588; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to; bh=ScGPgP09d6CTH8k9i/xxOIHmej0/0GzUcvmFJcghR1c=; b=OJIT+8tOs20p8c/6yTXF1Ai9ZpAI0VKoeCByoht5w6HHmU1dU7dwa7krT3rMqZFuTT /k8SiWXjrTpgDeJga4VqLWMvpDYsAz8HhnLZcXxrl3Mn6EKcIrI1zATYkwF6hPJV9fBc bbv8C6iqUcVxPfUMPkGxFtFJjQYhVDXGdLXghhOVcqJXOAhXwGeA5oKrYrE1Mun59LLm 6TPXRZqzlVOsE2Wquia+DZcfcZ32PgbBmXnhnZwgImk0DT95zZLWYmgxm4TA2rk8DOtl tMXkEzYExjDJwaKYMr9aZTGCYGhAf2jPwqYf/REAY6u0iUUQf44ny/1iSBrr8mX3SBie n8rQ== X-Forwarded-Encrypted: i=1; AFNElJ+3d8JE2JE7XkBBvxIR+sADMPbnw7t3kqD9iiH/++r5+a9nh20/g7cF9KajP/vYoYpoL++z4UlZ+MCUqyg=@vger.kernel.org X-Gm-Message-State: AOJu0Yz8pajjlJHyz9WqmCHB+Larjh21sAlB33ugNsiSdJrQ/VXDpQxl 98zh4nDCUt7783qfs4dKE1FBr4iRupFGL+dHw/pduKQV0KEA1L4MSJaU X-Gm-Gg: Acq92OFn4IueHvXmgUDbr4uUkevfvwpkhhP5WiCajkor8jC3SE7OJ0prR5WVxZKHwUl 00ud1/3dW00Zwr0DMdb897IxocIzy/pmmN+SQngLaLSI2K2sirtiaDVhFmPkuKrv5BYgCiaR0WJ bpIIqGsZRmj8rK4AeNsYKfR9BW1qhTz3VjegMj6ewS4skRWGBKEP8TpherPYLPXuqDK5PGVyWyX zAWuCxieBz4LPGQgRNMNVVP2GlzrTR7sHBS9oUkhAuc6+GG7jpiRLbh93Mn3CVq2BqNmUegHWW/ 2f3STim8pRp5CeoaeTWwtTBpojiDqfkKWCGu/eGGsem/iQERE6iP38t3ux9gETQJ7wS/6eGlXAU c2FKiq3cBC2sCOPRUrGRglhKKmTHWy1F8GHbBOc5tWyZ0JLWg7BiAb8Y+JNL+uBYByVvQwQl9rs NHJnmT+NbVxbCAkuNIeKo+hPfSmYusOY3j X-Received: by 2002:a17:902:d48f:b0:2c2:1982:5270 with SMTP id d9443c01a7336-2c6641e28f9mr8104325ad.21.1781287788002; Fri, 12 Jun 2026 11:09:48 -0700 (PDT) Received: from pve-server.rlab ([49.205.216.49]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2c42f7c70easm38956975ad.25.2026.06.12.11.09.42 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 12 Jun 2026 11:09:47 -0700 (PDT) From: "Ritesh Harjani (IBM)" To: linux-mm@kvack.org Cc: Madhavan Srinivasan , Michael Ellerman , Nicholas Piggin , Christophe Leroy , Andrew Morton , Chris Li , Kairui Song , Kemeng Shi , Nhat Pham , Baoquan He , Barry Song , Youngjun Park , David Hildenbrand , linuxppc-dev@lists.ozlabs.org, linux-kernel@vger.kernel.org, Sayali Patil , "Ritesh Harjani (IBM)" Subject: [PATCH v3 2/3] mm, swap: allow archs to override SWAP_NR_ORDERS via ARCH_MAX_PMD_ORDER Date: Fri, 12 Jun 2026 23:39:16 +0530 Message-Id: X-Mailer: git-send-email 2.39.5 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit SWAP_NR_ORDERS sizes a few small bounded arrays inside THP swap allocator code (nofull/frag cluster lists, percpu_swap_cluster's si/offset arrays, next array for rotational device). This currently expands to PMD_ORDER+1, which only works when PMD_ORDER is a compile time constant. However on architecture like PowerPC Book3S64, PMD_ORDER is a runtime variable which depends upon which MMU is selected (Radix / Hash), so in that case, PMD_ORDER cannot be used to size the static arrays. This patch provides an optional ARCH_MAX_PMD_ORDER (upper-bound) override for such architectures. The memory overhead on enabling this override is negligible. Even if we make SWAP_NR_ORDERS runtime alloc, default slab padding could cause some memory waste. Also we lose the per-cpu cacheline benefits (for percpu_swap_cluster) because it might cost an extra cacheline indirection overhead in swap_alloc_fast() for fetching si[order]/offset[order]. Note that a fully runtime SWAP_NR_ORDERS was considered in previous version but was dropped for this reason [1] [1]: https://lore.kernel.org/linuxppc-dev/pl1zdksc.ritesh.list@gmail.com/ Suggested-by: YoungJun Park Signed-off-by: Ritesh Harjani (IBM) --- arch/powerpc/include/asm/book3s/64/pgtable.h | 7 +++++++ include/linux/swap.h | 12 +++++++++++- 2 files changed, 18 insertions(+), 1 deletion(-) diff --git a/arch/powerpc/include/asm/book3s/64/pgtable.h b/arch/powerpc/include/asm/book3s/64/pgtable.h index e67e64ac6e8c..7f22d5d5fbdf 100644 --- a/arch/powerpc/include/asm/book3s/64/pgtable.h +++ b/arch/powerpc/include/asm/book3s/64/pgtable.h @@ -204,6 +204,13 @@ extern unsigned long __pmd_frag_size_shift; #define MAX_PTRS_PER_PGD (1 << (H_PGD_INDEX_SIZE > RADIX_PGD_INDEX_SIZE ? \ H_PGD_INDEX_SIZE : RADIX_PGD_INDEX_SIZE)) +/* + * Compile-time upper bound on PMD_ORDER across hash and radix MMUs. + * Used by THP SWAP code. Check include/linux/swap.h + */ +#define ARCH_MAX_PMD_ORDER ((H_PTE_INDEX_SIZE > RADIX_PTE_INDEX_SIZE) ? \ + H_PTE_INDEX_SIZE : RADIX_PTE_INDEX_SIZE) + /* PMD_SHIFT determines what a second-level page table entry can map */ #define PMD_SHIFT (PAGE_SHIFT + PTE_INDEX_SIZE) #define PMD_SIZE (1UL << PMD_SHIFT) diff --git a/include/linux/swap.h b/include/linux/swap.h index 8f0f68e245ba..317168aa2db5 100644 --- a/include/linux/swap.h +++ b/include/linux/swap.h @@ -229,11 +229,21 @@ enum { */ #define SWAP_ENTRY_INVALID 0 +/* + * ARCH_MAX_PMD_ORDER is an optional arch hook: a compile-time upper bound for + * PMD_ORDER across all possible MMU configurations of that arch. It is used to + * size SWAP_NR_ORDERS on architectures (e.g. powerpc book3s64) where PMD_ORDER + * is selected at boot rather than at compile time. + */ #ifdef CONFIG_THP_SWAP +#ifdef ARCH_MAX_PMD_ORDER +#define SWAP_NR_ORDERS (ARCH_MAX_PMD_ORDER + 1) +#else #define SWAP_NR_ORDERS (PMD_ORDER + 1) +#endif /* ARCH_MAX_PMD_ORDER */ #else #define SWAP_NR_ORDERS 1 -#endif +#endif /* CONFIG_THP_SWAP */ /* * We keep using same cluster for rotational device so IO will be sequential. -- 2.39.5