From: Luiz Capitulino <luizcap@redhat.com>
To: Usama Arif <usama.arif@linux.dev>
Cc: linux-kernel@vger.kernel.org, linux-mm@kvack.org,
david@kernel.org, baolin.wang@linux.alibaba.com, ziy@nvidia.com,
lance.yang@linux.dev, corbet@lwn.net, tsbogend@alpha.franken.de,
maddy@linux.ibm.com, mpe@ellerman.id.au, agordeev@linux.ibm.com,
gerald.schaefer@linux.ibm.com, hca@linux.ibm.com,
gor@linux.ibm.com, x86@kernel.org, tglx@kernel.org,
mingo@redhat.com, bp@alien8.de, hughd@google.com,
dave.hansen@linux.intel.com, djbw@kernel.org,
vishal.l.verma@intel.com, dave.jiang@intel.com,
akpm@linux-foundation.org, yintirui@huawei.com, dev.jain@arm.com
Subject: Re: [PATCH v8 14/14] mm: thp: always enable mTHP support
Date: Mon, 21 Sep 2026 22:06:15 -0400 [thread overview]
Message-ID: <3772f64c-de5f-40b3-914d-48f5e0b0a72a@redhat.com> (raw)
In-Reply-To: <20260921102913.2970139-1-usama.arif@linux.dev>
On 9/21/26 6:29 AM, Usama Arif wrote:
> On Thu, 17 Sep 2026 21:45:35 -0400 Luiz Capitulino <luizcap@redhat.com> wrote:
>
>> If PMD-sized pages are not supported on an architecture (ie. the
>> arch implements arch_has_pmd_leaves() and it returns false) then the
>> current code disables all THP, including mTHP.
>>
>> This commit fixes this by allowing mTHP to be always enabled for all
>> archs. When PMD-sized pages are not supported, its sysfs entry won't be
>> created and their mapping will be disallowed at page-fault time.
>>
>> Similarly, this commit implements the following changes for shmem in
>> shmem_allowable_huge_orders():
>>
>> - Drop the pgtable_has_pmd_leaves() check so that mTHP sizes are
>> considered
>> - Filter out PMD and PUD orders from allowable orders when
>> PMD-sized pages are not supported by the CPU
>>
>> Signed-off-by: Luiz Capitulino <luizcap@redhat.com>
>> ---
>> mm/huge_memory.c | 25 ++++++++++++++++++++-----
>> mm/shmem.c | 14 +++++++++-----
>> 2 files changed, 29 insertions(+), 10 deletions(-)
>>
>> diff --git a/mm/huge_memory.c b/mm/huge_memory.c
>> index a06025b87e7c..a2d6de3ea988 100644
>> --- a/mm/huge_memory.c
>> +++ b/mm/huge_memory.c
>> @@ -189,6 +189,15 @@ unsigned long __thp_vma_allowable_orders(struct vm_area_struct *vma,
>> else
>> supported_orders = THP_ORDERS_ALL_FILE_DEFAULT;
>>
>> + if (!pgtable_has_pmd_leaves()) {
>> + /*
>> + * If the CPU does not support PMD leaves, assume for
>> + * now that it does not support PUD leaves and disable
>> + * both folio orders.
>> + */
>> + supported_orders &= ~(BIT(PMD_ORDER) | BIT(PUD_ORDER));
>> + }
>> +
>> orders &= supported_orders;
>> if (!orders)
>> return 0;
>> @@ -196,7 +205,7 @@ unsigned long __thp_vma_allowable_orders(struct vm_area_struct *vma,
>> if (!vma->vm_mm) /* vdso */
>> return 0;
>>
>> - if (!pgtable_has_pmd_leaves() || vma_thp_disabled(vma, vm_flags, forced_collapse))
>> + if (vma_thp_disabled(vma, vm_flags, forced_collapse))
>> return 0;
>>
>> /* khugepaged doesn't collapse DAX vma, but page fault is fine. */
>> @@ -979,7 +988,7 @@ static int __init hugepage_init_sysfs(struct kobject **hugepage_kobj)
>> * disable all other sizes. powerpc's PMD_ORDER isn't a compile-time
>> * constant so we have to do this here.
>> */
>> - if (!anon_orders_configured)
>> + if (!anon_orders_configured && pgtable_has_pmd_leaves())
>> huge_anon_orders_inherit = BIT(PMD_ORDER);
>>
>> *hugepage_kobj = kobject_create_and_add("transparent_hugepage", mm_kobj);
>> @@ -1001,6 +1010,15 @@ static int __init hugepage_init_sysfs(struct kobject **hugepage_kobj)
>> }
>>
>> orders = THP_ORDERS_ALL_ANON | THP_ORDERS_ALL_FILE_DEFAULT;
>> + if (!pgtable_has_pmd_leaves()) {
>> + /*
>> + * If the CPU does not support PMD leaves, assume for
>> + * now that it does not support PUD leaves and disable
>> + * both folio orders.
>> + */
>> + orders &= ~(BIT(PMD_ORDER) | BIT(PUD_ORDER));
>> + }
>> +
>> order = highest_order(orders);
>> while (orders) {
>> thpsize = thpsize_create(order, *hugepage_kobj);
>> @@ -1091,9 +1109,6 @@ static int __init hugepage_init(void)
>> int err;
>> struct kobject *hugepage_kobj;
>>
>> - if (!pgtable_has_pmd_leaves())
>> - return -EINVAL;
>> -
>
> Removing this guard lets start_stop_khugepaged() run on a system
> without PMD leaves. All mTHP orders default to never, but the global always
> or madvise flag still makes hugepage_enabled() return true.
>
> That starts an idle khugepaged thread and can unnecessarily raise
> min_free_kbytes. Could hugepage_enabled() instead test whether
> an enabled order remains after masking PMD_ORDER when
> pgtable_has_pmd_leaves() is false?
You're right about the issue. Before this series, checking the global
configuration was sufficient to determine if THP was "enabled" as PMD
leaves would always be available (otherwise THP would be shut down).
With this series, we also need to check pgtable_has_pmd_leaves().
I prefer the simple fix below over masking orders in hugepage_enabled():
it's simple and adds the check only in the call site that requires it.
diff --git a/mm/khugepaged.c b/mm/khugepaged.c
index a3a9e4d93b46..175252a1d729 100644
--- a/mm/khugepaged.c
+++ b/mm/khugepaged.c
@@ -453,7 +453,7 @@ static bool hugepage_enabled(void)
* Shmem pmd-sized hugepages are also determined by its pmd-size control,
* except when the global shmem_huge is set to SHMEM_HUGE_DENY.
*/
- if (hugepage_global_enabled())
+ if (hugepage_global_enabled() && pgtable_has_pmd_leaves())
return true;
if (anon_hpage_enabled())
return true;
>
>> /*
>> * hugepages can't be allocated by the buddy allocator
>> */
>> diff --git a/mm/shmem.c b/mm/shmem.c
>> index bc2de3a7c1ea..8c0f7e3efeeb 100644
>> --- a/mm/shmem.c
>> +++ b/mm/shmem.c
>> @@ -2046,11 +2046,14 @@ unsigned long shmem_allowable_huge_orders(struct inode *inode,
>> unsigned long mask = READ_ONCE(huge_shmem_orders_always);
>> unsigned long within_size_orders = READ_ONCE(huge_shmem_orders_within_size);
>> vm_flags_t vm_flags = vma ? vma->vm_flags : 0;
>> - unsigned int global_orders;
>> + unsigned int global_orders, disabled_orders = 0;
>>
>> - if (!pgtable_has_pmd_leaves() || (vma && vma_thp_disabled(vma, vm_flags, shmem_huge_force)))
>> + if (vma && vma_thp_disabled(vma, vm_flags, shmem_huge_force))
>> return 0;
>>
>> + if (!pgtable_has_pmd_leaves())
>> + disabled_orders = BIT(PMD_ORDER);
>> +
>> global_orders = shmem_huge_global_enabled(inode, index, write_end,
>> shmem_huge_force, vma, vm_flags);
>> /*
>> @@ -2058,7 +2061,7 @@ unsigned long shmem_allowable_huge_orders(struct inode *inode,
>> * sysfs configs.
>> */
>> if (!vma || !vma_is_anon_shmem(vma) || shmem_huge_force)
>> - return global_orders;
>> + return global_orders & ~disabled_orders;
>>
>> /*
>> * Following the 'deny' semantics of the top level, force the huge
>> @@ -2072,7 +2075,7 @@ unsigned long shmem_allowable_huge_orders(struct inode *inode,
>> * means non-PMD sized THP can not override 'huge' mount option now.
>> */
>> if (shmem_huge == SHMEM_HUGE_FORCE)
>> - return READ_ONCE(huge_shmem_orders_inherit);
>> + return READ_ONCE(huge_shmem_orders_inherit) & ~disabled_orders;
>>
>> /* Allow mTHP that will be fully within i_size. */
>> mask |= shmem_get_orders_within_size(inode, within_size_orders, index, 0);
>> @@ -2083,6 +2086,7 @@ unsigned long shmem_allowable_huge_orders(struct inode *inode,
>> if (global_orders > 0)
>> mask |= READ_ONCE(huge_shmem_orders_inherit);
>>
>> + mask &= ~disabled_orders;
>> return THP_ORDERS_ALL_FILE_DEFAULT & mask;
>> }
>>
>> @@ -5630,7 +5634,7 @@ void __init shmem_init(void)
>> * Default to setting PMD-sized THP to inherit the global setting and
>> * disable all other multi-size THPs.
>> */
>> - if (!shmem_orders_configured)
>> + if (!shmem_orders_configured && pgtable_has_pmd_leaves())
>> huge_shmem_orders_inherit = BIT(HPAGE_PMD_ORDER);
>> #endif
>> return;
>> --
>> 2.55.0
>>
>>
>
next prev parent reply other threads:[~2026-09-22 2:06 UTC|newest]
Thread overview: 29+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-18 1:45 [PATCH v8 00/14] " Luiz Capitulino
2026-09-18 1:45 ` [PATCH v8 01/14] docs: tmpfs: remove implementation detail reference Luiz Capitulino
2026-09-18 1:45 ` [PATCH v8 02/14] mm: shmem: shmem_getattr(): set blksize to highest supported THP order Luiz Capitulino
2026-09-18 1:45 ` [PATCH v8 03/14] mm: introduce pgtable_has_pmd_leaves() Luiz Capitulino
2026-09-22 2:07 ` Zi Yan
2026-09-18 1:45 ` [PATCH v8 04/14] drivers: dax: use pgtable_has_pmd_leaves() Luiz Capitulino
2026-09-22 2:08 ` Zi Yan
2026-09-18 1:45 ` [PATCH v8 05/14] drivers: nvdimm: " Luiz Capitulino
2026-09-18 1:45 ` [PATCH v8 06/14] mm: debug_vm_pgtable: " Luiz Capitulino
2026-09-18 1:45 ` [PATCH v8 07/14] mm: shmem: allow THP support determination at folio allocation time Luiz Capitulino
2026-09-22 2:20 ` Zi Yan
2026-09-23 1:37 ` Luiz Capitulino
2026-09-18 1:45 ` [PATCH v8 08/14] s390: move has_transparent_hugepage() out of THP guard Luiz Capitulino
2026-09-22 2:21 ` Zi Yan
2026-09-18 1:45 ` [PATCH v8 09/14] powerpc: " Luiz Capitulino
2026-09-18 1:45 ` [PATCH v8 10/14] mips: " Luiz Capitulino
2026-09-18 1:45 ` [PATCH v8 11/14] x86: " Luiz Capitulino
2026-09-18 1:45 ` [PATCH v8 12/14] treewide: introduce arch_has_pmd_leaves() Luiz Capitulino
2026-09-22 2:26 ` Zi Yan
2026-09-18 1:45 ` [PATCH v8 13/14] mm: replace thp_disabled_by_hw() with pgtable_has_pmd_leaves() Luiz Capitulino
2026-09-18 1:45 ` [PATCH v8 14/14] mm: thp: always enable mTHP support Luiz Capitulino
2026-09-18 9:01 ` Baolin Wang
2026-09-21 10:29 ` Usama Arif
2026-09-22 2:06 ` Luiz Capitulino [this message]
2026-09-18 10:37 ` [PATCH v8 00/14] " Usama Arif
2026-09-18 14:01 ` Luiz Capitulino
2026-09-18 20:23 ` David Hildenbrand (Arm)
2026-09-21 10:36 ` Usama Arif
2026-09-22 2:11 ` Luiz Capitulino
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=3772f64c-de5f-40b3-914d-48f5e0b0a72a@redhat.com \
--to=luizcap@redhat.com \
--cc=agordeev@linux.ibm.com \
--cc=akpm@linux-foundation.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=bp@alien8.de \
--cc=corbet@lwn.net \
--cc=dave.hansen@linux.intel.com \
--cc=dave.jiang@intel.com \
--cc=david@kernel.org \
--cc=dev.jain@arm.com \
--cc=djbw@kernel.org \
--cc=gerald.schaefer@linux.ibm.com \
--cc=gor@linux.ibm.com \
--cc=hca@linux.ibm.com \
--cc=hughd@google.com \
--cc=lance.yang@linux.dev \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=maddy@linux.ibm.com \
--cc=mingo@redhat.com \
--cc=mpe@ellerman.id.au \
--cc=tglx@kernel.org \
--cc=tsbogend@alpha.franken.de \
--cc=usama.arif@linux.dev \
--cc=vishal.l.verma@intel.com \
--cc=x86@kernel.org \
--cc=yintirui@huawei.com \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®