From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-dl1-f44.google.com (mail-dl1-f44.google.com [74.125.82.44]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id BDA93261B96 for ; Tue, 3 Feb 2026 22:07:28 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.82.44 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1770156450; cv=none; b=sr8PWj3zhhNT9akaa/wqYmJbDp6Z/43A4lJ83l8pDCsGfqE9tPw2ET7dvA0po52z/oK81v+fztjBXijFVvXZ3uSs7xy0OkTvzKt0SIePnYVbAv8opGc8nWpUOrjBTusUItWkHMGoE3Pj5Hk0/W+99/t1ihF0WwP3WgXydnziR+o= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1770156450; c=relaxed/simple; bh=pqlOg/P1HPlw8PWuJxtKDE9vC/y1WuJ+XArnQfg0tQU=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=bWUq/X8D4nsc8g0SVgzAuuyq2OpWf3PezmVBt73/mY7Ij4Rv6jlYlqB70h6TyJC3wAXG9TLIkT/Lp7Q9p8Vdai8b4EuRROuf01imU4Xwis7GHVTk0yhdhla4spFfMmmqFxLhgepx1BP4j6VldKENPftkV+x0wnXZXKbHhfH1vGY= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=YMuZtAbC; arc=none smtp.client-ip=74.125.82.44 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="YMuZtAbC" Received: by mail-dl1-f44.google.com with SMTP id a92af1059eb24-12336c0a8b6so351889c88.1 for ; Tue, 03 Feb 2026 14:07:28 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20230601; t=1770156448; x=1770761248; darn=vger.kernel.org; h=content-transfer-encoding:in-reply-to:from:references:cc:to :content-language:subject:user-agent:mime-version:date:message-id :from:to:cc:subject:date:message-id:reply-to; bh=AOxus6Q+k+UFCDyHcbF/F78h7JC9Gdvh77//dMnEmMI=; b=YMuZtAbC0UB3mdBhrxIC1SfCrlg+qAlKrFKjn30huYW/oUKSiZwK23N8yBQrCje8NZ yghG/zuQLgxJo/Nm+th6XODcpoEMGeI7kQ08PuPc4HzzlGmv/+JK5L6Iv+DQ5fiLVuta Ete5Cz9kchHXOSCwXgEuIoxSbIzI2+ajXDU60jrf3Do4kc23OHnShcjANkzK3+bMdxBJ srppygmqdrqxfp8UI0prMaqAKhy7qdhMBWeD+FPoUEO0aZKv89geGGyK42s4vOr7F1se moLLw/W+X79Lhg4+V7ISZLaKaGjnMTftZvOm2brmr6tw5x5sK57CbZDscpUryXJCPU5z oyTQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1770156448; x=1770761248; h=content-transfer-encoding:in-reply-to:from:references:cc:to :content-language:subject:user-agent:mime-version:date:message-id :x-gm-gg:x-gm-message-state:from:to:cc:subject:date:message-id :reply-to; bh=AOxus6Q+k+UFCDyHcbF/F78h7JC9Gdvh77//dMnEmMI=; b=IEDGvRoDOW9sEtFBB8JodK+uDEg52qc79SLzPH+hsJrZTVV61EIhfD0/nIwF2h+6jt qnQ40doia98I61mPwwsGF7qjfaf9oxwBr2kjnnhDlJDV/U/gGpbtJGqMbkMNQXRCtYbq WDLTgi2QElVSIqNiG7xUrLzkIZAdVQnFJg97kuzQz2ByQ7uWt6sM0uKc1ho0B12TifKS Qn44EJl8Ohtg1Or3tJ3mvlprFMQKzgzwCYA1eDc7Lf7f7g+Qx73Tiku1i+WC/WzGOqSJ lWF7JAbr+NJt32Oh1WhzPQMccHgcres0prYY0I2oWlcElVaW4uGzwa+yC1EI/bOVVS8J SjaQ== X-Forwarded-Encrypted: i=1; AJvYcCV9XNjLz96abge53juDjugZcEwiUykKHTFJ8ec66NIjSsWDqaryXL7VDCVglDoXOGoNfT+Eourts0Ozbk4=@vger.kernel.org X-Gm-Message-State: AOJu0YzWBbN6Bj+NE2D/85VzuE6+JkwEG6Yy7KKId2CezIf19kr1UTrG 1HV8b3by3S+2KqZN/YMKKCd3Rh5GKHwQ4+CpGP7/DXY1hItVb1Zz6BzE X-Gm-Gg: AZuq6aLIox4sAiMjAAgrpbvtqWkyg7ZHidWL5zTi25a2dBME9qUS02BWpmdk69aMiEA jz9YhK1+bhlu7tq9XALbeO6qqDDArnLBxGuhTJeiZQFsxQQ0BJheqcT8MfGv8HFCY5WDDICZyiK S6T4vmcCSFxFjpEoGxJl0A1W6/FC8xrClkn4YA5qGUP31cj6Gzd4bWTeQ4L4NMCDoHxcJDTfHsT jsVDsmGzc7vbnGSsIb8njXd59sY57mnLLMYUwjQvrTxDX6Z3kth0GT22mDtgeVvF9fUfKh3uJx8 B0Kds81VGwqkUa+yZBVb/B6xK+cwLZfgckYSBtNuPVJ8q8xVfr7Mc+E6MPZTUAMY1Ed7iz7wBSO Hhv+8KVR9uSbNBxBq9tIpQIQtfum+kYF3kGunJWaTQq9triyVMzZumx7NS13w0AnLOMLGYlu5xh K8s/Tx/KREghMvhgadCGibnzk1+0G+ZpdepQ64vsFVIXsyqMbfngmpam9K8ijSRn0= X-Received: by 2002:a05:7022:6b9b:b0:119:e569:fbb2 with SMTP id a92af1059eb24-126f47cfbe2mr536040c88.33.1770156447603; Tue, 03 Feb 2026 14:07:27 -0800 (PST) Received: from ?IPV6:2a03:83e0:1151:15:1cc5:26fe:6b00:bcef? ([2620:10d:c090:500::21b0]) by smtp.gmail.com with ESMTPSA id a92af1059eb24-126f503d0fbsm454159c88.13.2026.02.03.14.07.26 (version=TLS1_3 cipher=TLS_AES_128_GCM_SHA256 bits=128/128); Tue, 03 Feb 2026 14:07:27 -0800 (PST) Message-ID: <05d5918f-b61b-4091-b8c6-20eebfffc3c4@gmail.com> Date: Tue, 3 Feb 2026 14:07:25 -0800 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [RFC 01/12] mm: add PUD THP ptdesc and rmap support Content-Language: en-GB To: Zi Yan , Kiryl Shutsemau , lorenzo.stoakes@oracle.com Cc: Andrew Morton , David Hildenbrand , linux-mm@kvack.org, hannes@cmpxchg.org, riel@surriel.com, shakeel.butt@linux.dev, baohua@kernel.org, dev.jain@arm.com, baolin.wang@linux.alibaba.com, npache@redhat.com, Liam.Howlett@oracle.com, ryan.roberts@arm.com, vbabka@suse.cz, lance.yang@linux.dev, linux-kernel@vger.kernel.org, kernel-team@meta.com References: <20260202005451.774496-1-usamaarif642@gmail.com> <20260202005451.774496-2-usamaarif642@gmail.com> <63D23D5F-AF35-4199-B52E-DFFC16DFDF91@nvidia.com> From: Usama Arif In-Reply-To: <63D23D5F-AF35-4199-B52E-DFFC16DFDF91@nvidia.com> Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 7bit On 02/02/2026 08:01, Zi Yan wrote: > On 2 Feb 2026, at 5:44, Kiryl Shutsemau wrote: > >> On Sun, Feb 01, 2026 at 04:50:18PM -0800, Usama Arif wrote: >>> For page table management, PUD THPs need to pre-deposit page tables >>> that will be used when the huge page is later split. When a PUD THP >>> is allocated, we cannot know in advance when or why it might need to >>> be split (COW, partial unmap, reclaim), but we need page tables ready >>> for that eventuality. Similar to how PMD THPs deposit a single PTE >>> table, PUD THPs deposit a PMD table which itself contains deposited >>> PTE tables - a two-level deposit. This commit adds the deposit/withdraw >>> infrastructure and a new pud_huge_pmd field in ptdesc to store the >>> deposited PMD. >>> >>> The deposited PMD tables are stored as a singly-linked stack using only >>> page->lru.next as the link pointer. A doubly-linked list using the >>> standard list_head mechanism would cause memory corruption: list_del() >>> poisons both lru.next (offset 8) and lru.prev (offset 16), but lru.prev >>> overlaps with ptdesc->pmd_huge_pte at offset 16. Since deposited PMD >>> tables have their own deposited PTE tables stored in pmd_huge_pte, >>> poisoning lru.prev would corrupt the PTE table list and cause crashes >>> when withdrawing PTE tables during split. PMD THPs don't have this >>> problem because their deposited PTE tables don't have sub-deposits. >>> Using only lru.next avoids the overlap entirely. >>> >>> For reverse mapping, PUD THPs need the same rmap support that PMD THPs >>> have. The page_vma_mapped_walk() function is extended to recognize and >>> handle PUD-mapped folios during rmap traversal. A new TTU_SPLIT_HUGE_PUD >>> flag tells the unmap path to split PUD THPs before proceeding, since >>> there is no PUD-level migration entry format - the split converts the >>> single PUD mapping into individual PTE mappings that can be migrated >>> or swapped normally. >>> >>> Signed-off-by: Usama Arif >>> --- >>> include/linux/huge_mm.h | 5 +++ >>> include/linux/mm.h | 19 ++++++++ >>> include/linux/mm_types.h | 5 ++- >>> include/linux/pgtable.h | 8 ++++ >>> include/linux/rmap.h | 7 ++- >>> mm/huge_memory.c | 8 ++++ >>> mm/internal.h | 3 ++ >>> mm/page_vma_mapped.c | 35 +++++++++++++++ >>> mm/pgtable-generic.c | 83 ++++++++++++++++++++++++++++++++++ >>> mm/rmap.c | 96 +++++++++++++++++++++++++++++++++++++--- >>> 10 files changed, 260 insertions(+), 9 deletions(-) >>> > > > >>> diff --git a/mm/pgtable-generic.c b/mm/pgtable-generic.c >>> index d3aec7a9926ad..2047558ddcd79 100644 >>> --- a/mm/pgtable-generic.c >>> +++ b/mm/pgtable-generic.c >>> @@ -195,6 +195,89 @@ pgtable_t pgtable_trans_huge_withdraw(struct mm_struct *mm, pmd_t *pmdp) >>> } >>> #endif >>> >>> +#ifdef CONFIG_HAVE_ARCH_TRANSPARENT_HUGEPAGE_PUD >>> +/* >>> + * Deposit page tables for PUD THP. >>> + * Called with PUD lock held. Stores PMD tables in a singly-linked stack >>> + * via pud_huge_pmd, using only pmd_page->lru.next as the link pointer. >>> + * >>> + * IMPORTANT: We use only lru.next (offset 8) for linking, NOT the full >>> + * list_head. This is because lru.prev (offset 16) overlaps with >>> + * ptdesc->pmd_huge_pte, which stores the PMD table's deposited PTE tables. >>> + * Using list_del() would corrupt pmd_huge_pte with LIST_POISON2. >> >> This is ugly. >> >> Sounds like you want to use llist_node/head instead of list_head for this. >> >> You might able to avoid taking the lock in some cases. Note that >> pud_lockptr() is mm->page_table_lock as of now. > > I agree. I used llist_node/head in my implementation[1] and it works. > I have an illustration at[2] to show the concept. Feel free to reuse the code. > > > [1] https://lore.kernel.org/all/20200928193428.GB30994@casper.infradead.org/ > [2] https://normal.zone/blog/2021-01-04-linux-1gb-thp-2/#new-mechanism > > Best Regards, > Yan, Zi Ah I should have looked at your patches more! I started working by just using lru and was using list_add/list_del which was ofcourse corrupting the list and took me way more time than I would like to admit to debug what was going on! The diagrams in your 2nd link are really useful. I ended up drawing by hand those to debug the corruption issue. I will point to that link in the next series :) How about something like the below diff over this patch? (Not included the comment changes that I will make everywhere) diff --git a/include/linux/mm_types.h b/include/linux/mm_types.h index 26a38490ae2e1..3653e24ce97d7 100644 --- a/include/linux/mm_types.h +++ b/include/linux/mm_types.h @@ -99,6 +99,9 @@ struct page { struct list_head buddy_list; struct list_head pcp_list; struct llist_node pcp_llist; + + /* PMD pagetable deposit head */ + struct llist_node pgtable_deposit_head; }; struct address_space *mapping; union { diff --git a/mm/pgtable-generic.c b/mm/pgtable-generic.c index 2047558ddcd79..764f14d0afcbb 100644 --- a/mm/pgtable-generic.c +++ b/mm/pgtable-generic.c @@ -215,9 +215,7 @@ void pgtable_trans_huge_pud_deposit(struct mm_struct *mm, pud_t *pudp, assert_spin_locked(pud_lockptr(mm, pudp)); - /* Push onto stack using only lru.next as the link */ - pmd_page->lru.next = (struct list_head *)pud_huge_pmd(pudp); - pud_huge_pmd(pudp) = pmd_page; + llist_add(&pmd_page->pgtable_deposit_head, (struct llist_head *)&pud_huge_pmd(pudp)); } /* @@ -227,16 +225,16 @@ void pgtable_trans_huge_pud_deposit(struct mm_struct *mm, pud_t *pudp, */ pmd_t *pgtable_trans_huge_pud_withdraw(struct mm_struct *mm, pud_t *pudp) { + struct llist_node *node; pgtable_t pmd_page; assert_spin_locked(pud_lockptr(mm, pudp)); - pmd_page = pud_huge_pmd(pudp); - if (!pmd_page) + node = llist_del_first((struct llist_head *)&pud_huge_pmd(pudp)); + if (!node) return NULL; - /* Pop from stack - lru.next points to next PMD page (or NULL) */ - pud_huge_pmd(pudp) = (pgtable_t)pmd_page->lru.next; + pmd_page = llist_entry(node, struct page, pgtable_deposit_head); return page_address(pmd_page); } Also, Zi is it ok if I add your Co-developed by on this patch in future revisions? I didn't want to do that without your explicit approval.