From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from out-189.mta0.migadu.com (out-189.mta0.migadu.com [91.218.175.189]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id BBB953ACF10 for ; Fri, 24 Jul 2026 10:05:40 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=91.218.175.189 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784887542; cv=none; b=Y8lyCsn5pVU0GbbODJZSc9dKWUYvpqwpJZObX/SzMSWz4latz/m/uvAqFdTbXvAMihFd/v+i+lhQUdy69mEPPpVwhV5UqHC3AGBvwFxpBouExfLJZbKwUgbV5rDLpa+8DsaoW7Vgu1Yq0AR34LSeAuo8DA0VrV4lkBbEcrt2k1c= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784887542; c=relaxed/simple; bh=5qw767GZDoCKryw8qanjrdeeYZKue27c3HjatjzszH4=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=cSI+8g04ATN8i54lNE0sD6bu3RZHmN053UIPm8JGMr+ToZPysCTcfbCD2mgapgBomjzSgKfRXWgBm3QQ03d+PY8u5oWlCTScicaRNJL6ny7MQ50Dc3VGVOBhozwOgVROtaOe4pc4G9cK5wbKICCK0903fbPlcDAKfQDsb5oBqls= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=a2wTIYut; arc=none smtp.client-ip=91.218.175.189 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="a2wTIYut" Message-ID: <740589e1-0c1e-4457-8781-05589151c06f@linux.dev> DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux.dev; s=key1; t=1784887538; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=lXYzAMJrIyhnjgtgqeRInUg03/n0tGgIY8htWWcGIJA=; b=a2wTIYutEZ0AFU+lwq6wQDQOLLm/VBDKtpQHds8WfHvu2c3a5ZIAo0qmSR6aCEjg5MMuDk AripeeEuk01hjqotkbVkaEAbRURYxkZR9JAcgsf1XfCOQD1Oexesd4Y8Y/pUqQL1fur7cY CTIIud+yQ5soiCnxdTkswW8xmFatxxw= Date: Fri, 24 Jul 2026 11:05:34 +0100 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Subject: Re: [PATCH v5 02/11] mm: add PMD swap entry splitting support To: Dev Jain , Andrew Morton , david@kernel.org, chrisl@kernel.org, kasong@tencent.com, ljs@kernel.org, ziy@nvidia.com, linux-mm@kvack.org Cc: ying.huang@linux.alibaba.com, Baoquan He , willy@infradead.org, youngjun.park@lge.com, hannes@cmpxchg.org, riel@surriel.com, shakeel.butt@linux.dev, alex@ghiti.fr, kas@kernel.org, baohua@kernel.org, baolin.wang@linux.alibaba.com, Nico Pache , "Liam R. Howlett" , ryan.roberts@arm.com, Vlastimil Babka , lance.yang@linux.dev, linux-kernel@vger.kernel.org, nphamcs@gmail.com, shikemeng@huaweicloud.com, yosry@kernel.org, kernel-team@meta.com References: <20260722152043.2273289-1-usama.arif@linux.dev> <20260722152043.2273289-3-usama.arif@linux.dev> <7c8b0e38-f13e-449d-a2a7-de1032b831dd@arm.com> Content-Language: en-US X-Report-Abuse: Please report any abuse attempt to abuse@migadu.com and include these headers. From: Usama Arif In-Reply-To: <7c8b0e38-f13e-449d-a2a7-de1032b831dd@arm.com> Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 7bit X-Migadu-Flow: FLOW_OUT On 24/07/2026 07:52, Dev Jain wrote: > > > On 22/07/26 8:49 pm, Usama Arif wrote: >> Add a swap branch in __split_huge_pmd_locked() that splits a PMD swap >> entry into 512 PTE swap entries. Unlike migration splits, no folio >> reference is needed because swap entries point to swap slots, not > > Confusing wording. swap entries need reference of swap slots instead > of folios. I would drop the mention of migration here. I added about migration because its the other if part of the if else and was a good comparison. Will remove it. Thanks! > >> pages. Each PTE inherits the correct sub-slot offset and preserves >> soft_dirty, uffd_wp, and exclusive flags. >> >> The folio_remove_rmap_pmd() gate at the end must inspect old_pmd >> rather than *pmd: for a present THP split, *pmd has already been >> cleared by pmdp_invalidate(), and that invalidated bit pattern can >> decode as a plausible swap entry. >> >> This branch is reached from the explicit __split_huge_pmd() callers >> that hit a non-present PMD: partial-range mprotect / munmap, the >> wp_huge_pmd() PMD-COW fallback, and the swap-in / swapoff fallbacks >> added in later patches when the cached folio is no longer PMD-sized. >> page_vma_mapped_walk() does not iterate PMD swap entries, so >> try_to_unmap_one() and try_to_migrate_one() do not reach this branch >> and freeze=true cannot occur in this branch today. page and folio >> are therefore left uninitialized in the swap branch; a >> VM_WARN_ON_ONCE(freeze) catches any future caller that breaks this >> invariant before the freeze path dereferences page_to_pfn(page + i) >> or put_page(page). >> >> Signed-off-by: Usama Arif >> --- >> mm/huge_memory.c | 29 ++++++++++++++++++++++++++++- >> 1 file changed, 28 insertions(+), 1 deletion(-) >> >> diff --git a/mm/huge_memory.c b/mm/huge_memory.c >> index 04e8a6b55343..9819c0ae228a 100644 >> --- a/mm/huge_memory.c >> +++ b/mm/huge_memory.c >> @@ -3210,6 +3210,14 @@ static void __split_huge_pmd_locked(struct vm_area_struct *vma, pmd_t *pmd, >> folio_add_anon_rmap_ptes(folio, page, HPAGE_PMD_NR, >> vma, haddr, rmap_flags); >> } >> + } else if (pmd_is_swap_entry(*pmd)) { >> + VM_WARN_ON_ONCE(freeze); >> + /* Swap entries have no page for the migration freeze path. */ >> + freeze = false; >> + old_pmd = *pmd; >> + soft_dirty = pmd_swp_soft_dirty(old_pmd); >> + uffd_wp = pmd_swp_uffd(old_pmd); >> + anon_exclusive = pmd_swp_exclusive(old_pmd); >> } else { >> /* >> * Up to this point the pmd is present and huge and userland has >> @@ -3346,6 +3354,25 @@ static void __split_huge_pmd_locked(struct vm_area_struct *vma, pmd_t *pmd, >> VM_WARN_ON(!pte_none(ptep_get(pte + i))); >> set_pte_at(mm, addr, pte + i, entry); >> } >> + } else if (pmd_is_swap_entry(old_pmd)) { >> + softleaf_t sl_entry = softleaf_from_pmd(old_pmd); >> + pte_t swp_pte; >> + swp_entry_t sub_entry; >> + >> + for (i = 0, addr = haddr; i < HPAGE_PMD_NR; >> + i++, addr += PAGE_SIZE) { >> + sub_entry = swp_entry(swp_type(sl_entry), >> + swp_offset(sl_entry) + i); >> + swp_pte = swp_entry_to_pte(sub_entry); >> + if (soft_dirty) >> + swp_pte = pte_swp_mksoft_dirty(swp_pte); >> + if (uffd_wp) >> + swp_pte = pte_swp_mkuffd(swp_pte); >> + if (anon_exclusive) >> + swp_pte = pte_swp_mkexclusive(swp_pte); >> + VM_WARN_ON(!pte_none(ptep_get(pte + i))); >> + set_pte_at(mm, addr, pte + i, swp_pte); >> + } > > All these for loops could later benefit from the set_softleaf_ptes() helper: > https://lore.kernel.org/all/20260723070905.3422276-6-dev.jain@arm.com/ > > >> } else { >> pte_t entry; >> >> @@ -3373,7 +3400,7 @@ static void __split_huge_pmd_locked(struct vm_area_struct *vma, pmd_t *pmd, >> } >> pte_unmap(pte); >> >> - if (!pmd_is_migration_entry(*pmd)) >> + if (!pmd_is_migration_entry(old_pmd) && !pmd_is_swap_entry(old_pmd)) >> folio_remove_rmap_pmd(folio, page, vma); >> if (freeze) >> put_page(page); >