mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Baolin Wang <baolin.wang@linux.alibaba.com>
To: Kiryl Shutsemau <kirill@shutemov.name>,
	Andrew Morton <akpm@linux-foundation.org>,
	David Hildenbrand <david@kernel.org>,
	Lorenzo Stoakes <ljs@kernel.org>, Zi Yan <ziy@nvidia.com>
Cc: "Kiryl Shutsemau (Meta)" <kas@kernel.org>,
	linux-mm@kvack.org, linux-kernel@vger.kernel.org,
	kernel-team@meta.com, "Liam R. Howlett" <liam@infradead.org>,
	Nico Pache <nico.pache@linux.dev>,
	Ryan Roberts <ryan.roberts@arm.com>, Dev Jain <dev.jain@arm.com>,
	Barry Song <baohua@kernel.org>, Lance Yang <lance.yang@linux.dev>,
	Usama Arif <usama.arif@linux.dev>,
	Vlastimil Babka <vbabka@kernel.org>, Jann Horn <jannh@google.com>
Subject: Re: [PATCH v3 08/12] mm/collapse: separate scanning a PTE table from collapsing it
Date: Fri, 18 Sep 2026 17:06:13 +0800	[thread overview]
Message-ID: <0230e861-8364-4302-a607-c34809a4a809@linux.alibaba.com> (raw)
In-Reply-To: <20260916093145.4022188-9-kirill@shutemov.name>



On 9/16/26 5:31 PM, Kiryl Shutsemau wrote:
> From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
> 
> A collapse is two jobs.  One reads a PTE table under mmap_lock and decides
> whether the range is worth collapsing.  The other allocates, isolates,
> copies and flushes, and wants the lock given up first.
> 
> collapse_single_pmd() did both, so the boundary between them was somewhere
> in the middle of a function.
> 
> Give each half its own function:
> 
>    - collapse_scan_pmd() scans one table and only reads.  The anonymous
>      scan that used to carry that name keeps its body as
>      collapse_scan_anon_pmd(), and collapse_scan_pmd() is now the entry
>      that picks the anonymous or the file side.
> 
>    - collapse_run_pmd() does the collapse the scan asked for, and is
>      handed what the scan returned.  SCAN_SUCCEED means there is
>      something to collapse.  SCAN_PTE_MAPPED_HUGEPAGE means the page
>      cache already holds the PMD folio and only the PTE table is left to
>      retract.  Both are work for the run; anything else is why there is
>      nothing to do.
> 
> collapse_single_pmd() is now the two of them with the mmap_lock drop in
> between, so its callers see what they saw before.
> 
> What the scan found and the run needs travels in collapse_control.  For
> an anonymous table that is the orders and the referenced and swapped-out
> counts, which mthp_collapse() and collapse_huge_page() now read from
> there instead of taking as arguments.  For a file it is the file itself
> and the offset in it.
> 
> The file side moves with the anonymous one.  collapse_scan_file() used to
> run with mmap_lock already given up, and called collapse_file() itself
> when the page cache looked worth it.  It now runs under the lock like the
> anonymous scan and only reads; the run does the collapse.  A file
> collapse works on the page cache and never sees a VMA, so the scan takes
> the file reference while it still has one and the run gives it back.
> 
> That changes what a refused file table costs khugepaged.  Every file
> table it scanned used to end its pass over that mm, because the lock had
> been dropped to scan it; now only a table it goes on to collapse does.
> 
> Two things on the file side stop being rescanned.  When the page cache
> already holds the PMD folio, the scan says so and the run goes straight
> to retracting the PTE table.  A run that refuses dirty pages and may
> write them back retries collapse_file() alone.  The checks the scan makes
> ahead of it are ones collapse_file() repeats under the page cache lock.
> 
> Tracing changes with it.  mm_khugepaged_scan_pmd and
> mm_khugepaged_scan_file used to fire after the collapse, so for an
> accepted table their status field carried what the collapse made of it.
> They now fire before it and read SCAN_SUCCEED for an accepted table.  What
> the collapse then made of it is for mm_collapse_huge_page and
> mm_khugepaged_collapse_file to report.
> 
> Assisted-by: LLM
> Reviewed-by: Zi Yan <ziy@nvidia.com>
> Signed-off-by: Kiryl Shutsemau (Meta) <kas@kernel.org>
> ---

LGTM.
Reviewed-by: Baolin Wang <baolin.wang@linux.alibaba.com>

  reply	other threads:[~2026-09-18  9:06 UTC|newest]

Thread overview: 18+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-16  9:31 [PATCH v3 00/12] mm/collapse: separate a collapse from its callers Kiryl Shutsemau
2026-09-16  9:31 ` [PATCH v3 01/12] mm/khugepaged: drop redundant mm_struct pin in madvise_collapse() Kiryl Shutsemau
2026-09-16  9:31 ` [PATCH v3 02/12] mm/khugepaged: count collapses where khugepaged makes them Kiryl Shutsemau
2026-09-16  9:31 ` [PATCH v3 03/12] mm/khugepaged: rename mthp_present_ptes bitmap to eligible_ptes Kiryl Shutsemau
2026-09-16  9:31 ` [PATCH v3 04/12] mm/collapse: add collapse.h for the collapse interface Kiryl Shutsemau
2026-09-16  9:31 ` [PATCH v3 05/12] mm/collapse: state what a collapse may do in the policy Kiryl Shutsemau
2026-09-16  9:31 ` [PATCH v3 06/12] mm/collapse: drop the collapse_possible() wrapper Kiryl Shutsemau
2026-09-16  9:31 ` [PATCH v3 07/12] mm/collapse: name the per-table scan reset for what it resets Kiryl Shutsemau
2026-09-16  9:31 ` [PATCH v3 08/12] mm/collapse: separate scanning a PTE table from collapsing it Kiryl Shutsemau
2026-09-18  9:06   ` Baolin Wang [this message]
2026-09-16  9:31 ` [PATCH v3 09/12] mm/collapse: open-code collapse_single_pmd() in its two callers Kiryl Shutsemau
2026-09-18  9:31   ` Baolin Wang
2026-09-16  9:31 ` [PATCH v3 10/12] mm/collapse: work out the orders a VMA allows once per VMA Kiryl Shutsemau
2026-09-16  9:31 ` [PATCH v3 11/12] mm/collapse: declare the collapse interface in collapse.h Kiryl Shutsemau
2026-09-16  9:31 ` [PATCH v3 12/12] mm/collapse: implement MADV_COLLAPSE in madvise.c Kiryl Shutsemau
2026-09-16 22:48 ` [PATCH v3 00/12] mm/collapse: separate a collapse from its callers Andrew Morton
2026-09-17 12:26   ` Kiryl Shutsemau
2026-09-17 22:11     ` Andrew Morton

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=0230e861-8364-4302-a607-c34809a4a809@linux.alibaba.com \
    --to=baolin.wang@linux.alibaba.com \
    --cc=akpm@linux-foundation.org \
    --cc=baohua@kernel.org \
    --cc=david@kernel.org \
    --cc=dev.jain@arm.com \
    --cc=jannh@google.com \
    --cc=kas@kernel.org \
    --cc=kernel-team@meta.com \
    --cc=kirill@shutemov.name \
    --cc=lance.yang@linux.dev \
    --cc=liam@infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=nico.pache@linux.dev \
    --cc=ryan.roberts@arm.com \
    --cc=usama.arif@linux.dev \
    --cc=vbabka@kernel.org \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®