mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: "Zi Yan" <ziy@nvidia.com>
To: "Kiryl Shutsemau" <kirill@shutemov.name>,
	"Andrew Morton" <akpm@linux-foundation.org>,
	"David Hildenbrand" <david@kernel.org>,
	"Lorenzo Stoakes" <ljs@kernel.org>,
	"Baolin Wang" <baolin.wang@linux.alibaba.com>
Cc: "Kiryl Shutsemau (Meta)" <kas@kernel.org>, <linux-mm@kvack.org>,
	<linux-kernel@vger.kernel.org>, <kernel-team@meta.com>,
	"Liam R . Howlett" <liam@infradead.org>,
	"Nico Pache" <nico.pache@linux.dev>,
	"Ryan Roberts" <ryan.roberts@arm.com>,
	"Dev Jain" <dev.jain@arm.com>, "Barry Song" <baohua@kernel.org>,
	"Lance Yang" <lance.yang@linux.dev>,
	"Usama Arif" <usama.arif@linux.dev>,
	"Vlastimil Babka" <vbabka@kernel.org>,
	"Jann Horn" <jannh@google.com>
Subject: Re: [PATCH v2 08/12] mm/collapse: separate scanning a PTE table from collapsing it
Date: Thu, 10 Sep 2026 22:38:13 -0400	[thread overview]
Message-ID: <DLC4ZVU8JXFG.3SS0W4STZKWBA@nvidia.com> (raw)
In-Reply-To: <20260910120238.2529819-9-kirill@shutemov.name>

On Thu Sep 10, 2026 at 8:02 AM EDT, Kiryl Shutsemau wrote:
> From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
>
> A collapse is two jobs.  One reads a PTE table under mmap_lock and decides
> whether the range is worth collapsing.  The other allocates, isolates,
> copies and flushes, and wants the lock given up first.
>
> collapse_single_pmd() did both, so the boundary between them was somewhere
> in the middle of a function.
>
> Give each half its own function:
>
>   - collapse_scan_pmd() scans one table and only reads.  The anonymous
>     scan that used to carry that name keeps its body as
>     collapse_scan_anon_pmd(), and collapse_scan_pmd() is now the entry
>     that picks the anonymous or the file side.
>
>   - collapse_run_pmd() does the collapse the scan asked for.
>     SCAN_SUCCEED from the scan means there is something to run; anything
>     else is why there is not.
>
> collapse_single_pmd() is now the two of them with the mmap_lock drop in
> between, so its callers see what they saw before.
>
> What the scan found and the run needs travels in collapse_control.  For
> an anonymous table that is the orders and the referenced and swapped-out
> counts.  For a file it is the file itself, the offset in it, and whether
> the PMD folio is already in the page cache.
>
> The file side moves with the anonymous one.  collapse_scan_file() used to
> run with mmap_lock already given up, and called collapse_file() itself
> when the page cache looked worth it.  It now runs under the lock like the
> anonymous scan and only judges; the run does the collapse.  A file
> collapse works on the page cache and never sees a VMA, so the scan takes
> the file reference while it still has one and the run gives it back.
>
> That changes what a refused file table costs khugepaged.  Every file
> table it scanned used to end its pass over that mm, because the lock had
> been dropped to scan it; now only a table it goes on to collapse does.
>
> Two things on the file side stop being rescanned.  When the page cache
> already holds the PMD folio, the scan says so and the run goes straight
> to retracting the PTE table.  A run that refuses dirty pages and may
> write them back retries collapse_file() alone.  The checks the scan makes
> ahead of it are ones collapse_file() repeats under the page cache lock.
>
> Tracing changes with it.  mm_khugepaged_scan_pmd and
> mm_khugepaged_scan_file used to fire after the collapse, so for an
> accepted table their status field carried what the collapse made of it.
> They now fire before it and read SCAN_SUCCEED for an accepted table.  What
> the collapse then made of it is for mm_collapse_huge_page and
> mm_khugepaged_collapse_file to report.
>
> Assisted-by: LLM
> Signed-off-by: Kiryl Shutsemau (Meta) <kas@kernel.org>
> ---
>  mm/collapse.h   |  16 ++++++
>  mm/khugepaged.c | 147 ++++++++++++++++++++++++++++++++++++------------
>  2 files changed, 128 insertions(+), 35 deletions(-)
>
> diff --git a/mm/collapse.h b/mm/collapse.h
> index 7044dc71c7c2..346859a2184f 100644
> --- a/mm/collapse.h
> +++ b/mm/collapse.h
> @@ -88,6 +88,22 @@ struct collapse_control {
>  
>  	/* Each bit marks a PTE the scan accepted as a collapse source */
>  	DECLARE_BITMAP(eligible_ptes, MAX_PTRS_PER_PTE);
> +
> +	/*
> +	 * What a scan found and the run after it needs.  Live only between the
> +	 * two, and read by nobody else.
> +	 *
> +	 * The file side takes a reference while it still has the VMA, since a
> +	 * file collapse works on the page cache and never sees one; the run is
> +	 * what gives it back.  A scan that found the PMD folio already in the
> +	 * cache leaves only the PTE table to retract.
> +	 */
> +	unsigned long scan_orders;
> +	int scan_referenced;
> +	int scan_unmapped;
> +	struct file *scan_file;
> +	pgoff_t scan_pgoff;
> +	bool scan_retract_only;

scan_retract_pte_only ?

>  };
>  

<snip>
>  
> -/*
> - * Try to collapse a single PMD starting at a PMD aligned addr, and return
> - * the results.
> - */
> -static enum scan_result collapse_single_pmd(unsigned long addr,
> -		struct vm_area_struct *vma, bool *lock_dropped,
> -		struct collapse_control *cc)
> +static void collapse_control_init(struct collapse_control *cc)
> +{
> +	cc->progress = 0;
> +	cc->scan_file = NULL;
> +}
> +
> +static void collapse_control_release(struct collapse_control *cc)
> +{
> +	/* A scan that took a file reference should have been run */
> +	if (WARN_ON_ONCE(cc->scan_file)) {
> +		fput(cc->scan_file);
> +		cc->scan_file = NULL;
> +	}
> +}
> +
> +static enum scan_result collapse_scan_pmd(struct vm_area_struct *vma,
> +		unsigned long addr, struct collapse_control *cc)
>  {
> -	struct mm_struct *mm = vma->vm_mm;
> -	bool triggered_wb = false;
>  	enum scan_result result;
> -	struct file *file;
>  	pgoff_t pgoff;
>  
> -	mmap_assert_locked(mm);
> +	mmap_assert_locked(vma->vm_mm);
> +	/* Whatever the last scan found has to have been run by now */
> +	if (WARN_ON_ONCE(cc->scan_file)) {
> +		fput(cc->scan_file);
> +		cc->scan_file = NULL;
> +	}

scan_file should be set to NULL by collapse_control_init(). Anyway, the
code is duplicated here and in collapse_control_release(), maybe add a
helper.

>  
>  	if (vma_is_anonymous(vma))
> -		return collapse_scan_pmd(mm, vma, addr, lock_dropped, cc);
> +		return collapse_scan_anon_pmd(vma, addr, cc);
>  
> -	file = get_file(vma->vm_file);
>  	pgoff = linear_page_index(vma, addr);
> +	result = collapse_scan_file(vma->vm_mm, addr, vma->vm_file, pgoff, cc);
> +	switch (result) {
> +	case SCAN_SUCCEED:
> +		cc->scan_retract_only = false;
> +		break;
> +	case SCAN_PTE_MAPPED_HUGEPAGE:
> +		/*
> +		 * The page cache already holds the PMD folio; what is left is
> +		 * to retract the PTE table, which is the run's job.
> +		 */
> +		cc->scan_retract_only = true;
> +		result = SCAN_SUCCEED;

<snip>

> +
> +	if (cc->scan_retract_only) {
> +		result = SCAN_PTE_MAPPED_HUGEPAGE;
> +		goto retract;
> +	}

<snip>

> +retract:
>  	fput(file);
>  
> +	/*
> +	 * A PMD folio is in the page cache, whether the collapse just put it
> +	 * there or found it: retract the PTE table, and map the PMD if asked.
> +	 */
>  	if (result == SCAN_PTE_MAPPED_HUGEPAGE) {
>  		mmap_read_lock(mm);
>  		if (collapse_test_exit_or_disable(mm))

result is changed from SCAN_PTE_MAPPED_HUGEPAGE to SCAN_SUCCEED to
SCAN_PTE_MAPPED_HUGEPAGE to get here. Is there a way of avoiding this
result churn?


> @@ -2805,6 +2857,28 @@ static enum scan_result collapse_single_pmd(unsigned long addr,
>  	return result;
>  }
>  
> +/*
> + * Try to collapse a single PMD starting at a PMD aligned addr, and return
> + * the results.
> + */
> +static enum scan_result collapse_single_pmd(unsigned long addr,
> +		struct vm_area_struct *vma, bool *lock_dropped,
> +		struct collapse_control *cc)
> +{
> +	struct mm_struct *mm = vma->vm_mm;
> +	enum scan_result result;
> +
> +	result = collapse_scan_pmd(vma, addr, cc);
> +	if (result != SCAN_SUCCEED)
> +		return result;

Can it be changed to?

if (result != SCAN_SUCCEED && result != SCAN_PTE_MAPPED_HUGEPAGE)
	return result;


-- 
Best Regards,
Yan, Zi


  reply	other threads:[~2026-09-11  2:38 UTC|newest]

Thread overview: 27+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-10 12:02 [PATCH v2 00/12] mm/collapse: separate a collapse from its callers Kiryl Shutsemau
2026-09-10 12:02 ` [PATCH v2 01/12] mm/khugepaged: drop redundant mm_struct pin in madvise_collapse() Kiryl Shutsemau
2026-09-10 12:02 ` [PATCH v2 02/12] mm/khugepaged: count collapses where khugepaged makes them Kiryl Shutsemau
2026-09-10 12:02 ` [PATCH v2 03/12] mm/khugepaged: rename mthp_present_ptes bitmap to eligible_ptes Kiryl Shutsemau
2026-09-10 12:02 ` [PATCH v2 04/12] mm/collapse: add collapse.h for the collapse interface Kiryl Shutsemau
2026-09-10 12:02 ` [PATCH v2 05/12] mm/collapse: state what a collapse may do in the policy Kiryl Shutsemau
2026-09-11  2:06   ` Zi Yan
2026-09-10 12:02 ` [PATCH v2 06/12] mm/collapse: drop the collapse_possible() wrapper Kiryl Shutsemau
2026-09-10 12:02 ` [PATCH v2 07/12] mm/collapse: name the per-table scan reset for what it resets Kiryl Shutsemau
2026-09-10 12:02 ` [PATCH v2 08/12] mm/collapse: separate scanning a PTE table from collapsing it Kiryl Shutsemau
2026-09-11  2:38   ` Zi Yan [this message]
2026-09-11 13:37     ` Kiryl Shutsemau
2026-09-11 14:40       ` Zi Yan
2026-09-10 12:02 ` [PATCH v2 09/12] mm/collapse: open-code collapse_single_pmd() in its two callers Kiryl Shutsemau
2026-09-11 14:57   ` Zi Yan
2026-09-11 15:24     ` Kiryl Shutsemau
2026-09-11 15:26       ` Zi Yan
2026-09-11 22:09   ` Zi Yan
2026-09-10 12:02 ` [PATCH v2 10/12] mm/collapse: work out the orders a VMA allows once per VMA Kiryl Shutsemau
2026-09-11 15:56   ` Zi Yan
2026-09-10 12:02 ` [PATCH v2 11/12] mm/collapse: declare the collapse interface in collapse.h Kiryl Shutsemau
2026-09-11 19:02   ` Zi Yan
2026-09-10 12:02 ` [PATCH v2 12/12] mm/collapse: implement MADV_COLLAPSE in madvise.c Kiryl Shutsemau
2026-09-11 15:06 ` [PATCH v2 00/12] mm/collapse: separate a collapse from its callers David Hildenbrand (Arm)
2026-09-11 15:56   ` Kiryl Shutsemau
2026-09-11 15:58     ` Kiryl Shutsemau
2026-09-11 18:35     ` David Hildenbrand (Arm)

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=DLC4ZVU8JXFG.3SS0W4STZKWBA@nvidia.com \
    --to=ziy@nvidia.com \
    --cc=akpm@linux-foundation.org \
    --cc=baohua@kernel.org \
    --cc=baolin.wang@linux.alibaba.com \
    --cc=david@kernel.org \
    --cc=dev.jain@arm.com \
    --cc=jannh@google.com \
    --cc=kas@kernel.org \
    --cc=kernel-team@meta.com \
    --cc=kirill@shutemov.name \
    --cc=lance.yang@linux.dev \
    --cc=liam@infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=nico.pache@linux.dev \
    --cc=ryan.roberts@arm.com \
    --cc=usama.arif@linux.dev \
    --cc=vbabka@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®