From: "Zi Yan" <ziy@nvidia.com>
To: "Kiryl Shutsemau" <kirill@shutemov.name>,
"Andrew Morton" <akpm@linux-foundation.org>,
"David Hildenbrand" <david@kernel.org>,
"Lorenzo Stoakes" <ljs@kernel.org>,
"Baolin Wang" <baolin.wang@linux.alibaba.com>
Cc: "Kiryl Shutsemau (Meta)" <kas@kernel.org>, <linux-mm@kvack.org>,
<linux-kernel@vger.kernel.org>, <kernel-team@meta.com>,
"Liam R . Howlett" <liam@infradead.org>,
"Nico Pache" <nico.pache@linux.dev>,
"Ryan Roberts" <ryan.roberts@arm.com>,
"Dev Jain" <dev.jain@arm.com>, "Barry Song" <baohua@kernel.org>,
"Lance Yang" <lance.yang@linux.dev>,
"Usama Arif" <usama.arif@linux.dev>,
"Vlastimil Babka" <vbabka@kernel.org>,
"Jann Horn" <jannh@google.com>
Subject: Re: [PATCH v2 08/12] mm/collapse: separate scanning a PTE table from collapsing it
Date: Thu, 10 Sep 2026 22:38:13 -0400 [thread overview]
Message-ID: <DLC4ZVU8JXFG.3SS0W4STZKWBA@nvidia.com> (raw)
In-Reply-To: <20260910120238.2529819-9-kirill@shutemov.name>
On Thu Sep 10, 2026 at 8:02 AM EDT, Kiryl Shutsemau wrote:
> From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
>
> A collapse is two jobs. One reads a PTE table under mmap_lock and decides
> whether the range is worth collapsing. The other allocates, isolates,
> copies and flushes, and wants the lock given up first.
>
> collapse_single_pmd() did both, so the boundary between them was somewhere
> in the middle of a function.
>
> Give each half its own function:
>
> - collapse_scan_pmd() scans one table and only reads. The anonymous
> scan that used to carry that name keeps its body as
> collapse_scan_anon_pmd(), and collapse_scan_pmd() is now the entry
> that picks the anonymous or the file side.
>
> - collapse_run_pmd() does the collapse the scan asked for.
> SCAN_SUCCEED from the scan means there is something to run; anything
> else is why there is not.
>
> collapse_single_pmd() is now the two of them with the mmap_lock drop in
> between, so its callers see what they saw before.
>
> What the scan found and the run needs travels in collapse_control. For
> an anonymous table that is the orders and the referenced and swapped-out
> counts. For a file it is the file itself, the offset in it, and whether
> the PMD folio is already in the page cache.
>
> The file side moves with the anonymous one. collapse_scan_file() used to
> run with mmap_lock already given up, and called collapse_file() itself
> when the page cache looked worth it. It now runs under the lock like the
> anonymous scan and only judges; the run does the collapse. A file
> collapse works on the page cache and never sees a VMA, so the scan takes
> the file reference while it still has one and the run gives it back.
>
> That changes what a refused file table costs khugepaged. Every file
> table it scanned used to end its pass over that mm, because the lock had
> been dropped to scan it; now only a table it goes on to collapse does.
>
> Two things on the file side stop being rescanned. When the page cache
> already holds the PMD folio, the scan says so and the run goes straight
> to retracting the PTE table. A run that refuses dirty pages and may
> write them back retries collapse_file() alone. The checks the scan makes
> ahead of it are ones collapse_file() repeats under the page cache lock.
>
> Tracing changes with it. mm_khugepaged_scan_pmd and
> mm_khugepaged_scan_file used to fire after the collapse, so for an
> accepted table their status field carried what the collapse made of it.
> They now fire before it and read SCAN_SUCCEED for an accepted table. What
> the collapse then made of it is for mm_collapse_huge_page and
> mm_khugepaged_collapse_file to report.
>
> Assisted-by: LLM
> Signed-off-by: Kiryl Shutsemau (Meta) <kas@kernel.org>
> ---
> mm/collapse.h | 16 ++++++
> mm/khugepaged.c | 147 ++++++++++++++++++++++++++++++++++++------------
> 2 files changed, 128 insertions(+), 35 deletions(-)
>
> diff --git a/mm/collapse.h b/mm/collapse.h
> index 7044dc71c7c2..346859a2184f 100644
> --- a/mm/collapse.h
> +++ b/mm/collapse.h
> @@ -88,6 +88,22 @@ struct collapse_control {
>
> /* Each bit marks a PTE the scan accepted as a collapse source */
> DECLARE_BITMAP(eligible_ptes, MAX_PTRS_PER_PTE);
> +
> + /*
> + * What a scan found and the run after it needs. Live only between the
> + * two, and read by nobody else.
> + *
> + * The file side takes a reference while it still has the VMA, since a
> + * file collapse works on the page cache and never sees one; the run is
> + * what gives it back. A scan that found the PMD folio already in the
> + * cache leaves only the PTE table to retract.
> + */
> + unsigned long scan_orders;
> + int scan_referenced;
> + int scan_unmapped;
> + struct file *scan_file;
> + pgoff_t scan_pgoff;
> + bool scan_retract_only;
scan_retract_pte_only ?
> };
>
<snip>
>
> -/*
> - * Try to collapse a single PMD starting at a PMD aligned addr, and return
> - * the results.
> - */
> -static enum scan_result collapse_single_pmd(unsigned long addr,
> - struct vm_area_struct *vma, bool *lock_dropped,
> - struct collapse_control *cc)
> +static void collapse_control_init(struct collapse_control *cc)
> +{
> + cc->progress = 0;
> + cc->scan_file = NULL;
> +}
> +
> +static void collapse_control_release(struct collapse_control *cc)
> +{
> + /* A scan that took a file reference should have been run */
> + if (WARN_ON_ONCE(cc->scan_file)) {
> + fput(cc->scan_file);
> + cc->scan_file = NULL;
> + }
> +}
> +
> +static enum scan_result collapse_scan_pmd(struct vm_area_struct *vma,
> + unsigned long addr, struct collapse_control *cc)
> {
> - struct mm_struct *mm = vma->vm_mm;
> - bool triggered_wb = false;
> enum scan_result result;
> - struct file *file;
> pgoff_t pgoff;
>
> - mmap_assert_locked(mm);
> + mmap_assert_locked(vma->vm_mm);
> + /* Whatever the last scan found has to have been run by now */
> + if (WARN_ON_ONCE(cc->scan_file)) {
> + fput(cc->scan_file);
> + cc->scan_file = NULL;
> + }
scan_file should be set to NULL by collapse_control_init(). Anyway, the
code is duplicated here and in collapse_control_release(), maybe add a
helper.
>
> if (vma_is_anonymous(vma))
> - return collapse_scan_pmd(mm, vma, addr, lock_dropped, cc);
> + return collapse_scan_anon_pmd(vma, addr, cc);
>
> - file = get_file(vma->vm_file);
> pgoff = linear_page_index(vma, addr);
> + result = collapse_scan_file(vma->vm_mm, addr, vma->vm_file, pgoff, cc);
> + switch (result) {
> + case SCAN_SUCCEED:
> + cc->scan_retract_only = false;
> + break;
> + case SCAN_PTE_MAPPED_HUGEPAGE:
> + /*
> + * The page cache already holds the PMD folio; what is left is
> + * to retract the PTE table, which is the run's job.
> + */
> + cc->scan_retract_only = true;
> + result = SCAN_SUCCEED;
<snip>
> +
> + if (cc->scan_retract_only) {
> + result = SCAN_PTE_MAPPED_HUGEPAGE;
> + goto retract;
> + }
<snip>
> +retract:
> fput(file);
>
> + /*
> + * A PMD folio is in the page cache, whether the collapse just put it
> + * there or found it: retract the PTE table, and map the PMD if asked.
> + */
> if (result == SCAN_PTE_MAPPED_HUGEPAGE) {
> mmap_read_lock(mm);
> if (collapse_test_exit_or_disable(mm))
result is changed from SCAN_PTE_MAPPED_HUGEPAGE to SCAN_SUCCEED to
SCAN_PTE_MAPPED_HUGEPAGE to get here. Is there a way of avoiding this
result churn?
> @@ -2805,6 +2857,28 @@ static enum scan_result collapse_single_pmd(unsigned long addr,
> return result;
> }
>
> +/*
> + * Try to collapse a single PMD starting at a PMD aligned addr, and return
> + * the results.
> + */
> +static enum scan_result collapse_single_pmd(unsigned long addr,
> + struct vm_area_struct *vma, bool *lock_dropped,
> + struct collapse_control *cc)
> +{
> + struct mm_struct *mm = vma->vm_mm;
> + enum scan_result result;
> +
> + result = collapse_scan_pmd(vma, addr, cc);
> + if (result != SCAN_SUCCEED)
> + return result;
Can it be changed to?
if (result != SCAN_SUCCEED && result != SCAN_PTE_MAPPED_HUGEPAGE)
return result;
--
Best Regards,
Yan, Zi
next prev parent reply other threads:[~2026-09-11 2:38 UTC|newest]
Thread overview: 27+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-10 12:02 [PATCH v2 00/12] mm/collapse: separate a collapse from its callers Kiryl Shutsemau
2026-09-10 12:02 ` [PATCH v2 01/12] mm/khugepaged: drop redundant mm_struct pin in madvise_collapse() Kiryl Shutsemau
2026-09-10 12:02 ` [PATCH v2 02/12] mm/khugepaged: count collapses where khugepaged makes them Kiryl Shutsemau
2026-09-10 12:02 ` [PATCH v2 03/12] mm/khugepaged: rename mthp_present_ptes bitmap to eligible_ptes Kiryl Shutsemau
2026-09-10 12:02 ` [PATCH v2 04/12] mm/collapse: add collapse.h for the collapse interface Kiryl Shutsemau
2026-09-10 12:02 ` [PATCH v2 05/12] mm/collapse: state what a collapse may do in the policy Kiryl Shutsemau
2026-09-11 2:06 ` Zi Yan
2026-09-10 12:02 ` [PATCH v2 06/12] mm/collapse: drop the collapse_possible() wrapper Kiryl Shutsemau
2026-09-10 12:02 ` [PATCH v2 07/12] mm/collapse: name the per-table scan reset for what it resets Kiryl Shutsemau
2026-09-10 12:02 ` [PATCH v2 08/12] mm/collapse: separate scanning a PTE table from collapsing it Kiryl Shutsemau
2026-09-11 2:38 ` Zi Yan [this message]
2026-09-11 13:37 ` Kiryl Shutsemau
2026-09-11 14:40 ` Zi Yan
2026-09-10 12:02 ` [PATCH v2 09/12] mm/collapse: open-code collapse_single_pmd() in its two callers Kiryl Shutsemau
2026-09-11 14:57 ` Zi Yan
2026-09-11 15:24 ` Kiryl Shutsemau
2026-09-11 15:26 ` Zi Yan
2026-09-11 22:09 ` Zi Yan
2026-09-10 12:02 ` [PATCH v2 10/12] mm/collapse: work out the orders a VMA allows once per VMA Kiryl Shutsemau
2026-09-11 15:56 ` Zi Yan
2026-09-10 12:02 ` [PATCH v2 11/12] mm/collapse: declare the collapse interface in collapse.h Kiryl Shutsemau
2026-09-11 19:02 ` Zi Yan
2026-09-10 12:02 ` [PATCH v2 12/12] mm/collapse: implement MADV_COLLAPSE in madvise.c Kiryl Shutsemau
2026-09-11 15:06 ` [PATCH v2 00/12] mm/collapse: separate a collapse from its callers David Hildenbrand (Arm)
2026-09-11 15:56 ` Kiryl Shutsemau
2026-09-11 15:58 ` Kiryl Shutsemau
2026-09-11 18:35 ` David Hildenbrand (Arm)
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=DLC4ZVU8JXFG.3SS0W4STZKWBA@nvidia.com \
--to=ziy@nvidia.com \
--cc=akpm@linux-foundation.org \
--cc=baohua@kernel.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=david@kernel.org \
--cc=dev.jain@arm.com \
--cc=jannh@google.com \
--cc=kas@kernel.org \
--cc=kernel-team@meta.com \
--cc=kirill@shutemov.name \
--cc=lance.yang@linux.dev \
--cc=liam@infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=nico.pache@linux.dev \
--cc=ryan.roberts@arm.com \
--cc=usama.arif@linux.dev \
--cc=vbabka@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®