From: Kiryl Shutsemau <kirill@shutemov.name>
To: Andrew Morton <akpm@linux-foundation.org>,
David Hildenbrand <david@kernel.org>,
Lorenzo Stoakes <ljs@kernel.org>
Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org,
kernel-team@meta.com, Zi Yan <ziy@nvidia.com>,
Baolin Wang <baolin.wang@linux.alibaba.com>,
"Liam R . Howlett" <liam@infradead.org>,
Nico Pache <nico.pache@linux.dev>,
Ryan Roberts <ryan.roberts@arm.com>, Dev Jain <dev.jain@arm.com>,
Barry Song <baohua@kernel.org>, Lance Yang <lance.yang@linux.dev>,
Usama Arif <usama.arif@linux.dev>,
Vlastimil Babka <vbabka@kernel.org>, Jann Horn <jannh@google.com>,
"Kiryl Shutsemau (Meta)" <kas@kernel.org>
Subject: [PATCH 08/12] mm/collapse: separate scanning a PTE table from collapsing it
Date: Fri, 4 Sep 2026 16:10:22 +0100 [thread overview]
Message-ID: <1a1bc537850bd7ef73bed5ac4985634bb8dd95e1.1788533997.git.kas@kernel.org> (raw)
In-Reply-To: <cover.1788533997.git.kas@kernel.org>
From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
A collapse is two jobs. One reads a PTE table under mmap_lock and decides
whether the range is worth collapsing. The other allocates, isolates,
copies and flushes, and wants the lock given up first.
collapse_single_pmd() did both, so the boundary between them was somewhere
in the middle of a function.
Give each half its own function:
- collapse_scan_pmd() scans one table and only reads. The anonymous
scan that used to carry that name keeps its body as
collapse_scan_anon_pmd(), and collapse_scan_pmd() is now the entry
that picks the anonymous or the file side.
- collapse_run_pmd() does the collapse the scan asked for.
SCAN_SUCCEED from the scan means there is something to run; anything
else is why there is not.
collapse_single_pmd() is now the two of them with the mmap_lock drop in
between, so its callers see what they saw before.
Scan results (beyond SCAN_SUCCEED) communicated via collapse_control
structure: the orders, the referenced and swapped-out counts, and for a
file the file itself and the offset in it.
A file collapse works on the page cache and never sees a VMA. The scan
takes the file reference while it still has VMA and the run unpins it
when it is done.
Tracing changes with it. mm_khugepaged_scan_pmd now fires before
mm_collapse_huge_page instead of after it. Its status field already reads
SCAN_SUCCEED for an accepted table, so what the collapse then made of that
table is mm_collapse_huge_page's to report, per order.
The two calls to that tracepoint become one. They differed in what the
collapse between them changed; with the collapse no longer here, both
carry the same arguments. failed_pfn is set only where a PTE was refused,
so it is -1 exactly when the result is SCAN_SUCCEED.
Assisted-by: Claude-Code:claude-opus-5
Signed-off-by: Kiryl Shutsemau (Meta) <kas@kernel.org>
---
mm/collapse.h | 14 ++++++
mm/khugepaged.c | 121 ++++++++++++++++++++++++++++++++++--------------
2 files changed, 100 insertions(+), 35 deletions(-)
diff --git a/mm/collapse.h b/mm/collapse.h
index 05282eed9a35..f03cad8ed40e 100644
--- a/mm/collapse.h
+++ b/mm/collapse.h
@@ -100,6 +100,20 @@ struct collapse_control {
/* Each bit marks a PTE the scan accepted as a collapse source */
DECLARE_BITMAP(eligible_ptes, MAX_PTRS_PER_PTE);
+
+ /*
+ * What a scan found and the run after it needs. Live only between the
+ * two, and read by nobody else.
+ *
+ * The file side takes a reference while it still has the VMA, since a
+ * file collapse works on the page cache and never sees one; the run is
+ * what gives it back.
+ */
+ unsigned long scan_orders;
+ int scan_referenced;
+ int scan_unmapped;
+ struct file *scan_file;
+ pgoff_t scan_pgoff;
};
#endif /* __MM_COLLAPSE_H */
diff --git a/mm/khugepaged.c b/mm/khugepaged.c
index 511ffb381fe9..120af57540df 100644
--- a/mm/khugepaged.c
+++ b/mm/khugepaged.c
@@ -1550,14 +1550,14 @@ static enum scan_result mthp_collapse(struct mm_struct *mm,
return last_result;
}
-static enum scan_result collapse_scan_pmd(struct mm_struct *mm,
- struct vm_area_struct *vma, unsigned long start_addr,
- bool *lock_dropped, struct collapse_control *cc)
+static enum scan_result collapse_scan_anon_pmd(struct vm_area_struct *vma,
+ unsigned long start_addr, struct collapse_control *cc)
{
const unsigned int max_ptes_shared = collapse_max_ptes_shared(cc, HPAGE_PMD_ORDER);
const unsigned int max_ptes_swap = collapse_max_ptes_swap(cc, HPAGE_PMD_ORDER);
unsigned int max_ptes_none = collapse_max_ptes_none(cc, vma, HPAGE_PMD_ORDER);
enum tva_type tva_flags = cc->policy.tva_type;
+ struct mm_struct *mm = vma->vm_mm;
pmd_t *pmd;
pte_t *pte, *_pte, pteval;
int i;
@@ -1737,19 +1737,17 @@ static enum scan_result collapse_scan_pmd(struct mm_struct *mm,
out_unmap:
pte_unmap_unlock(pte, ptl);
if (result == SCAN_SUCCEED) {
- /* collapse_huge_page() expects the lock to be dropped before calling */
- mmap_read_unlock(mm);
- result = mthp_collapse(mm, start_addr, referenced,
- unmapped, cc, enabled_orders);
- /* mmap_lock was released above, set lock_dropped */
- *lock_dropped = true;
- trace_mm_khugepaged_scan_pmd(mm, -1, referenced, none_or_zero,
- SCAN_SUCCEED, unmapped);
- } else {
-out:
- trace_mm_khugepaged_scan_pmd(mm, failed_pfn, referenced,
- none_or_zero, result, unmapped);
+ cc->scan_orders = enabled_orders;
+ cc->scan_referenced = referenced;
+ cc->scan_unmapped = unmapped;
}
+out:
+ /*
+ * failed_pfn is only set where a PTE was refused, so it is -1 on the
+ * path that returns SCAN_SUCCEED.
+ */
+ trace_mm_khugepaged_scan_pmd(mm, failed_pfn, referenced,
+ none_or_zero, result, unmapped);
return result;
}
@@ -2759,30 +2757,58 @@ static enum scan_result collapse_scan_file(struct mm_struct *mm,
return result;
}
-/*
- * Try to collapse a single PMD starting at a PMD aligned addr, and return
- * the results.
- */
-static enum scan_result collapse_single_pmd(unsigned long addr,
- struct vm_area_struct *vma, bool *lock_dropped,
- struct collapse_control *cc)
+static void collapse_control_init(struct collapse_control *cc)
{
- struct mm_struct *mm = vma->vm_mm;
- bool triggered_wb = false;
- enum scan_result result;
- struct file *file;
- pgoff_t pgoff;
+ cc->progress = 0;
+ cc->scan_file = NULL;
+}
- mmap_assert_locked(mm);
+static void collapse_control_release(struct collapse_control *cc)
+{
+ /* A scan that took a file reference should have been run */
+ if (WARN_ON_ONCE(cc->scan_file)) {
+ fput(cc->scan_file);
+ cc->scan_file = NULL;
+ }
+}
+
+static enum scan_result collapse_scan_pmd(struct vm_area_struct *vma,
+ unsigned long addr, struct collapse_control *cc)
+{
+ mmap_assert_locked(vma->vm_mm);
+ /* Whatever the last scan found has to have been run by now */
+ if (WARN_ON_ONCE(cc->scan_file)) {
+ fput(cc->scan_file);
+ cc->scan_file = NULL;
+ }
if (vma_is_anonymous(vma))
- return collapse_scan_pmd(mm, vma, addr, lock_dropped, cc);
+ return collapse_scan_anon_pmd(vma, addr, cc);
- file = get_file(vma->vm_file);
- pgoff = linear_page_index(vma, addr);
+ /*
+ * A file collapse works on the page cache and never sees a VMA, so take
+ * what it needs from this one while it is still here. Judging the
+ * range needs the page cache and no lock, so it happens in the run.
+ */
+ cc->scan_file = get_file(vma->vm_file);
+ cc->scan_pgoff = linear_page_index(vma, addr);
+ return SCAN_SUCCEED;
+}
- mmap_read_unlock(mm);
- *lock_dropped = true;
+static enum scan_result collapse_run_pmd(struct mm_struct *mm,
+ unsigned long addr, struct collapse_control *cc)
+{
+ struct file *file = cc->scan_file;
+ bool triggered_wb = false;
+ enum scan_result result;
+ pgoff_t pgoff;
+
+ if (!file)
+ return mthp_collapse(mm, addr, cc->scan_referenced,
+ cc->scan_unmapped, cc, cc->scan_orders);
+
+ cc->scan_file = NULL;
+ pgoff = cc->scan_pgoff;
retry:
result = collapse_scan_file(mm, addr, file, pgoff, cc);
@@ -2812,6 +2838,28 @@ static enum scan_result collapse_single_pmd(unsigned long addr,
return result;
}
+/*
+ * Try to collapse a single PMD starting at a PMD aligned addr, and return
+ * the results.
+ */
+static enum scan_result collapse_single_pmd(unsigned long addr,
+ struct vm_area_struct *vma, bool *lock_dropped,
+ struct collapse_control *cc)
+{
+ struct mm_struct *mm = vma->vm_mm;
+ enum scan_result result;
+
+ result = collapse_scan_pmd(vma, addr, cc);
+ if (result != SCAN_SUCCEED)
+ return result;
+
+ /* The collapse takes its own locks, so give this up */
+ mmap_read_unlock(mm);
+ *lock_dropped = true;
+
+ return collapse_run_pmd(mm, addr, cc);
+}
+
static void collapse_scan_mm_slot(unsigned int progress_max,
enum scan_result *result, struct collapse_control *cc)
__releases(&khugepaged_mm_lock)
@@ -2954,10 +3002,10 @@ static void khugepaged_do_scan(struct collapse_control *cc)
lru_add_drain_all();
+ collapse_control_init(cc);
/* One policy for the whole pass, so every table is judged the same */
collapse_policy_khugepaged(&cc->policy);
- cc->progress = 0;
while (true) {
cond_resched();
@@ -2988,6 +3036,8 @@ static void khugepaged_do_scan(struct collapse_control *cc)
khugepaged_alloc_sleep();
}
}
+
+ collapse_control_release(cc);
}
static bool khugepaged_should_wakeup(void)
@@ -3184,8 +3234,8 @@ int madvise_collapse(struct vm_area_struct *vma, unsigned long start,
cc = kmalloc_obj(*cc);
if (!cc)
return -ENOMEM;
+ collapse_control_init(cc);
collapse_policy_forced(&cc->policy);
- cc->progress = 0;
lru_add_drain_all();
@@ -3242,6 +3292,7 @@ int madvise_collapse(struct vm_area_struct *vma, unsigned long start,
}
out_nolock:
mmap_assert_locked(mm);
+ collapse_control_release(cc);
kfree(cc);
return thps == ((hend - hstart) >> HPAGE_PMD_SHIFT) ? 0
--
2.54.0
next prev parent reply other threads:[~2026-09-04 15:10 UTC|newest]
Thread overview: 20+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-04 15:10 [PATCH 00/12] mm/collapse: separate a collapse from its callers Kiryl Shutsemau
2026-09-04 15:10 ` [PATCH 01/12] mm/khugepaged: drop redundant mm_struct pin in madvise_collapse() Kiryl Shutsemau
2026-09-04 15:58 ` Zi Yan
2026-09-04 15:10 ` [PATCH 02/12] mm/khugepaged: count collapses where khugepaged makes them Kiryl Shutsemau
2026-09-05 2:25 ` Zi Yan
2026-09-04 15:10 ` [PATCH 03/12] mm/khugepaged: rename mthp_present_ptes bitmap to eligible_ptes Kiryl Shutsemau
2026-09-05 2:28 ` Zi Yan
2026-09-04 15:10 ` [PATCH 04/12] mm/collapse: add collapse.h for the collapse interface Kiryl Shutsemau
2026-09-05 2:36 ` Zi Yan
2026-09-04 15:10 ` [PATCH 05/12] mm/collapse: state what a collapse may do in the policy Kiryl Shutsemau
2026-09-05 2:44 ` Zi Yan
2026-09-04 15:10 ` [PATCH 06/12] mm/collapse: drop the collapse_possible() wrapper Kiryl Shutsemau
2026-09-05 2:45 ` Zi Yan
2026-09-04 15:10 ` [PATCH 07/12] mm/collapse: name the per-table scan reset for what it resets Kiryl Shutsemau
2026-09-05 18:05 ` Zi Yan
2026-09-04 15:10 ` Kiryl Shutsemau [this message]
2026-09-04 15:10 ` [PATCH 09/12] mm/collapse: open-code collapse_single_pmd() in its two callers Kiryl Shutsemau
2026-09-04 15:10 ` [PATCH 10/12] mm/collapse: work out the orders a VMA allows once per VMA Kiryl Shutsemau
2026-09-04 15:10 ` [PATCH 11/12] mm/collapse: declare the collapse interface in collapse.h Kiryl Shutsemau
2026-09-04 15:10 ` [PATCH 12/12] mm/collapse: implement MADV_COLLAPSE in madvise.c Kiryl Shutsemau
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=1a1bc537850bd7ef73bed5ac4985634bb8dd95e1.1788533997.git.kas@kernel.org \
--to=kirill@shutemov.name \
--cc=akpm@linux-foundation.org \
--cc=baohua@kernel.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=david@kernel.org \
--cc=dev.jain@arm.com \
--cc=jannh@google.com \
--cc=kas@kernel.org \
--cc=kernel-team@meta.com \
--cc=lance.yang@linux.dev \
--cc=liam@infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=nico.pache@linux.dev \
--cc=ryan.roberts@arm.com \
--cc=usama.arif@linux.dev \
--cc=vbabka@kernel.org \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®