From: "David Hildenbrand (Arm)" <david@kernel.org>
To: Kiryl Shutsemau <kirill@shutemov.name>
Cc: akpm@linux-foundation.org, ljs@kernel.org, nico.pache@linux.dev,
baolin.wang@linux.alibaba.com, baohua@kernel.org,
dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev,
liam@infradead.org, mhocko@suse.com, rppt@kernel.org,
ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com,
usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com,
usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org,
linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org,
jannh@google.com, willy@infradead.org, pfalcato@suse.de,
rostedt@goodmis.org, mhiramat@kernel.org,
linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org
Subject: Re: [RFC PATCH 00/57] mm/collapse: rebuild collapse on migration primitives
Date: Mon, 14 Sep 2026 17:07:37 +0200 [thread overview]
Message-ID: <fd179a26-fd69-4b5f-9af1-b3755ba9030a@kernel.org> (raw)
In-Reply-To: <aoXRi3ngBhO54Akk@thinkstation>
On 8/19/26 19:09, Kiryl Shutsemau wrote:
> On Tue, Aug 18, 2026 at 03:55:55PM +0200, David Hildenbrand (Arm) wrote:
>>> This replaces khugepaged's anonymous collapse with an engine that
>>> can collapse sub-PMD ranges. It is built around migration entries and
>>> frozen folios instead of heavy locking and isolation, aiming for better
>>> scalability and less disruption to the workload being collapsed.
>>
>> I recall us discussing something around using some PTE/PMD markers (e.g.,
>> migration entries) in the past.
>>
>> One thing that needed care is handling concurrent MADV_DONTNEED + faultin after
>> dropping relevant locks.
>
> Handled at install time.
>
> I drop the PTL after the freeze to allow allocation and copy, but sample
> the PTE values (see saved_ptes) at freeze time. If something changed
> under us by the time we install the new page table entries, we give up
> on that candidate and roll it back; the rest of the round still
> installs. We allow harmless transitions: zero page to none.
I think there was more to it, and Jann also hinted at some examples in his
reply. But we'll get to that part once it's no longer buried in 39 patches ;)
>
>>> Which is why hugepage_vma_revalidate() demands that the VMA span the
>>> whole PMD even for an mTHP order -- "we'd need to lock all VMAs in the
>>> PMD range to support this", as the comment there puts it. A PMD-granular
>>> operation is only safe when one VMA owns the PMD, and that is exactly the
>>> restriction in the way. The alignment is the symptom; the PMD is the
>>> design.
>>
>> I disagree with "A PMD-granular operation is only safe when one VMA owns the
>> PMD". It's safe when all page table walkers can be stopped (see above).
>
> Fair, the sentence is too strong. mmap_write_lock plus a VMA write lock on
> every VMA the PMD covers, plus their rmap locks, would make it safe.
Ack.
>
> But the mechanism still clears the whole PMD, flushes, IPIs and
> repopulates it to collapse each 16-page window. That's very noisy to the
> workload.
>
Right. khugepaged itself is pretty noise already, though. So one would have to
understand "how much more noisy and who cares".
I can understand why one would want to make khugepaged less noisy, though.
> And I am not sure how to deal with rmap locking here. Nothing in mm
> holds two unrelated anon_vma rwsems: vma_prepare() takes one for both VMAs
> it touches, because a merge requires them to share the anon_vma, and
> anon_vma_clone() takes one because "all anon_vma's share the same root".
> The lock ordering in mm/rmap.c has a single anon_vma->rwsem level, so a PMD
> spanning unrelated mappings would need an ordering rule that does not exist
> today.
That's a good point. Try-locking would likely work but might have other effects.
Nobody tried this so far.
Lorenzo is on his way to simplify a lot on that anon_vma front (scalable cow),
and IIRC it would also simplify that case.
>
>>> Between the two, nothing can reach a source, so the copy runs with no
>>> lock held at all -- and the address space is left alone while it does.
>>
>> Right. Concurrent MADV_DONTNEED can zap migration PTEs and other faults even
>> re-fault fresh anon folios. So that must be detected before replacing migration
>> entries again I guess.
>
> Yes, that is the install-time check above.
>
> The copy itself is safe: a zap of a migration entry only clears the slot
> and adjusts rss -- see zap_nonpresent_ptes(), which neither puts the
> folio nor drops its rmap -- so a frozen, locked source cannot go away
> under the copy.
I'll have to think about the impact of having these folios frozen for a longer
time, instead of only very briefly during migration.
E.g., these folios will then be unmovable for the entirety of the collapse
operation, because folio_try_get() by memory offlining/cma/compaction will just
fail.
[...]
>
> There are no PMD-level migration entries here.
That's good.
>
> The freeze is always at PTE level, so the pmd keeps pointing at the
> table until the last step, and the PMD leaf goes in as the terminal
> layer: verify, pmdp_collapse_flush(), deposit a fresh table, set the
> leaf, all in one section under the pmd lock with the pte ptl nested
> inside.
>
> A pmd-level walker sees the old table or the leaf and never pmd_none,
> and faults stay held at pte level by the migration entries throughout,
> which is what lets PMD collapse run under a VMA read lock like
> everything else.
>
>>
>>> Working in windows rather than whole PMDs takes care of the other root.
>>> A sub-PMD window is collapsed under the page table lock, so a collapse
>>> disturbs only the window it collapses, and each candidate is validated
>>
>> I recall us discussing that holding the PT lock for a longer collapse operation
>> (especially on 64k) is problematic. But I don't get all the details from your
>> description here.
>
> As I mentioned above, we drop the ptl after the freeze. And take it a
> second time for the install. Allocation and the copy run in between with
> no lock held -- on 64K the copy at PMD order is 512M of it, which is why
> it cannot sit under either lock.
Hold on, are freezing all folios to collapse? That cannot possibly work with
COW-shared folios that are mapped into other address spaces.
We must only freeze a folio if we are sure that no frozen reference can go away
concurrently.
But maybe I misunderstood or you handle this in a special way elsewhere.
--
Cheers,
David
next prev parent reply other threads:[~2026-09-14 15:07 UTC|newest]
Thread overview: 119+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-16 22:45 Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 01/57] mm: add pte_folio() Kiryl Shutsemau
2026-08-18 16:38 ` Rik van Riel
2026-08-18 18:13 ` David Hildenbrand (Arm)
2026-08-18 20:04 ` Rik van Riel
2026-08-19 7:57 ` David Hildenbrand (Arm)
2026-08-18 17:09 ` David Hildenbrand (Arm)
2026-08-18 18:30 ` Lorenzo Stoakes (ARM)
2026-08-20 10:52 ` Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 02/57] mm: add pte_none_or_zero() Kiryl Shutsemau
2026-08-17 17:57 ` David Hildenbrand (Arm)
2026-08-20 11:03 ` Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 03/57] mm/collapse: add collapse.h for the shared collapse state Kiryl Shutsemau
2026-08-18 10:50 ` Lorenzo Stoakes (ARM)
2026-08-20 11:06 ` Kiryl Shutsemau
2026-08-19 14:19 ` David Hildenbrand (Arm)
2026-08-20 11:11 ` Kiryl Shutsemau
2026-08-24 11:47 ` David Hildenbrand (Arm)
2026-08-24 12:10 ` Kiryl Shutsemau
2026-08-24 12:28 ` David Hildenbrand (Arm)
2026-08-24 12:36 ` Kiryl Shutsemau
2026-08-24 14:09 ` Lorenzo Stoakes (ARM)
2026-08-24 15:18 ` Kiryl Shutsemau
2026-08-24 15:23 ` Lorenzo Stoakes (ARM)
2026-08-24 15:51 ` Lorenzo Stoakes (ARM)
2026-08-16 22:45 ` [RFC PATCH 04/57] mm/collapse: rename mthp_present_ptes to eligible_ptes Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 05/57] mm/collapse: state what a collapse may do in the policy Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 06/57] mm/collapse: move the smallest collapse order to collapse.h Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 07/57] mm/collapse: sketch the new anonymous collapse engine Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 08/57] mm/collapse: scan a table for what a collapse could use Kiryl Shutsemau
2026-08-24 8:39 ` Lance Yang
2026-08-24 9:36 ` Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 09/57] mm/collapse: collect candidate windows into a round Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 10/57] mm/collapse: run a round and feed the outcomes back Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 11/57] mm/collapse: sketch the passes of a round Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 12/57] mm/collapse: allocate a destination per candidate Kiryl Shutsemau
2026-08-24 11:20 ` Lance Yang
2026-08-24 12:37 ` Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 13/57] mm/collapse: revalidate a round against the VMA Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 14/57] mm/collapse: fault the sources in before the freeze Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 15/57] mm/collapse: check what a candidate would freeze Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 16/57] mm/collapse: freeze the sources behind migration entries Kiryl Shutsemau
2026-08-24 13:12 ` Lance Yang
2026-08-24 14:13 ` Kiryl Shutsemau
2026-08-24 15:49 ` Johannes Weiner
2026-08-24 16:13 ` Usama Arif
2026-08-25 2:44 ` Lance Yang
2026-08-16 22:45 ` [RFC PATCH 17/57] mm/collapse: copy the sources into the destinations Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 18/57] mm/collapse: install the destinations at PTE level Kiryl Shutsemau
2026-08-25 6:48 ` Lance Yang
2026-08-26 16:57 ` Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 19/57] mm/collapse: install a PMD leaf as the terminal layer Kiryl Shutsemau
2026-08-25 12:23 ` Lance Yang
2026-08-26 17:54 ` Kiryl Shutsemau
2026-08-25 16:52 ` Jann Horn
2026-08-26 18:36 ` Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 20/57] mm/collapse: put the sources back Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 21/57] mm/collapse: settle whatever the round reached Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 22/57] mm/collapse: walk a table with a selection cursor Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 23/57] mm/collapse: give a refused region a second chance Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 24/57] mm/collapse: report each candidate's outcome to tracing Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 25/57] mm/collapse: collapse anonymous memory with the new engine Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 26/57] mm/collapse: give collapse_single_pmd() the range to work on Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 27/57] mm/collapse: scan the windows a VMA can hold Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 28/57] mm/collapse: remove the mechanism the engine replaces Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 29/57] mm/collapse: move what a collapse is judged on into collapse.c Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 30/57] mm/collapse: name the max_ptes ceiling after collapse Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 31/57] mm/khugepaged: count collapses where khugepaged makes them Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 32/57] mm/collapse: move the file collapse into collapse.c Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 33/57] mm/collapse: split collapse into a scan and a run Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 34/57] mm/collapse: implement MADV_COLLAPSE in madvise.c Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 35/57] mm/madvise: drop MADV_COLLAPSE's redundant mm reference Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 36/57] mm/collapse: report what the scan found Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 37/57] mm/collapse: report what the fault-in pass paid Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 38/57] mm/collapse: report the round, and what it made faulters wait Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 39/57] mm/collapse: name the file collapse's tracepoints after collapse Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 40/57] mm/collapse: remove the tracepoints of the mechanism that is gone Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 41/57] mm/collapse: give collapse its own trace header Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 42/57] mm/collapse: allow error injection into the freeze Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 43/57] mm/khugepaged: check the scan budget before the work, not after Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 44/57] mm/khugepaged: hold the address space open across a scan Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 45/57] mm/collapse: take a per-VMA read lock for the round Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 46/57] mm/khugepaged: scan under a per-VMA read lock Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 47/57] mm/madvise: collapse " Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 48/57] mm/collapse: assert the mm reference the engine relies on Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 49/57] mm/khugepaged: drop the mmap_lock barrier from __khugepaged_exit() Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 50/57] selftests/mm: attribute collapses by candidate event alone Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 51/57] selftests/mm: cover collapse inside a sub-PMD VMA Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 52/57] selftests/mm: cover a hole-y window in " Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 53/57] selftests/mm: cover collapse of mlocked ranges Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 54/57] selftests/mm: cover collapse beside a MADV_FREE'd page Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 55/57] selftests/mm: cover collapse beside a pinned page Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 56/57] selftests/mm: cover the scaled max_ptes_shared limit Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 57/57] MAINTAINERS: add an entry for collapse Kiryl Shutsemau
2026-08-17 8:04 ` Lorenzo Stoakes (ARM)
2026-08-17 8:08 ` David Hildenbrand (Arm)
2026-08-17 10:12 ` Kiryl Shutsemau
2026-08-17 2:02 ` [RFC PATCH 00/57] mm/collapse: rebuild collapse on migration primitives Zi Yan
2026-08-17 10:07 ` Kiryl Shutsemau
2026-08-17 8:52 ` Lorenzo Stoakes (ARM)
2026-08-17 13:38 ` Kiryl Shutsemau
2026-08-18 13:06 ` Lorenzo Stoakes (ARM)
2026-08-18 14:12 ` David Hildenbrand (Arm)
2026-08-18 14:33 ` Lorenzo Stoakes (ARM)
2026-08-19 18:08 ` Kiryl Shutsemau
2026-08-24 14:22 ` Lorenzo Stoakes (ARM)
2026-08-24 14:59 ` Kiryl Shutsemau
2026-08-24 17:00 ` Lorenzo Stoakes (ARM)
2026-09-14 14:54 ` David Hildenbrand (Arm)
2026-09-15 10:40 ` Kiryl Shutsemau
2026-08-18 14:15 ` David Hildenbrand (Arm)
2026-08-18 14:41 ` Lorenzo Stoakes (ARM)
2026-08-19 18:22 ` Kiryl Shutsemau
2026-08-19 18:14 ` Kiryl Shutsemau
2026-08-18 13:55 ` David Hildenbrand (Arm)
2026-08-19 17:09 ` Kiryl Shutsemau
2026-09-14 15:07 ` David Hildenbrand (Arm) [this message]
2026-09-15 10:46 ` Kiryl Shutsemau
2026-09-15 15:21 ` David Hildenbrand (Arm)
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=fd179a26-fd69-4b5f-9af1-b3755ba9030a@kernel.org \
--to=david@kernel.org \
--cc=agordeev@linux.ibm.com \
--cc=akpm@linux-foundation.org \
--cc=baohua@kernel.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=bpf@vger.kernel.org \
--cc=dev.jain@arm.com \
--cc=hughd@google.com \
--cc=jannh@google.com \
--cc=kirill@shutemov.name \
--cc=lance.yang@linux.dev \
--cc=liam@infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=linux-trace-kernel@vger.kernel.org \
--cc=ljs@kernel.org \
--cc=mhiramat@kernel.org \
--cc=mhocko@suse.com \
--cc=nico.pache@linux.dev \
--cc=pfalcato@suse.de \
--cc=rostedt@goodmis.org \
--cc=rppt@kernel.org \
--cc=ryan.roberts@arm.com \
--cc=shuah@kernel.org \
--cc=surenb@google.com \
--cc=usama.anjum@arm.com \
--cc=usama.arif@linux.dev \
--cc=vbabka@kernel.org \
--cc=willy@infradead.org \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®