From: Kiryl Shutsemau <kirill@shutemov.name>
To: "David Hildenbrand (Arm)" <david@kernel.org>
Cc: Andrew Morton <akpm@linux-foundation.org>,
Lorenzo Stoakes <ljs@kernel.org>, Zi Yan <ziy@nvidia.com>,
Baolin Wang <baolin.wang@linux.alibaba.com>,
linux-mm@kvack.org, linux-kernel@vger.kernel.org,
kernel-team@meta.com, "Liam R. Howlett" <liam@infradead.org>,
Nico Pache <nico.pache@linux.dev>,
Ryan Roberts <ryan.roberts@arm.com>, Dev Jain <dev.jain@arm.com>,
Barry Song <baohua@kernel.org>,
Lance Yang <lance.yang@linux.dev>,
Usama Arif <usama.arif@linux.dev>,
Vlastimil Babka <vbabka@kernel.org>,
Jann Horn <jannh@google.com>
Subject: Re: [PATCH v4 10/13] mm/collapse: open-code collapse_single_pmd() in its two callers
Date: Fri, 2 Oct 2026 15:27:18 +0100 [thread overview]
Message-ID: <ar-nKHZ5hlbsXeFP@thinkstation> (raw)
In-Reply-To: <b2133cff-ce69-4a45-bf33-c44b774dee7b@kernel.org>
On Thu, Oct 01, 2026 at 11:37:31AM +0200, David Hildenbrand (Arm) wrote:
> On 9/28/26 12:06, Kiryl Shutsemau wrote:
> > From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
> >
> > A scan and a collapse want different things from mmap_lock. The scan
> > reads one PTE table under the lock the caller holds, refuses most of the
> > time, and the caller moves on to the next table without letting go. The
> > collapse allocates, may sleep in writeback and takes the lock for write
> > itself, so the lock it is handed is of no use to it.
> >
> > collapse_single_pmd() kept that boundary inside itself. It dropped the
> > lock on some paths and not others, and reported which by way of a bool
> > its callers had to carry along and then act on.
>
> Let me think this through. collapse_scan_file: it doesn't actually need the MM at
> all. The only reason is to do tracing. Rather stupid, it just should not consume
> the MM at all anymore. Consequently it doesn't even need the mmap lock. But the
> caller needs the mmap lock to figure out the file + range from the vma (the
> per-vma lock would also be sufficient for that).
Right on both.
As I mentioned before, virtual address scanning might not be the best
way to collapse page cache.
We might want to start from inodes on superblocks that have large folios
enabled.
But it is out of scope for the patchset.
> Also, I guess we can convert some of the scanning to use per-vma locks in the
> future, whereby we would actually want to scan with the per-vma lock held.
That is the reason the caller owns the lock here.
On top of this series I have khugepaged scanning under the per-VMA lock:
the scan asserts whatever lock the caller holds, the caller does
vma_end_read() before the run, and the run takes the locks it needs
itself.
Neither the scan nor the run changed for it.
With collapse_single_pmd() dropping the lock, it has to know which lock
the caller took. process_madvise() on another process's mm cannot move
to the VMA lock: untagged_addr_remote() asserts mmap_lock, so a remote
MADV_COLLAPSE keeps it while a local one and khugepaged move on.
The single call would need a flag saying which lock to drop. That is
the lock_dropped bool again, pointing the other way.
> I do wonder about one thing: should we really care so much about keeping the
> mmap lock locked? Meaning, why not provide a single collapse_single_pmd() that
>
> * Is always called without the mmap lock (as is)
> * Always returns with the mmap lock unlocked (change)
Yes, we should.
Walking the tables under one hold and giving the lock up only to
collapse is how khugepaged has worked since ba76149f47d8 ("thp:
khugepaged").
The bool is just how that signal got plumbed when MADV_COLLAPSE arrived
in 50ad2f24b3b4, and it has already cost one bug, 5a62019807da
("mm/khugepaged: fix issue with tracking lock").
This series removes the bool and keeps the behaviour.
> Sure, we drop+re-acquire the mmap lock a couple of times and lookup the vma, but
> isn't that actually being nice to the other parts of the system? In the future
> it would simply get called with the per-vma lock and would return with it unlocked.
>
> In the good old days, looking up VMAs was expensive, but nowadays ... not sure
> if it still matters?
The rwsem and the VMA lookup are not the only cost.
In your version every table is a visit to the mm: unlock,
khugepaged_mm_lock, back through khugepaged_do_scan(),
khugepaged_mm_lock again, a trylock and a VMA walk, for a scan that on
memory already huge is one pmd read.
I measured your prototype against patch 9 as posted, both on a production
config, 8G of anonymous memory already collapsed, scan_sleep_millisecs 0,
pages_to_scan 65536, PMD order only, five runs each, medians:
as posted yours
khugepaged, nothing to collapse:
full passes over 8G in 20s 251,180 55,477
khugepaged CPU 100% 100%
cost per refused table 19 ns 88 ns
MADV_COLLAPSE, already huge:
512M per call, p50 3.8 us 17.4 us
2M per call, p50 703 ns 703 ns
tables walked in full then refused
(max_ptes_none 0), passes in 20s 803 836
On being nice to the rest of the system: the scan is already bounded.
The read hold ends after pages_to_scan worth of progress. We already
have properly sized scan batching in place.
> khugepaged? Not sure if this matters. madvise? I suspect many real users operate
> on a single PMD only (e.g., tcmalloc, jemalloc). For the other ones, not sure if
> dropping the lock every PMD is really a problem?
khugepaged matters. We run it across the fleet, and its scan cost is
what decides how aggressive we can afford to be with it. I am not
giving up scan rate to keep one entry point into the engine.
For a single PMD there is no difference either way. For a range, today a
MADV_COLLAPSE over memory that is already huge walks it without letting
go of the lock at all; in your version it relocks and revalidates for
every table, and tells madvise_walk_vmas() to look the VMA up again.
> IOW, how bad would the following simplification be (prototype that needs more
> work and thought):
>
...
>
> Based on that, I'd rather want to see collapse_single_pmd() to just inline the file
> and anon paths, and see how we can further optimize the locking internally (e.g., perform
> the pagecache scanning without the mmap lock).
I see the appeal of the single entry point for the collapse engine. I do.
But it costs us on both the locking picture and the scan rate. That
does not work for me.
--
Kiryl Shutsemau / Kirill A. Shutemov
next prev parent reply other threads:[~2026-10-02 14:27 UTC|newest]
Thread overview: 29+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-28 10:06 [PATCH v4 00/13] mm/collapse: separate a collapse from its callers Kiryl Shutsemau
2026-09-28 10:06 ` [PATCH v4 01/13] mm/khugepaged: drop redundant mm_struct pin in madvise_collapse() Kiryl Shutsemau
2026-09-28 10:06 ` [PATCH v4 02/13] mm/khugepaged: count collapses where khugepaged makes them Kiryl Shutsemau
2026-09-28 10:06 ` [PATCH v4 03/13] mm/khugepaged: rename mthp_present_ptes bitmap to eligible_ptes Kiryl Shutsemau
2026-09-28 10:06 ` [PATCH v4 04/13] mm/collapse: add collapse.h for the collapse interface Kiryl Shutsemau
2026-09-28 10:06 ` [PATCH v4 05/13] mm/collapse: state what a collapse may do in the policy Kiryl Shutsemau
2026-09-28 19:26 ` David Hildenbrand (Arm)
2026-09-29 1:32 ` Zi Yan
2026-09-29 8:12 ` Baolin Wang
2026-09-28 10:06 ` [PATCH v4 06/13] mm/collapse: drop the collapse_possible() wrapper Kiryl Shutsemau
2026-09-28 10:06 ` [PATCH v4 07/13] mm/collapse: name the per-table scan reset for what it resets Kiryl Shutsemau
2026-09-28 10:06 ` [PATCH v4 08/13] mm/collapse: call collapse_file() from collapse_single_pmd() Kiryl Shutsemau
2026-09-29 1:41 ` Zi Yan
2026-10-02 11:04 ` Kiryl Shutsemau
2026-10-01 8:08 ` David Hildenbrand (Arm)
2026-10-02 11:08 ` Kiryl Shutsemau
2026-09-28 10:06 ` [PATCH v4 09/13] mm/collapse: separate scanning a PTE table from collapsing it Kiryl Shutsemau
2026-09-29 1:52 ` Zi Yan
2026-10-01 8:34 ` David Hildenbrand (Arm)
2026-10-02 11:16 ` Kiryl Shutsemau
2026-09-28 10:06 ` [PATCH v4 10/13] mm/collapse: open-code collapse_single_pmd() in its two callers Kiryl Shutsemau
2026-10-01 9:37 ` David Hildenbrand (Arm)
2026-10-02 14:27 ` Kiryl Shutsemau [this message]
2026-09-28 10:06 ` [PATCH v4 11/13] mm/collapse: work out the orders a VMA allows once per VMA Kiryl Shutsemau
2026-09-28 10:06 ` [PATCH v4 12/13] mm/collapse: declare the collapse interface in collapse.h Kiryl Shutsemau
2026-09-29 2:02 ` Zi Yan
2026-09-28 10:06 ` [PATCH v4 13/13] mm/collapse: implement MADV_COLLAPSE in madvise.c Kiryl Shutsemau
2026-09-29 2:03 ` Zi Yan
2026-09-28 21:59 ` [PATCH v4 00/13] mm/collapse: separate a collapse from its callers Andrew Morton
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=ar-nKHZ5hlbsXeFP@thinkstation \
--to=kirill@shutemov.name \
--cc=akpm@linux-foundation.org \
--cc=baohua@kernel.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=david@kernel.org \
--cc=dev.jain@arm.com \
--cc=jannh@google.com \
--cc=kernel-team@meta.com \
--cc=lance.yang@linux.dev \
--cc=liam@infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=nico.pache@linux.dev \
--cc=ryan.roberts@arm.com \
--cc=usama.arif@linux.dev \
--cc=vbabka@kernel.org \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®