From: Yuan-Hao Hsu <aa9736195201@gmail.com>
To: David Hildenbrand <david@kernel.org>
Cc: Andrew Morton <akpm@linux-foundation.org>,
Lorenzo Stoakes <ljs@kernel.org>,
liam@infradead.org, Vlastimil Babka <vbabka@kernel.org>,
Mike Rapoport <rppt@kernel.org>,
Suren Baghdasaryan <surenb@google.com>,
Michal Hocko <mhocko@suse.com>, Barry Song <baohua@kernel.org>,
Ryan Roberts <ryan.roberts@arm.com>, Dev Jain <dev.jain@arm.com>,
linux-mm@kvack.org, linux-kernel@vger.kernel.org
Subject: Re: [PATCH v2 1/2] mm/memory: reuse 16 PTEs of an exclusive large folio on a write fault
Date: Wed, 30 Sep 2026 06:45:52 +0800 [thread overview]
Message-ID: <20260929224552.467-1-aa9736195201@gmail.com> (raw)
In-Reply-To: <82738d2b-9c2c-474e-b90a-60a1140b5bd6@kernel.org>
On Thu, 24 Sep 2026 20:46:06 +0200, David Hildenbrand (Arm) wrote:
> The patch needs work. I disagree with various decisions
> either you or the LLM came up with like
[...]
> I tried to see how to implement it cleaner. I think we should definitely
> start with:
[...]
> Entirely untested:
Thanks for writing it out. I took your two patches as they are, on
238650ef6c7c, and ran them through what I had for v1/v2. Nothing
broke:
- builds on x86-64, arm64 4K/16K/64K, i386 (+PAE), x86 without THP,
arm, arm nommu, riscv64, powerpc64le, s390x; sparse and W=1 clean
- DEBUG_VM + DEBUG_VM_PGTABLE + PROVE_LOCKING + PAGE_TABLE_CHECK_ENFORCED,
with the mm selftests, my 12 COW scenarios (vmsplice, PROT_READ VMA
inside the folio, soft-dirty/uffd-wp counts, mremap, holes, pageout,
FOLL_FORCE) and a fork/pageout/mprotect/vmsplice stress: no warnings
- arm64 under QEMU: 64K folios go from 8,200 to 520 faults per pass,
contpte_convert() from 512 to 0, and contpte_ptep_set_access_flags()
from 512 to 0
- x86-64 (i7-12700KF): same fault counts and pass times as my 16-PTE
version; the fault that handles the block takes 720-730 ns instead
of 740-750
- Redis on 64K mTHP (BGSAVE, then 1M SETs): faults 262,100 -> 18,900,
CPU 1.53 -> 1.38 s, 640k -> 715k req/s; THP-off control unchanged
> * Hardcoding WP_REUSE_MAX_NR_PTES, likely should be determine differently.
Same patch, only the constant changed; 256 MiB of PTE-mapped 2M
folios after fork() + child exit:
WP_REUSE_MAX_NR_PTES 16 32 64 128 512
one byte per page (seq)
faults 4,161 2,113 1,089 577 193
ms 5.6 4.7 4.2 4.1 3.9
one byte per page, random 7.0 6.0 5.4 5.2 5.0
one store per 64K 3.2 2.3 1.9 1.7 1.6
one store per 2M: the fault
that walks the block (us) 0.84 1.13 1.71 3.01 10.32
So roughly 0.35 us per block plus 20 ns per PTE, and the gain is flat
from 64 on. For where the number could come from: arm64's CONT_PTES
is 16/128/32 for 4K/16K/64K pages, and fault_around_pages defaults to
64K worth, which is 16 on 4K pages but 4 and 1 on the larger ones. An
arch-overridable default of 16, with arm64 using CONT_PTES, would fit
the numbers. Your call.
> * Marking all 16 PTEs young+dirty. It's somewhat the same thing as we do in
> map_anon_folio_pte_pf(). On arm64 it's already fuzzy with cont-pte. With
> transparent coalescing we'd actually allow it directly. So it does feel like the right thing.
I think it's right, and for a simpler reason: nothing looks at young
or dirty per PTE within a folio. try_to_unmap_one() marks the folio
dirty from the merged pteval of the batch, ttu_anon_lazyfree_folio()
checks folio_test_dirty(), folio_referenced_one() adds the young bits
up. The faulting PTE is young+dirty anyway, so the other 15 can't
change any decision. I tried the most sensitive case I could think
of, a MADV_FREE'd folio with one store per folio after fork(): two
MADV_PAGEOUT passes keep and swap the whole 64K/2M folio on today's
kernel exactly as with your patch.
> I also wonder whether some part of the function could be factored out as helpers for
> other code to use in the future. I also suspect that there are more cleanups to be had.
One candidate: numa_rebuild_large_mapping() does the same
folio/VMA/page-table bounded walk by hand, so a "batch around this
PTE within the folio" helper could serve both.
Thanks,
Yuan-Hao
next prev parent reply other threads:[~2026-09-29 22:45 UTC|newest]
Thread overview: 18+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-18 6:42 [PATCH] mm/memory: reuse the whole " Yuan-Hao Hsu
2026-09-18 12:14 ` David Hildenbrand (Arm)
2026-09-18 18:28 ` Yuan-Hao Hsu
2026-09-18 23:48 ` Barry Song
2026-09-19 7:24 ` Yuan-Hao Hsu
2026-09-18 13:54 ` Lorenzo Stoakes (ARM)
2026-09-19 7:31 ` [PATCH v2 0/2] " Yuan-Hao Hsu
2026-09-19 7:31 ` [PATCH v2 1/2] mm/memory: reuse 16 PTEs of an " Yuan-Hao Hsu
2026-09-24 18:46 ` David Hildenbrand (Arm)
2026-09-29 22:45 ` Yuan-Hao Hsu [this message]
2026-09-19 7:31 ` [PATCH v2 2/2] mm/memory: reuse the whole " Yuan-Hao Hsu
2026-09-19 10:10 ` David Hildenbrand (Arm)
2026-09-19 11:18 ` Yuan-Hao Hsu
2026-09-24 19:54 ` David Hildenbrand (Arm)
2026-09-19 10:08 ` [PATCH v2 0/2] " David Hildenbrand (Arm)
2026-09-19 11:18 ` Yuan-Hao Hsu
2026-09-21 12:36 ` Lorenzo Stoakes (ARM)
2026-09-21 19:59 ` Yuan-Hao Hsu
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260929224552.467-1-aa9736195201@gmail.com \
--to=aa9736195201@gmail.com \
--cc=akpm@linux-foundation.org \
--cc=baohua@kernel.org \
--cc=david@kernel.org \
--cc=dev.jain@arm.com \
--cc=liam@infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=mhocko@suse.com \
--cc=rppt@kernel.org \
--cc=ryan.roberts@arm.com \
--cc=surenb@google.com \
--cc=vbabka@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®