* [REGRESSION] mm/mprotect: shared-dirty base-page toggle slower since v6.17
@ 2026-06-24 17:28 Chengfeng Lin
2026-07-01 10:50 ` Chengfeng Lin
0 siblings, 1 reply; 8+ messages in thread
From: Chengfeng Lin @ 2026-06-24 17:28 UTC (permalink / raw)
To: Pedro Falcato, Andrew Morton, linux-mm
Cc: Liam R. Howlett, Lorenzo Stoakes, Vlastimil Babka, Jann Horn,
linux-kernel, regressions
Hi,
I have a refreshed bare-metal result for the shared-dirty mprotect()
slowdown I reported earlier from QEMU/lab testing.
The reproducer is intentionally narrow:
- MAP_SHARED | MAP_ANONYMOUS mapping
- 64 MiB range, write-prefaulted before timing
- state check: 4 KiB base pages, no THP backing
- repeated full-range mprotect(PROT_READ)
- restore with mprotect(PROT_READ | PROT_WRITE)
- write-touch after each protect/restore cycle
So this is not a generic mprotect() regression claim. The scope is the
shared-dirty base-page PTE permission-change path.
The bare-metal machine is an Intel Core i7-14700 system. The workload is
single-threaded and pinned to one logical CPU with `taskset -c 2`. The primary
metric is `iteration_ns_per_page`, lower is better. It is the wall-clock time
for one full protect/restore/write-touch iteration, divided by the number of
4 KiB pages in the range. Each benchmark step used 9 external rounds, 1000
iterations, and 10 warmup iterations.
First, the v6.12 -> v6.19 result still reproduces on bare metal:
kernel iteration_ns_per_page
v6.12.77 26
v6.19.9 37
I then narrowed the release window with 3 interleaved boot/run steps per
kernel:
kernel values mean
v6.16 25 25 25 25.000
v6.17 37 37 37 37.000
v6.18 38 38 38 38.000
v6.18.19 38 38 38 38.000
v6.19.9 37 36 37 36.667
I also checked later context with the same standalone command:
kernel values mean
v7.0.9 36 36 36 36.000
v6.19.9 + Pedro v3 patch-only 39 39 39 39.000
v7.1.0-rc3 mm-unstable/Pedro 39 39 39 39.000
I do not treat the mm-unstable result as a clean release-kernel comparison.
It is only a follow-up check, and in this workload it did not improve the
standalone result.
All of these runs reported `expected_match_ratio=100` and
`unexpected_results=0`. The state check in the standalone output stays in the
same shape: 4 KiB pages, no THP.
This puts the slowdown in the v6.16 -> v6.17 release window.
As an attribution check, I also built a v6.17 probe kernel that only changes
the present-PTE path in `mm/mprotect.c::change_pte_range()` for this workload
back to a single-PTE start/commit/flush shape. That is not an upstream patch
and not a clean release-kernel comparison; it is only a hot-path probe.
The result was:
kernel values mean
v6.16 25 25 25 25.000
v6.17 37 37 37 37.000
v6.17 single-PTE probe 25 25 25 25.000
So the targeted probe brings v6.17 back to the v6.16 range for this workload.
That points at the v6.17 PTE-batching shape in `change_pte_range()` as the
main cost for this shared-dirty 4 KiB base-page case.
I do not want to overstate the attribution. I tried reversing the official
`cac1db8c3aad ("mm: optimize mprotect() by PTE batching")` patch onto my
linux-6.17 tree, but it did not apply cleanly. That means this is not an
exact revert result. I can only say that the slowdown appears in the
v6.16 -> v6.17 window, and that this focused probe brings the v6.17 result
back to the v6.16 range.
Evidence bundle:
https://github.com/lcf0399/linux-mm-regression-evidence/tree/acd7fef0e0276ac361971b0960e6611811edf5b3/mprotect-shared-dirty-toggle
Standalone reproducer:
https://github.com/lcf0399/linux-mm-regression-evidence/tree/acd7fef0e0276ac361971b0960e6611811edf5b3/mprotect-shared-dirty-toggle/reproducer
For each installed kernel, the standalone reproducer was run as:
taskset -c 2 env MAPPING_MB=64 ITERATIONS=1000 WARMUP=10 \
EXTERNAL_ROUNDS=9 ./run_mprotect_shared_dirty_reproducer.sh
For the release-window check, a small systemd/GRUB queue booted each target
kernel before running the same command.
Bare-metal summaries and raw run logs:
https://github.com/lcf0399/linux-mm-regression-evidence/tree/acd7fef0e0276ac361971b0960e6611811edf5b3/mprotect-shared-dirty-toggle/bare-metal
Release-window narrowing:
https://github.com/lcf0399/linux-mm-regression-evidence/tree/acd7fef0e0276ac361971b0960e6611811edf5b3/mprotect-shared-dirty-toggle/bare-metal/20260623-narrow-6.16-6.19-3rounds
v6.17 single-PTE probe:
https://github.com/lcf0399/linux-mm-regression-evidence/tree/acd7fef0e0276ac361971b0960e6611811edf5b3/mprotect-shared-dirty-toggle/bare-metal/20260624-6.17-singlepte-probe
Probe patch used for that attribution run:
https://github.com/lcf0399/linux-mm-regression-evidence/blob/acd7fef0e0276ac361971b0960e6611811edf5b3/mprotect-shared-dirty-toggle/bare-metal/20260624-6.17-singlepte-probe/0001-mm-mprotect-probe-6.17-single-pte-hotpath.patch
#regzbot introduced: v6.16..v6.17
Does this scope look useful to investigate further? If yes, I can try a more
exact commit-level check or test a patch you think is the right direction.
Thanks,
Chengfeng
^ permalink raw reply [flat|nested] 8+ messages in thread* Re: [REGRESSION] mm/mprotect: shared-dirty base-page toggle slower since v6.17 2026-06-24 17:28 [REGRESSION] mm/mprotect: shared-dirty base-page toggle slower since v6.17 Chengfeng Lin @ 2026-07-01 10:50 ` Chengfeng Lin 2026-07-21 12:39 ` Chengfeng Lin 0 siblings, 1 reply; 8+ messages in thread From: Chengfeng Lin @ 2026-07-01 10:50 UTC (permalink / raw) To: Pedro Falcato, Andrew Morton, linux-mm Cc: Liam R. Howlett, Lorenzo Stoakes, Vlastimil Babka, Jann Horn, linux-kernel, regressions Hi, I wanted to add one small follow-up to make the scope clearer. The original reproducer repeatedly does: mprotect(PROT_READ) mprotect(PROT_READ | PROT_WRITE) write-touch I also checked whether the signal only comes from that repeated protect/restore loop. In the follow-up test, each timed iteration creates a fresh 64 MiB shared-dirty mapping, checks that it is still 4 KiB/no THP, and then times exactly one full-range: mprotect(PROT_READ) This is still the same narrow shared-dirty base-page workload, not a separate regression claim. I also checked a MADV_HUGEPAGE control on the current kernel. With the default shmem THP policy, the shared mapping did not actually form THP. After temporarily switching shmem THP to `advise`, the same shared mapping did form THP and the single mprotect(PROT_READ) cost was much lower. So the numbers below are intentionally for the 4 KiB base-page shape. The bare-metal result is: kernel single_protect_ns_per_page v6.16 8 8 8 mean 8.000 v6.17 14 14 14 mean 14.000 v7.1 18 15 18 mean 17.000 All steps reported `expected_match_ratio=100` and `unexpected_results=0`. So a single `mprotect(PROT_READ)` on the prepared shared-dirty range already shows the v6.16 -> v6.17 slowdown. The earlier toggle result is therefore not only measuring steady-state restore/write-touch effects. Follow-up evidence: https://github.com/lcf0399/linux-regression-evidence/tree/8a19614b08d8d2a4c2698c1cfac5a747f12aa021/mprotect-shared-dirty-toggle/bare-metal/20260630-single-protect-followup Updated evidence root: https://github.com/lcf0399/linux-regression-evidence/tree/8a19614b08d8d2a4c2698c1cfac5a747f12aa021/mprotect-shared-dirty-toggle If this scope looks useful, I can provide more runs or narrower measurements around the v6.16..v6.17 window. Thanks, Chengfeng > -----原始邮件----- > 发件人: "Chengfeng Lin" <chengfenglin@stu.xmu.edu.cn> > 发送时间:2026-06-25 01:28:04 (星期四) > 收件人: "Pedro Falcato" <pfalcato@suse.de>, "Andrew Morton" <akpm@linux-foundation.org>, linux-mm@kvack.org > 抄送: "Liam R. Howlett" <liam@infradead.org>, "Lorenzo Stoakes" <ljs@kernel.org>, "Vlastimil Babka" <vbabka@kernel.org>, "Jann Horn" <jannh@google.com>, linux-kernel@vger.kernel.org, regressions@lists.linux.dev > 主题: [REGRESSION] mm/mprotect: shared-dirty base-page toggle slower since v6.17 > > Hi, > > I have a refreshed bare-metal result for the shared-dirty mprotect() > slowdown I reported earlier from QEMU/lab testing. > > The reproducer is intentionally narrow: > > - MAP_SHARED | MAP_ANONYMOUS mapping > - 64 MiB range, write-prefaulted before timing > - state check: 4 KiB base pages, no THP backing > - repeated full-range mprotect(PROT_READ) > - restore with mprotect(PROT_READ | PROT_WRITE) > - write-touch after each protect/restore cycle > > So this is not a generic mprotect() regression claim. The scope is the > shared-dirty base-page PTE permission-change path. > > The bare-metal machine is an Intel Core i7-14700 system. The workload is > single-threaded and pinned to one logical CPU with `taskset -c 2`. The primary > metric is `iteration_ns_per_page`, lower is better. It is the wall-clock time > for one full protect/restore/write-touch iteration, divided by the number of > 4 KiB pages in the range. Each benchmark step used 9 external rounds, 1000 > iterations, and 10 warmup iterations. > > First, the v6.12 -> v6.19 result still reproduces on bare metal: > > kernel iteration_ns_per_page > v6.12.77 26 > v6.19.9 37 > > I then narrowed the release window with 3 interleaved boot/run steps per > kernel: > > kernel values mean > v6.16 25 25 25 25.000 > v6.17 37 37 37 37.000 > v6.18 38 38 38 38.000 > v6.18.19 38 38 38 38.000 > v6.19.9 37 36 37 36.667 > > I also checked later context with the same standalone command: > > kernel values mean > v7.0.9 36 36 36 36.000 > v6.19.9 + Pedro v3 patch-only 39 39 39 39.000 > v7.1.0-rc3 mm-unstable/Pedro 39 39 39 39.000 > > I do not treat the mm-unstable result as a clean release-kernel comparison. > It is only a follow-up check, and in this workload it did not improve the > standalone result. > > All of these runs reported `expected_match_ratio=100` and > `unexpected_results=0`. The state check in the standalone output stays in the > same shape: 4 KiB pages, no THP. > > This puts the slowdown in the v6.16 -> v6.17 release window. > > As an attribution check, I also built a v6.17 probe kernel that only changes > the present-PTE path in `mm/mprotect.c::change_pte_range()` for this workload > back to a single-PTE start/commit/flush shape. That is not an upstream patch > and not a clean release-kernel comparison; it is only a hot-path probe. > > The result was: > > kernel values mean > v6.16 25 25 25 25.000 > v6.17 37 37 37 37.000 > v6.17 single-PTE probe 25 25 25 25.000 > > So the targeted probe brings v6.17 back to the v6.16 range for this workload. > That points at the v6.17 PTE-batching shape in `change_pte_range()` as the > main cost for this shared-dirty 4 KiB base-page case. > > I do not want to overstate the attribution. I tried reversing the official > `cac1db8c3aad ("mm: optimize mprotect() by PTE batching")` patch onto my > linux-6.17 tree, but it did not apply cleanly. That means this is not an > exact revert result. I can only say that the slowdown appears in the > v6.16 -> v6.17 window, and that this focused probe brings the v6.17 result > back to the v6.16 range. > > Evidence bundle: > > https://github.com/lcf0399/linux-mm-regression-evidence/tree/acd7fef0e0276ac361971b0960e6611811edf5b3/mprotect-shared-dirty-toggle > > Standalone reproducer: > > https://github.com/lcf0399/linux-mm-regression-evidence/tree/acd7fef0e0276ac361971b0960e6611811edf5b3/mprotect-shared-dirty-toggle/reproducer > > For each installed kernel, the standalone reproducer was run as: > > taskset -c 2 env MAPPING_MB=64 ITERATIONS=1000 WARMUP=10 \ > EXTERNAL_ROUNDS=9 ./run_mprotect_shared_dirty_reproducer.sh > > For the release-window check, a small systemd/GRUB queue booted each target > kernel before running the same command. > > Bare-metal summaries and raw run logs: > > https://github.com/lcf0399/linux-mm-regression-evidence/tree/acd7fef0e0276ac361971b0960e6611811edf5b3/mprotect-shared-dirty-toggle/bare-metal > > Release-window narrowing: > > https://github.com/lcf0399/linux-mm-regression-evidence/tree/acd7fef0e0276ac361971b0960e6611811edf5b3/mprotect-shared-dirty-toggle/bare-metal/20260623-narrow-6.16-6.19-3rounds > > v6.17 single-PTE probe: > > https://github.com/lcf0399/linux-mm-regression-evidence/tree/acd7fef0e0276ac361971b0960e6611811edf5b3/mprotect-shared-dirty-toggle/bare-metal/20260624-6.17-singlepte-probe > > Probe patch used for that attribution run: > > https://github.com/lcf0399/linux-mm-regression-evidence/blob/acd7fef0e0276ac361971b0960e6611811edf5b3/mprotect-shared-dirty-toggle/bare-metal/20260624-6.17-singlepte-probe/0001-mm-mprotect-probe-6.17-single-pte-hotpath.patch > > #regzbot introduced: v6.16..v6.17 > > Does this scope look useful to investigate further? If yes, I can try a more > exact commit-level check or test a patch you think is the right direction. > > Thanks, > Chengfeng > ^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [REGRESSION] mm/mprotect: shared-dirty base-page toggle slower since v6.17 2026-07-01 10:50 ` Chengfeng Lin @ 2026-07-21 12:39 ` Chengfeng Lin 2026-07-22 5:48 ` Dev Jain 0 siblings, 1 reply; 8+ messages in thread From: Chengfeng Lin @ 2026-07-21 12:39 UTC (permalink / raw) To: Pedro Falcato, Andrew Morton, linux-mm Cc: Liam R. Howlett, Lorenzo Stoakes, Vlastimil Babka, Jann Horn, linux-kernel, regressions Hi, Just a gentle ping, with results from an exact commit-level check. I tested cac1db8c3aad ("mm: optimize mprotect() by PTE batching") against its direct parent 45199f715b74. Each point used a fresh boot: 45199f715b74 parent A -> cac1db8c3aad child -> 45199f715b74 parent B I used the same narrow 64 MiB shared-dirty 4 KiB base-page toggle workload from the report. Each point had 3 warm-up processes and 15 measured processes. Each measured process ran 1,000 iterations. The primary result, iteration_ns_per_page (lower is better), was: point mean ns/page SD parent A 38.133 0.743 child 53.533 0.516 parent B 38.467 0.640 The parent midpoint was 38.300 ns/page. The child was 39.77% slower, while parent B differed from parent A by only 0.87%. Dropping the first measured process from every point left a 39.53% child delta. All 45 measured processes passed the semantic checks and used 4 KiB pages with no THP. The protect and restore components each increased from 13 ns/page in the parent controls to 20 ns/page in the child. The post-touch component stayed close to 12 ns/page. The exact commit changes only mm/mprotect.c. Both kernels were built from the exact commits with the same normalized config, GCC 15.2.0, and Kbuild metadata. Every point used the same CPU settings and runtime preempt=none. For this narrow workload, the exact parent/child result identifies cac1db8c3aad as the source of the measured slowdown. I am not claiming a generic mprotect() or application-level regression. Evidence: https://github.com/lcf0399/linux-regression-evidence/tree/a3ae7bf1a50e4d88b572c6005a037679530c9b1f/mprotect-shared-dirty-toggle/bare-metal/20260721-cac1db8c3aad-exact-ab #regzbot introduced: cac1db8c3aad97d6ffb56ced8868d6cbbbd2bfbe If useful, I can still collect a perf profile of the child around change_pte_range(). Thanks, Chengfeng ^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [REGRESSION] mm/mprotect: shared-dirty base-page toggle slower since v6.17 2026-07-21 12:39 ` Chengfeng Lin @ 2026-07-22 5:48 ` Dev Jain 2026-07-22 10:37 ` Chengfeng Lin 0 siblings, 1 reply; 8+ messages in thread From: Dev Jain @ 2026-07-22 5:48 UTC (permalink / raw) To: Chengfeng Lin, Pedro Falcato, Andrew Morton, linux-mm Cc: Liam R. Howlett, Lorenzo Stoakes, Vlastimil Babka, Jann Horn, linux-kernel, regressions On 21/07/26 6:09 pm, Chengfeng Lin wrote: > Hi, > > Just a gentle ping, with results from an exact commit-level check. > > I tested cac1db8c3aad ("mm: optimize mprotect() by PTE batching") > against its direct parent 45199f715b74. Each point used a fresh boot: > > 45199f715b74 parent A -> cac1db8c3aad child -> > 45199f715b74 parent B > > I used the same narrow 64 MiB shared-dirty 4 KiB base-page toggle > workload from the report. Each point had 3 warm-up processes and > 15 measured processes. Each measured process ran 1,000 iterations. > The primary result, iteration_ns_per_page (lower is better), was: > > point mean ns/page SD > parent A 38.133 0.743 > child 53.533 0.516 > parent B 38.467 0.640 > > The parent midpoint was 38.300 ns/page. The child was 39.77% slower, > while parent B differed from parent A by only 0.87%. Dropping the first > measured process from every point left a 39.53% child delta. > > All 45 measured processes passed the semantic checks and used 4 KiB > pages with no THP. The protect and restore components each increased > from 13 ns/page in the parent controls to 20 ns/page in the child. > The post-touch component stayed close to 12 ns/page. > > The exact commit changes only mm/mprotect.c. Both kernels were built from > the exact commits with the same normalized config, GCC 15.2.0, and Kbuild > metadata. Every point used the same CPU settings and runtime preempt=none. > > For this narrow workload, the exact parent/child result identifies > cac1db8c3aad as the source of the measured slowdown. I am not claiming > a generic mprotect() or application-level regression. > > Evidence: > > https://github.com/lcf0399/linux-regression-evidence/tree/a3ae7bf1a50e4d88b572c6005a037679530c9b1f/mprotect-shared-dirty-toggle/bare-metal/20260721-cac1db8c3aad-exact-ab > > #regzbot introduced: cac1db8c3aad97d6ffb56ced8868d6cbbbd2bfbe > > If useful, I can still collect a perf profile of the child around > change_pte_range(). > Sorry it is difficult for me to parse your report. So I didn't understand whether Pedro's patches helped you or not? > Thanks, > Chengfeng ^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [REGRESSION] mm/mprotect: shared-dirty base-page toggle slower since v6.17 2026-07-22 5:48 ` Dev Jain @ 2026-07-22 10:37 ` Chengfeng Lin 2026-07-22 13:54 ` Dev Jain 0 siblings, 1 reply; 8+ messages in thread From: Chengfeng Lin @ 2026-07-22 10:37 UTC (permalink / raw) To: Dev Jain Cc: Pedro Falcato, Andrew Morton, linux-mm, Liam R. Howlett, Lorenzo Stoakes, Vlastimil Babka, Jann Horn, linux-kernel, regressions Hi, Sorry that my earlier message was difficult to follow. The previous exact parent/child test showed that cac1db8c3aad slowed this workload by 39.77%. To answer your question directly: no. In a separate matched v7.1.3 patch-on/patch-off test, Pedro's full v3 series was 6.20% slower than the no-v3 controls, so it did not reduce the slowdown in this workload. I also checked for an update. The latest public revision I found is still v3, and it is already included in Linux v7.1.3. To isolate the series, I used the same v7.1.3 source for both kernels. The no-v3 kernel replaces only mm/mprotect.c with the exact pre-series file at 19999e479c2a. The full-v3 kernel is unmodified v7.1.3 and matches the series tip at 89e613bc0b2d. The normalized configs and all other build conditions were the same. I used a fresh boot for every point: point mean ns/page no-v3 A 55.333 full v3 58.200 no-v3 B 54.267 The no-v3 midpoint was 54.800 ns/page. Full v3 was 6.20% slower than that midpoint. The two controls drifted by -1.93%. Dropping the first measured run from every point left the v3 delta at +6.18%. This was the same narrow standalone workload: a 64 MiB shared anonymous mapping, dirty 4 KiB pages, and repeated full-range read-only, restore, and write-touch cycles. Each point had 3 warm-up processes and 15 measured processes; each measured process ran 1,000 cycles. The benchmark was pinned to one P-core with Turbo disabled and preempt=none. All 45 measured processes passed the returned-value checks and stayed on 4 KiB pages with no THP. The protect and restore phases were each about 23 ns/page with v3 and 21 ns/page without v3. The later write-touch phase was unchanged. So, for this specific microbenchmark on this machine, Pedro v3 did not reduce the slowdown. It produced a small slowdown instead. I am not making a claim about other mprotect patterns or application-level performance. My earlier QEMU result that suggested a partial improvement should not be used. Thanks, Chengfeng ^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [REGRESSION] mm/mprotect: shared-dirty base-page toggle slower since v6.17 2026-07-22 10:37 ` Chengfeng Lin @ 2026-07-22 13:54 ` Dev Jain 2026-07-22 14:11 ` Pedro Falcato 2026-07-23 17:43 ` Chengfeng Lin 0 siblings, 2 replies; 8+ messages in thread From: Dev Jain @ 2026-07-22 13:54 UTC (permalink / raw) To: Chengfeng Lin Cc: Pedro Falcato, Andrew Morton, linux-mm, Liam R. Howlett, Lorenzo Stoakes, Vlastimil Babka, Jann Horn, linux-kernel, regressions On 22/07/26 4:07 pm, Chengfeng Lin wrote: > Hi, > > Sorry that my earlier message was difficult to follow. > > The previous exact parent/child test showed that cac1db8c3aad slowed this > workload by 39.77%. To answer your question directly: no. In a separate > matched v7.1.3 patch-on/patch-off test, Pedro's full v3 series was 6.20% > slower than the no-v3 controls, so it did not reduce the slowdown in this > workload. > > I also checked for an update. The latest public revision I found is still v3, > and it is already included in Linux v7.1.3. > > To isolate the series, I used the same v7.1.3 source for both kernels. The > no-v3 kernel replaces only mm/mprotect.c with the exact pre-series file at > 19999e479c2a. The full-v3 kernel is unmodified v7.1.3 and matches the series > tip at 89e613bc0b2d. The normalized configs and all other build conditions > were the same. > > I used a fresh boot for every point: > > point mean ns/page > no-v3 A 55.333 > full v3 58.200 > no-v3 B 54.267 > > The no-v3 midpoint was 54.800 ns/page. Full v3 was 6.20% slower than that > midpoint. The two controls drifted by -1.93%. Dropping the first measured run > from every point left the v3 delta at +6.18%. > > This was the same narrow standalone workload: a 64 MiB shared anonymous > mapping, dirty 4 KiB pages, and repeated full-range read-only, restore, and > write-touch cycles. Each point had 3 warm-up processes and 15 measured > processes; each measured process ran 1,000 cycles. The benchmark was pinned to > one P-core with Turbo disabled and preempt=none. > > All 45 measured processes passed the returned-value checks and stayed on 4 KiB > pages with no THP. The protect and restore phases were each about 23 ns/page > with v3 and 21 ns/page without v3. The later write-touch phase was unchanged. > > So, for this specific microbenchmark on this machine, Pedro v3 did not reduce > the slowdown. It produced a small slowdown instead. I am not making a claim > about other mprotect patterns or application-level performance. My earlier > QEMU result that suggested a partial improvement should not be used. I don't see why Pedro's patches should cause a slowdown. The only diff I can think of is the if (likely(nr_ptes == 1)) check. But again, we have had these microbench regression discussions before. I think it is clear by now that vm_normal_folio() is bad on some systems. Best we can do here IMO is avoid that by doing something like https://lore.kernel.org/all/20250506050056.59250-3-dev.jain@arm.com/ > > Thanks, > Chengfeng > ^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [REGRESSION] mm/mprotect: shared-dirty base-page toggle slower since v6.17 2026-07-22 13:54 ` Dev Jain @ 2026-07-22 14:11 ` Pedro Falcato 2026-07-23 17:43 ` Chengfeng Lin 1 sibling, 0 replies; 8+ messages in thread From: Pedro Falcato @ 2026-07-22 14:11 UTC (permalink / raw) To: Dev Jain Cc: Chengfeng Lin, Andrew Morton, linux-mm, Liam R. Howlett, Lorenzo Stoakes, Vlastimil Babka, Jann Horn, linux-kernel, regressions On Wed, Jul 22, 2026 at 07:24:31PM +0530, Dev Jain wrote: > > > On 22/07/26 4:07 pm, Chengfeng Lin wrote: > > Hi, > > > > Sorry that my earlier message was difficult to follow. > > > > The previous exact parent/child test showed that cac1db8c3aad slowed this > > workload by 39.77%. To answer your question directly: no. In a separate > > matched v7.1.3 patch-on/patch-off test, Pedro's full v3 series was 6.20% > > slower than the no-v3 controls, so it did not reduce the slowdown in this > > workload. > > > > I also checked for an update. The latest public revision I found is still v3, > > and it is already included in Linux v7.1.3. > > > > To isolate the series, I used the same v7.1.3 source for both kernels. The > > no-v3 kernel replaces only mm/mprotect.c with the exact pre-series file at > > 19999e479c2a. The full-v3 kernel is unmodified v7.1.3 and matches the series > > tip at 89e613bc0b2d. The normalized configs and all other build conditions > > were the same. > > > > I used a fresh boot for every point: > > > > point mean ns/page > > no-v3 A 55.333 > > full v3 58.200 > > no-v3 B 54.267 > > > > The no-v3 midpoint was 54.800 ns/page. Full v3 was 6.20% slower than that > > midpoint. The two controls drifted by -1.93%. Dropping the first measured run > > from every point left the v3 delta at +6.18%. > > > > This was the same narrow standalone workload: a 64 MiB shared anonymous > > mapping, dirty 4 KiB pages, and repeated full-range read-only, restore, and > > write-touch cycles. Each point had 3 warm-up processes and 15 measured > > processes; each measured process ran 1,000 cycles. The benchmark was pinned to > > one P-core with Turbo disabled and preempt=none. > > > > All 45 measured processes passed the returned-value checks and stayed on 4 KiB > > pages with no THP. The protect and restore phases were each about 23 ns/page > > with v3 and 21 ns/page without v3. The later write-touch phase was unchanged. > > > > So, for this specific microbenchmark on this machine, Pedro v3 did not reduce > > the slowdown. It produced a small slowdown instead. I am not making a claim > > about other mprotect patterns or application-level performance. My earlier > > QEMU result that suggested a partial improvement should not be used. > > I don't see why Pedro's patches should cause a slowdown. The only diff I can think > of is the if (likely(nr_ptes == 1)) check. Agree. I don't really understand what this slowdown is about (nor have I had a proper look, to be fair). > > But again, we have had these microbench regression discussions before. I think it > is clear by now that vm_normal_folio() is bad on some systems. Best we can do > here IMO is avoid that by doing something like On my desktop the results were very clear: vm_normal_folio() didn't really make a difference. Perhaps we're looking at microarchitectural differences here. I dunno. -- Pedro ^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [REGRESSION] mm/mprotect: shared-dirty base-page toggle slower since v6.17 2026-07-22 13:54 ` Dev Jain 2026-07-22 14:11 ` Pedro Falcato @ 2026-07-23 17:43 ` Chengfeng Lin 1 sibling, 0 replies; 8+ messages in thread From: Chengfeng Lin @ 2026-07-23 17:43 UTC (permalink / raw) To: Dev Jain Cc: Pedro Falcato, Andrew Morton, linux-mm, Liam R. Howlett, Lorenzo Stoakes, Vlastimil Babka, Jann Horn, linux-kernel, regressions Hi, Thanks for the comments. First, the exact direct parent/child test showed that cac1db8c3aad ("mm: optimize mprotect() by PTE batching") made this workload 39.77% slower. A separate matched Pedro v3 test did not help: on v7.1.3, full-v3 was 6.20% slower than the midpoint of the no-v3 controls. I do not treat that small difference as a separate regression. For context, this remains the same narrow synthetic microbenchmark: a 64 MiB shared anonymous mapping, dirty 4 KiB base pages with no THP, and repeated full-range protect, restore, and write-touch cycles. Exact-commit decomposition -------------------------- I ran a nine-point, fresh-boot nested ablation around the exact parent/child pair, returning through matched controls to limit time drift. It reproduced a 36.51% child slowdown. Measured as a fraction of that gap: diagnostic change gap recovered parent-style single-PTE update/flush path 43.06% skip normal-path batch discovery no measurable effect additionally skip the normal-path folio lookup 44.47% all three diagnostic changes combined 87.29% The first diagnostic kept the child's lookup and batch discovery unchanged, restoring the parent-style update/flush path only when discovery returned one PTE. It recovered 43.06% of the gap, placing a substantial cost after discovery in the generic single-PTE handling, though not at a specific instruction. In the nested sequence, the intermediate discovery bypass had no measurable effect; additionally skipping the normal-path folio lookup recovered another 44.47%, so lookup contributes a similarly large share. These costs occur in sequence, and both the protect and restore walks pay both. The combined probe was still 4.64% slower than the parent, leaving 12.71% of the original gap unexplained. All 135 measured processes passed the semantic and 4 KiB/no-THP checks; the maximum CV was 1.56%. v7.1.3 lookup safety check -------------------------- The exact sequence above uses the cac1db8c3aad child, which calls vm_normal_folio(). This separate check uses unmodified v7.1.3, where the corresponding path calls vm_normal_page() followed by page_folio(); its percentages are not part of the exact-commit decomposition. For a shared, non-NUMA PTE with pte_batch_hint() == 1, a diagnostic that skips the current page/folio lookup and batch discovery was 17.36% faster on the 4 KiB workload. A paired trace removed exactly 2,048 vm_normal_page() calls, but this is not a standalone timing of that function. The same diagnostic was 65.80% slower for a verified shared 2 MiB folio mapped by 512 PTEs, because on x86 the hint is also one and the diagnostic disables folio batching. All tests used fresh boots on an i7-12700KF with 32 GiB RAM. Within each comparison, builds used matched configs and GCC 15.2.0. The benchmark used P-core CPU 2, performance governor and EPP, Turbo off, and preempt=none. Thus, the generic single-PTE update/flush path and the lookup account for about 87% of the exact-commit gap. The v7.1.3 bypass is unsafe for PTE-mapped large folios, so I am not proposing it as a fix. Pedro did not observe the same lookup cost on his desktop. That difference may be machine-specific, but I have not tested enough systems to establish why. I have also not measured application-level impact. Compact evidence bundle: https://github.com/lcf0399/linux-regression-evidence/tree/ea3ec2d41702502336b9854b9da78c43840b86e8/mprotect-shared-dirty-toggle/bare-metal Standalone reproducers: https://github.com/lcf0399/linux-regression-evidence/tree/ea3ec2d41702502336b9854b9da78c43840b86e8/mprotect-shared-dirty-toggle/reproducer Thanks, Chengfeng ^ permalink raw reply [flat|nested] 8+ messages in thread
end of thread, other threads:[~2026-07-23 17:43 UTC | newest] Thread overview: 8+ messages (download: mbox.gz / follow: Atom feed) -- links below jump to the message on this page -- 2026-06-24 17:28 [REGRESSION] mm/mprotect: shared-dirty base-page toggle slower since v6.17 Chengfeng Lin 2026-07-01 10:50 ` Chengfeng Lin 2026-07-21 12:39 ` Chengfeng Lin 2026-07-22 5:48 ` Dev Jain 2026-07-22 10:37 ` Chengfeng Lin 2026-07-22 13:54 ` Dev Jain 2026-07-22 14:11 ` Pedro Falcato 2026-07-23 17:43 ` Chengfeng Lin
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox
Powered by JetHome