From: "David Hildenbrand (Arm)" <david@kernel.org>
To: Dev Jain <dev.jain@arm.com>,
akpm@linux-foundation.org, ljs@kernel.org, hughd@google.com,
chrisl@kernel.org, kasong@tencent.com, davem@davemloft.net,
andreas@gaisler.com
Cc: riel@surriel.com, liam@infradead.org, vbabka@kernel.org,
harry@kernel.org, jannh@google.com, lance.yang@linux.dev,
baolin.wang@linux.alibaba.com, shikemeng@huaweicloud.com,
nphamcs@gmail.com, baoquan.he@linux.dev, baohua@kernel.org,
youngjun.park@lge.com, linux-mm@kvack.org,
linux-kernel@vger.kernel.org, rppt@kernel.org, surenb@google.com,
mhocko@suse.com, pfalcato@suse.de, jgg@ziepe.ca,
thuth@redhat.com, sparclinux@vger.kernel.org,
ryan.roberts@arm.com, anshuman.khandual@arm.com
Subject: Re: [PATCH v3 0/9] Optimize anonymous swapbacked large folio unmapping
Date: Fri, 2 Oct 2026 12:15:10 +0200 [thread overview]
Message-ID: <54a7252b-ab2d-4230-9a1a-d4d1b670b7c8@kernel.org> (raw)
In-Reply-To: <20260924131106.1730494-1-dev.jain@arm.com>
On 9/24/26 15:09, Dev Jain wrote:
> Speed up unmapping of anonymous swapbacked large folios by clearing
> the ptes, and setting swap ptes, in one go.
>
> The following benchmark (stolen from Barry) is used to measure the
> time taken to swapout 256M worth of memory backed by 64K large folios:
>
> #define _GNU_SOURCE
> #include <stdio.h>
> #include <stdlib.h>
> #include <sys/mman.h>
> #include <string.h>
> #include <time.h>
> #include <unistd.h>
> #include <errno.h>
>
> #define SIZE_MB 256
> #define SIZE_BYTES (SIZE_MB * 1024 * 1024)
>
> int main() {
> void *addr = mmap(NULL, SIZE_BYTES, PROT_READ | PROT_WRITE,
> MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
> if (addr == MAP_FAILED) {
> perror("mmap failed");
> return 1;
> }
>
> memset(addr, 0, SIZE_BYTES);
>
> struct timespec start, end;
> clock_gettime(CLOCK_MONOTONIC, &start);
>
> if (madvise(addr, SIZE_BYTES, MADV_PAGEOUT) != 0) {
> perror("madvise(MADV_PAGEOUT) failed");
> munmap(addr, SIZE_BYTES);
> return 1;
> }
>
> clock_gettime(CLOCK_MONOTONIC, &end);
>
> long duration_ns = (end.tv_sec - start.tv_sec) * 1e9 +
> (end.tv_nsec - start.tv_nsec);
> printf("madvise(MADV_PAGEOUT) took %ld ns (%.3f ms)\n",
> duration_ns, duration_ns / 1e6);
>
> munmap(addr, SIZE_BYTES);
> return 0;
> }
>
> Performance as measured on a Linux VM on Apple M3 (arm64):
>
> Vanilla - Mean: 37401913 ns, std dev: 12%
> Patched - Mean: 17420282 ns, std dev: 11%
>
> resulting in more than 2x speedup.
>
> No regression observed on 4K folios.
>
> Performance as measured on bare metal x86:
>
> Vanilla - mean: 54986286 ns, std dev: 1.5%
> Patched - mean: 51930795 ns, std dev: 3%
>
> I tried magnifying the difference on x86 by using 1M large folios, but
> can't spot an obvious improvement (looks like my system is too fast to
> benefit from batched atomic operations!), hinting that the benefit lies
> mainly in the reduction of ptep_get() calls and the reduction of TLB
> flushes during contpte-unfolding, on arm64.
>
> No regression is observed on 4K folios on x86 too.
>
> ---
> Applies on mm-unstable.
Given that we are approaching rc6 and are flooded with stuff, this will target 7.5.
I have this on my todo list.
--
Cheers,
David
prev parent reply other threads:[~2026-10-02 12:08 UTC|newest]
Thread overview: 21+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-24 13:09 Dev Jain
2026-09-24 13:09 ` [PATCH v3 1/9] mm/swapfile: add batched version of folio_dup_swap Dev Jain
2026-09-29 9:20 ` Kairui Song
2026-09-24 13:09 ` [PATCH v3 2/9] mm/swapfile: add batched version of folio_put_swap Dev Jain
2026-09-29 9:22 ` Kairui Song
2026-09-24 13:09 ` [PATCH v3 3/9] mm: move anon-exclusive batch helper to rmap.h Dev Jain
2026-09-24 21:02 ` Barry Song
2026-09-25 10:18 ` Dev Jain
2026-09-24 13:09 ` [PATCH v3 4/9] mm/rmap: Add batched version of folio_try_share_anon_rmap_pte Dev Jain
2026-09-25 5:30 ` Barry Song
2026-09-25 11:09 ` Dev Jain
2026-09-25 12:03 ` Barry Song
2026-09-24 13:09 ` [PATCH v3 5/9] mm/internal: rename swap offset helpers to softleaf offset Dev Jain
2026-09-24 13:09 ` [PATCH v3 6/9] mm/internal: add set_softleaf_ptes Dev Jain
2026-09-24 13:09 ` [PATCH v3 7/9] mm/memory: use set_softleaf_ptes for uffd-wp markers Dev Jain
2026-09-24 13:09 ` [PATCH v3 8/9] mm/rmap: batch unmap anonymous swap-backed large folios Dev Jain
2026-09-25 6:15 ` Barry Song
2026-09-25 11:25 ` Dev Jain
2026-09-26 12:37 ` Dev Jain
2026-09-24 13:09 ` [PATCH v3 9/9] mm, sparc: batch arch_unmap_one() Dev Jain
2026-10-02 10:15 ` David Hildenbrand (Arm) [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=54a7252b-ab2d-4230-9a1a-d4d1b670b7c8@kernel.org \
--to=david@kernel.org \
--cc=akpm@linux-foundation.org \
--cc=andreas@gaisler.com \
--cc=anshuman.khandual@arm.com \
--cc=baohua@kernel.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=baoquan.he@linux.dev \
--cc=chrisl@kernel.org \
--cc=davem@davemloft.net \
--cc=dev.jain@arm.com \
--cc=harry@kernel.org \
--cc=hughd@google.com \
--cc=jannh@google.com \
--cc=jgg@ziepe.ca \
--cc=kasong@tencent.com \
--cc=lance.yang@linux.dev \
--cc=liam@infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=mhocko@suse.com \
--cc=nphamcs@gmail.com \
--cc=pfalcato@suse.de \
--cc=riel@surriel.com \
--cc=rppt@kernel.org \
--cc=ryan.roberts@arm.com \
--cc=shikemeng@huaweicloud.com \
--cc=sparclinux@vger.kernel.org \
--cc=surenb@google.com \
--cc=thuth@redhat.com \
--cc=vbabka@kernel.org \
--cc=youngjun.park@lge.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®