mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Matthew Wilcox <willy@infradead.org>
To: Barry Song <baohua@kernel.org>
Cc: "Lorenzo Stoakes (ARM)" <ljs@kernel.org>,
	Hongru Zhang <zhanghongru06@gmail.com>,
	akpm@linux-foundation.org, linux-mm@kvack.org,
	linux-kernel@vger.kernel.org, surenb@google.com,
	david@kernel.org, liam@infradead.org, mhocko@suse.com,
	rppt@kernel.org, shakeel.butt@linux.dev, vbabka@kernel.org,
	zhaonanzhe@xiaomi.com, linux@armlinux.org.uk,
	catalin.marinas@arm.com, will@kernel.org, mark.rutland@arm.com,
	linux-arm-kernel@lists.infradead.org, chenhuacai@kernel.org,
	kernel@xen0n.name, loongarch@lists.linux.dev,
	maddy@linux.ibm.com, mpe@ellerman.id.au, npiggin@gmail.com,
	chleroy@kernel.org, linuxppc-dev@lists.ozlabs.org,
	pjw@kernel.org, palmer@dabbelt.com, aou@eecs.berkeley.edu,
	alex@ghiti.fr, linux-riscv@lists.infradead.org,
	agordeev@linux.ibm.com, gerald.schaefer@linux.ibm.com,
	hca@linux.ibm.com, gor@linux.ibm.com, borntraeger@linux.ibm.com,
	svens@linux.ibm.com, linux-s390@vger.kernel.org,
	dave.hansen@linux.intel.com, luto@kernel.org,
	peterz@infradead.org, tglx@kernel.org, mingo@redhat.com,
	bp@alien8.de, x86@kernel.org, hpa@zytor.com,
	Hongru Zhang <zhanghongru@xiaomi.com>
Subject: Re: [PATCH v6] mm: retry page faults once under the per-VMA lock
Date: Thu, 1 Oct 2026 22:12:16 +0100	[thread overview]
Message-ID: <ar7MsOb6IJLt4eLo@casper.infradead.org> (raw)
In-Reply-To: <CAGsJ_4xuFi7K+crAcFQ998SxXS0xPM6QeA34JV+6+Sy005rXLg@mail.gmail.com>

On Mon, Sep 28, 2026 at 10:48:41AM +0800, Barry Song wrote:
> On Mon, Sep 28, 2026 at 6:55 AM Matthew Wilcox <willy@infradead.org> wrote:
> > So while doing my slides, I realised that what we need to avoid doing
> > is (a) sleeping while holding the mmap_lock (b) returning RETRY while
> > holding the VMA lock
> >
> > And that turns out to be as simple as this patch:
> >
> > diff --git a/include/linux/mm.h b/include/linux/mm.h
> > index dd09c438fa23..94ed2333f8d8 100644
> > --- a/include/linux/mm.h
> > +++ b/include/linux/mm.h
> > @@ -723,6 +723,8 @@ enum {
> >   */
> >  static inline bool fault_flag_allow_retry_first(enum fault_flag flags)
> >  {
> > +       if (flags & FAULT_FLAG_VMA_LOCK)
> > +               return false;
> >         return (flags & FAULT_FLAG_ALLOW_RETRY) &&
> >             (!(flags & FAULT_FLAG_TRIED));
> >  }
> >
> > OK, this is a hack.  The function is spectacularly badly named, and
> > needs to be renamed before a patch can go upstream.  But this should
> > fix the contention on mmap_lock.
> 
> Thanks for your suggestion.
> This is exactly what we did in Android Common Kernel before we had
> Lorenzo's proposal (bypassing `fault_flag_allow_retry_first()`):
> 
> https://android.googlesource.com/kernel/common/+/1b9b045a586245cc1c29b2747c6586234c7f5bad%5E%21/#F2

Looks like that one didn't cover __folio_lock_or_retry(), but that
doesn't invalidate your point.

> Note that Lorenzo's proposal avoids mmap_lock contention without
> introducing any new VMA lock contention. It also doesn't require a new
> flag that would break KMI. So this is clearly the preferred approach.

But it does retry multiple times in cases where we know the fault
will always fail (eg the fault is on a device-private VMA)

> > Could somebody try it?  I've verified it boots and runs some userspace
> > fine, but I don't have the workload to test the contention.
> 
> Both Nanzhe and Hongru tested it before and reported the fork issue.

So what I didn't realise is that fork() waits for page faults to finish.
I don't think that's necessary, so we can just stop doing that (whitespace
damaged):

diff --git a/mm/mmap.c b/mm/mmap.c
index 4bf26b0f1e6e..e79555247d3a 100644
--- a/mm/mmap.c
+++ b/mm/mmap.c
@@ -1739,9 +1739,6 @@ __latent_entropy int dup_mmap(struct mm_struct *mm, struct mm_struct *oldmm)
        for_each_vma(vmi, mpnt) {
                struct file *file;

-               retval = vma_start_write_killable(mpnt);
-               if (retval < 0)
-                       goto loop_out;
                if (vma_test(mpnt, VMA_DONTCOPY_BIT)) {
                        retval = vma_iter_clear_gfp(&vmi, mpnt->vm_start,
                                                    mpnt->vm_end, GFP_KERNEL);

I think this is safe.  I've booted a kernel with this change, and
everything seems to run fine.  Of course I don't have any multithreaded
applications which call fork() because that's a stupid way to write an
application, so it's not really tested.

My argument for why it's safe is that a thread which takes a page
fault during fork() might have taken the page fault either before or
after fork().  The faults will definitely happen in the parent process.
They may or may not have happened in the child process, which can't
possibly care whether or not they've happened.

The only difference I can think of being observable is that the child
may observe some later faults to have occurred, while some earlier faults
to have not occurred.  I have a hard time believing any application can
possibly depend on it.

  reply	other threads:[~2026-10-01 21:12 UTC|newest]

Thread overview: 10+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-11  2:56 Hongru Zhang
2026-09-21 17:54 ` Lorenzo Stoakes (ARM)
2026-09-21 18:57   ` Matthew Wilcox
2026-09-22  8:55     ` Lorenzo Stoakes (ARM)
2026-09-27 22:55       ` Matthew Wilcox
2026-09-28  2:48         ` Barry Song
2026-10-01 21:12           ` Matthew Wilcox [this message]
2026-10-02 13:12             ` Lorenzo Stoakes (ARM)
2026-10-02 19:30             ` Barry Song
2026-09-23  9:53   ` Hongru Zhang

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=ar7MsOb6IJLt4eLo@casper.infradead.org \
    --to=willy@infradead.org \
    --cc=agordeev@linux.ibm.com \
    --cc=akpm@linux-foundation.org \
    --cc=alex@ghiti.fr \
    --cc=aou@eecs.berkeley.edu \
    --cc=baohua@kernel.org \
    --cc=borntraeger@linux.ibm.com \
    --cc=bp@alien8.de \
    --cc=catalin.marinas@arm.com \
    --cc=chenhuacai@kernel.org \
    --cc=chleroy@kernel.org \
    --cc=dave.hansen@linux.intel.com \
    --cc=david@kernel.org \
    --cc=gerald.schaefer@linux.ibm.com \
    --cc=gor@linux.ibm.com \
    --cc=hca@linux.ibm.com \
    --cc=hpa@zytor.com \
    --cc=kernel@xen0n.name \
    --cc=liam@infradead.org \
    --cc=linux-arm-kernel@lists.infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=linux-riscv@lists.infradead.org \
    --cc=linux-s390@vger.kernel.org \
    --cc=linux@armlinux.org.uk \
    --cc=linuxppc-dev@lists.ozlabs.org \
    --cc=ljs@kernel.org \
    --cc=loongarch@lists.linux.dev \
    --cc=luto@kernel.org \
    --cc=maddy@linux.ibm.com \
    --cc=mark.rutland@arm.com \
    --cc=mhocko@suse.com \
    --cc=mingo@redhat.com \
    --cc=mpe@ellerman.id.au \
    --cc=npiggin@gmail.com \
    --cc=palmer@dabbelt.com \
    --cc=peterz@infradead.org \
    --cc=pjw@kernel.org \
    --cc=rppt@kernel.org \
    --cc=shakeel.butt@linux.dev \
    --cc=surenb@google.com \
    --cc=svens@linux.ibm.com \
    --cc=tglx@kernel.org \
    --cc=vbabka@kernel.org \
    --cc=will@kernel.org \
    --cc=x86@kernel.org \
    --cc=zhanghongru06@gmail.com \
    --cc=zhanghongru@xiaomi.com \
    --cc=zhaonanzhe@xiaomi.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®