mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: "Lorenzo Stoakes (ARM)" <ljs@kernel.org>
To: Pedro Falcato <pfalcato@suse.de>
Cc: Andrew Morton <akpm@linux-foundation.org>,
	 David Hildenbrand <david@kernel.org>,
	"Liam R. Howlett" <liam@infradead.org>,
	 Vlastimil Babka <vbabka@kernel.org>,
	Mike Rapoport <rppt@kernel.org>,
	 Suren Baghdasaryan <surenb@google.com>,
	Michal Hocko <mhocko@suse.com>, Rik van Riel <riel@surriel.com>,
	 Harry Yoo <harry@kernel.org>, Jann Horn <jannh@google.com>,
	Lance Yang <lance.yang@linux.dev>,
	 linux-mm@kvack.org, linux-kernel@vger.kernel.org,
	Pan Deng <pan.deng@intel.com>
Subject: Re: [PATCH] mm/vma: don't remove VMA from rmap if pgoff unchanged
Date: Mon, 28 Sep 2026 18:01:11 +0100	[thread overview]
Message-ID: <arqczL8K39qBoqgV@gremlin> (raw)
In-Reply-To: <arqI5QB5tUr8dd_r@pedro-suse.tail5790ac.ts.net>

On Mon, Sep 28, 2026 at 04:37:52PM +0100, Pedro Falcato wrote:
> On Fri, Sep 25, 2026 at 07:32:20PM +0100, Lorenzo Stoakes (ARM) wrote:
> > When updating a VMA, vma_prepare() unconditionally removes it from its rmap
> > interval trees under the rmap lock, and vma_complete() reinserts it before
> > releasing the lock.
> >
> > This is wholly unnecessary if its page offset (file rmap) or anonymous page
> > offset (anon rmap) is unchanged.
> >
> > So, track whether they will change in the newly introduced
> > vp->anon_pgoff_unchanged and vp->pgoff_unchanged fields, and use them to
> > determine whether to remove the VMA or not.
> >
> > The rmap lock keeps things safe as no rmap walks can concurrently occur
> > during the operation.
> >
> > Additionally, some architectures (arm, parisc, nios2, csky) have dcache
> > flush rmap walkers which take only flush_dcache_mmap_lock(), which is
> > likewise held across the operation.
> >
> > It's also necessary to keep the rb_subtree_last field updated in the
> > interval tree so implement anon_rmap_tree_update_inplace() and
> > mapping_rmap_tree_update_inplace() to do that.
> >
> > This is done in vma_complete(), after the VMA's range has been updated, so
> > in the interim the field may be invalid. However, given the locks described
> > above, this cannot be observed until after the state is corrected.
> >
> > The anonymous rmap is keyed on anon_vma_chains not VMAs, so in those
> > instances anon_rmap_tree_update_vma_inplace() iterates over
> > vma->anon_vma_chain, invoking anon_rmap_tree_update_inplace() on each one.
> >
> > For the anon rmap case, with CONFIG_DEBUG_VM_RB set, avc->cached_vma_last
> > is also updated in anon_rmap_tree_update_inplace().
> >
> > When performing a VMA shrink or a split where the VMA is the lower one, the
> > page offset cannot change, so set the flags unconditionally in these cases.
> >
> > When merging VMAs the page offset is unchanged only in some cases, so
> > update init_multi_vma_prep() to set the flags only if the page offsets
> > remain the same.
> >
> > These changes ultimately result in less rmap lock contention.
>
> I think this asks for numbers?

Well I don't have any :)

It logically reduces the contention, and that can only be a good thing.

Pan had some numbers, I've asked him to re-run against this one.

>
> >
> > Reported-by: Pan Deng <pan.deng@intel.com>
> > Closes: https://lore.kernel.org/linux-mm/20260924054301.2330822-1-pan.deng@intel.com/
> > Signed-off-by: Lorenzo Stoakes <ljs@kernel.org>
> > ---
> > Signed-off-by: Lorenzo Stoakes (ARM) <ljs@kernel.org>
> > ---
> >  include/linux/mm.h                |  3 +++
> >  mm/interval_tree.c                | 33 ++++++++++++++++++++++++++++++++
> >  mm/vma.c                          | 40 ++++++++++++++++++++++++++++++++++-----
> >  mm/vma.h                          |  2 ++
> >  tools/testing/vma/include/stubs.h |  8 ++++++++
> >  5 files changed, 81 insertions(+), 5 deletions(-)
> >
> > diff --git a/include/linux/mm.h b/include/linux/mm.h
> > index 6e71eaa4af3f..94c2eb055716 100644
> > --- a/include/linux/mm.h
> > +++ b/include/linux/mm.h
> > @@ -4357,6 +4357,8 @@ void mapping_rmap_tree_insert_after(struct vm_area_struct *vma,
> >  				    struct address_space *mapping);
> >  void mapping_rmap_tree_remove(struct vm_area_struct *vma,
> >  			      struct address_space *mapping);
> > +void mapping_rmap_tree_update_inplace(struct vm_area_struct *vma);
> > +
> >  struct vm_area_struct *
> >  mapping_rmap_tree_iter_first(struct address_space *mapping,
> >  			     pgoff_t pgoff_start, pgoff_t pgoff_last);
> > @@ -4374,6 +4376,7 @@ void anon_rmap_tree_insert(struct anon_vma_chain *avc,
> >  			   struct anon_vma *anon_vma);
> >  void anon_rmap_tree_remove(struct anon_vma_chain *avc,
> >  			   struct anon_vma *anon_vma);
> > +void anon_rmap_tree_update_inplace(struct anon_vma_chain *avc);
> >  struct anon_vma_chain *
> >  anon_rmap_tree_iter_first(struct anon_vma *anon_vma,
> >  			  pgoff_t pgoff_start, pgoff_t pgoff_last);
> > diff --git a/mm/interval_tree.c b/mm/interval_tree.c
> > index 7bbbf15cfbf0..eafde5d12ef5 100644
> > --- a/mm/interval_tree.c
> > +++ b/mm/interval_tree.c
> > @@ -64,6 +64,21 @@ void mapping_rmap_tree_remove(struct vm_area_struct *vma,
> >  	__mapping_rmap_tree_remove(vma, &mapping->i_mmap);
> >  }
> >
> > +/**
> > + * mapping_rmap_tree_update_inplace() - Update file rmap tree to reflect an
> > + * in-place change in a VMA's size.
> > + * @vma: The VMA whose size has changed.
> > + *
> > + * The file rmap lock must be held.
> > + *
> > + * Invalid to do so if @vma->vm_pgoff has changed.
> > + */
> > +void mapping_rmap_tree_update_inplace(struct vm_area_struct *vma)
> > +{
> > +	/* Propagate all the way up the tree. */
> > +	__mapping_rmap_tree_augment.propagate(&vma->shared.rb, NULL);
> > +}
> > +
> >  struct vm_area_struct *
> >  mapping_rmap_tree_iter_first(struct address_space *mapping,
> >  			     pgoff_t pgoff_start, pgoff_t pgoff_last)
> > @@ -111,6 +126,24 @@ void anon_rmap_tree_remove(struct anon_vma_chain *avc,
> >  	__anon_rmap_tree_remove(avc, &anon_vma->rb_root);
> >  }
> >
> > +/**
> > + * anon_rmap_tree_update_inplace() - Update anon rmap tree to reflect an
> > + * in-place change in the size of @avc's VMA.
> > + * @avc: The anon_vma_chain whose VMA's size has changed.
> > + *
> > + * The anon rmap root lock must be held.
> > + *
> > + * Invalid to do so if the VMA's anonymous pgoff has changed.
> > + */
> > +void anon_rmap_tree_update_inplace(struct anon_vma_chain *avc)
> > +{
> > +#ifdef CONFIG_DEBUG_VM_RB
> > +	avc->cached_vma_last = avc_last_pgoff(avc);
> > +#endif
> > +	/* Propagate all the way up the tree. */
> > +	__anon_rmap_tree_augment.propagate(&avc->rb, NULL);
> > +}
> > +
> >  struct anon_vma_chain *
> >  anon_rmap_tree_iter_first(struct anon_vma *anon_vma,
> >  			  pgoff_t pgoff_start, pgoff_t pgoff_last)
> > diff --git a/mm/vma.c b/mm/vma.c
> > index 077e23694143..8b333ec0c958 100644
> > --- a/mm/vma.c
> > +++ b/mm/vma.c
> > @@ -201,8 +201,15 @@ static void init_multi_vma_prep(struct vma_prepare *vp,
> >  	if (vp->file)
> >  		vp->mapping = vma->vm_file->f_mapping;
> >
> > -	if (vmg && vmg->skip_vma_uprobe)
> > +	if (!vmg)
> > +		return;
> > +
> > +	if (vmg->skip_vma_uprobe)
> >  		vp->skip_vma_uprobe = true;
> > +	if (vma_start_pgoff(vma) == vmg_start_pgoff(vmg))
> > +		vp->pgoff_unchanged = true;
>
> file_pgoff_unchanged perhaps? since we're distinguishing.

Yeah makes sense.

>
> In any case, I would prefer if we moved this logic to mm/interval_tree.c, or
> any rmap-related header. I don't love opencoding rmap tree assumptions this
> deep in vma.c code. WDYT?

Sure that's fair, will put it over there!

>
>
> Rest obviously looks great to me :)

Thanks :)

>
> > +	if (vma_start_anon_pgoff(vma) == vmg_start_anon_pgoff(vmg))
> > +		vp->anon_pgoff_unchanged = true;
> >  }
> >
> >  /*
> > @@ -331,6 +338,15 @@ anon_rmap_tree_post_update_vma(struct vm_area_struct *vma)
> >  		anon_rmap_tree_insert(avc, avc->anon_vma);
> >  }
> >
> > +static void
> > +anon_rmap_tree_update_vma_inplace(struct vm_area_struct *vma)
> > +{
> > +	struct anon_vma_chain *avc;
> > +
> > +	list_for_each_entry(avc, &vma->anon_vma_chain, same_vma)
> > +		anon_rmap_tree_update_inplace(avc);
> > +}
> > +
> >  /*
> >   * vma_prepare() - Helper function for handling locking VMAs prior to altering
> >   * @vp: The initialized vma_prepare struct
> > @@ -359,14 +375,16 @@ static void vma_prepare(struct vma_prepare *vp)
> >
> >  	if (vp->anon_vma) {
> >  		anon_vma_lock_write(vp->anon_vma);
> > -		anon_rmap_tree_pre_update_vma(vp->vma);
> > +		if (!vp->anon_pgoff_unchanged)
> > +			anon_rmap_tree_pre_update_vma(vp->vma);
> >  		if (vp->adj_next)
> >  			anon_rmap_tree_pre_update_vma(vp->adj_next);
> >  	}
> >
> >  	if (vp->file) {
> >  		flush_dcache_mmap_lock(vp->mapping);
> > -		mapping_rmap_tree_remove(vp->vma, vp->mapping);
> > +		if (!vp->pgoff_unchanged)
> > +			mapping_rmap_tree_remove(vp->vma, vp->mapping);
> >  		if (vp->adj_next)
> >  			mapping_rmap_tree_remove(vp->adj_next, vp->mapping);
> >  	}
> > @@ -387,7 +405,11 @@ static void vma_complete(struct vma_prepare *vp, struct vma_iterator *vmi,
> >  	if (vp->file) {
> >  		if (vp->adj_next)
> >  			mapping_rmap_tree_insert(vp->adj_next, vp->mapping);
> > -		mapping_rmap_tree_insert(vp->vma, vp->mapping);
> > +		/* Need only propagate the change inplace. */
> > +		if (vp->pgoff_unchanged)
> > +			mapping_rmap_tree_update_inplace(vp->vma);
> > +		else
> > +			mapping_rmap_tree_insert(vp->vma, vp->mapping);
>
> And perhaps similar for this, hiding update vs re-insert in interval tree
> code (via a helper) sounds cleaner.

OK will see how that looks!

>
> --
> Pedro

--
Cheers, Lorenzo

  reply	other threads:[~2026-09-28 17:01 UTC|newest]

Thread overview: 7+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-25 18:32 Lorenzo Stoakes (ARM)
2026-09-25 18:59 ` Rik van Riel
2026-09-28 15:10 ` Lorenzo Stoakes (ARM)
2026-09-28 15:37 ` Pedro Falcato
2026-09-28 17:01   ` Lorenzo Stoakes (ARM) [this message]
2026-09-29 13:45     ` Deng, Pan
2026-09-29 14:08       ` Lorenzo Stoakes (ARM)

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=arqczL8K39qBoqgV@gremlin \
    --to=ljs@kernel.org \
    --cc=akpm@linux-foundation.org \
    --cc=david@kernel.org \
    --cc=harry@kernel.org \
    --cc=jannh@google.com \
    --cc=lance.yang@linux.dev \
    --cc=liam@infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=mhocko@suse.com \
    --cc=pan.deng@intel.com \
    --cc=pfalcato@suse.de \
    --cc=riel@surriel.com \
    --cc=rppt@kernel.org \
    --cc=surenb@google.com \
    --cc=vbabka@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®