mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Rik van Riel <riel@surriel.com>
To: Yuan-Hao Hsu <aa9736195201@gmail.com>,
	Andrew Morton <akpm@linux-foundation.org>,
	David Hildenbrand <david@kernel.org>
Cc: Jason Gunthorpe <jgg@ziepe.ca>,
	John Hubbard <jhubbard@nvidia.com>,  Peter Xu <peterx@redhat.com>,
	Lorenzo Stoakes <ljs@kernel.org>,
	liam@infradead.org, Vlastimil Babka	 <vbabka@kernel.org>,
	Mike Rapoport <rppt@kernel.org>,
	Suren Baghdasaryan	 <surenb@google.com>,
	Michal Hocko <mhocko@suse.com>,
	Aristeu Rozanski	 <aris@ruivo.org>,
	Ryan Roberts <ryan.roberts@arm.com>, Dev Jain <dev.jain@arm.com>,
		linux-mm@kvack.org, linux-kernel@vger.kernel.org
Subject: Re: [PATCH] mm/gup: batch PTE-mapped large folios in gup_fast_pte_range()
Date: Wed, 07 Oct 2026 06:47:06 -0400	[thread overview]
Message-ID: <7dd4ec994502233cdc09fe5c4369adc89168d738.camel@surriel.com> (raw)
In-Reply-To: <20260917083747.786-1-aa9736195201@gmail.com>

On Thu, 2026-09-17 at 16:37 +0800, Yuan-Hao Hsu wrote:
> GUP-fast grabs a PTE-mapped large folio one page at a time.  Every
> PTE
> costs a try_grab_folio_fast() (a refcount cmpxchg, plus the pincount
> atomic, a full barrier and a node stat update for FOLL_PIN), a
> gup_fast_folio_allowed() and a folio_set_referenced(), so a 64 kB
> mTHP
> pays sixteen of each and a PTE-mapped 2 MB THP pays 512.  The PMD and
> PUD leaf paths already take the whole range with one
> try_grab_folio_fast() call, and the slow path is getting the same
> treatment for PTEs in Rik's follow_page_mask() series.
> 

> fio O_DIRECT randread of a null_blk device (CPU bound) with the I/O
> buffer in 64 kB mTHP, median of 15 five-second runs:
> 
>   bs=1M psync                 126.2/125.6 GB/s  ->  220.6 GB/s
>   bs=1M io_uring, iodepth 16   71.2/71.6 GB/s   ->  112.1 GB/s
>   bs=64k io_uring, iodepth 16  50.2/50.7 GB/s   ->   60.1 GB/s
> 
That's huge! Nice.

> 
> +static unsigned int gup_fast_pte_batch(struct folio *folio,
> +		struct page *page, pmd_t pmd, pmd_t *pmdp, pte_t
> *ptep,
> +		pte_t pte, unsigned int flags, unsigned int max_nr,
> +		struct page **pages)
> +{
> +	unsigned int nr, i;
> +
> +	if (max_nr == 1)
> +		return 1;
> +
> +	/*
> +	 * The reference on @folio keeps folio_nr_pages() stable. 
> The PTEs
> +	 * are read without the PTL, hence FPB_LOCKLESS.
> +	 */
> +	nr = folio_pte_batch_flags(folio, NULL, ptep, &pte, max_nr,
> +				   FPB_RESPECT_WRITE |
> FPB_LOCKLESS);
> 
...
> +	if (try_grab_folio(folio, nr - 1, flags))
> +		return 1;
> 

Since this is lockless, the PTEs could change under
us between when we walk them in folio_pte_batch_flags()
and the batched try_grab_folio().

Is that race condition safe, or do we need to re-check
afterward that:
1) the ptes still point at the same folio, and/or
2) the folio still spans the entire range, and was
   not split before we grabbed the refcount?

The equivalent changes to GUP do not need that check,
because that all happens under the pte_lock:

https://lore.kernel.org/all/20260811025157.1632867-1-riel@surriel.com/

-- 
All Rights Reversed.

      parent reply	other threads:[~2026-10-07 10:47 UTC|newest]

Thread overview: 7+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-17  8:37 Yuan-Hao Hsu
2026-09-17  8:42 ` David Hildenbrand (Arm)
2026-09-17  8:43   ` David Hildenbrand (Arm)
2026-09-17  9:01     ` Yuan-Hao Hsu
2026-10-07 10:07   ` Rik van Riel
2026-10-07 10:14     ` David Hildenbrand (Arm)
2026-10-07 10:47 ` Rik van Riel [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=7dd4ec994502233cdc09fe5c4369adc89168d738.camel@surriel.com \
    --to=riel@surriel.com \
    --cc=aa9736195201@gmail.com \
    --cc=akpm@linux-foundation.org \
    --cc=aris@ruivo.org \
    --cc=david@kernel.org \
    --cc=dev.jain@arm.com \
    --cc=jgg@ziepe.ca \
    --cc=jhubbard@nvidia.com \
    --cc=liam@infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=mhocko@suse.com \
    --cc=peterx@redhat.com \
    --cc=rppt@kernel.org \
    --cc=ryan.roberts@arm.com \
    --cc=surenb@google.com \
    --cc=vbabka@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®