From: "Lorenzo Stoakes (ARM)" <ljs@kernel.org>
To: Gregory Price <gourry@gourry.net>
Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org,
kernel-team@meta.com, akpm@linux-foundation.org,
liam@infradead.org, david@kernel.org, vbabka@kernel.org,
jannh@google.com, wangjiexun@tinylab.org,
sashiko-bot <sashiko-bot@kernel.org>,
stable@vger.kernel.org
Subject: Re: [PATCH] mm/madvise: reclaim isolated folios if PTE restart fails
Date: Tue, 15 Sep 2026 16:55:47 +0100 [thread overview]
Message-ID: <aqloB0zm3RjJ47m1@gremlin> (raw)
In-Reply-To: <20260912110832.3203902-1-gourry@gourry.net>
On Sat, Sep 12, 2026 at 07:08:32AM -0400, Gregory Price wrote:
> MADV_PAGEOUT collects isolated folios on a local list before reclaiming
> them after the PTE walk. The reschedule path drops the PTE lock and then
> restarts the mapping with pte_offset_map_lock().
Is this reschedule path even necessary? It's pretty bloody sketchy.
I thought the modern approach (TM) was to not do cond_resched() and friends so
can we actually look at removing this?
>
> A concurrent operation can remove or replace the PTE table while the lock
> is dropped, causing pte_offset_map_lock() to return NULL. Returning directly
> in that case bypasses reclaim_pages(), leaving the collected folios off the
> LRU with elevated references.
>
> Route the failure through the existing cleanup path so any isolated folios
> are reclaimed or put back.
>
> Simplest userland pseudo-code reproducer:
>
> p = mmap(PMD_SIZE, ANONYMOUS);
> touch_every_page(p, PMD_SIZE);
> parallel {
> while (1) madvise(p, PMD_SIZE, MADV_PAGEOUT);
> while (1) {
> madvise(p, PMD_SIZE, MADV_DONTNEED);
> touch_every_page(p, PMD_SIZE);
> }
> }
>
> Reproduced in qemu trivially with some explicit widening of the race window.
>
> Fixes: b2f557a21bc8 ("mm/madvise: add cond_resched() in madvise_cold_or_pageout_pte_range()")
> Reported-by: sashiko-bot <sashiko-bot@kernel.org>
> Closes: https://sashiko.dev/#/patchset/20260821150912.183976-1-gourry@gourry.net
> Cc: <stable@vger.kernel.org>
> Assisted-by: LLM
> Signed-off-by: Gregory Price (Meta) <gourry@gourry.net>
> ---
> mm/madvise.c | 3 ++-
> 1 file changed, 2 insertions(+), 1 deletion(-)
>
> diff --git a/mm/madvise.c b/mm/madvise.c
> index 574aa2bb7c7e..ba5a3d77241a 100644
> --- a/mm/madvise.c
> +++ b/mm/madvise.c
> @@ -464,7 +464,7 @@ static int madvise_cold_or_pageout_pte_range(pmd_t *pmd,
This function is truly some of the absolute worst code I've read in
mm. Shocking.
> restart:
> start_pte = pte = pte_offset_map_lock(vma->vm_mm, pmd, addr, &ptl);
> if (!start_pte)
> - return 0;
> + goto out;
> flush_tlb_batched_pending(mm);
> lazy_mmu_mode_enable();
> for (; addr < end; pte += nr, addr += nr * PAGE_SIZE) {
> @@ -568,6 +568,7 @@ static int madvise_cold_or_pageout_pte_range(pmd_t *pmd,
> folio_deactivate(folio);
> }
>
> +out:
Seems odd to put this here for a condition that explicitly guarantees
!start_pte? It should be before the if (pageout) reclaim_pages(...);
surely?
I honestly wonder whether, rather than adding yet another label/goto into this
absolute bloody mess, whether we should just live with a bit of duplication and do:
if (!start_pte) {
if (pageout)
reclaim_pages(&folio_list);
return 0;
}
I actually think that'd be clearer at this point than throwing in some more
indirection.
That's for the backport but somebody needs to rework this entire bloody
function going forwards...
> if (start_pte) {
> lazy_mmu_mode_disable();
> pte_unmap_unlock(start_pte, ptl);
> --
> 2.55.0
>
--
Cheers, Lorenzo
next prev parent reply other threads:[~2026-09-15 15:55 UTC|newest]
Thread overview: 22+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-12 11:08 Gregory Price
2026-09-15 6:32 ` Andrew Morton
2026-09-15 14:13 ` Gregory Price
2026-09-15 15:55 ` Lorenzo Stoakes (ARM) [this message]
2026-09-15 17:00 ` Gregory Price
2026-09-16 13:18 ` Lorenzo Stoakes (ARM)
2026-09-16 14:02 ` Gregory Price
2026-09-16 14:08 ` Lorenzo Stoakes (ARM)
2026-09-15 17:12 ` Gregory Price
2026-09-16 13:03 ` Lorenzo Stoakes (ARM)
2026-09-16 13:57 ` Gregory Price
2026-09-16 14:12 ` Lorenzo Stoakes (ARM)
2026-09-16 14:42 ` Gregory Price
2026-09-16 14:48 ` David Hildenbrand (Arm)
2026-09-16 14:58 ` Gregory Price
2026-09-16 14:59 ` David Hildenbrand (Arm)
2026-09-16 15:19 ` Gregory Price
2026-09-16 15:21 ` Lorenzo Stoakes (ARM)
2026-09-16 15:24 ` David Hildenbrand (Arm)
2026-09-16 15:37 ` Gregory Price
2026-09-16 15:42 ` David Hildenbrand (Arm)
2026-09-16 16:10 ` Gregory Price
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=aqloB0zm3RjJ47m1@gremlin \
--to=ljs@kernel.org \
--cc=akpm@linux-foundation.org \
--cc=david@kernel.org \
--cc=gourry@gourry.net \
--cc=jannh@google.com \
--cc=kernel-team@meta.com \
--cc=liam@infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=sashiko-bot@kernel.org \
--cc=stable@vger.kernel.org \
--cc=vbabka@kernel.org \
--cc=wangjiexun@tinylab.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®