mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: JP Kobryn <jp.kobryn@linux.dev>
To: Hugh Dickins <hughd@google.com>,
	Andrew Morton <akpm@linux-foundation.org>
Cc: Ackerley Tng <ackerleytng@google.com>,
	Alexander Viro <viro@zeniv.linux.org.uk>,
	Alexandre Ghiti <alex@ghiti.fr>,
	Baolin Wang <baolin.wang@linux.alibaba.com>,
	Barry Song <baohua@kernel.org>,
	Binbin Wu <binbin.wu@linux.intel.com>,
	Christian Brauner <brauner@kernel.org>,
	Christoph Hellwig <hch@lst.de>, Christoph Lameter <cl@gentwo.org>,
	Claudio Imbrenda <imbrenda@linux.ibm.com>,
	David Hildenbrand <david@kernel.org>, Jan Kara <jack@suse.cz>,
	Jens Axboe <axboe@kernel.dk>,
	Johannes Weiner <hannes@cmpxchg.org>,
	Kairui Song <ryncsn@gmail.com>, Kiryl Shutsemau <kas@kernel.org>,
	Lance Yang <lance.yang@linux.dev>,
	Leonardo Bras <leobras.c@gmail.com>,
	Lorenzo Stoakes <ljs@kernel.org>,
	Marcelo Tosatti <mtosatti@redhat.com>,
	Matthew Wilcox <willy@infradead.org>,
	Mel Gorman <mgorman@techsingularity.net>,
	Miaohe Lin <linmiaohe@huawei.com>, Michal Hocko <mhocko@suse.com>,
	Minchan Kim <minchan@kernel.org>,
	Muchun Song <muchun.song@linux.dev>,
	Oscar Salvador <osalvador@suse.de>,
	Peter Zijlstra <peterz@infradead.org>,
	Qi Zheng <qi.zheng@linux.dev>, Rik van Riel <riel@surriel.com>,
	Sebastian Andrzej Siewior <bigeasy@linutronix.de>,
	Shakeel Butt <shakeel.butt@linux.dev>,
	Suren Baghdasaryan <surenb@google.com>,
	Vlastimil Babka <vbabka@kernel.org>,
	Yang Shi <yang@os.amperecomputing.com>,
	Yu Zhao <yuzhao@google.com>, Zach O'Keefe <zokeefe@google.com>,
	Zi Yan <ziy@nvidia.com>,
	linux-block@vger.kernel.org, linux-fsdevel@vger.kernel.org,
	linux-kernel@vger.kernel.org, linux-mm@kvack.org
Subject: Re: [PATCH v2 00/26] mm/fbatch: drain lru_add_drain() and _all()
Date: Fri, 25 Sep 2026 19:07:35 -0700	[thread overview]
Message-ID: <064255ee-c8be-4a82-8be0-fa209eac8637@linux.dev> (raw)
In-Reply-To: <e28f9a94-4339-f8ac-8301-6be3c9b5b7ce@google.com>

Hi Hugh,

On 9/9/26 2:39 AM, Hugh Dickins wrote:
> [PATCH v2 00/26] mm/fbatch: drain lru_add_drain() and _all()
> 
> This series was prompted by lru_add_drain_all() appearing in watchdog
> backtraces: not to blame, but blocked on an unresponsive CPU to run
> its workqueue. lru_add_drain_all() is a heavyweight operation, which
> we often hope to avoid by a much lighter lru_add_drain(); but even
> those local drains can aggravate the lruvec lock contention which
> per-cpu fbatches are intended to ease.
> 
> Now let the folios on the lru_add and other per-cpu fbatches remain
> there, isolatable, with PG_lru set and refcount unraised, just like
> when on an actual LRU. Then many calls to lru_add_drain() and _all()
> can be removed.
> 
> It's something I've wanted to do for years, tried several times, but
> only hit on a good way to do it a few weeks ago: now it's uncomfortable
> to watch others wrestling with the drainage, while I'm sitting on this.
> 
> The key idea came from the "if (!folio_test_clear_lru(folio)) continue;".
> If we don't mind missing to take an action on rare occasions, then maybe
> we won't mind taking an action on the wrong folio on rare occasions, so
> long as it is a folio consenting to PG_lru rules.
> 
> No new locking, but relies on folio_try_get() and folio_test_clear_lru()
> even on the lru_add fbatch; with use of bits not set in aligned pointers,
> and some try_cmpxchg()ing.  Speculative references to folios are already
> accepted: this adds another source of them.
> 
> Performance? You (and the bots) tell me. I'm considering this as a
> cleanup, to make life easier for developers. I expect that some loads
> will show improvement, but also expect some disappointments (perhaps
> I go too far against lru_cache_disable()? or not far enough).
> 
I found one source of regression but I think it's fixable.

Using a test of 4 processes repeatedly write faulting anon pages then
discarding with MADV_DONTNEED, the perf A/B below shows that the series
(minus patch 27) spends more time in __page_cache_release due to new
lru_lock traffic.

baseline (7891fbb9512f):
  0.03%  folio-churn  [kernel.kallsyms]  [k] queued_spin_lock_slowpath
  0.26%  folio-churn  [kernel.kallsyms]  [k] _raw_spin_lock_irqsave
  0.01%  folio-churn  [kernel.kallsyms]  [k] folio_lruvec_lock_irqsave
  0.00%  folio-churn  [kernel.kallsyms]  [k] folio_lruvec_relock_irqsave
  0.03%  folio-churn  [kernel.kallsyms]  [k] __page_cache_release

series (without patch 27):
      6.97%  folio-churn  [kernel.kallsyms]  [k] _raw_spin_lock_irqsave
             |
             ---_raw_spin_lock_irqsave
                |
                 --6.96%--folio_lruvec_lock_irqsave
                           folio_lruvec_relock_irqsave
                           |
                            --6.72%--__page_cache_release
                                      folios_put_refs
                                      free_pages_and_swap_cache
                                      tlb_flush_mmu
                                      tlb_finish_mmu
                                      do_madvise
                                      __x64_sys_madvise
                                      do_syscall_64
                                      entry_SYSCALL_64_after_hwframe
                                      __madvise

      6.62%  folio-churn  [kernel.kallsyms] [k] queued_spin_lock_slowpath
             |
             ---queued_spin_lock_slowpath
                _raw_spin_lock_irqsave
                folio_lruvec_lock_irqsave
                folio_lruvec_relock_irqsave
                |
                 --6.43%--__page_cache_release
                           folios_put_refs
                           free_pages_and_swap_cache
                           tlb_flush_mmu
                           tlb_finish_mmu
                           do_madvise
                           __x64_sys_madvise
                           do_syscall_64
                           entry_SYSCALL_64_after_hwframe
                           __madvise

There was a pre-existing dead folio optimization [0] which allowed
caller-released folios remaining on batches to be cleaned up at drain
time without lru_lock acquisition. This series removes it which makes
sense since the batch reference concept is gone. But now that callers
effectively invoke the cleanup when refcounts go directly to zero,
there's a point in which an lru_lock will be taken for pending folios
that are dead (refcount zero but still on batch). That's seen in the
call graphs above and is unnecessary since these folios haven't been
inserted into the LRU. Luckily It looks like the new API call
lru_add_del_folio() can be used to avoid this, since it returns true
after cleaning up the pending LRU batch state. I applied this diff and
re-ran the same benchmark. The new locking overhead was nearly all
eliminated. The perf output afterward shows the fractional amount that
remains.

@@ -64,8 +64,10 @@ static void __page_cache_release(struct folio *folio, 
struct lruvec **lruvecp,
                 unsigned long *flagsp)
  {
         if (folio_test_lru(folio)) {
-               folio_lruvec_relock_irqsave(folio, lruvecp, flagsp);
-               lruvec_del_folio(*lruvecp, folio);
+               if (!lru_add_del_folio(folio)) {
+                       folio_lruvec_relock_irqsave(folio, lruvecp, flagsp);
+                       lruvec_del_folio(*lruvecp, folio);
+               }
                 __folio_clear_lru_flags(folio);
         }
  }

series modified with patch above (without patch 27):
      0.77%  folio-churn  [kernel.kallsyms]  [k] queued_spin_lock_slowpath
             |
             ---queued_spin_lock_slowpath
                _raw_spin_lock_irqsave
                |
                 --0.77%--folio_lruvec_lock_irqsave
                           folio_lruvec_relock_irqsave
                           |
                            --0.52%--__page_cache_release
                                      folios_put_refs
                                      free_pages_and_swap_cache
                                      tlb_flush_mmu
                                      tlb_finish_mmu
                                      do_madvise
                                      __x64_sys_madvise
                                      do_syscall_64
                                      entry_SYSCALL_64_after_hwframe
                                      __madvise

      0.37% folio-churn [kernel.kallsyms] [k] _raw_spin_lock_irqsave
      0.04% folio-churn [kernel.kallsyms] [k] folio_lruvec_lock_irqsave
      0.01% folio-churn [kernel.kallsyms] [k] folio_lruvec_relock_irqsave

[0] 
https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/mm?id=9669b87065a6fe96198f3df2c3d125c5f5c1f210

      parent reply	other threads:[~2026-09-26  2:07 UTC|newest]

Thread overview: 46+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-09  9:39 Hugh Dickins
2026-09-09  9:42 ` [PATCH v2 01/26] mm/fbatch: remove !CONFIG_SMP special case of folio_activate() Hugh Dickins
2026-09-09  9:44 ` [PATCH v2 02/26] mm/fbatch: allow folios_put_refs() to skip xa_is_value() entries Hugh Dickins
2026-09-09  9:46 ` [PATCH v2 03/26] mm/fbatch: temporarily disable lazyfree and mlock+munlock batching Hugh Dickins
2026-09-09  9:49 ` [PATCH v2 04/26] mm/fbatch: lru bit set, no extra ref, while folio on per-cpu fbatch Hugh Dickins
2026-09-09 12:25   ` Vlastimil Babka (SUSE)
2026-09-12 19:30     ` Hugh Dickins
2026-09-09  9:51 ` [PATCH v2 05/26] mm/fbatch: lru_add_del_folio()+folio_add_lru() after clear_lru() Hugh Dickins
2026-09-09 15:04   ` Vlastimil Babka (SUSE)
2026-09-12 21:59     ` Hugh Dickins
2026-09-09  9:53 ` [PATCH v2 06/26] mm/fbatch: fbatch_drain_lazyfree(onstack fbatch) before ptl unlock Hugh Dickins
2026-09-09 21:02   ` Vlastimil Babka (SUSE)
2026-09-12 22:07     ` Hugh Dickins
2026-09-09  9:55 ` [PATCH v2 07/26] mm/fbatch: LRU_NEXT_ACTIVATE bit to optimize folio_activate() Hugh Dickins
2026-09-10 12:01   ` Vlastimil Babka (SUSE)
2026-09-12 22:35     ` Hugh Dickins
2026-09-14 19:55       ` Hugh Dickins
2026-09-10 16:42   ` Kiryl Shutsemau
2026-09-12 23:33     ` Hugh Dickins
2026-09-14 20:19       ` Hugh Dickins
2026-09-15 13:08         ` Kiryl Shutsemau
2026-09-09  9:57 ` [PATCH v2 08/26] mm/fbatch: replace mlock_new_folio() by __folio_add_lru(,mlockit) Hugh Dickins
2026-09-10 17:47   ` Vlastimil Babka (SUSE)
2026-09-09  9:59 ` [PATCH v2 09/26] mm/fbatch: restore mlock+munlock batching, without extra ref Hugh Dickins
2026-09-10 21:05   ` Vlastimil Babka (SUSE)
2026-09-12 23:46     ` Hugh Dickins
2026-09-14  8:16       ` Vlastimil Babka (SUSE)
2026-09-09 10:01 ` [PATCH v2 10/26] mm/fbatch: remove several uses of mlock_drain_local() Hugh Dickins
2026-09-09 10:03 ` [PATCH v2 11/26] mm/fbatch: remove migration's PAGE_WAS_MLOCKED lru_add_drain() Hugh Dickins
2026-09-09 10:05 ` [PATCH v2 12/26] mm/fbatch: remove percpu_pvec_drained and folios_put() Hugh Dickins
2026-09-09 10:08 ` [PATCH v2 13/26] mm/fbatch: no lru_add_drain to collect_longterm_unpinnable_folios() Hugh Dickins
2026-09-09 10:10 ` [PATCH v2 14/26] mm/fbatch: no lru_add_drain() nor _all() for memfd_wait_for_pins() Hugh Dickins
2026-09-09 10:12 ` [PATCH v2 15/26] mm/fbatch: remove shake_folio() shake_page() from memory-failure Hugh Dickins
2026-09-09 10:14 ` [PATCH v2 16/26] mm/fbatch: remove lru_cache_disable() from NUMA folio migration Hugh Dickins
2026-09-09 10:16 ` [PATCH v2 17/26] mm/fbatch: no lru_cache_disable() in __alloc_contig_migrate_range() Hugh Dickins
2026-09-09 10:18 ` [PATCH v2 18/26] mm/fbatch: remove lru_add_drain() and _all() calls from various Hugh Dickins
2026-09-09 10:20 ` [PATCH v2 19/26] mm/fbatch: vm/stat_refresh include lru_add_drain() on each cpu Hugh Dickins
2026-09-09 10:23 ` [PATCH v2 20/26] s390/fbatch: no lru_add_drain_all() in s390_wiggle_split_folio() Hugh Dickins
2026-09-09 10:25 ` [PATCH v2 21/26] block/fbatch: no lru_add_drain_all() in invalidate_bdev() Hugh Dickins
2026-09-09 10:27 ` [PATCH v2 22/26] fs/fbatch: drop_caches invalidate_bh_lrus() not lru_add_drain_all() Hugh Dickins
2026-09-09 10:30 ` [PATCH v2 23/26] fs,mm/fbatch: use invalidate_bh_lrus() not invalidate_bh_lrus_cpu() Hugh Dickins
2026-09-09 10:33 ` [PATCH v2 24/26] fs,mm/fbatch: lru_cache_disable() keep off buffer_head lrus only Hugh Dickins
2026-09-09 10:35 ` [PATCH v2 25/26] mm/fbatch: move lru_add_drain_all() declaration to mm/internal.h Hugh Dickins
2026-09-09 10:37 ` [PATCH v2 26/26] mm/fbatch: drop reference inside the loop when draining Hugh Dickins
2026-09-09 10:41 ` [PATCH v2 27/26] mm/fbatch: paranoid folio vmstats in folio_batch_move_lru() Hugh Dickins
2026-09-26  2:07 ` JP Kobryn [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=064255ee-c8be-4a82-8be0-fa209eac8637@linux.dev \
    --to=jp.kobryn@linux.dev \
    --cc=ackerleytng@google.com \
    --cc=akpm@linux-foundation.org \
    --cc=alex@ghiti.fr \
    --cc=axboe@kernel.dk \
    --cc=baohua@kernel.org \
    --cc=baolin.wang@linux.alibaba.com \
    --cc=bigeasy@linutronix.de \
    --cc=binbin.wu@linux.intel.com \
    --cc=brauner@kernel.org \
    --cc=cl@gentwo.org \
    --cc=david@kernel.org \
    --cc=hannes@cmpxchg.org \
    --cc=hch@lst.de \
    --cc=hughd@google.com \
    --cc=imbrenda@linux.ibm.com \
    --cc=jack@suse.cz \
    --cc=kas@kernel.org \
    --cc=lance.yang@linux.dev \
    --cc=leobras.c@gmail.com \
    --cc=linmiaohe@huawei.com \
    --cc=linux-block@vger.kernel.org \
    --cc=linux-fsdevel@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=mgorman@techsingularity.net \
    --cc=mhocko@suse.com \
    --cc=minchan@kernel.org \
    --cc=mtosatti@redhat.com \
    --cc=muchun.song@linux.dev \
    --cc=osalvador@suse.de \
    --cc=peterz@infradead.org \
    --cc=qi.zheng@linux.dev \
    --cc=riel@surriel.com \
    --cc=ryncsn@gmail.com \
    --cc=shakeel.butt@linux.dev \
    --cc=surenb@google.com \
    --cc=vbabka@kernel.org \
    --cc=viro@zeniv.linux.org.uk \
    --cc=willy@infradead.org \
    --cc=yang@os.amperecomputing.com \
    --cc=yuzhao@google.com \
    --cc=ziy@nvidia.com \
    --cc=zokeefe@google.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®