mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v3] jbd2: fix shrinker scan budget accounting in jbd2_journal_shrink_scan()
@ 2026-09-28 14:44 Qiliang Yuan
  2026-09-29 21:05 ` Jan Kara
  2026-09-30  3:31 ` Zhang Yi
  0 siblings, 2 replies; 3+ messages in thread
From: Qiliang Yuan @ 2026-09-28 14:44 UTC (permalink / raw)
  To: Theodore Ts'o, Jan Kara
  Cc: linux-ext4, linux-kernel, stable, Qiliang Yuan

jbd2_journal_shrink_checkpoint_list() already tracks its examined
checkpoint buffers accurately: it decrements its nr_to_scan in/out
parameter for every journal_head walked, whether or not it gets freed.
jbd2_journal_shrink_scan(), the shrinker's scan_objects() callback,
never surfaces that into sc->nr_scanned.

Because sc->nr_scanned defaults to sc->nr_to_scan before every call,
do_shrink_slab() always assumes a full batch was examined and keeps
calling scan_objects() until the one-shot budget derived from the
(possibly stale) percpu checkpoint count is drained, even when every
buffer in the checkpoint list is busy and nothing gets freed.

Derive sc->nr_scanned from the difference between the nr_to_scan value
passed in and the value left behind by
jbd2_journal_shrink_checkpoint_list(). Trigger SHRINK_STOP on
nr_shrunk == 0 (nothing freed) instead of leaving do_shrink_slab() to
keep calling us: a checkpoint list full of buffers that are still busy
being written back can honestly report a full batch scanned while
freeing nothing, and retrying immediately within the same synchronous
reclaim pass will not make any of that in-flight writeback complete
sooner.

Tested by creating 20000 small files without an explicit sync (to let
jbd2's normal 5-second commit timer move them onto the checkpoint list
while their buffers are still busy being written back) and then
triggering "echo 2 > /proc/sys/vm/drop_caches", tracing jbd2_shrink_*:

              scan_objects() calls   calls with nr_shrunk == 0
  before             176                    176 (100%)
  after                8                      8 (100%)

Fixes: 4ba3fcdde7e3 ("jbd2,ext4: add a shrinker to release checkpointed buffers")
Cc: stable@vger.kernel.org
Signed-off-by: Qiliang Yuan <odys.yuan@gmail.com>
---
V2 -> V3:
- Rebase onto v7.3-rc5. It already includes Max Kellermann's
  "jbd2: bound shrinker scans by examined checkpoint buffers"
  (15cb16496446), which independently fixes the busy-buffer
  accounting in journal_shrink_one_cp_list()/
  jbd2_journal_shrink_checkpoint_list() that v2 also touched. Drop
  the now-redundant checkpoint.c changes; this revision only touches
  journal.c, reusing the accurate nr_to_scan tracking Kellermann's
  fix already provides.
- Add a Fixes: tag for the commit that introduced
  jbd2_journal_shrink_scan() without ever setting sc->nr_scanned.
- Re-measure test data against the new baseline (176 -> 8 calls,
  vs the old baseline's 168 -> 7).

V1 -> V2:
- Count examined buffers, not just freed ones, in
  journal_shrink_one_cp_list() (Sashiko AI review finding).
- Trigger SHRINK_STOP on nr_shrunk == 0 instead of
  sc->nr_scanned == 0.

v2: https://lore.kernel.org/r/20260928-fix-jbd2-shrink-scan-nr-scanned-v2-1-7e1efec3afdf@gmail.com
v1: https://lore.kernel.org/r/20260928-fix-jbd2-shrink-scan-nr-scanned-v1-1-e6f4016ec699@gmail.com
---
 fs/jbd2/journal.c | 17 +++++++++++++++++
 1 file changed, 17 insertions(+)

diff --git a/fs/jbd2/journal.c b/fs/jbd2/journal.c
index 00f5a98f3d4fe..a61b7a59c0a0e 100644
--- a/fs/jbd2/journal.c
+++ b/fs/jbd2/journal.c
@@ -1263,10 +1263,27 @@ static unsigned long jbd2_journal_shrink_scan(struct shrinker *shrink,
 	trace_jbd2_shrink_scan_enter(journal, sc->nr_to_scan, count);
 
 	nr_shrunk = jbd2_journal_shrink_checkpoint_list(journal, &nr_to_scan);
+	sc->nr_scanned = sc->nr_to_scan - nr_to_scan;
 
 	count = percpu_counter_read_positive(&journal->j_checkpoint_jh_count);
 	trace_jbd2_shrink_scan_exit(journal, nr_to_scan, nr_shrunk, count);
 
+	/*
+	 * Give up on this reclaim pass if this call didn't manage to free
+	 * anything. This is deliberately based on nr_shrunk, not on
+	 * sc->nr_scanned: a checkpoint list can be full of buffers that are
+	 * still busy being written back, in which case a call can
+	 * legitimately scan (and correctly report through sc->nr_scanned)
+	 * a full batch of them without freeing a single one. Retrying
+	 * immediately within the same synchronous reclaim pass is not going
+	 * to let any of that in-flight writeback complete any sooner, so
+	 * there is nothing to gain from letting do_shrink_slab() keep
+	 * calling us against the same stale freeable count until its
+	 * one-shot scan budget for this priority level is exhausted.
+	 */
+	if (nr_shrunk == 0)
+		return SHRINK_STOP;
+
 	return nr_shrunk;
 }
 

---
base-commit: 72d3fcf802c45d00b300f25b848a93c3a2bd7c7e
change-id: 20260928-fix-jbd2-shrink-scan-nr-scanned-c3098cd901e2

Best regards,
-- 
Qiliang Yuan <odys.yuan@gmail.com>


^ permalink raw reply	[flat|nested] 3+ messages in thread

* Re: [PATCH v3] jbd2: fix shrinker scan budget accounting in jbd2_journal_shrink_scan()
  2026-09-28 14:44 [PATCH v3] jbd2: fix shrinker scan budget accounting in jbd2_journal_shrink_scan() Qiliang Yuan
@ 2026-09-29 21:05 ` Jan Kara
  2026-09-30  3:31 ` Zhang Yi
  1 sibling, 0 replies; 3+ messages in thread
From: Jan Kara @ 2026-09-29 21:05 UTC (permalink / raw)
  To: Qiliang Yuan
  Cc: Theodore Ts'o, Jan Kara, linux-ext4, linux-kernel, stable

On Mon 28-09-26 22:44:21, Qiliang Yuan wrote:
> jbd2_journal_shrink_checkpoint_list() already tracks its examined
> checkpoint buffers accurately: it decrements its nr_to_scan in/out
> parameter for every journal_head walked, whether or not it gets freed.
> jbd2_journal_shrink_scan(), the shrinker's scan_objects() callback,
> never surfaces that into sc->nr_scanned.
> 
> Because sc->nr_scanned defaults to sc->nr_to_scan before every call,
> do_shrink_slab() always assumes a full batch was examined and keeps
> calling scan_objects() until the one-shot budget derived from the
> (possibly stale) percpu checkpoint count is drained, even when every
> buffer in the checkpoint list is busy and nothing gets freed.
> 
> Derive sc->nr_scanned from the difference between the nr_to_scan value
> passed in and the value left behind by
> jbd2_journal_shrink_checkpoint_list(). Trigger SHRINK_STOP on
> nr_shrunk == 0 (nothing freed) instead of leaving do_shrink_slab() to
> keep calling us: a checkpoint list full of buffers that are still busy
> being written back can honestly report a full batch scanned while
> freeing nothing, and retrying immediately within the same synchronous
> reclaim pass will not make any of that in-flight writeback complete
> sooner.
> 
> Tested by creating 20000 small files without an explicit sync (to let
> jbd2's normal 5-second commit timer move them onto the checkpoint list
> while their buffers are still busy being written back) and then
> triggering "echo 2 > /proc/sys/vm/drop_caches", tracing jbd2_shrink_*:
> 
>               scan_objects() calls   calls with nr_shrunk == 0
>   before             176                    176 (100%)
>   after                8                      8 (100%)
> 
> Fixes: 4ba3fcdde7e3 ("jbd2,ext4: add a shrinker to release checkpointed buffers")
> Cc: stable@vger.kernel.org
> Signed-off-by: Qiliang Yuan <odys.yuan@gmail.com>

And you're back at v1 so my reviewed-by tag still holds :). Feel free to
add:

Reviewed-by: Jan Kara <jack@suse.cz>

								Honza

> ---
> V2 -> V3:
> - Rebase onto v7.3-rc5. It already includes Max Kellermann's
>   "jbd2: bound shrinker scans by examined checkpoint buffers"
>   (15cb16496446), which independently fixes the busy-buffer
>   accounting in journal_shrink_one_cp_list()/
>   jbd2_journal_shrink_checkpoint_list() that v2 also touched. Drop
>   the now-redundant checkpoint.c changes; this revision only touches
>   journal.c, reusing the accurate nr_to_scan tracking Kellermann's
>   fix already provides.
> - Add a Fixes: tag for the commit that introduced
>   jbd2_journal_shrink_scan() without ever setting sc->nr_scanned.
> - Re-measure test data against the new baseline (176 -> 8 calls,
>   vs the old baseline's 168 -> 7).
> 
> V1 -> V2:
> - Count examined buffers, not just freed ones, in
>   journal_shrink_one_cp_list() (Sashiko AI review finding).
> - Trigger SHRINK_STOP on nr_shrunk == 0 instead of
>   sc->nr_scanned == 0.
> 
> v2: https://lore.kernel.org/r/20260928-fix-jbd2-shrink-scan-nr-scanned-v2-1-7e1efec3afdf@gmail.com
> v1: https://lore.kernel.org/r/20260928-fix-jbd2-shrink-scan-nr-scanned-v1-1-e6f4016ec699@gmail.com
> ---
>  fs/jbd2/journal.c | 17 +++++++++++++++++
>  1 file changed, 17 insertions(+)
> 
> diff --git a/fs/jbd2/journal.c b/fs/jbd2/journal.c
> index 00f5a98f3d4fe..a61b7a59c0a0e 100644
> --- a/fs/jbd2/journal.c
> +++ b/fs/jbd2/journal.c
> @@ -1263,10 +1263,27 @@ static unsigned long jbd2_journal_shrink_scan(struct shrinker *shrink,
>  	trace_jbd2_shrink_scan_enter(journal, sc->nr_to_scan, count);
>  
>  	nr_shrunk = jbd2_journal_shrink_checkpoint_list(journal, &nr_to_scan);
> +	sc->nr_scanned = sc->nr_to_scan - nr_to_scan;
>  
>  	count = percpu_counter_read_positive(&journal->j_checkpoint_jh_count);
>  	trace_jbd2_shrink_scan_exit(journal, nr_to_scan, nr_shrunk, count);
>  
> +	/*
> +	 * Give up on this reclaim pass if this call didn't manage to free
> +	 * anything. This is deliberately based on nr_shrunk, not on
> +	 * sc->nr_scanned: a checkpoint list can be full of buffers that are
> +	 * still busy being written back, in which case a call can
> +	 * legitimately scan (and correctly report through sc->nr_scanned)
> +	 * a full batch of them without freeing a single one. Retrying
> +	 * immediately within the same synchronous reclaim pass is not going
> +	 * to let any of that in-flight writeback complete any sooner, so
> +	 * there is nothing to gain from letting do_shrink_slab() keep
> +	 * calling us against the same stale freeable count until its
> +	 * one-shot scan budget for this priority level is exhausted.
> +	 */
> +	if (nr_shrunk == 0)
> +		return SHRINK_STOP;
> +
>  	return nr_shrunk;
>  }
>  
> 
> ---
> base-commit: 72d3fcf802c45d00b300f25b848a93c3a2bd7c7e
> change-id: 20260928-fix-jbd2-shrink-scan-nr-scanned-c3098cd901e2
> 
> Best regards,
> -- 
> Qiliang Yuan <odys.yuan@gmail.com>
> 
-- 
Jan Kara <jack@suse.com>
SUSE Labs, CR

^ permalink raw reply	[flat|nested] 3+ messages in thread

* Re: [PATCH v3] jbd2: fix shrinker scan budget accounting in jbd2_journal_shrink_scan()
  2026-09-28 14:44 [PATCH v3] jbd2: fix shrinker scan budget accounting in jbd2_journal_shrink_scan() Qiliang Yuan
  2026-09-29 21:05 ` Jan Kara
@ 2026-09-30  3:31 ` Zhang Yi
  1 sibling, 0 replies; 3+ messages in thread
From: Zhang Yi @ 2026-09-30  3:31 UTC (permalink / raw)
  To: Qiliang Yuan
  Cc: Theodore Ts'o, Jan Kara, linux-ext4, linux-kernel, stable

On 9/28/2026 10:44 PM, Qiliang Yuan wrote:
> jbd2_journal_shrink_checkpoint_list() already tracks its examined
> checkpoint buffers accurately: it decrements its nr_to_scan in/out
> parameter for every journal_head walked, whether or not it gets freed.
> jbd2_journal_shrink_scan(), the shrinker's scan_objects() callback,
> never surfaces that into sc->nr_scanned.
> 
> Because sc->nr_scanned defaults to sc->nr_to_scan before every call,
> do_shrink_slab() always assumes a full batch was examined and keeps
> calling scan_objects() until the one-shot budget derived from the
> (possibly stale) percpu checkpoint count is drained, even when every
> buffer in the checkpoint list is busy and nothing gets freed.
> 
> Derive sc->nr_scanned from the difference between the nr_to_scan value
> passed in and the value left behind by
> jbd2_journal_shrink_checkpoint_list(). Trigger SHRINK_STOP on
> nr_shrunk == 0 (nothing freed) instead of leaving do_shrink_slab() to
> keep calling us: a checkpoint list full of buffers that are still busy
> being written back can honestly report a full batch scanned while
> freeing nothing, and retrying immediately within the same synchronous
> reclaim pass will not make any of that in-flight writeback complete
> sooner.
> 
> Tested by creating 20000 small files without an explicit sync (to let
> jbd2's normal 5-second commit timer move them onto the checkpoint list
> while their buffers are still busy being written back) and then
> triggering "echo 2 > /proc/sys/vm/drop_caches", tracing jbd2_shrink_*:
> 
>               scan_objects() calls   calls with nr_shrunk == 0
>   before             176                    176 (100%)
>   after                8                      8 (100%)
> 
> Fixes: 4ba3fcdde7e3 ("jbd2,ext4: add a shrinker to release checkpointed buffers")
> Cc: stable@vger.kernel.org
> Signed-off-by: Qiliang Yuan <odys.yuan@gmail.com>
> ---
> V2 -> V3:
> - Rebase onto v7.3-rc5. It already includes Max Kellermann's
>   "jbd2: bound shrinker scans by examined checkpoint buffers"
>   (15cb16496446), which independently fixes the busy-buffer
>   accounting in journal_shrink_one_cp_list()/
>   jbd2_journal_shrink_checkpoint_list() that v2 also touched. Drop
>   the now-redundant checkpoint.c changes; this revision only touches
>   journal.c, reusing the accurate nr_to_scan tracking Kellermann's
>   fix already provides.
> - Add a Fixes: tag for the commit that introduced
>   jbd2_journal_shrink_scan() without ever setting sc->nr_scanned.
> - Re-measure test data against the new baseline (176 -> 8 calls,
>   vs the old baseline's 168 -> 7).
> 
> V1 -> V2:
> - Count examined buffers, not just freed ones, in
>   journal_shrink_one_cp_list() (Sashiko AI review finding).
> - Trigger SHRINK_STOP on nr_shrunk == 0 instead of
>   sc->nr_scanned == 0.
> 
> v2: https://lore.kernel.org/r/20260928-fix-jbd2-shrink-scan-nr-scanned-v2-1-7e1efec3afdf@gmail.com
> v1: https://lore.kernel.org/r/20260928-fix-jbd2-shrink-scan-nr-scanned-v1-1-e6f4016ec699@gmail.com
> ---
>  fs/jbd2/journal.c | 17 +++++++++++++++++
>  1 file changed, 17 insertions(+)
> 
> diff --git a/fs/jbd2/journal.c b/fs/jbd2/journal.c
> index 00f5a98f3d4fe..a61b7a59c0a0e 100644
> --- a/fs/jbd2/journal.c
> +++ b/fs/jbd2/journal.c
> @@ -1263,10 +1263,27 @@ static unsigned long jbd2_journal_shrink_scan(struct shrinker *shrink,
>  	trace_jbd2_shrink_scan_enter(journal, sc->nr_to_scan, count);
>  
>  	nr_shrunk = jbd2_journal_shrink_checkpoint_list(journal, &nr_to_scan);
> +	sc->nr_scanned = sc->nr_to_scan - nr_to_scan;

Similar to your patch "ext4: fix shrinker scan budget accounting in
ext4_es_scan()", I'd tend to leave sc->nr_scanned alone. Your issue
should be resolved by simply returning SHRINK_STOP when nr_shrunk is
0, right or is there some other benefit to modifying nr_scanned?

Thanks,
Yi.
>  
>  	count = percpu_counter_read_positive(&journal->j_checkpoint_jh_count);
>  	trace_jbd2_shrink_scan_exit(journal, nr_to_scan, nr_shrunk, count);
>  
> +	/*
> +	 * Give up on this reclaim pass if this call didn't manage to free
> +	 * anything. This is deliberately based on nr_shrunk, not on
> +	 * sc->nr_scanned: a checkpoint list can be full of buffers that are
> +	 * still busy being written back, in which case a call can
> +	 * legitimately scan (and correctly report through sc->nr_scanned)
> +	 * a full batch of them without freeing a single one. Retrying
> +	 * immediately within the same synchronous reclaim pass is not going
> +	 * to let any of that in-flight writeback complete any sooner, so
> +	 * there is nothing to gain from letting do_shrink_slab() keep
> +	 * calling us against the same stale freeable count until its
> +	 * one-shot scan budget for this priority level is exhausted.
> +	 */
> +	if (nr_shrunk == 0)
> +		return SHRINK_STOP;
> +
>  	return nr_shrunk;
>  }
>  
> 
> ---
> base-commit: 72d3fcf802c45d00b300f25b848a93c3a2bd7c7e
> change-id: 20260928-fix-jbd2-shrink-scan-nr-scanned-c3098cd901e2
> 
> Best regards,


^ permalink raw reply	[flat|nested] 3+ messages in thread

end of thread, other threads:[~2026-09-30  3:31 UTC | newest]

Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-28 14:44 [PATCH v3] jbd2: fix shrinker scan budget accounting in jbd2_journal_shrink_scan() Qiliang Yuan
2026-09-29 21:05 ` Jan Kara
2026-09-30  3:31 ` Zhang Yi

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®