From: Andrew Morton <akpm@linux-foundation.org>
To: Breno Leitao <leitao@debian.org>
Cc: David Hildenbrand <david@kernel.org>, Zi Yan <ziy@nvidia.com>,
Matthew Brost <matthew.brost@intel.com>,
Joshua Hahn <joshua.hahnjy@gmail.com>,
Rakie Kim <rakie.kim@sk.com>, Byungchul Park <byungchul@sk.com>,
Gregory Price <gourry@gourry.net>,
Ying Huang <ying.huang@linux.alibaba.com>,
Alistair Popple <apopple@nvidia.com>,
paulmck@kernel.org, linux-mm@kvack.org,
linux-kernel@vger.kernel.org, kernel-team@meta.com
Subject: Re: [PATCH] mm/migrate: report RCU-tasks quiescent states in migrate_pages_batch()
Date: Mon, 27 Jul 2026 13:49:50 -0700 [thread overview]
Message-ID: <20260727134950.2d091094dbb9bb3ecb7a3eb6@linux-foundation.org> (raw)
In-Reply-To: <20260727-kcompact-v1-1-bdfefddd6874@debian.org>
On Mon, 27 Jul 2026 06:50:19 -0700 Breno Leitao <leitao@debian.org> wrote:
> migrate_pages_batch() unmaps each folio before moving it, and every
> unmap runs the mmu_notifier invalidate callbacks. On KVM hosts
> try_to_migrate() ends up in kvm_mmu_notifier_invalidate_range_start() ->
> tdp_mmu_zap_leafs(), which is expensive, so unmapping a large batch keeps
> the CPU busy for a long time.
>
> The loop already calls cond_resched(), but on PREEMPTION kernels that is
> a no-op, and involuntary preemption is not a Tasks-RCU quiescent state.
>
> A long batch therefore never reports a quiescent state, and the
> migrating task (e.g. kcompactd) becomes a Tasks-RCU holdout, stalling the
> Tasks-RCU grace period for minutes, which is common at Meta fleet:
>
> INFO: rcu_tasks detected stalls on tasks:
> 0000000055349ecc: .. nvcsw: 1157401/1157401 holdout: 1 idle_cpu: -1/56 task:kcompactd0 state:R running task
> Call Trace:
> tdp_mmu_zap_leafs
> tdp_mmu_next_root
> gfn_to_pfn_cache_invalidate_start
> kvm_mmu_notifier_invalidate_range_start
> __mmu_notifier_invalidate_range_start
> try_to_migrate_one
> try_to_migrate
> migrate_pages_batch
> migrate_pages
> compact_zone
> compact_node
> kcompactd
> kthread
Well I doubt if users of 7.1 kernels and earlier want to see this. So
a Fixes: and a cc:stable are needed. The affected code is quite old
and might even predate the addition of cond_resched_tasks_rcu_qs(). So
I can't begin to suggest a Fixes: target. Maybe omit it and let the
-stable team figure it out ;)
> --- a/mm/migrate.c
> +++ b/mm/migrate.c
> @@ -1843,7 +1843,7 @@ static int migrate_pages_batch(struct list_head *from,
> is_thp = folio_test_pmd_mappable(folio);
> nr_pages = folio_nr_pages(folio);
>
> - cond_resched();
> + cond_resched_tasks_rcu_qs();
>
Totally off-topic but why the heck was that implemented as a macro.
Which invokes another macro and another and another and turtles all the
way down. End result:
do { do { if (!((false)) && ({ do { __attribute__((__noreturn__)) extern void __compiletime_assert_606(void) __attribute__((__error__("Unsupported access size for {READ,WRITE}_ONCE()."))); if (!((sizeof(((current))->rcu_tasks_holdout) == sizeof(char) || sizeof(((current))->rcu_tasks_holdout) == sizeof(short) || sizeof(((current))->rcu_tasks_holdout) == sizeof(int) || sizeof(((current))->rcu_tasks_holdout) == sizeof(long)) || sizeof(((current))->rcu_tasks_holdout) == sizeof(long long))) __compiletime_assert_606(); } while (0); (*(const volatile __typeof_unqual__(((current))->rcu_tasks_holdout) *)&(((current))->rcu_tasks_holdout)); })) do { do { __attribute__((__noreturn__)) extern void __compiletime_assert_607(void) __attribute__((__error__("Unsupported access size for {READ,WRITE}_ONCE()."))); if (!((sizeof(((current))->rcu_tasks_holdout) == sizeof(char) || sizeof(((current))->rcu_tasks_holdout) == sizeof(short) || sizeof(((current))->rcu_tasks_holdout) == sizeof(int) || sizeof(((
current))->rcu_tasks_holdout) == sizeof(long)) || sizeof(((current))->rcu_tasks_holdout) == sizeof(long long))) __compiletime_assert_607(); } while (0); do { *(volatile typeof(((current))->rcu_tasks_holdout) *)&(((current))->rcu_tasks_holdout) = (false); } while (0); } while (0); } while (0); ({ __might_resched("mm/migrate.c", 1846, 0); _cond_resched(); }); } while (0);
How much nicer would it be to have
static inline void cond_resched_tasks_rcu_qs(void)
{
if (some brief and efficient test)
some_slow_uninlined_thing();
}
next prev parent reply other threads:[~2026-07-27 20:49 UTC|newest]
Thread overview: 7+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-27 13:50 Breno Leitao
2026-07-27 14:24 ` Zi Yan
2026-07-27 15:57 ` Gregory Price
2026-07-27 18:39 ` Paul E. McKenney
2026-07-27 20:49 ` Andrew Morton [this message]
2026-07-27 21:14 ` Paul E. McKenney
2026-07-29 9:21 ` David Hildenbrand (Arm)
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260727134950.2d091094dbb9bb3ecb7a3eb6@linux-foundation.org \
--to=akpm@linux-foundation.org \
--cc=apopple@nvidia.com \
--cc=byungchul@sk.com \
--cc=david@kernel.org \
--cc=gourry@gourry.net \
--cc=joshua.hahnjy@gmail.com \
--cc=kernel-team@meta.com \
--cc=leitao@debian.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=matthew.brost@intel.com \
--cc=paulmck@kernel.org \
--cc=rakie.kim@sk.com \
--cc=ying.huang@linux.alibaba.com \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
Powered by JetHome