From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CDF7B3DD520 for ; Mon, 27 Jul 2026 20:49:51 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785185393; cv=none; b=GV9be9AaiS6ajNoThWwj7gRFMW4LRr0lH7KiYG5jKprPZesXx00us2axlZxX5EleZE892VeEIqlOqlnTRq0LF3R1KUTR5vhKp7Q2sWuayBJsoN9BNcRkd3V9CnmOBoatbR9RJ9kIYQ8ev3sBA3mS1qWHvhKsy3ST7tq8g/afAJc= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785185393; c=relaxed/simple; bh=5W174lbsIuFDjgrHjyOyteoaQx6u6McqYdEldcDtCTc=; h=Date:From:To:Cc:Subject:Message-Id:In-Reply-To:References: Mime-Version:Content-Type; b=XsJf4n+dV+7RrODsaaYNyeImdb632YyBTK0kVa6PqsMbx4HBPL5VmnEKXYqRprTOLpmdLcYVdIZMfIxLpdOpAc4YQCQATD44DXvc7Fbc6BbSf9VK7caB+gfNHsOBtPTVvdy/dJE3dqnuniwDJL1/SiORxDQlxLKBFlK4p6eS2u4= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dkim=fail (1024-bit key) header.d=linux-foundation.org header.i=@linux-foundation.org header.b=Km6SraPs reason="signature verification failed"; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=fail reason="signature verification failed" (1024-bit key) header.d=linux-foundation.org header.i=@linux-foundation.org header.b="Km6SraPs" Received: by smtp.kernel.org (Postfix) with ESMTPSA id E70ED1F000E9; Mon, 27 Jul 2026 20:49:50 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux-foundation.org; s=korg; t=1785185391; bh=fggn7tx/E/h6j/YgcVlmy7J3VJhfZLSKiSzgGEbZLCM=; h=Date:From:To:Cc:Subject:In-Reply-To:References; b=Km6SraPsRfrDyw3Tupf/G1B+Rq9uKup57Mq7c1UEyR2qDg2EhTvX8CcADk9i6AIjR p02VPtt5X7SGe6OmcWFjJ67nZ03Q+eMPzr8n9E7IzWDMgIy3YwiETmc4cKbzBhsBO6 AQ1rpKVKwyKuoByJBoYU/JOcfV2JrYqTNZWeo3E4= Date: Mon, 27 Jul 2026 13:49:50 -0700 From: Andrew Morton To: Breno Leitao Cc: David Hildenbrand , Zi Yan , Matthew Brost , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , paulmck@kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, kernel-team@meta.com Subject: Re: [PATCH] mm/migrate: report RCU-tasks quiescent states in migrate_pages_batch() Message-Id: <20260727134950.2d091094dbb9bb3ecb7a3eb6@linux-foundation.org> In-Reply-To: <20260727-kcompact-v1-1-bdfefddd6874@debian.org> References: <20260727-kcompact-v1-1-bdfefddd6874@debian.org> X-Mailer: Sylpheed 3.8.0beta1 (GTK+ 2.24.33; x86_64-pc-linux-gnu) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 Content-Type: text/plain; charset=US-ASCII Content-Transfer-Encoding: 7bit On Mon, 27 Jul 2026 06:50:19 -0700 Breno Leitao wrote: > migrate_pages_batch() unmaps each folio before moving it, and every > unmap runs the mmu_notifier invalidate callbacks. On KVM hosts > try_to_migrate() ends up in kvm_mmu_notifier_invalidate_range_start() -> > tdp_mmu_zap_leafs(), which is expensive, so unmapping a large batch keeps > the CPU busy for a long time. > > The loop already calls cond_resched(), but on PREEMPTION kernels that is > a no-op, and involuntary preemption is not a Tasks-RCU quiescent state. > > A long batch therefore never reports a quiescent state, and the > migrating task (e.g. kcompactd) becomes a Tasks-RCU holdout, stalling the > Tasks-RCU grace period for minutes, which is common at Meta fleet: > > INFO: rcu_tasks detected stalls on tasks: > 0000000055349ecc: .. nvcsw: 1157401/1157401 holdout: 1 idle_cpu: -1/56 task:kcompactd0 state:R running task > Call Trace: > tdp_mmu_zap_leafs > tdp_mmu_next_root > gfn_to_pfn_cache_invalidate_start > kvm_mmu_notifier_invalidate_range_start > __mmu_notifier_invalidate_range_start > try_to_migrate_one > try_to_migrate > migrate_pages_batch > migrate_pages > compact_zone > compact_node > kcompactd > kthread Well I doubt if users of 7.1 kernels and earlier want to see this. So a Fixes: and a cc:stable are needed. The affected code is quite old and might even predate the addition of cond_resched_tasks_rcu_qs(). So I can't begin to suggest a Fixes: target. Maybe omit it and let the -stable team figure it out ;) > --- a/mm/migrate.c > +++ b/mm/migrate.c > @@ -1843,7 +1843,7 @@ static int migrate_pages_batch(struct list_head *from, > is_thp = folio_test_pmd_mappable(folio); > nr_pages = folio_nr_pages(folio); > > - cond_resched(); > + cond_resched_tasks_rcu_qs(); > Totally off-topic but why the heck was that implemented as a macro. Which invokes another macro and another and another and turtles all the way down. End result: do { do { if (!((false)) && ({ do { __attribute__((__noreturn__)) extern void __compiletime_assert_606(void) __attribute__((__error__("Unsupported access size for {READ,WRITE}_ONCE()."))); if (!((sizeof(((current))->rcu_tasks_holdout) == sizeof(char) || sizeof(((current))->rcu_tasks_holdout) == sizeof(short) || sizeof(((current))->rcu_tasks_holdout) == sizeof(int) || sizeof(((current))->rcu_tasks_holdout) == sizeof(long)) || sizeof(((current))->rcu_tasks_holdout) == sizeof(long long))) __compiletime_assert_606(); } while (0); (*(const volatile __typeof_unqual__(((current))->rcu_tasks_holdout) *)&(((current))->rcu_tasks_holdout)); })) do { do { __attribute__((__noreturn__)) extern void __compiletime_assert_607(void) __attribute__((__error__("Unsupported access size for {READ,WRITE}_ONCE()."))); if (!((sizeof(((current))->rcu_tasks_holdout) == sizeof(char) || sizeof(((current))->rcu_tasks_holdout) == sizeof(short) || sizeof(((current))->rcu_tasks_holdout) == sizeof(int) || sizeof((( current))->rcu_tasks_holdout) == sizeof(long)) || sizeof(((current))->rcu_tasks_holdout) == sizeof(long long))) __compiletime_assert_607(); } while (0); do { *(volatile typeof(((current))->rcu_tasks_holdout) *)&(((current))->rcu_tasks_holdout) = (false); } while (0); } while (0); } while (0); ({ __might_resched("mm/migrate.c", 1846, 0); _cond_resched(); }); } while (0); How much nicer would it be to have static inline void cond_resched_tasks_rcu_qs(void) { if (some brief and efficient test) some_slow_uninlined_thing(); }