mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Qiliang Yuan <odys.yuan@gmail.com>
To: Ingo Molnar <mingo@redhat.com>,
	Peter Zijlstra <peterz@infradead.org>,
	 Juri Lelli <juri.lelli@redhat.com>,
	 Vincent Guittot <vincent.guittot@linaro.org>,
	 Dietmar Eggemann <dietmar.eggemann@arm.com>,
	 Steven Rostedt <rostedt@goodmis.org>,
	Ben Segall <bsegall@google.com>,  Mel Gorman <mgorman@suse.de>,
	Valentin Schneider <vschneid@redhat.com>,
	 "Paul E. McKenney" <paulmck@kernel.org>,
	 Frederic Weisbecker <frederic@kernel.org>,
	 Neeraj Upadhyay <neeraj.upadhyay@kernel.org>,
	 Joel Fernandes <joelagnelf@nvidia.com>,
	 Josh Triplett <josh@joshtriplett.org>,
	Boqun Feng <boqun@kernel.org>,
	 Uladzislau Rezki <urezki@gmail.com>,
	 Mathieu Desnoyers <mathieu.desnoyers@efficios.com>,
	 Lai Jiangshan <jiangshanlai@gmail.com>,
	Zqiang <qiang.zhang@linux.dev>,
	 Anna-Maria Behnsen <anna-maria@linutronix.de>,
	Tejun Heo <tj@kernel.org>,  Jonathan Corbet <corbet@lwn.net>,
	Shuah Khan <skhan@linuxfoundation.org>,
	 Shuah Khan <shuah@kernel.org>, Thomas Gleixner <tglx@kernel.org>
Cc: Waiman Long <longman@redhat.com>,
	linux-kernel@vger.kernel.org,  rcu@vger.kernel.org,
	cgroups@vger.kernel.org, linux-doc@vger.kernel.org,
	 linux-kselftest@vger.kernel.org,
	Qiliang Yuan <odys.yuan@gmail.com>
Subject: [PATCH v5 10/12] cpuset: Drive kernel-noise isolation via per-CPU hotplug cycling
Date: Fri, 02 Oct 2026 21:10:30 +0800	[thread overview]
Message-ID: <20261002-wujing-dhm-v5-10-78a6996d87ad@gmail.com> (raw)
In-Reply-To: <20261002-wujing-dhm-v5-0-78a6996d87ad@gmail.com>

Track A (cpuset_update_sd_hk_unlock) updates the HK_TYPE_KERNEL_NOISE
and HK_TYPE_MANAGED_IRQ cpumasks but performs no per-CPU reconfiguration.
Tick suppression, RCU callback offloading and managed-IRQ remapping only
take effect when the affected CPUs pass through the CPU hotplug machinery.

Implement dhm_cycle_isolated_cpus() and call it from
cpuset_update_sd_hk_unlock() with cpuset_top_mutex still held: per the
locking convention at the top of this file, cpuset_top_mutex is the
outermost lock, so remove_cpu()/add_cpu() may freely acquire
cpus_write_lock() underneath it.  Keeping the mutex held across the
whole cycle also matches dhm_prev_isolated's existing "protected by
cpuset_top_mutex" comment and prevents a second isolated-partition
update from entering cpuset_update_sd_hk_unlock() while a cycle for
the previous one is still in flight.

On isolation, for each newly-isolated CPU:
  1. remove_cpu()              - offline; dying callbacks migrate IRQs
  2. housekeeping_update_types() - publish this CPU as kernel-noise
                                 isolated now that it is actually offline
  3. tick_nohz_cpu_isolate()  - enable context tracking so the tick
                                 is suppressed for this CPU (B0/B3)
  4. rcu_nocb_cpu_isolate()   - lazy nocb init, spawn kthreads, offload
                                 callbacks (B1)
  5. add_cpu()                 - online; tick and IRQ online callbacks
                                 reconfigure against the now-published
                                 HK masks

On de-isolation, the reverse order is applied, publishing the CPU as
no longer kernel-noise isolated right after its own remove_cpu()
succeeds instead of before any CPU in the batch is touched.

The managed-IRQ remapping requires no explicit call:
irq_migrate_all_off_this_cpu() (dying callback) and
irq_affinity_online_cpu() (online callback) already consult the
updated HK_TYPE_MANAGED_IRQ mask.

dhm_prev_isolated tracks the previous isolation set so that only CPUs
whose state changed are cycled rather than the full isolation set.
lockup_detector_hk_update() (B2) is called once after all CPUs are
cycled to update the watchdog mask.

Some CPUs cannot be taken offline: cpu_is_hotpluggable() rejects a
CPU with hotplug disabled in the architecture (e.g. the x86-64 boot
CPU), and remove_cpu() itself can still fail for a CPU that passed
that filter (e.g. it turns out to be the last CPU in its sched
domain, or it is the current tick_do_timer_cpu holder and
tick_nohz_cpu_hotpluggable() rejects it).  Publishing
housekeeping_update_types() per CPU, strictly after that CPU's own
remove_cpu() has already succeeded, makes both cases handle
themselves: a CPU this function ends up skipping is simply never
added to the published mask, so HK_TYPE_KERNEL_NOISE never claims a
CPU is isolated while it keeps ticking and running RCU callbacks
normally, without needing a separate pre-filter pass or a
republish-on-failure step afterwards.  Also free
newly_isolated/newly_deisolated/cur_isolated on the early return for
a no-op update, which previously leaked the allocations.

Signed-off-by: Qiliang Yuan <odys.yuan@gmail.com>
---
 kernel/cgroup/cpuset.c | 155 ++++++++++++++++++++++++++++++++++++++++++++-----
 1 file changed, 140 insertions(+), 15 deletions(-)

diff --git a/kernel/cgroup/cpuset.c b/kernel/cgroup/cpuset.c
index 6213e63cf348d..9fa4bf504c3d9 100644
--- a/kernel/cgroup/cpuset.c
+++ b/kernel/cgroup/cpuset.c
@@ -20,6 +20,8 @@
  */
 #include "cpuset-internal.h"
 
+#include <linux/cpu.h>
+#include <linux/cpuhplock.h>
 #include <linux/init.h>
 #include <linux/interrupt.h>
 #include <linux/kernel.h>
@@ -33,7 +35,9 @@
 #include <linux/sched/task.h>
 #include <linux/security.h>
 #include <linux/oom.h>
+#include <linux/nmi.h>
 #include <linux/sched/isolation.h>
+#include <linux/tick.h>
 #include <linux/wait.h>
 #include <linux/workqueue.h>
 #include <linux/task_work.h>
@@ -165,6 +169,14 @@ static cpumask_var_t	isolated_hk_cpus;	/* T */
 static DEFINE_SPINLOCK(dhm_cycling_lock);
 static cpumask_var_t	dhm_cycling_cpus;
 
+/*
+ * Snapshot of the isolated CPUs from the previous housekeeping update.
+ * Used to compute the delta (newly isolated / newly de-isolated) so that
+ * only the changed CPUs are cycled rather than the full isolation set.
+ * Protected by cpuset_top_mutex.
+ */
+static cpumask_var_t	dhm_prev_isolated;
+
 /*
  * A flag to force sched domain rebuild at the end of an operation.
  * It can be set in
@@ -1423,6 +1435,108 @@ static bool prstate_housekeeping_conflict(int prstate, struct cpumask *new_cpus)
 	return false;
 }
 
+/*
+ * dhm_cycle_isolated_cpus - Apply kernel-noise isolation via hotplug cycling
+ *
+ * For each CPU newly entering isolation: cycle it offline, configure tick
+ * suppression and RCU callback offloading while it is offline, then bring
+ * it back online.  The managed-IRQ state is handled automatically by the
+ * existing irq_migrate_all_off_this_cpu() dying callback and the
+ * irq_affinity_online_cpu() online callback which both consult the
+ * already-updated HK_TYPE_MANAGED_IRQ mask.
+ *
+ * For each CPU leaving isolation: cycle it offline, de-offload RCU and
+ * restore the tick, then bring it back online.
+ *
+ * Called with cpuset_top_mutex held and no other cpuset or hotplug locks
+ * held: cpuset_top_mutex is the outermost lock, so remove_cpu()/add_cpu()
+ * may freely take cpus_write_lock() underneath it.
+ */
+static void dhm_cycle_isolated_cpus(const struct cpumask *new_isolated)
+{
+	static const unsigned long noise_types =
+		BIT(HK_TYPE_KERNEL_NOISE) | BIT(HK_TYPE_MANAGED_IRQ);
+	cpumask_var_t newly_isolated, newly_deisolated, cur_isolated;
+	int cpu;
+
+	if (!alloc_cpumask_var(&newly_isolated, GFP_KERNEL) ||
+	    !alloc_cpumask_var(&newly_deisolated, GFP_KERNEL) ||
+	    !alloc_cpumask_var(&cur_isolated, GFP_KERNEL)) {
+		free_cpumask_var(newly_isolated);
+		free_cpumask_var(newly_deisolated);
+		return;
+	}
+
+	cpumask_andnot(newly_isolated, new_isolated, dhm_prev_isolated);
+	cpumask_andnot(newly_deisolated, dhm_prev_isolated, new_isolated);
+	/* cur_isolated tracks the mask actually published so far. */
+	cpumask_copy(cur_isolated, dhm_prev_isolated);
+	cpumask_copy(dhm_prev_isolated, new_isolated);
+
+	if (cpumask_empty(newly_isolated) && cpumask_empty(newly_deisolated))
+		goto out_free;
+
+	/* Mark cycling CPUs so cpuset_hotplug_update_tasks skips invalidation */
+	spin_lock(&dhm_cycling_lock);
+	cpumask_or(dhm_cycling_cpus, newly_isolated, newly_deisolated);
+	spin_unlock(&dhm_cycling_lock);
+
+	/*
+	 * Publish each CPU's kernel-noise/managed-IRQ housekeeping state
+	 * strictly between its own remove_cpu() succeeding and add_cpu()
+	 * bringing it back, never as a batch before the whole cycle.
+	 * housekeeping_update_types() then always describes exactly which
+	 * CPUs are actually ticking and running RCU callbacks normally at
+	 * that instant, including the current tick_do_timer_cpu holder,
+	 * which tick_nohz_cpu_hotpluggable() must see as still a
+	 * housekeeping CPU for as long as it is still online and un-cycled.
+	 */
+	for_each_cpu(cpu, newly_isolated) {
+		if (!cpu_is_hotpluggable(cpu)) {
+			pr_warn_once("cpuset: CPU%d cannot be isolated (hotplug disabled)\n",
+				     cpu);
+			cpumask_clear_cpu(cpu, dhm_prev_isolated);
+			continue;
+		}
+		if (remove_cpu(cpu)) {
+			pr_warn_once("cpuset: failed to offline CPU%d for isolation\n",
+				     cpu);
+			cpumask_clear_cpu(cpu, dhm_prev_isolated);
+			continue;
+		}
+		cpumask_set_cpu(cpu, cur_isolated);
+		WARN_ON_ONCE(housekeeping_update_types(noise_types, cur_isolated));
+		WARN_ON_ONCE(tick_nohz_cpu_isolate(cpu));
+		WARN_ON_ONCE(rcu_nocb_cpu_isolate(cpu));
+		WARN_ON_ONCE(add_cpu(cpu));
+	}
+
+	for_each_cpu(cpu, newly_deisolated) {
+		if (remove_cpu(cpu)) {
+			pr_warn_once("cpuset: failed to offline CPU%d for de-isolation\n",
+				     cpu);
+			cpumask_set_cpu(cpu, dhm_prev_isolated);
+			continue;
+		}
+		cpumask_clear_cpu(cpu, cur_isolated);
+		WARN_ON_ONCE(housekeeping_update_types(noise_types, cur_isolated));
+		WARN_ON_ONCE(rcu_nocb_cpu_deoffload(cpu));
+		tick_nohz_cpu_deisolate(cpu);
+		WARN_ON_ONCE(add_cpu(cpu));
+	}
+
+	spin_lock(&dhm_cycling_lock);
+	cpumask_clear(dhm_cycling_cpus);
+	spin_unlock(&dhm_cycling_lock);
+
+	lockup_detector_hk_update();
+
+out_free:
+	free_cpumask_var(newly_isolated);
+	free_cpumask_var(newly_deisolated);
+	free_cpumask_var(cur_isolated);
+}
+
 /*
  * cpuset_update_sd_hk_unlock - Rebuild sched domains, update HK & unlock
  *
@@ -1439,8 +1553,6 @@ static void cpuset_update_sd_hk_unlock(void)
 		rebuild_sched_domains_locked();
 
 	if (update_housekeeping) {
-		static const unsigned long noise_types =
-			BIT(HK_TYPE_KERNEL_NOISE) | BIT(HK_TYPE_MANAGED_IRQ);
 		int ret;
 
 		update_housekeeping = false;
@@ -1450,9 +1562,16 @@ static void cpuset_update_sd_hk_unlock(void)
 		cpus_read_unlock();
 
 		/*
-		 * housekeeping_update() is now called without holding
-		 * cpus_read_lock and cpuset_mutex. Only cpuset_top_mutex
-		 * is still being held for mutual exclusion.
+		 * housekeeping_update() and housekeeping_update_types() are
+		 * now called without holding cpus_read_lock and cpuset_mutex.
+		 * cpuset_top_mutex stays held all the way through
+		 * dhm_cycle_isolated_cpus() below: per the locking convention
+		 * at the top of this file it is the outermost lock, so it may
+		 * legally nest around the cpus_write_lock() that remove_cpu()/
+		 * add_cpu() take. Holding it here is also what makes
+		 * dhm_prev_isolated's "protected by cpuset_top_mutex" comment
+		 * true, and keeps a second isolated-partition update from
+		 * entering this function while a cycle is still in flight.
 		 */
 
 		/*
@@ -1464,18 +1583,23 @@ static void cpuset_update_sd_hk_unlock(void)
 		WARN_ON_ONCE(ret);
 
 		/*
-		 * Only touch the kernel-noise housekeeping masks
-		 * (HK_TYPE_KERNEL_NOISE and HK_TYPE_MANAGED_IRQ) once the
-		 * sched domain update above actually succeeded: HK_TYPE_
-		 * KERNEL_NOISE must stay a subset of HK_TYPE_DOMAIN, so a CPU
-		 * can never end up tick-suppressed while still scheduled as
-		 * part of the normal (non-isolated) sched domain.  The tick,
-		 * RCU and managed-interrupt state is reconfigured as the
-		 * affected CPUs are cycled through the CPU hotplug machinery.
+		 * Only cycle CPUs through hotplug, applying the kernel-noise
+		 * types along the way, when the sched domain update above
+		 * actually succeeded.  HK_TYPE_KERNEL_NOISE must stay a
+		 * subset of HK_TYPE_DOMAIN, so a CPU can never end up
+		 * tick-suppressed while still scheduled as part of the
+		 * normal (non-isolated) sched domain.  The tick, RCU and
+		 * managed-interrupt state is reconfigured as the affected
+		 * CPUs are cycled through the CPU hotplug machinery;
+		 * dhm_cycle_isolated_cpus() publishes HK_TYPE_KERNEL_NOISE
+		 * and HK_TYPE_MANAGED_IRQ itself, one CPU at a time, strictly
+		 * between that CPU's own remove_cpu() and add_cpu(), so a
+		 * CPU that cpu_is_hotpluggable() rejects (e.g. the current
+		 * tick_do_timer_cpu) is simply never published as isolated.
 		 */
 		if (!ret)
-			WARN_ON_ONCE(housekeeping_update_types(noise_types,
-							       isolated_hk_cpus));
+			dhm_cycle_isolated_cpus(isolated_hk_cpus);
+
 		mutex_unlock(&cpuset_top_mutex);
 	} else {
 		cpuset_full_unlock();
@@ -3915,6 +4039,7 @@ int __init cpuset_init(void)
 	BUG_ON(!zalloc_cpumask_var(&isolated_cpus, GFP_KERNEL));
 	BUG_ON(!zalloc_cpumask_var(&isolated_hk_cpus, GFP_KERNEL));
 	BUG_ON(!zalloc_cpumask_var(&dhm_cycling_cpus, GFP_KERNEL));
+	BUG_ON(!zalloc_cpumask_var(&dhm_prev_isolated, GFP_KERNEL));
 
 	cpumask_setall(top_cpuset.cpus_allowed);
 	nodes_setall(top_cpuset.mems_allowed);

-- 
2.43.0


  parent reply	other threads:[~2026-10-02 13:11 UTC|newest]

Thread overview: 14+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-02 13:10 [PATCH v5 00/12] Dynamic Housekeeping Management (DHM) via CPUSets Qiliang Yuan
2026-10-02 13:10 ` [PATCH v5 01/12] sched/isolation: Enforce nohz_full as a subset of isolcpus=domain at boot Qiliang Yuan
2026-10-02 13:10 ` [PATCH v5 02/12] sched/isolation: Add runtime housekeeping mask updates with boot snapshots Qiliang Yuan
2026-10-02 13:10 ` [PATCH v5 03/12] sched/isolation: RCU-protect runtime-mutable housekeeping cpumask readers Qiliang Yuan
2026-10-02 13:10 ` [PATCH v5 04/12] cpuset: Drive kernel-noise housekeeping from isolated partitions Qiliang Yuan
2026-10-02 13:10 ` [PATCH v5 05/12] context_tracking: Allow runtime per-CPU user tracking enable/disable Qiliang Yuan
2026-10-02 13:10 ` [PATCH v5 06/12] rcu/nocb: Support lazy init for runtime CPU isolation Qiliang Yuan
2026-10-02 13:10 ` [PATCH v5 07/12] watchdog: Sync watchdog_cpumask with HK_TYPE_KERNEL_NOISE on isolation Qiliang Yuan
2026-10-02 15:26   ` Bradley Morgan
2026-10-02 13:10 ` [PATCH v5 08/12] tick/nohz: Derive full-dynticks state from HK_TYPE_KERNEL_NOISE Qiliang Yuan
2026-10-02 13:10 ` [PATCH v5 09/12] cpuset: Add dhm_cycling_cpus mask to suppress transient invalidation Qiliang Yuan
2026-10-02 13:10 ` Qiliang Yuan [this message]
2026-10-02 13:10 ` [PATCH v5 11/12] docs: cgroup-v2: Document kernel-noise isolation via isolated partitions Qiliang Yuan
2026-10-02 13:10 ` [PATCH v5 12/12] selftests/cgroup: Add kernel-noise isolation test to cpuset selftest Qiliang Yuan

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261002-wujing-dhm-v5-10-78a6996d87ad@gmail.com \
    --to=odys.yuan@gmail.com \
    --cc=anna-maria@linutronix.de \
    --cc=boqun@kernel.org \
    --cc=bsegall@google.com \
    --cc=cgroups@vger.kernel.org \
    --cc=corbet@lwn.net \
    --cc=dietmar.eggemann@arm.com \
    --cc=frederic@kernel.org \
    --cc=jiangshanlai@gmail.com \
    --cc=joelagnelf@nvidia.com \
    --cc=josh@joshtriplett.org \
    --cc=juri.lelli@redhat.com \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-kselftest@vger.kernel.org \
    --cc=longman@redhat.com \
    --cc=mathieu.desnoyers@efficios.com \
    --cc=mgorman@suse.de \
    --cc=mingo@redhat.com \
    --cc=neeraj.upadhyay@kernel.org \
    --cc=paulmck@kernel.org \
    --cc=peterz@infradead.org \
    --cc=qiang.zhang@linux.dev \
    --cc=rcu@vger.kernel.org \
    --cc=rostedt@goodmis.org \
    --cc=shuah@kernel.org \
    --cc=skhan@linuxfoundation.org \
    --cc=tglx@kernel.org \
    --cc=tj@kernel.org \
    --cc=urezki@gmail.com \
    --cc=vincent.guittot@linaro.org \
    --cc=vschneid@redhat.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®