mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v5 00/12] Dynamic Housekeeping Management (DHM) via CPUSets
@ 2026-10-02 13:10 Qiliang Yuan
  2026-10-02 13:10 ` [PATCH v5 01/12] sched/isolation: Enforce nohz_full as a subset of isolcpus=domain at boot Qiliang Yuan
                   ` (11 more replies)
  0 siblings, 12 replies; 14+ messages in thread
From: Qiliang Yuan @ 2026-10-02 13:10 UTC (permalink / raw)
  To: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot,
	Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, Paul E. McKenney, Frederic Weisbecker,
	Neeraj Upadhyay, Joel Fernandes, Josh Triplett, Boqun Feng,
	Uladzislau Rezki, Mathieu Desnoyers, Lai Jiangshan, Zqiang,
	Anna-Maria Behnsen, Tejun Heo, Jonathan Corbet, Shuah Khan,
	Shuah Khan, Thomas Gleixner
  Cc: Waiman Long, linux-kernel, rcu, cgroups, linux-doc,
	linux-kselftest, Qiliang Yuan

This series introduces Dynamic Housekeeping Management (DHM) to the Linux
kernel, enabling runtime reconfiguration of kernel-noise housekeeping
(nohz_full tick suppression, RCU NOCB offloading, and managed IRQ
migration) through the existing cgroup v2 cpuset isolated partition
mechanism — no new kernel ABI required.

When a cpuset partition is set to isolated mode, DHM cycles each CPU in
that partition offline, reconfigures the housekeeping masks (removing the
CPU from HK_TYPE_KERNEL_NOISE and related types), and brings it back
online.  The subsystems (tick/nohz, RCU NOCB, genirq) pick up the new
masks through their existing CPU hotplug callbacks.  Destroying the
partition reverses the cycle: each CPU is taken offline, restored to all
housekeeping masks, and brought back online.

Housekeeping cpumask pointers are RCU-protected to allow lock-free readers
during updates.  A global dhm_cycling_cpus mask suppresses transient
cpuset partition invalidation while CPUs are being cycled.

This work is related to Waiman Long's [PATCH-next 00/23] series
(20260421030351.281436-1-longman@redhat.com), which also targets runtime
cpuset housekeeping control.  DHM's distinguishing feature is zero-boot-
parameter activation: nohz_full+nocb isolation can be achieved at runtime
without any boot-time nohz_full= or rcu_nocbs= parameters.

Signed-off-by: Qiliang Yuan <odys.yuan@gmail.com>
---
V4 -> V5:
- Rebase onto current mainline (was based on a v7.1-era snapshot).
- Gate the HK_TYPE_KERNEL_NOISE/HK_TYPE_MANAGED_IRQ mask update on the
  HK_TYPE_DOMAIN update succeeding, so a CPU can never end up
  tick-suppressed while still part of the normal sched domain.
  Frederic asked for HK_TYPE_KERNEL_NOISE to stay a subset of
  HK_TYPE_DOMAIN.
- Keep cpuset_top_mutex held across dhm_cycle_isolated_cpus() instead
  of releasing it before the hotplug cycle, so dhm_prev_isolated's
  "protected by cpuset_top_mutex" invariant actually holds and a second
  isolated-partition update cannot race with one still in flight.
  Frederic asked whether cpus_read_lock should stay held here.
- Remove tick_nohz_full_mask and tick_nohz_full_running; derive
  tick_nohz_full_enabled()/tick_nohz_full_cpu() from
  housekeeping_enabled(HK_TYPE_KERNEL_NOISE) and
  housekeeping_cpumask_rcu(HK_TYPE_KERNEL_NOISE) instead, removing the
  now-redundant lazy-allocation and locking in
  tick_nohz_cpu_isolate()/deisolate().  Frederic pointed out the
  duplication.
- Publish each CPU's HK_TYPE_KERNEL_NOISE/HK_TYPE_MANAGED_IRQ mask
  update from dhm_cycle_isolated_cpus() itself, strictly between that
  CPU's own remove_cpu() succeeding and add_cpu() bringing it back,
  instead of batching the whole isolation set before or after the
  cycle.  A CPU the cycling loop skips (not hotpluggable, or
  remove_cpu() fails) is then simply never published as isolated,
  with no separate pre-filter or republish-on-failure step needed.
- Install tick_do_timer_cpu's hotplug protection from
  housekeeping_update_types()'s HK_TYPE_KERNEL_NOISE first-enable
  path too, via a new tick_nohz_full_hotplug_init(): tick_nohz_init()
  never runs on DHM's zero-boot-param path, so the current duty
  holder had no protection against being pulled down by remove_cpu().
- Keep tick_nohz_cpu_hotpluggable()'s duty-holder check on the plain
  housekeeping_enabled() flag rather than a mask-aware variant:
  housekeeping_update_types() only publishes a CPU's isolation after
  its own remove_cpu() already succeeded, so a mask-aware check would
  see every single-CPU isolation's own live mask as still empty and
  leave the duty holder unprotected each time, not just the first.
  Use the mask-aware variant only for tick_sched_do_timer()'s
  diagnostic WARN, since HK_TYPE_KERNEL_NOISE can stay enabled with an
  empty mask after a full runtime de-isolation, an ordinary
  NO_HZ_IDLE handover rather than a violation.
- Close a context_tracking_key activation race: a CPU whose
  kernel<->user transition lands in the window before every CPU
  observes static_branch_inc()'s patched code sees
  context_tracking_enabled() as still false and silently drops the
  transition, tripping CT_WARN_ON() on its next kernel entry.  Add
  context_tracking_activating, a plain per-CPU bool with no
  code-patching delay of its own, set via IPI before
  static_branch_inc() and cleared via IPI only after it returns, and
  route user_enter_irqoff()/user_exit_irqoff()/CT_WARN_ON()/ct_state()
  through it alongside the static key.
- Fix the boot-time nohz_full=/isolcpus=domain invariant check
  itself: it ran before static_branch_enable(&housekeeping_overridden),
  so housekeeping_cpumask() always returned cpu_possible_mask for
  both sides of the subset comparison and the check never actually
  rejected anything.  Enable housekeeping_overridden first.
- Fix a cpumask_var_t leak in dhm_cycle_isolated_cpus() on the no-op
  path (a partition update that nets to no CPU delta).
- Let HK_TYPE_MANAGED_IRQ use the same runtime first-enable path as
  HK_TYPE_KERNEL_NOISE; it was unconditionally skipped without a
  boot-time isolcpus=managed_irq, so managed-IRQ migration silently
  never activated on a zero-boot-parameter system.
- Finish converting HK_TYPE_MANAGED_IRQ/HK_TYPE_KERNEL_NOISE readers
  to housekeeping_cpumask_rcu(): vmbus_channel_set_cpu() in
  drivers/hv/vmbus_drv.c, rps_cpumask_housekeeping() in
  net/core/net-sysfs.c, and tmigr_isolated_exclude_cpumask() in
  kernel/time/timer_migration.c were still reading the mask
  without RCU protection.
- Register the CONFIG_RCU_LAZY shrinker from the runtime nocb
  first-enable path too; previously only the boot path did, so a
  system with no rcu_nocbs=/nohz_full= at boot never got lazy-callback
  reclaim once DHM enabled NOCB offload at runtime.
- Add a standalone patch enforcing nohz_full=/isolcpus=nohz as a
  subset of isolcpus=domain at boot: disable nohz_full entirely
  (including when it is passed alone, with no isolcpus=domain at
  all) if the invariant does not hold, checked in housekeeping_init()
  so the outcome does not depend on cmdline argument order.
  Frederic's superset invariant above only covers the cpuset-driven
  runtime path; this closes the same gap for the two boot
  parameters, which he separately said he'd enforce "if he could do
  it again".

V3 -> V4:
- Drop apply() callback architecture entirely (struct housekeeping_cbs,
  pre_validate/apply hooks, housekeeping_update_types() notification).
  Thomas Gleixner identified concurrent in-kernel mask apply as
  fundamentally unsafe ("broken beyond repair") and endorsed CPU-by-CPU
  hotplug cycling as the correct approach.
- Track A prerequisites (housekeeping boot type isolation, RCU reader
  protection, cpuset trigger) submitted upstream separately; v4 builds
  on top of those three committed patches.
- Fix rcu/nocb lazy_init: remove __init from rcu_organize_nocb_kthreads()
  forward declaration (tree.h) to prevent GCC from placing the function in
  .init.text, which caused an NX-protected page fault when
  rcu_nocb_cpu_isolate() was called at runtime.  Add noinline to
  rcu_nocb_lazy_init() to preserve the GCC IPA call chain.
- Fix cpuset remote partition cycling suppression: extend dhm_cycling_cpus
  guard to the remote-partition disable path in cpuset_hotplug_update_tasks(),
  preventing false partition invalidation during hotplug cycling steps.
  Fixes selftest TEST_MATRIX[69] (A2: expected 1-2, got 1-3).

V2 -> V3:
- Replace notifier chain with explicit per-type callback interface
  (struct housekeeping_cbs with .name, .pre_validate, .apply fields).
- RCU-protect all housekeeping cpumask pointers; callers must hold
  rcu_read_lock() or use housekeeping_cpumask_rcu() in apply() callbacks.
- Drop 5 patches from v2: HK_TYPE enum separation (upstream aliases are
  already correct), no-op timer/hrtimer patches, kthread dead code, and
  workqueue double-update.
- Fix deadlock in rcu_hk_workfn(): remove cpus_read_lock() wrapper around
  remove_cpu()/add_cpu() which take cpu_hotplug_lock write side.
- Fix UAF in rcu_hk_apply(): snapshot the housekeeping cpumask inside the
  work function under rcu_read_lock(), not at apply() time where the old
  pointer may be freed by synchronize_rcu() before the work runs.
- Fix tick apply(): snapshot housekeeping_cpumask_rcu() under
  rcu_read_lock() as required by lockdep for runtime-mutable types.
- Activate context_tracking dynamically via ct_cpu_track_user() /
  ct_cpu_untrack_user() in tick apply(), eliminating the dependency on
  CONFIG_CONTEXT_TRACKING_USER_FORCE flagged by tglx.
- Fix genirq apply(): snapshot HK_TYPE_MANAGED_IRQ mask under
  rcu_read_lock() before the IRQ iteration loop.
- Simplify cpuset noise_types to BIT(HK_TYPE_KERNEL_NOISE) |
  BIT(HK_TYPE_MANAGED_IRQ), replacing the redundant per-alias bitmask.
- housekeeping_update_types(): always use cpu_possible_mask as base
  for HK_TYPE_KERNEL_NOISE, so de-isolation restores the mask to all
  possible CPUs rather than leaving it at its last non-trivial value.
- Initialize watchdog_cpumask from HK_TYPE_KERNEL_NOISE (not
  HK_TYPE_TIMER) at boot; keep it in sync at runtime via a new
  housekeeping_cbs callback.
- Add kernel-noise selftest to test_cpuset_prs.sh, including
  cpu_in_cpulist() for correct cpulist range membership detection and
  nohz_full sysfs verification when CONFIG_NO_HZ_FULL is active.
- Add RCU caller fixes: sched/core (HK_TYPE_KERNEL_NOISE) and
  drivers/hv (HK_TYPE_MANAGED_IRQ) are required because those types
  are updated at runtime; hrtimer (HK_TYPE_TIMER) and arm64/topology
  (HK_TYPE_TICK) are defensive fixes.
- Reorder patches so all subsystem callbacks are registered before the
  cpuset patch that triggers housekeeping_update_types().

V1 -> V2:
- Rebrand series from DHEI to DHM (Dynamic Housekeeping Management).
- Drop custom sysfs interface entirely.
- Integrate housekeeping control into cgroup v2 cpuset isolated partition
  mechanism.
- Add SMT-aware isolation constraints to prevent splitting SMT siblings.
- Add comprehensive documentation and cgroup functional selftests.
- Refactor mask transition logic to use RCU-safe handover.

v4: https://lore.kernel.org/r/20260710-wujing-dhm-v4-0-2e912e5d9645@gmail.com
v3: https://lore.kernel.org/r/20260618-wujing-dhm-v3-0-28f1a4d83b68@gmail.com
v2: https://lore.kernel.org/r/20260413-wujing-dhm-v2-0-06df21caba5d@gmail.com
v1: https://lore.kernel.org/all/20260325-dhei-v12-final-v1-0-919cca23cadf@gmail.com

---
Qiliang Yuan (12):
      sched/isolation: Enforce nohz_full as a subset of isolcpus=domain at boot
      sched/isolation: Add runtime housekeeping mask updates with boot snapshots
      sched/isolation: RCU-protect runtime-mutable housekeeping cpumask readers
      cpuset: Drive kernel-noise housekeeping from isolated partitions
      context_tracking: Allow runtime per-CPU user tracking enable/disable
      rcu/nocb: Support lazy init for runtime CPU isolation
      watchdog: Sync watchdog_cpumask with HK_TYPE_KERNEL_NOISE on isolation
      tick/nohz: Derive full-dynticks state from HK_TYPE_KERNEL_NOISE
      cpuset: Add dhm_cycling_cpus mask to suppress transient invalidation
      cpuset: Drive kernel-noise isolation via per-CPU hotplug cycling
      docs: cgroup-v2: Document kernel-noise isolation via isolated partitions
      selftests/cgroup: Add kernel-noise isolation test to cpuset selftest

 Documentation/admin-guide/cgroup-v2.rst           |  17 +
 Documentation/admin-guide/kernel-parameters.txt   |  12 +
 arch/arm64/kernel/topology.c                      |   9 +-
 drivers/base/cpu.c                                |  42 +-
 drivers/hv/channel_mgmt.c                         |  50 +-
 drivers/hv/vmbus_drv.c                            |  13 +-
 fs/resctrl/internal.h                             |   6 +-
 include/linux/context_tracking.h                  |  16 +-
 include/linux/context_tracking_state.h            |  25 +-
 include/linux/nmi.h                               |   2 +
 include/linux/rcupdate.h                          |   2 +
 include/linux/sched/isolation.h                   |  34 +-
 include/linux/tick.h                              |  40 +-
 kernel/cgroup/cpuset.c                            | 203 +++++++-
 kernel/context_tracking.c                         |  91 +++-
 kernel/rcu/tree.h                                 |   2 +-
 kernel/rcu/tree_nocb.h                            | 105 +++-
 kernel/sched/core.c                               |   7 +-
 kernel/sched/isolation.c                          | 253 +++++++++-
 kernel/sched/sched.h                              |   2 +-
 kernel/time/hrtimer.c                             |   5 +-
 kernel/time/tick-sched.c                          | 158 +++++-
 kernel/time/timer_migration.c                     |   9 +-
 kernel/watchdog.c                                 |  26 +-
 net/core/net-sysfs.c                              |  10 +-
 tools/testing/selftests/cgroup/test_cpuset_prs.sh | 578 ++++++++++++++++++++++
 26 files changed, 1558 insertions(+), 159 deletions(-)
---
base-commit: 551c722f40809618230001baccf219193e22fc5a
change-id: 20260408-wujing-dhm-8f43e2d49cd8

Best regards,
-- 
Qiliang Yuan <odys.yuan@gmail.com>


^ permalink raw reply	[flat|nested] 14+ messages in thread

* [PATCH v5 01/12] sched/isolation: Enforce nohz_full as a subset of isolcpus=domain at boot
  2026-10-02 13:10 [PATCH v5 00/12] Dynamic Housekeeping Management (DHM) via CPUSets Qiliang Yuan
@ 2026-10-02 13:10 ` Qiliang Yuan
  2026-10-02 13:10 ` [PATCH v5 02/12] sched/isolation: Add runtime housekeeping mask updates with boot snapshots Qiliang Yuan
                   ` (10 subsequent siblings)
  11 siblings, 0 replies; 14+ messages in thread
From: Qiliang Yuan @ 2026-10-02 13:10 UTC (permalink / raw)
  To: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot,
	Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, Paul E. McKenney, Frederic Weisbecker,
	Neeraj Upadhyay, Joel Fernandes, Josh Triplett, Boqun Feng,
	Uladzislau Rezki, Mathieu Desnoyers, Lai Jiangshan, Zqiang,
	Anna-Maria Behnsen, Tejun Heo, Jonathan Corbet, Shuah Khan,
	Shuah Khan, Thomas Gleixner
  Cc: Waiman Long, linux-kernel, rcu, cgroups, linux-doc,
	linux-kselftest, Qiliang Yuan

HK_TYPE_KERNEL_NOISE and HK_TYPE_DOMAIN are configured independently
when nohz_full=/isolcpus=nohz and isolcpus=domain are passed as
separate boot parameters, or when nohz_full=/isolcpus=nohz is passed
without isolcpus=domain at all.  Nothing stops a CPU from ending up
tick-suppressed (nohz_full) while still scheduled as part of the
normal, non-isolated sched domain, which defeats the point of
isolating it from kernel noise in the first place.

A CPU with the tick stopped must always be excluded from the normal
scheduler domain: HK_TYPE_KERNEL_NOISE's housekeeping set has to stay
a superset of HK_TYPE_DOMAIN's.  Checking this while __setup()
parameters are still being parsed is order-dependent: rejecting
nohz_full= the moment it is seen, before a later isolcpus=domain on
the same command line has been parsed yet, would discard a
perfectly valid combination just because of argument order.

Check the invariant once in housekeeping_init() instead, which runs
after every __setup() cmdline parameter has been parsed regardless of
order.  Disable HK_TYPE_KERNEL_NOISE rather than panicking or leaving
the inconsistency in place: this gives the same end state as not
having passed nohz_full=/isolcpus=nohz at all.

housekeeping_cpumask() only dereferences the real per-type masks once
housekeeping_overridden is enabled; before that every call returns
cpu_possible_mask regardless of housekeeping.flags.  Enable the static
key before the subset check runs, not after: checking first would
compare cpu_possible_mask against itself for both sides and never
reject anything.

Signed-off-by: Qiliang Yuan <odys.yuan@gmail.com>
---
 Documentation/admin-guide/kernel-parameters.txt | 12 ++++++++++++
 kernel/sched/isolation.c                        | 24 ++++++++++++++++++++++++
 2 files changed, 36 insertions(+)

diff --git a/Documentation/admin-guide/kernel-parameters.txt b/Documentation/admin-guide/kernel-parameters.txt
index e75344f4e0cde..cd4fb6d752a6c 100644
--- a/Documentation/admin-guide/kernel-parameters.txt
+++ b/Documentation/admin-guide/kernel-parameters.txt
@@ -2810,6 +2810,11 @@ Kernel parameters
 			  so to protect individual CPUs the 'cpumask' file has to
 			  be configured manually after bootup.
 
+			  When 'domain' is configured through a separate
+			  isolcpus= or nohz_full= invocation, every 'nohz'
+			  CPU must also be a 'domain' CPU; a combination that
+			  violates this is rejected at boot with a warning.
+
 			domain
 			  Isolate from the general SMP balancing and scheduling
 			  algorithms. Note that performing domain isolation this way
@@ -4564,6 +4569,13 @@ Kernel parameters
 			Note that this argument takes precedence over
 			the CONFIG_RCU_NOCB_CPU_DEFAULT_ALL option.
 
+			When isolcpus=domain is also given, every CPU in
+			this list (or in isolcpus=nohz) must also be in
+			the isolcpus=domain list: a tick-suppressed CPU
+			must always be excluded from the normal scheduler
+			domain.  A combination that violates this is
+			rejected at boot with a warning.
+
 	noinitrd	[Deprecated,RAM] Tells the kernel not to load any configured
 			initial RAM disk. Currently this parameter applies to
 			initrd only, not to initramfs. But it applies to both
diff --git a/kernel/sched/isolation.c b/kernel/sched/isolation.c
index 156025ef81b75..c6a41095a9002 100644
--- a/kernel/sched/isolation.c
+++ b/kernel/sched/isolation.c
@@ -171,8 +171,32 @@ void __init housekeeping_init(void)
 	if (!housekeeping.flags)
 		return;
 
+	/*
+	 * housekeeping_cpumask() only dereferences the real per-type masks
+	 * once housekeeping_overridden is live; before that it always
+	 * returns cpu_possible_mask regardless of housekeeping.flags, which
+	 * would make the subset check below vacuously pass for every type.
+	 * Enable it first so the check actually sees the parsed masks.
+	 */
 	static_branch_enable(&housekeeping_overridden);
 
+	/*
+	 * A nohz_full CPU must always run on an isolated sched domain:
+	 * reject nohz_full=/isolcpus=nohz unless isolcpus=domain covers
+	 * at least the same CPUs.  Checked here, after all __setup()
+	 * cmdline parsing has run, so the outcome does not depend on
+	 * which of the two parameters came first on the command line.
+	 */
+	if ((housekeeping.flags & HK_FLAG_KERNEL_NOISE) &&
+	    !cpumask_subset(housekeeping_cpumask(HK_TYPE_DOMAIN),
+			    housekeeping_cpumask(HK_TYPE_KERNEL_NOISE))) {
+		pr_warn("Housekeeping: nohz_full=/isolcpus=nohz must be a subset "
+			"of isolcpus=domain, disabling nohz_full\n");
+		housekeeping.flags &= ~HK_FLAG_KERNEL_NOISE;
+		if (!housekeeping.flags)
+			return;
+	}
+
 	if (housekeeping.flags & HK_FLAG_KERNEL_NOISE)
 		sched_tick_offload_init();
 	/*

-- 
2.43.0


^ permalink raw reply	[flat|nested] 14+ messages in thread

* [PATCH v5 02/12] sched/isolation: Add runtime housekeeping mask updates with boot snapshots
  2026-10-02 13:10 [PATCH v5 00/12] Dynamic Housekeeping Management (DHM) via CPUSets Qiliang Yuan
  2026-10-02 13:10 ` [PATCH v5 01/12] sched/isolation: Enforce nohz_full as a subset of isolcpus=domain at boot Qiliang Yuan
@ 2026-10-02 13:10 ` Qiliang Yuan
  2026-10-02 13:10 ` [PATCH v5 03/12] sched/isolation: RCU-protect runtime-mutable housekeeping cpumask readers Qiliang Yuan
                   ` (9 subsequent siblings)
  11 siblings, 0 replies; 14+ messages in thread
From: Qiliang Yuan @ 2026-10-02 13:10 UTC (permalink / raw)
  To: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot,
	Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, Paul E. McKenney, Frederic Weisbecker,
	Neeraj Upadhyay, Joel Fernandes, Josh Triplett, Boqun Feng,
	Uladzislau Rezki, Mathieu Desnoyers, Lai Jiangshan, Zqiang,
	Anna-Maria Behnsen, Tejun Heo, Jonathan Corbet, Shuah Khan,
	Shuah Khan, Thomas Gleixner
  Cc: Waiman Long, linux-kernel, rcu, cgroups, linux-doc,
	linux-kselftest, Qiliang Yuan

The housekeeping cpumasks for kernel noise (HK_TYPE_KERNEL_NOISE) and
managed interrupts (HK_TYPE_MANAGED_IRQ) are computed once at boot from
the nohz_full= and isolcpus=managed_irq arguments.  Dynamic CPU
isolation driven by cpuset isolated partitions needs to update these
masks after boot.

Add housekeeping_update_types() to recompute one or more housekeeping
masks as (boot snapshot & ~isolated), publish them via RCU and free the
old masks after a grace period.

Introduce HK_TYPE_KERNEL_NOISE_BOOT and HK_TYPE_MANAGED_IRQ_BOOT to
record the immutable boot configuration.  Anchoring every update on the
boot snapshot keeps the runtime mask a subset of the boot set and lets
de-isolation restore the boot configuration exactly.

When neither nohz_full=/isolcpus=nohz nor isolcpus=managed_irq was
given at boot, no boot snapshot exists for the corresponding type, so
cpu_possible_mask is the implicit boot set and the type is enabled on
the first runtime call.  housekeeping_cpumask() already falls back to
cpu_possible_mask for a type with no flag bit set, so this needs no
special-casing in the trial-mask computation; it only needs the
type's housekeeping_enabled() check not to gate out the first call,
which the HK_TYPE_KERNEL_NOISE check already did but the
HK_TYPE_MANAGED_IRQ one did not, silently skipping managed-IRQ
isolation whenever no boot parameter was present.  HK_TYPE_KERNEL_NOISE
additionally calls sched_tick_offload_init() on first enable to
allocate the tick offload percpu data that would otherwise only be
allocated by housekeeping_init() when nohz_full= is present at boot.
Remove __init from sched_tick_offload_init() and guard it with a
tick_work_cpu check so it is safe to call at runtime.

housekeeping_init() may now reject nohz_full=/isolcpus=nohz and clear
HK_FLAG_KERNEL_NOISE for violating the isolcpus=domain subset
invariant.  Clear HK_FLAG_KERNEL_NOISE_BOOT in that same path: DHM's
runtime first-enable anchors on the boot snapshot whenever that flag
is set, and the rejected configuration must not become the anchor.

Signed-off-by: Qiliang Yuan <odys.yuan@gmail.com>
---
 include/linux/sched/isolation.h |  32 +++++-
 kernel/sched/core.c             |   4 +-
 kernel/sched/isolation.c        | 209 +++++++++++++++++++++++++++++++++++-----
 kernel/sched/sched.h            |   2 +-
 4 files changed, 220 insertions(+), 27 deletions(-)

diff --git a/include/linux/sched/isolation.h b/include/linux/sched/isolation.h
index cf0fd03dd7a24..70602a74c1410 100644
--- a/include/linux/sched/isolation.h
+++ b/include/linux/sched/isolation.h
@@ -14,10 +14,26 @@ enum hk_type {
 	 * is always a subset of HK_TYPE_DOMAIN_BOOT.
 	 */
 	HK_TYPE_DOMAIN,
-	/* Inverse of boot-time isolcpus=managed_irq argument */
-	HK_TYPE_MANAGED_IRQ,
-	/* Inverse of boot-time nohz_full= or isolcpus=nohz arguments */
+
+	/*
+	 * Inverse of the boot-time nohz_full= or isolcpus=nohz arguments.
+	 * When neither is given, DHM still records cpu_possible_mask here so
+	 * that kernel-noise isolation can be enabled purely at runtime.
+	 */
+	HK_TYPE_KERNEL_NOISE_BOOT,
+	/*
+	 * A subset of HK_TYPE_KERNEL_NOISE_BOOT: it may exclude additional
+	 * CPUs isolated at runtime via cpuset isolated partitions.
+	 */
 	HK_TYPE_KERNEL_NOISE,
+
+	/* Inverse of the boot-time isolcpus=managed_irq argument */
+	HK_TYPE_MANAGED_IRQ_BOOT,
+	/*
+	 * A subset of HK_TYPE_MANAGED_IRQ_BOOT: it may exclude additional
+	 * CPUs isolated at runtime via cpuset isolated partitions.
+	 */
+	HK_TYPE_MANAGED_IRQ,
 	HK_TYPE_MAX,
 
 	/*
@@ -40,10 +56,13 @@ enum hk_type {
 DECLARE_STATIC_KEY_FALSE(housekeeping_overridden);
 extern int housekeeping_any_cpu(enum hk_type type);
 extern const struct cpumask *housekeeping_cpumask(enum hk_type type);
+extern const struct cpumask *housekeeping_cpumask_rcu(enum hk_type type);
 extern bool housekeeping_enabled(enum hk_type type);
 extern void housekeeping_affine(struct task_struct *t, enum hk_type type);
 extern bool housekeeping_test_cpu(int cpu, enum hk_type type);
 extern int housekeeping_update(struct cpumask *isol_mask);
+extern int housekeeping_update_types(unsigned long type_mask,
+				     struct cpumask *isol_mask);
 extern void __init housekeeping_init(void);
 
 #else
@@ -58,6 +77,11 @@ static inline const struct cpumask *housekeeping_cpumask(enum hk_type type)
 	return cpu_possible_mask;
 }
 
+static inline const struct cpumask *housekeeping_cpumask_rcu(enum hk_type type)
+{
+	return cpu_possible_mask;
+}
+
 static inline bool housekeeping_enabled(enum hk_type type)
 {
 	return false;
@@ -72,6 +96,8 @@ static inline bool housekeeping_test_cpu(int cpu, enum hk_type type)
 }
 
 static inline int housekeeping_update(struct cpumask *isol_mask) { return 0; }
+static inline int housekeeping_update_types(unsigned long type_mask,
+					    struct cpumask *isol_mask) { return 0; }
 static inline void housekeeping_init(void) { }
 #endif /* CONFIG_CPU_ISOLATION */
 
diff --git a/kernel/sched/core.c b/kernel/sched/core.c
index 1fe40de6ebe3c..6c67874e639a5 100644
--- a/kernel/sched/core.c
+++ b/kernel/sched/core.c
@@ -5960,8 +5960,10 @@ static void sched_tick_stop(int cpu)
 }
 #endif /* CONFIG_HOTPLUG_CPU */
 
-int __init sched_tick_offload_init(void)
+int sched_tick_offload_init(void)
 {
+	if (tick_work_cpu)
+		return 0;
 	tick_work_cpu = alloc_percpu(struct tick_work);
 	BUG_ON(!tick_work_cpu);
 	return 0;
diff --git a/kernel/sched/isolation.c b/kernel/sched/isolation.c
index c6a41095a9002..7725514ac290e 100644
--- a/kernel/sched/isolation.c
+++ b/kernel/sched/isolation.c
@@ -13,10 +13,12 @@
 #include "sched.h"
 
 enum hk_flags {
-	HK_FLAG_DOMAIN_BOOT	= BIT(HK_TYPE_DOMAIN_BOOT),
-	HK_FLAG_DOMAIN		= BIT(HK_TYPE_DOMAIN),
-	HK_FLAG_MANAGED_IRQ	= BIT(HK_TYPE_MANAGED_IRQ),
-	HK_FLAG_KERNEL_NOISE	= BIT(HK_TYPE_KERNEL_NOISE),
+	HK_FLAG_DOMAIN_BOOT	  = BIT(HK_TYPE_DOMAIN_BOOT),
+	HK_FLAG_DOMAIN		  = BIT(HK_TYPE_DOMAIN),
+	HK_FLAG_KERNEL_NOISE_BOOT = BIT(HK_TYPE_KERNEL_NOISE_BOOT),
+	HK_FLAG_KERNEL_NOISE	  = BIT(HK_TYPE_KERNEL_NOISE),
+	HK_FLAG_MANAGED_IRQ_BOOT  = BIT(HK_TYPE_MANAGED_IRQ_BOOT),
+	HK_FLAG_MANAGED_IRQ	  = BIT(HK_TYPE_MANAGED_IRQ),
 };
 
 DEFINE_STATIC_KEY_FALSE(housekeeping_overridden);
@@ -36,25 +38,40 @@ bool housekeeping_enabled(enum hk_type type)
 }
 EXPORT_SYMBOL_GPL(housekeeping_enabled);
 
+/*
+ * Types that can change at runtime via cpuset isolated partitions.
+ * Boot-only types (DOMAIN_BOOT) are always safe to read without lockdep.
+ */
+static bool housekeeping_type_can_change(enum hk_type type)
+{
+	switch (type) {
+	case HK_TYPE_DOMAIN:
+	case HK_TYPE_KERNEL_NOISE:
+	case HK_TYPE_MANAGED_IRQ:
+		return true;
+	default:
+		return false;
+	}
+}
+
 static bool housekeeping_dereference_check(enum hk_type type)
 {
-	if (IS_ENABLED(CONFIG_LOCKDEP) && type == HK_TYPE_DOMAIN) {
-		/* Cpuset isn't even writable yet? */
-		if (system_state <= SYSTEM_SCHEDULING)
-			return true;
+	if (!IS_ENABLED(CONFIG_LOCKDEP) || !housekeeping_type_can_change(type))
+		return true;
 
-		/* CPU hotplug write locked, so cpuset partition can't be overwritten */
-		if (IS_ENABLED(CONFIG_HOTPLUG_CPU) && lockdep_is_cpus_write_held())
-			return true;
+	/* Cpuset isn't even writable yet? */
+	if (system_state <= SYSTEM_SCHEDULING)
+		return true;
 
-		/* Cpuset lock held, partitions not writable */
-		if (IS_ENABLED(CONFIG_CPUSETS) && lockdep_is_cpuset_held())
-			return true;
+	/* CPU hotplug write locked, so cpuset partition can't be overwritten */
+	if (IS_ENABLED(CONFIG_HOTPLUG_CPU) && lockdep_is_cpus_write_held())
+		return true;
 
-		return false;
-	}
+	/* Cpuset lock held, partitions not writable */
+	if (IS_ENABLED(CONFIG_CPUSETS) && lockdep_is_cpuset_held())
+		return true;
 
-	return true;
+	return false;
 }
 
 static inline struct cpumask *housekeeping_cpumask_dereference(enum hk_type type)
@@ -77,12 +94,26 @@ const struct cpumask *housekeeping_cpumask(enum hk_type type)
 }
 EXPORT_SYMBOL_GPL(housekeeping_cpumask);
 
+const struct cpumask *housekeeping_cpumask_rcu(enum hk_type type)
+{
+	const struct cpumask *mask = NULL;
+
+	if (static_branch_unlikely(&housekeeping_overridden)) {
+		if (READ_ONCE(housekeeping.flags) & BIT(type))
+			mask = rcu_dereference(housekeeping.cpumasks[type]);
+	}
+	if (!mask)
+		mask = cpu_possible_mask;
+	return mask;
+}
+EXPORT_SYMBOL_GPL(housekeeping_cpumask_rcu);
+
 int housekeeping_any_cpu(enum hk_type type)
 {
 	int cpu;
 
 	if (static_branch_unlikely(&housekeeping_overridden)) {
-		if (housekeeping.flags & BIT(type)) {
+		if (READ_ONCE(housekeeping.flags) & BIT(type)) {
 			cpu = sched_numa_find_closest(housekeeping_cpumask(type), smp_processor_id());
 			if (cpu < nr_cpu_ids)
 				return cpu;
@@ -164,6 +195,132 @@ int housekeeping_update(struct cpumask *isol_mask)
 	return 0;
 }
 
+/**
+ * housekeeping_update_types - Update housekeeping masks for specified types
+ * @type_mask: Bitmask of housekeeping types to update
+ * @isol_mask: CPUs being added to the isolation set
+ *
+ * For each type in @type_mask, compute the trial mask as
+ * (boot snapshot & ~@isol_mask), validate it against @cpu_online_mask,
+ * then swap the RCU mask pointer and free the old mask after
+ * synchronize_rcu().  Anchoring on the immutable boot snapshot
+ * (HK_TYPE_*_BOOT) keeps the runtime mask a subset of the boot set and
+ * lets de-isolation restore the boot configuration exactly.
+ *
+ * The updated mask only takes effect for subsystems as CPUs cycle
+ * through hotplug; callers isolate CPUs via the CPU hotplug machinery so
+ * that tick, RCU and interrupt state is reconfigured by the existing
+ * online/offline callbacks rather than reconfigured in place.
+ *
+ * HK_TYPE_KERNEL_NOISE and HK_TYPE_MANAGED_IRQ also support runtime
+ * first-enable: when neither nohz_full=/isolcpus=nohz nor
+ * isolcpus=managed_irq was given at boot, no boot snapshot exists for
+ * the type, so cpu_possible_mask is the implicit boot set and the type
+ * flag is set in housekeeping.flags on the first call.
+ *
+ * Return: 0 on success, -ENOMEM on allocation failure, -EINVAL if
+ * a trial mask has no online CPUs.
+ */
+int housekeeping_update_types(unsigned long type_mask,
+			      struct cpumask *isol_mask)
+{
+	struct cpumask *trials[HK_TYPE_MAX] = {};
+	struct cpumask *old_masks[HK_TYPE_MAX] = {};
+	enum hk_type type;
+	int ret = 0;
+
+	for_each_set_bit(type, &type_mask, HK_TYPE_MAX) {
+		const struct cpumask *base;
+
+		if (type == HK_TYPE_DOMAIN_BOOT)
+			continue;
+		if (!housekeeping_enabled(type)) {
+			/*
+			 * HK_TYPE_KERNEL_NOISE and HK_TYPE_MANAGED_IRQ support
+			 * runtime first-enable for DHM isolated partitions
+			 * created without nohz_full=/isolcpus=managed_irq at
+			 * boot.  All other types must be boot-enabled.
+			 */
+			if (type != HK_TYPE_KERNEL_NOISE && type != HK_TYPE_MANAGED_IRQ)
+				continue;
+		}
+
+		/*
+		 * Compute the trial mask relative to the immutable boot
+		 * snapshot, never relative to the current (already shrunk)
+		 * mask.  Using the current mask would let it shrink
+		 * monotonically across isolation/de-isolation cycles and would
+		 * never recover CPUs once de-isolated.  Anchoring on the boot
+		 * snapshot keeps the runtime mask a subset of the boot set and
+		 * lets de-isolation restore exactly the boot configuration.
+		 *
+		 * HK_TYPE_KERNEL_NOISE additionally supports runtime
+		 * first-enable: when no nohz_full=/isolcpus=nohz was given at
+		 * boot, no boot snapshot exists, so cpu_possible_mask is the
+		 * implicit boot set.
+		 */
+		if (type == HK_TYPE_KERNEL_NOISE &&
+		    !(housekeeping.flags & HK_FLAG_KERNEL_NOISE_BOOT))
+			base = cpu_possible_mask;
+		else if (type == HK_TYPE_KERNEL_NOISE)
+			base = housekeeping_cpumask(HK_TYPE_KERNEL_NOISE_BOOT);
+		else if (type == HK_TYPE_MANAGED_IRQ)
+			base = housekeeping_cpumask(HK_TYPE_MANAGED_IRQ_BOOT);
+		else
+			base = housekeeping_cpumask(type);
+		trials[type] = kmalloc(cpumask_size(), GFP_KERNEL);
+		if (!trials[type]) {
+			ret = -ENOMEM;
+			goto err_free;
+		}
+		cpumask_andnot(trials[type], base, isol_mask);
+		if (!cpumask_intersects(trials[type], cpu_online_mask)) {
+			ret = -EINVAL;
+			goto err_free;
+		}
+	}
+
+	if (!housekeeping.flags) {
+		ret = -EINVAL;
+		goto err_free;
+	}
+
+	for_each_set_bit(type, &type_mask, HK_TYPE_MAX) {
+		if (!trials[type])
+			continue;
+		old_masks[type] = housekeeping_cpumask_dereference(type);
+		/* First-time runtime enable: register the type now. */
+		if (!housekeeping_enabled(type)) {
+			WRITE_ONCE(housekeeping.flags,
+				   housekeeping.flags | BIT(type));
+			/*
+			 * HK_TYPE_KERNEL_NOISE first-enable at runtime
+			 * (zero-boot-param path): tick offload percpu data
+			 * was never allocated at boot since nohz_full= was
+			 * absent.  Allocate it now before CPUs cycle through
+			 * hotplug and sched_tick_stop() dereferences
+			 * tick_work_cpu.
+			 */
+			if (type == HK_TYPE_KERNEL_NOISE)
+				WARN_ON_ONCE(sched_tick_offload_init());
+		}
+		rcu_assign_pointer(housekeeping.cpumasks[type], trials[type]);
+		trials[type] = NULL;
+	}
+
+	synchronize_rcu();
+
+	for_each_set_bit(type, &type_mask, HK_TYPE_MAX)
+		kfree(old_masks[type]);
+
+	return 0;
+
+err_free:
+	for_each_set_bit(type, &type_mask, HK_TYPE_MAX)
+		kfree(trials[type]);
+	return ret;
+}
+
 void __init housekeeping_init(void)
 {
 	enum hk_type type;
@@ -192,7 +349,15 @@ void __init housekeeping_init(void)
 			    housekeeping_cpumask(HK_TYPE_KERNEL_NOISE))) {
 		pr_warn("Housekeeping: nohz_full=/isolcpus=nohz must be a subset "
 			"of isolcpus=domain, disabling nohz_full\n");
-		housekeeping.flags &= ~HK_FLAG_KERNEL_NOISE;
+		/*
+		 * Clear HK_TYPE_KERNEL_NOISE_BOOT too, not just the live
+		 * type: DHM's housekeeping_update_types() anchors runtime
+		 * isolation on this boot snapshot when it is set, and the
+		 * rejected boot configuration must not become that anchor.
+		 * Clearing both leaves the system in the same state as if
+		 * nohz_full=/isolcpus=nohz had never been passed.
+		 */
+		housekeeping.flags &= ~(HK_FLAG_KERNEL_NOISE | HK_FLAG_KERNEL_NOISE_BOOT);
 		if (!housekeeping.flags)
 			return;
 	}
@@ -343,7 +508,7 @@ static int __init housekeeping_nohz_full_setup(char *str)
 {
 	unsigned long flags;
 
-	flags = HK_FLAG_KERNEL_NOISE;
+	flags = HK_FLAG_KERNEL_NOISE | HK_FLAG_KERNEL_NOISE_BOOT;
 
 	return housekeeping_setup(str, flags);
 }
@@ -362,7 +527,7 @@ static int __init housekeeping_isolcpus_setup(char *str)
 		 */
 		if (!strncmp(str, "nohz,", 5)) {
 			str += 5;
-			flags |= HK_FLAG_KERNEL_NOISE;
+			flags |= HK_FLAG_KERNEL_NOISE | HK_FLAG_KERNEL_NOISE_BOOT;
 			continue;
 		}
 
@@ -374,7 +539,7 @@ static int __init housekeeping_isolcpus_setup(char *str)
 
 		if (!strncmp(str, "managed_irq,", 12)) {
 			str += 12;
-			flags |= HK_FLAG_MANAGED_IRQ;
+			flags |= HK_FLAG_MANAGED_IRQ | HK_FLAG_MANAGED_IRQ_BOOT;
 			continue;
 		}
 
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index e656c7059bf86..a640accf1d890 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -3018,7 +3018,7 @@ extern void post_init_entity_util_avg(struct task_struct *p);
 
 #ifdef CONFIG_NO_HZ_FULL
 extern bool sched_can_stop_tick(struct rq *rq);
-extern int __init sched_tick_offload_init(void);
+extern int sched_tick_offload_init(void);
 
 /*
  * Tick may be needed by tasks in the runqueue depending on their policy and

-- 
2.43.0


^ permalink raw reply	[flat|nested] 14+ messages in thread

* [PATCH v5 03/12] sched/isolation: RCU-protect runtime-mutable housekeeping cpumask readers
  2026-10-02 13:10 [PATCH v5 00/12] Dynamic Housekeeping Management (DHM) via CPUSets Qiliang Yuan
  2026-10-02 13:10 ` [PATCH v5 01/12] sched/isolation: Enforce nohz_full as a subset of isolcpus=domain at boot Qiliang Yuan
  2026-10-02 13:10 ` [PATCH v5 02/12] sched/isolation: Add runtime housekeeping mask updates with boot snapshots Qiliang Yuan
@ 2026-10-02 13:10 ` Qiliang Yuan
  2026-10-02 13:10 ` [PATCH v5 04/12] cpuset: Drive kernel-noise housekeeping from isolated partitions Qiliang Yuan
                   ` (8 subsequent siblings)
  11 siblings, 0 replies; 14+ messages in thread
From: Qiliang Yuan @ 2026-10-02 13:10 UTC (permalink / raw)
  To: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot,
	Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, Paul E. McKenney, Frederic Weisbecker,
	Neeraj Upadhyay, Joel Fernandes, Josh Triplett, Boqun Feng,
	Uladzislau Rezki, Mathieu Desnoyers, Lai Jiangshan, Zqiang,
	Anna-Maria Behnsen, Tejun Heo, Jonathan Corbet, Shuah Khan,
	Shuah Khan, Thomas Gleixner
  Cc: Waiman Long, linux-kernel, rcu, cgroups, linux-doc,
	linux-kselftest, Qiliang Yuan

Now that HK_TYPE_KERNEL_NOISE and HK_TYPE_MANAGED_IRQ can be updated at
runtime, their cpumask pointers are swapped and the old masks freed after
an RCU grace period.  Readers that dereference these masks must do so
inside an RCU read-side critical section, otherwise the mask can be freed
while it is still in use.

Convert the runtime-mutable readers to housekeeping_cpumask_rcu() under
rcu_read_lock():

  - get_nohz_timer_target() (HK_TYPE_KERNEL_NOISE)
  - hrtimer target selection (HK_TYPE_TIMER)
  - arm64 topology (HK_TYPE_TICK)
  - Hyper-V channel management, both channel_mgmt.c and
    vmbus_channel_set_cpu() in vmbus_drv.c (HK_TYPE_MANAGED_IRQ)
  - the housekeeping sysfs attribute (HK_TYPE_KERNEL_NOISE)
  - rps_cpumask_housekeeping() (HK_TYPE_WQ, an alias of
    HK_TYPE_KERNEL_NOISE)
  - tmigr_isolated_exclude_cpumask() (HK_TYPE_KERNEL_NOISE)

The watchdog boot-time cpumask initialisation is switched from the
HK_TYPE_TIMER alias to HK_TYPE_KERNEL_NOISE for consistency; both alias
the same value.

Signed-off-by: Qiliang Yuan <odys.yuan@gmail.com>
---
 arch/arm64/kernel/topology.c  |  9 ++++++--
 drivers/base/cpu.c            | 20 ++++++++++++-----
 drivers/hv/channel_mgmt.c     | 50 +++++++++++++++++++++++++++++--------------
 drivers/hv/vmbus_drv.c        | 13 ++++++++++-
 kernel/sched/core.c           |  3 +--
 kernel/time/hrtimer.c         |  5 ++++-
 kernel/time/timer_migration.c |  9 +++++++-
 kernel/watchdog.c             |  2 +-
 net/core/net-sysfs.c          | 10 ++++++++-
 9 files changed, 91 insertions(+), 30 deletions(-)

diff --git a/arch/arm64/kernel/topology.c b/arch/arm64/kernel/topology.c
index d28438f8b83f1..1a7badffa45d4 100644
--- a/arch/arm64/kernel/topology.c
+++ b/arch/arm64/kernel/topology.c
@@ -212,8 +212,13 @@ int arch_freq_get_on_cpu(int cpu)
 			if (!policy)
 				return -EINVAL;
 
-			if (!cpumask_intersects(policy->related_cpus,
-						housekeeping_cpumask(HK_TYPE_TICK))) {
+			bool no_hk_in_policy;
+
+			rcu_read_lock();
+			no_hk_in_policy = !cpumask_intersects(policy->related_cpus,
+							      housekeeping_cpumask_rcu(HK_TYPE_TICK));
+			rcu_read_unlock();
+			if (no_hk_in_policy) {
 				cpufreq_cpu_put(policy);
 				return -EOPNOTSUPP;
 			}
diff --git a/drivers/base/cpu.c b/drivers/base/cpu.c
index 69e52fed42415..1f85fcbba867d 100644
--- a/drivers/base/cpu.c
+++ b/drivers/base/cpu.c
@@ -303,13 +303,23 @@ static DEVICE_ATTR(isolated, 0444, print_cpus_isolated, NULL);
 static ssize_t housekeeping_show(struct device *dev,
 			     struct device_attribute *attr, char *buf)
 {
-	const struct cpumask *hk_mask;
+	ssize_t len;
 
-	hk_mask = housekeeping_cpumask(HK_TYPE_KERNEL_NOISE);
+	if (!housekeeping_enabled(HK_TYPE_KERNEL_NOISE))
+		return sysfs_emit(buf, "\n");
 
-	if (housekeeping_enabled(HK_TYPE_KERNEL_NOISE))
-		return sysfs_emit(buf, "%*pbl\n", cpumask_pr_args(hk_mask));
-	return sysfs_emit(buf, "\n");
+	/*
+	 * HK_TYPE_KERNEL_NOISE is runtime-mutable: the mask pointer can be
+	 * swapped and the old mask freed after an RCU grace period.  Hold the
+	 * RCU read lock across the dereference and the format so the mask
+	 * cannot be freed while it is being printed.
+	 */
+	rcu_read_lock();
+	len = sysfs_emit(buf, "%*pbl\n",
+			 cpumask_pr_args(housekeeping_cpumask_rcu(HK_TYPE_KERNEL_NOISE)));
+	rcu_read_unlock();
+
+	return len;
 }
 static DEVICE_ATTR_RO(housekeeping);
 
diff --git a/drivers/hv/channel_mgmt.c b/drivers/hv/channel_mgmt.c
index a044fd3b3c4e7..c2aa01205c2cc 100644
--- a/drivers/hv/channel_mgmt.c
+++ b/drivers/hv/channel_mgmt.c
@@ -750,26 +750,43 @@ static void init_vp_index(struct vmbus_channel *channel)
 {
 	bool perf_chn = hv_is_perf_channel(channel);
 	u32 i, ncpu = num_online_cpus();
-	cpumask_var_t available_mask;
+	cpumask_var_t available_mask, hk_snap;
 	struct cpumask *allocated_mask;
-	const struct cpumask *hk_mask = housekeeping_cpumask(HK_TYPE_MANAGED_IRQ);
 	u32 target_cpu;
 	int numa_node;
 
-	if (!perf_chn ||
-	    !alloc_cpumask_var(&available_mask, GFP_KERNEL) ||
-	    cpumask_empty(hk_mask)) {
-		/*
-		 * If the channel is not a performance critical
-		 * channel, bind it to VMBUS_CONNECT_CPU.
-		 * In case alloc_cpumask_var() fails, bind it to
-		 * VMBUS_CONNECT_CPU.
-		 * If all the cpus are isolated, bind it to
-		 * VMBUS_CONNECT_CPU.
-		 */
+	if (!perf_chn) {
+		channel->target_cpu = VMBUS_CONNECT_CPU;
+		return;
+	}
+
+	if (!alloc_cpumask_var(&available_mask, GFP_KERNEL)) {
+		channel->target_cpu = VMBUS_CONNECT_CPU;
+		hv_set_allocated_cpu(VMBUS_CONNECT_CPU);
+		return;
+	}
+
+	/*
+	 * Snapshot HK_TYPE_MANAGED_IRQ cpumask under RCU read lock.
+	 * housekeeping_update_types() frees the old cpumask after
+	 * synchronize_rcu(), so we must not hold the pointer beyond an
+	 * RCU read-side critical section.
+	 */
+	if (!alloc_cpumask_var(&hk_snap, GFP_KERNEL)) {
+		free_cpumask_var(available_mask);
+		channel->target_cpu = VMBUS_CONNECT_CPU;
+		hv_set_allocated_cpu(VMBUS_CONNECT_CPU);
+		return;
+	}
+	rcu_read_lock();
+	cpumask_copy(hk_snap, housekeeping_cpumask_rcu(HK_TYPE_MANAGED_IRQ));
+	rcu_read_unlock();
+
+	if (cpumask_empty(hk_snap)) {
+		free_cpumask_var(hk_snap);
+		free_cpumask_var(available_mask);
 		channel->target_cpu = VMBUS_CONNECT_CPU;
-		if (perf_chn)
-			hv_set_allocated_cpu(VMBUS_CONNECT_CPU);
+		hv_set_allocated_cpu(VMBUS_CONNECT_CPU);
 		return;
 	}
 
@@ -788,7 +805,7 @@ static void init_vp_index(struct vmbus_channel *channel)
 
 retry:
 		cpumask_xor(available_mask, allocated_mask, cpumask_of_node(numa_node));
-		cpumask_and(available_mask, available_mask, hk_mask);
+		cpumask_and(available_mask, available_mask, hk_snap);
 
 		if (cpumask_empty(available_mask)) {
 			/*
@@ -809,6 +826,7 @@ static void init_vp_index(struct vmbus_channel *channel)
 
 	channel->target_cpu = target_cpu;
 
+	free_cpumask_var(hk_snap);
 	free_cpumask_var(available_mask);
 }
 
diff --git a/drivers/hv/vmbus_drv.c b/drivers/hv/vmbus_drv.c
index 5ebdbe24b5a1e..ac9b4800ed9fd 100644
--- a/drivers/hv/vmbus_drv.c
+++ b/drivers/hv/vmbus_drv.c
@@ -1734,6 +1734,7 @@ int vmbus_channel_set_cpu(struct vmbus_channel *channel, u32 target_cpu)
 {
 	u32 origin_cpu;
 	int ret = 0;
+	bool on_housekeeping_cpu;
 
 	lockdep_assert_cpus_held();
 	lockdep_assert_held(&vmbus_connection.channel_mutex);
@@ -1745,7 +1746,17 @@ int vmbus_channel_set_cpu(struct vmbus_channel *channel, u32 target_cpu)
 	if (target_cpu >= nr_cpumask_bits)
 		return -EINVAL;
 
-	if (!cpumask_test_cpu(target_cpu, housekeeping_cpumask(HK_TYPE_MANAGED_IRQ)))
+	/*
+	 * Snapshot the HK_TYPE_MANAGED_IRQ test under RCU read lock:
+	 * housekeeping_update_types() frees the old cpumask after
+	 * synchronize_rcu(), so the pointer must not be dereferenced
+	 * outside an RCU read-side critical section.
+	 */
+	rcu_read_lock();
+	on_housekeeping_cpu = cpumask_test_cpu(target_cpu,
+						housekeeping_cpumask_rcu(HK_TYPE_MANAGED_IRQ));
+	rcu_read_unlock();
+	if (!on_housekeeping_cpu)
 		return -EINVAL;
 
 	if (!cpu_online(target_cpu))
diff --git a/kernel/sched/core.c b/kernel/sched/core.c
index 6c67874e639a5..03a791a1dde9a 100644
--- a/kernel/sched/core.c
+++ b/kernel/sched/core.c
@@ -1301,9 +1301,8 @@ int get_nohz_timer_target(void)
 		default_cpu = cpu;
 	}
 
-	hk_mask = housekeeping_cpumask(HK_TYPE_KERNEL_NOISE);
-
 	guard(rcu)();
+	hk_mask = housekeeping_cpumask_rcu(HK_TYPE_KERNEL_NOISE);
 
 	for_each_domain(cpu, sd) {
 		for_each_cpu_and(i, sched_domain_span(sd), hk_mask) {
diff --git a/kernel/time/hrtimer.c b/kernel/time/hrtimer.c
index cbf1693c86b38..ac9d4eba7380e 100644
--- a/kernel/time/hrtimer.c
+++ b/kernel/time/hrtimer.c
@@ -243,8 +243,11 @@ static bool hrtimer_suitable_target(struct hrtimer *timer, struct hrtimer_clock_
 static inline struct hrtimer_cpu_base *get_target_base(struct hrtimer_cpu_base *base, bool pinned)
 {
 	if (!hrtimer_base_is_online(base)) {
-		int cpu = cpumask_any_and(cpu_online_mask, housekeeping_cpumask(HK_TYPE_TIMER));
+		int cpu;
 
+		rcu_read_lock();
+		cpu = cpumask_any_and(cpu_online_mask, housekeeping_cpumask_rcu(HK_TYPE_TIMER));
+		rcu_read_unlock();
 		return &per_cpu(hrtimer_bases, cpu);
 	}
 
diff --git a/kernel/time/timer_migration.c b/kernel/time/timer_migration.c
index 059d43355e650..f56748e9981b0 100644
--- a/kernel/time/timer_migration.c
+++ b/kernel/time/timer_migration.c
@@ -1631,7 +1631,14 @@ int tmigr_isolated_exclude_cpumask(struct cpumask *exclude_cpumask)
 	 * There cannot be overlap with the newly available ones.
 	 */
 	cpumask_and(cpumask, exclude_cpumask, tmigr_available_cpumask);
-	cpumask_and(cpumask, cpumask, housekeeping_cpumask(HK_TYPE_KERNEL_NOISE));
+	/*
+	 * HK_TYPE_KERNEL_NOISE is runtime-mutable: housekeeping_update_types()
+	 * frees the old cpumask after synchronize_rcu(), so dereference it
+	 * only under rcu_read_lock().
+	 */
+	rcu_read_lock();
+	cpumask_and(cpumask, cpumask, housekeeping_cpumask_rcu(HK_TYPE_KERNEL_NOISE));
+	rcu_read_unlock();
 	/*
 	 * Handle this here and not in the cpuset code because exclude_cpumask
 	 * might include also the tick CPU if included in isolcpus.
diff --git a/kernel/watchdog.c b/kernel/watchdog.c
index e5134ad7b6634..e567fbb0d4692 100644
--- a/kernel/watchdog.c
+++ b/kernel/watchdog.c
@@ -1389,7 +1389,7 @@ void __init lockup_detector_init(void)
 		pr_info("Disabling watchdog on nohz_full cores by default\n");
 
 	cpumask_copy(&watchdog_cpumask,
-		     housekeeping_cpumask(HK_TYPE_TIMER));
+		     housekeeping_cpumask(HK_TYPE_KERNEL_NOISE));
 
 	if (!watchdog_hardlockup_probe())
 		watchdog_hardlockup_available = true;
diff --git a/net/core/net-sysfs.c b/net/core/net-sysfs.c
index 352173df75785..15f19ac0f6df5 100644
--- a/net/core/net-sysfs.c
+++ b/net/core/net-sysfs.c
@@ -1017,7 +1017,15 @@ int rps_cpumask_housekeeping(struct cpumask *mask)
 {
 	if (!cpumask_empty(mask)) {
 		cpumask_and(mask, mask, housekeeping_cpumask(HK_TYPE_DOMAIN_BOOT));
-		cpumask_and(mask, mask, housekeeping_cpumask(HK_TYPE_WQ));
+		/*
+		 * HK_TYPE_WQ aliases HK_TYPE_KERNEL_NOISE, which is
+		 * runtime-mutable: housekeeping_update_types() frees the old
+		 * cpumask after synchronize_rcu(), so dereference it only
+		 * under rcu_read_lock().
+		 */
+		rcu_read_lock();
+		cpumask_and(mask, mask, housekeeping_cpumask_rcu(HK_TYPE_WQ));
+		rcu_read_unlock();
 		if (cpumask_empty(mask))
 			return -EINVAL;
 	}

-- 
2.43.0


^ permalink raw reply	[flat|nested] 14+ messages in thread

* [PATCH v5 04/12] cpuset: Drive kernel-noise housekeeping from isolated partitions
  2026-10-02 13:10 [PATCH v5 00/12] Dynamic Housekeeping Management (DHM) via CPUSets Qiliang Yuan
                   ` (2 preceding siblings ...)
  2026-10-02 13:10 ` [PATCH v5 03/12] sched/isolation: RCU-protect runtime-mutable housekeeping cpumask readers Qiliang Yuan
@ 2026-10-02 13:10 ` Qiliang Yuan
  2026-10-02 13:10 ` [PATCH v5 05/12] context_tracking: Allow runtime per-CPU user tracking enable/disable Qiliang Yuan
                   ` (7 subsequent siblings)
  11 siblings, 0 replies; 14+ messages in thread
From: Qiliang Yuan @ 2026-10-02 13:10 UTC (permalink / raw)
  To: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot,
	Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, Paul E. McKenney, Frederic Weisbecker,
	Neeraj Upadhyay, Joel Fernandes, Josh Triplett, Boqun Feng,
	Uladzislau Rezki, Mathieu Desnoyers, Lai Jiangshan, Zqiang,
	Anna-Maria Behnsen, Tejun Heo, Jonathan Corbet, Shuah Khan,
	Shuah Khan, Thomas Gleixner
  Cc: Waiman Long, linux-kernel, rcu, cgroups, linux-doc,
	linux-kselftest, Qiliang Yuan

An isolated cpuset partition already updates the HK_TYPE_DOMAIN
housekeeping mask.  Extend it to also update the kernel-noise masks
(HK_TYPE_KERNEL_NOISE and HK_TYPE_MANAGED_IRQ) so that creating or
destroying an isolated partition reconfigures the full set of
housekeeping cpumasks.

HK_TYPE_KERNEL_NOISE must stay a subset of HK_TYPE_DOMAIN: a CPU with
the tick suppressed still needs to be excluded from the normal sched
domain, or scheduler load balancing can keep nominating it as a
target.  Update the sched domain mask first and only touch the
kernel-noise types when that update succeeds, so the two masks can
never diverge for the same isolation request.

housekeeping_update() and housekeeping_update_types() are called after
dropping cpus_read_lock and cpuset_mutex, with only cpuset_top_mutex held
for mutual exclusion.

Signed-off-by: Qiliang Yuan <odys.yuan@gmail.com>
---
 kernel/cgroup/cpuset.c | 32 +++++++++++++++++++++++++++++---
 1 file changed, 29 insertions(+), 3 deletions(-)

diff --git a/kernel/cgroup/cpuset.c b/kernel/cgroup/cpuset.c
index 3f52717c19654..8b3bb034adf62 100644
--- a/kernel/cgroup/cpuset.c
+++ b/kernel/cgroup/cpuset.c
@@ -1431,17 +1431,43 @@ static void cpuset_update_sd_hk_unlock(void)
 		rebuild_sched_domains_locked();
 
 	if (update_housekeeping) {
+		static const unsigned long noise_types =
+			BIT(HK_TYPE_KERNEL_NOISE) | BIT(HK_TYPE_MANAGED_IRQ);
+		int ret;
+
 		update_housekeeping = false;
 		cpumask_copy(isolated_hk_cpus, isolated_cpus);
 
+		mutex_unlock(&cpuset_mutex);
+		cpus_read_unlock();
+
 		/*
 		 * housekeeping_update() is now called without holding
 		 * cpus_read_lock and cpuset_mutex. Only cpuset_top_mutex
 		 * is still being held for mutual exclusion.
 		 */
-		mutex_unlock(&cpuset_mutex);
-		cpus_read_unlock();
-		WARN_ON_ONCE(housekeeping_update(isolated_hk_cpus));
+
+		/*
+		 * Update the sched domain mask first; it must succeed
+		 * before the kernel-noise types because workqueue flush
+		 * and timer migration depend on the sched domain mask.
+		 */
+		ret = housekeeping_update(isolated_hk_cpus);
+		WARN_ON_ONCE(ret);
+
+		/*
+		 * Only touch the kernel-noise housekeeping masks
+		 * (HK_TYPE_KERNEL_NOISE and HK_TYPE_MANAGED_IRQ) once the
+		 * sched domain update above actually succeeded: HK_TYPE_
+		 * KERNEL_NOISE must stay a subset of HK_TYPE_DOMAIN, so a CPU
+		 * can never end up tick-suppressed while still scheduled as
+		 * part of the normal (non-isolated) sched domain.  The tick,
+		 * RCU and managed-interrupt state is reconfigured as the
+		 * affected CPUs are cycled through the CPU hotplug machinery.
+		 */
+		if (!ret)
+			WARN_ON_ONCE(housekeeping_update_types(noise_types,
+							       isolated_hk_cpus));
 		mutex_unlock(&cpuset_top_mutex);
 	} else {
 		cpuset_full_unlock();

-- 
2.43.0


^ permalink raw reply	[flat|nested] 14+ messages in thread

* [PATCH v5 05/12] context_tracking: Allow runtime per-CPU user tracking enable/disable
  2026-10-02 13:10 [PATCH v5 00/12] Dynamic Housekeeping Management (DHM) via CPUSets Qiliang Yuan
                   ` (3 preceding siblings ...)
  2026-10-02 13:10 ` [PATCH v5 04/12] cpuset: Drive kernel-noise housekeeping from isolated partitions Qiliang Yuan
@ 2026-10-02 13:10 ` Qiliang Yuan
  2026-10-02 13:10 ` [PATCH v5 06/12] rcu/nocb: Support lazy init for runtime CPU isolation Qiliang Yuan
                   ` (6 subsequent siblings)
  11 siblings, 0 replies; 14+ messages in thread
From: Qiliang Yuan @ 2026-10-02 13:10 UTC (permalink / raw)
  To: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot,
	Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, Paul E. McKenney, Frederic Weisbecker,
	Neeraj Upadhyay, Joel Fernandes, Josh Triplett, Boqun Feng,
	Uladzislau Rezki, Mathieu Desnoyers, Lai Jiangshan, Zqiang,
	Anna-Maria Behnsen, Tejun Heo, Jonathan Corbet, Shuah Khan,
	Shuah Khan, Thomas Gleixner
  Cc: Waiman Long, linux-kernel, rcu, cgroups, linux-doc,
	linux-kselftest, Qiliang Yuan

ct_cpu_track_user() and the context_tracking_key static key are currently
restricted to boot-time use: the key is __ro_after_init and the function
is __init with __initdata state.  This prevents enabling nohz_full context
tracking for CPUs isolated at runtime via cpuset partitions.

Split ct_cpu_track_user() into three functions:

  ct_cpu_track_user(cpu)      - sets per_cpu(context_tracking.active) and
                                increments context_tracking_key; callable
                                at runtime with the CPU offline.

  ct_cpu_untrack_user(cpu)    - reverses the above; for de-isolation.

  ct_cpu_track_user_init(cpu) - __init wrapper; calls ct_cpu_track_user()
                                and handles TIF_NOHZ / tasklist setup.

Change context_tracking_key from DEFINE_STATIC_KEY_FALSE_RO to
DEFINE_STATIC_KEY_FALSE so that static_branch_inc/dec() can be called
after the __ro_after_init window closes.

Update tick_nohz_init() to call ct_cpu_track_user_init() so boot
behaviour is unchanged.

This is a prerequisite for DHM (Dynamic Housekeeping Management) runtime
CPU noise isolation without boot parameters.

context_tracking_key is a single systemwide static branch, not a
per-CPU gate: once live, __ct_user_enter()/__ct_user_exit() run
unconditionally on every CPU, regardless of that CPU's own
context_tracking.active.  Going live happens through code patching,
and other CPUs only observe the patched code some time after
static_branch_inc() returns.  A CPU whose kernel<->user transition
lands in that window sees context_tracking_enabled() as still false
and silently skips recording it, leaving context_tracking.state
stuck, so the next traced kernel entry on that CPU wrongly trips
CT_WARN_ON(__ct_state() != CT_STATE_USER).  Reordering the enable and
a fixup sweep around each other cannot close this: whichever runs
last still has its own propagation delay to every other CPU.

Add context_tracking_activating, a plain per-CPU bool with no
code-patching delay of its own, and
context_tracking_enabled_or_activating() to test it alongside the
static key.  ct_cpu_track_user() sets it on every CPU via IPI
strictly before static_branch_inc(), and clears it via another IPI
only after static_branch_inc() returns, so every relevant call site
sees it in place for the whole window during which the static key
might not have propagated yet.  Route user_enter_irqoff(),
user_exit_irqoff(), the guest variants, CT_WARN_ON() and ct_state()
through it instead of the raw static key.  Also directly bootstrap
CT_STATE_USER for a CPU caught sitting in user mode by its
interrupted pt_regs, rather than leaving it to self-correct on its
own next transition.

Signed-off-by: Qiliang Yuan <odys.yuan@gmail.com>
---
 include/linux/context_tracking.h       | 16 +++---
 include/linux/context_tracking_state.h | 25 +++++++++-
 kernel/context_tracking.c              | 91 ++++++++++++++++++++++++++++++++--
 kernel/time/tick-sched.c               |  2 +-
 4 files changed, 122 insertions(+), 12 deletions(-)

diff --git a/include/linux/context_tracking.h b/include/linux/context_tracking.h
index af9fe87a09225..a83a2f1f9f9a9 100644
--- a/include/linux/context_tracking.h
+++ b/include/linux/context_tracking.h
@@ -12,6 +12,8 @@
 
 #ifdef CONFIG_CONTEXT_TRACKING_USER
 extern void ct_cpu_track_user(int cpu);
+extern void ct_cpu_untrack_user(int cpu);
+extern void __init ct_cpu_track_user_init(int cpu);
 
 /* Called with interrupts disabled.  */
 extern void __ct_user_enter(enum ctx_state state);
@@ -25,26 +27,26 @@ extern void user_exit_callable(void);
 
 static inline void user_enter(void)
 {
-	if (context_tracking_enabled())
+	if (context_tracking_enabled_or_activating())
 		ct_user_enter(CT_STATE_USER);
 
 }
 static inline void user_exit(void)
 {
-	if (context_tracking_enabled())
+	if (context_tracking_enabled_or_activating())
 		ct_user_exit(CT_STATE_USER);
 }
 
 /* Called with interrupts disabled.  */
 static __always_inline void user_enter_irqoff(void)
 {
-	if (context_tracking_enabled())
+	if (context_tracking_enabled_or_activating())
 		__ct_user_enter(CT_STATE_USER);
 
 }
 static __always_inline void user_exit_irqoff(void)
 {
-	if (context_tracking_enabled())
+	if (context_tracking_enabled_or_activating())
 		__ct_user_exit(CT_STATE_USER);
 }
 
@@ -74,7 +76,7 @@ static inline void exception_exit(enum ctx_state prev_ctx)
 
 static __always_inline bool context_tracking_guest_enter(void)
 {
-	if (context_tracking_enabled())
+	if (context_tracking_enabled_or_activating())
 		__ct_user_enter(CT_STATE_GUEST);
 
 	return context_tracking_enabled_this_cpu();
@@ -82,13 +84,13 @@ static __always_inline bool context_tracking_guest_enter(void)
 
 static __always_inline bool context_tracking_guest_exit(void)
 {
-	if (context_tracking_enabled())
+	if (context_tracking_enabled_or_activating())
 		__ct_user_exit(CT_STATE_GUEST);
 
 	return context_tracking_enabled_this_cpu();
 }
 
-#define CT_WARN_ON(cond) WARN_ON(context_tracking_enabled() && (cond))
+#define CT_WARN_ON(cond) WARN_ON(context_tracking_enabled_or_activating() && (cond))
 
 #else
 static inline void user_enter(void) { }
diff --git a/include/linux/context_tracking_state.h b/include/linux/context_tracking_state.h
index 0b81248aa03e2..25f87a9763313 100644
--- a/include/linux/context_tracking_state.h
+++ b/include/linux/context_tracking_state.h
@@ -138,6 +138,28 @@ static __always_inline bool context_tracking_enabled(void)
 	return static_branch_unlikely(&context_tracking_key);
 }
 
+/*
+ * context_tracking_key goes live via code patching, which other CPUs only
+ * observe some time after ct_cpu_track_user() calls static_branch_inc().
+ * A CPU whose kernel<->user transition lands in that window would see
+ * context_tracking_enabled() as still false and silently skip recording
+ * it, leaving context_tracking.state stale.  ct_cpu_track_user() sets
+ * context_tracking_activating on every CPU with an IPI strictly before
+ * calling static_branch_inc(), and clears it again with another IPI only
+ * after static_branch_inc() returns (so only once every CPU is
+ * guaranteed to already observe the branch as enabled).  Checking it
+ * here closes that window: every transition in between is recorded via
+ * the normal __ct_user_enter()/__ct_user_exit() path instead of being
+ * silently dropped.
+ */
+DECLARE_PER_CPU(bool, context_tracking_activating);
+
+static __always_inline bool context_tracking_enabled_or_activating(void)
+{
+	return context_tracking_enabled() ||
+	       unlikely(__this_cpu_read(context_tracking_activating));
+}
+
 static __always_inline bool context_tracking_enabled_cpu(int cpu)
 {
 	return context_tracking_enabled() && per_cpu(context_tracking.active, cpu);
@@ -159,7 +181,7 @@ static __always_inline int ct_state(void)
 {
 	int ret;
 
-	if (!context_tracking_enabled())
+	if (!context_tracking_enabled_or_activating())
 		return CT_STATE_DISABLED;
 
 	preempt_disable();
@@ -171,6 +193,7 @@ static __always_inline int ct_state(void)
 
 #else
 static __always_inline bool context_tracking_enabled(void) { return false; }
+static __always_inline bool context_tracking_enabled_or_activating(void) { return false; }
 static __always_inline bool context_tracking_enabled_cpu(int cpu) { return false; }
 static __always_inline bool context_tracking_enabled_this_cpu(void) { return false; }
 #endif /* CONFIG_CONTEXT_TRACKING_USER */
diff --git a/kernel/context_tracking.c b/kernel/context_tracking.c
index a743e7ffa6c00..326862a2679e3 100644
--- a/kernel/context_tracking.c
+++ b/kernel/context_tracking.c
@@ -23,6 +23,9 @@
 #include <linux/hardirq.h>
 #include <linux/export.h>
 #include <linux/kprobes.h>
+#include <linux/smp.h>
+#include <linux/ptrace.h>
+#include <asm/irq_regs.h>
 #include <trace/events/rcu.h>
 
 
@@ -411,9 +414,12 @@ static __always_inline void ct_kernel_enter(bool user, int offset) { }
 #define CREATE_TRACE_POINTS
 #include <trace/events/context_tracking.h>
 
-DEFINE_STATIC_KEY_FALSE_RO(context_tracking_key);
+DEFINE_STATIC_KEY_FALSE(context_tracking_key);
 EXPORT_SYMBOL_GPL(context_tracking_key);
 
+DEFINE_PER_CPU(bool, context_tracking_activating);
+EXPORT_SYMBOL_GPL(context_tracking_activating);
+
 static noinstr bool context_tracking_recursion_enter(void)
 {
 	int recursion;
@@ -674,14 +680,93 @@ void user_exit_callable(void)
 }
 NOKPROBE_SYMBOL(user_exit_callable);
 
-void __init ct_cpu_track_user(int cpu)
+/*
+ * context_tracking_key is a single systemwide static branch, not a per-CPU
+ * gate: once live, __ct_user_enter()/__ct_user_exit() run unconditionally
+ * on every CPU (so that a task migrating between a tracked and an
+ * untracked CPU always sees consistent state), regardless of that CPU's
+ * own context_tracking.active.  Going live happens through code patching,
+ * though, and other CPUs only observe the patched code some time after
+ * static_branch_inc() is called on this one.  A CPU whose kernel<->user
+ * transition lands in that window would see context_tracking_enabled()
+ * as still false and silently skip recording it, leaving
+ * context_tracking.state stuck at whatever it was, so the first traced
+ * kernel entry on that CPU afterwards would wrongly trip
+ * CT_WARN_ON(__ct_state() != CT_STATE_USER).  Reordering the two calls
+ * below cannot close this: whichever runs last still has its own
+ * propagation delay to every other CPU.
+ *
+ * context_tracking_enabled_or_activating() closes the window instead:
+ * every user_enter_irqoff()/user_exit_irqoff()/CT_WARN_ON() site treats
+ * a CPU as tracking once context_tracking_activating is set on it, with
+ * no code-patching delay of its own, since it is a plain per-CPU bool
+ * set directly by the interrupting IPI handler rather than inferred
+ * from a jump label.  Set it on every CPU before static_branch_inc(),
+ * and only clear it once static_branch_inc() has returned, so there is
+ * no gap during which a CPU observes neither signal: every transition
+ * in between is recorded through the normal path instead of being
+ * silently dropped.  Also directly bootstrap CT_STATE_USER for a CPU
+ * caught sitting in user mode (via its interrupted pt_regs), rather
+ * than leaving it to self-correct on its own next transition.
+ */
+static void ct_activate_set_pending_ipi(void *unused)
 {
-	static __initdata bool initialized = false;
+	struct pt_regs *regs = get_irq_regs();
 
+	__this_cpu_write(context_tracking_activating, true);
+	if (regs && user_mode(regs))
+		__ct_user_enter(CT_STATE_USER);
+}
+
+static void ct_activate_clear_pending_ipi(void *unused)
+{
+	__this_cpu_write(context_tracking_activating, false);
+}
+
+/**
+ * ct_cpu_track_user - enable context tracking for a CPU
+ * @cpu: target CPU (must be offline when called at runtime)
+ *
+ * Marks @cpu as actively tracking user/kernel transitions and increments
+ * the context_tracking_key refcount.  Safe to call at runtime provided
+ * the CPU is offline so no context-tracking readers are active on it.
+ */
+void ct_cpu_track_user(int cpu)
+{
 	if (!per_cpu(context_tracking.active, cpu)) {
+		bool first_activation = !context_tracking_enabled();
+
 		per_cpu(context_tracking.active, cpu) = true;
+		if (first_activation)
+			on_each_cpu(ct_activate_set_pending_ipi, NULL, 1);
 		static_branch_inc(&context_tracking_key);
+		if (first_activation)
+			on_each_cpu(ct_activate_clear_pending_ipi, NULL, 1);
 	}
+}
+EXPORT_SYMBOL_GPL(ct_cpu_track_user);
+
+/**
+ * ct_cpu_untrack_user - disable context tracking for a CPU
+ * @cpu: target CPU (must be offline when called)
+ *
+ * Reverses ct_cpu_track_user().  The CPU must be offline so that no
+ * context-tracking readers are active on it.
+ */
+void ct_cpu_untrack_user(int cpu)
+{
+	if (per_cpu(context_tracking.active, cpu)) {
+		per_cpu(context_tracking.active, cpu) = false;
+		static_branch_dec(&context_tracking_key);
+	}
+}
+EXPORT_SYMBOL_GPL(ct_cpu_untrack_user);
+
+void __init ct_cpu_track_user_init(int cpu)
+{
+	static __initdata bool initialized = false;
+
+	ct_cpu_track_user(cpu);
 
 	if (initialized)
 		return;
diff --git a/kernel/time/tick-sched.c b/kernel/time/tick-sched.c
index 6c3fea3867139..8c53754c4ac4a 100644
--- a/kernel/time/tick-sched.c
+++ b/kernel/time/tick-sched.c
@@ -675,7 +675,7 @@ void __init tick_nohz_init(void)
 	}
 
 	for_each_cpu(cpu, tick_nohz_full_mask)
-		ct_cpu_track_user(cpu);
+		ct_cpu_track_user_init(cpu);
 
 	ret = cpuhp_setup_state_nocalls(CPUHP_AP_ONLINE_DYN,
 					"kernel/nohz:predown", NULL,

-- 
2.43.0


^ permalink raw reply	[flat|nested] 14+ messages in thread

* [PATCH v5 06/12] rcu/nocb: Support lazy init for runtime CPU isolation
  2026-10-02 13:10 [PATCH v5 00/12] Dynamic Housekeeping Management (DHM) via CPUSets Qiliang Yuan
                   ` (4 preceding siblings ...)
  2026-10-02 13:10 ` [PATCH v5 05/12] context_tracking: Allow runtime per-CPU user tracking enable/disable Qiliang Yuan
@ 2026-10-02 13:10 ` Qiliang Yuan
  2026-10-02 13:10 ` [PATCH v5 07/12] watchdog: Sync watchdog_cpumask with HK_TYPE_KERNEL_NOISE on isolation Qiliang Yuan
                   ` (5 subsequent siblings)
  11 siblings, 0 replies; 14+ messages in thread
From: Qiliang Yuan @ 2026-10-02 13:10 UTC (permalink / raw)
  To: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot,
	Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, Paul E. McKenney, Frederic Weisbecker,
	Neeraj Upadhyay, Joel Fernandes, Josh Triplett, Boqun Feng,
	Uladzislau Rezki, Mathieu Desnoyers, Lai Jiangshan, Zqiang,
	Anna-Maria Behnsen, Tejun Heo, Jonathan Corbet, Shuah Khan,
	Shuah Khan, Thomas Gleixner
  Cc: Waiman Long, linux-kernel, rcu, cgroups, linux-doc,
	linux-kselftest, Qiliang Yuan

When a cpuset isolated partition first requests kernel-noise isolation,
it needs to offload RCU callbacks from the affected CPUs. The existing
rcu_nocb_cpu_offload() requires rcu_nocbs= or nohz_full= at boot so
that rcu_organize_nocb_kthreads() has already set rdp->nocb_gp_rdp for
each CPU. Without a boot parameter, nocb_gp_rdp is NULL and offload
fails immediately.

Introduce lazy nocb initialization so that the first call into the
isolation path triggers the one-time setup automatically:

  rcu_nocb_lazy_init() - allocates rcu_nocb_mask and calls
    rcu_organize_nocb_kthreads() (now without __init) to set
    nocb_gp_rdp for every possible CPU. Uses a dedicated mutex
    for serialization with a fast-path read of nocb_is_setup.
    Registers the CONFIG_RCU_LAZY shrinker on this first-ever
    setup too: rcu_init_nohz() only does so when nocb_is_setup was
    already true at boot, so the lazy-callback reclaim path would
    otherwise never exist on a system with no rcu_nocbs=/nohz_full=
    boot parameter.  Factor the registration into
    rcu_nocb_register_lazy_shrinker(), shared with rcu_init_nohz().

  rcu_nocb_cpu_isolate() - exported entry point called per-CPU
    while the CPU is offline. Calls rcu_nocb_lazy_init() for the
    one-time setup, spawns the GP and CB kthreads via the existing
    rcu_spawn_cpu_nocb_kthread(), then finalizes offload through
    rcu_nocb_cpu_offload(). Adding the CPU to the GP kthread's
    nocb_head_rdp list is handled by nocb_gp_toggle_rdp() in the
    GP kthread, so rcu_organize_nocb_kthreads() can run with an
    empty mask.

Remove __init from rcu_organize_nocb_kthreads() to allow this
runtime call path; the function itself has no __initdata dependencies.

Signed-off-by: Qiliang Yuan <odys.yuan@gmail.com>
---
 include/linux/rcupdate.h |  2 ++
 kernel/rcu/tree.h        |  2 +-
 kernel/rcu/tree_nocb.h   | 81 ++++++++++++++++++++++++++++++++++++++++--------
 3 files changed, 71 insertions(+), 14 deletions(-)

diff --git a/include/linux/rcupdate.h b/include/linux/rcupdate.h
index 44c07a66edfff..1425f04ceb4e9 100644
--- a/include/linux/rcupdate.h
+++ b/include/linux/rcupdate.h
@@ -158,6 +158,7 @@ static __always_inline void rcu_irq_work_resched(void) { }
 void rcu_init_nohz(void);
 int rcu_nocb_cpu_offload(int cpu);
 int rcu_nocb_cpu_deoffload(int cpu);
+int rcu_nocb_cpu_isolate(int cpu);
 void rcu_nocb_flush_deferred_wakeup(void);
 
 #define RCU_NOCB_LOCKDEP_WARN(c, s) RCU_LOCKDEP_WARN(c, s)
@@ -167,6 +168,7 @@ void rcu_nocb_flush_deferred_wakeup(void);
 static inline void rcu_init_nohz(void) { }
 static inline int rcu_nocb_cpu_offload(int cpu) { return -EINVAL; }
 static inline int rcu_nocb_cpu_deoffload(int cpu) { return 0; }
+static inline int rcu_nocb_cpu_isolate(int cpu) { return -EINVAL; }
 static inline void rcu_nocb_flush_deferred_wakeup(void) { }
 
 #define RCU_NOCB_LOCKDEP_WARN(c, s)
diff --git a/kernel/rcu/tree.h b/kernel/rcu/tree.h
index eedfa43059e80..60bcf52817f3e 100644
--- a/kernel/rcu/tree.h
+++ b/kernel/rcu/tree.h
@@ -522,7 +522,7 @@ static void rcu_nocb_unlock_irqrestore(struct rcu_data *rdp,
 				       unsigned long flags);
 static void rcu_lockdep_assert_cblist_protected(struct rcu_data *rdp);
 #ifdef CONFIG_RCU_NOCB_CPU
-static void __init rcu_organize_nocb_kthreads(void);
+static void rcu_organize_nocb_kthreads(void);
 
 /*
  * Disable IRQs before checking offloaded state so that local
diff --git a/kernel/rcu/tree_nocb.h b/kernel/rcu/tree_nocb.h
index 19bb42672baf8..99f3e4cec34e3 100644
--- a/kernel/rcu/tree_nocb.h
+++ b/kernel/rcu/tree_nocb.h
@@ -1344,12 +1344,30 @@ lazy_rcu_shrink_scan(struct shrinker *shrink, struct shrink_control *sc)
 }
 #endif // #ifdef CONFIG_RCU_LAZY
 
+#ifdef CONFIG_RCU_LAZY
+static void rcu_nocb_register_lazy_shrinker(void)
+{
+	struct shrinker *lazy_rcu_shrinker = shrinker_alloc(0, "rcu-lazy");
+
+	if (!lazy_rcu_shrinker) {
+		pr_err("Failed to allocate lazy_rcu shrinker!\n");
+		return;
+	}
+
+	lazy_rcu_shrinker->count_objects = lazy_rcu_shrink_count;
+	lazy_rcu_shrinker->scan_objects = lazy_rcu_shrink_scan;
+
+	shrinker_register(lazy_rcu_shrinker);
+}
+#else
+static void rcu_nocb_register_lazy_shrinker(void) { }
+#endif // #ifdef CONFIG_RCU_LAZY
+
 void __init rcu_init_nohz(void)
 {
 	int cpu;
 	struct rcu_data *rdp;
 	const struct cpumask *cpumask = NULL;
-	struct shrinker * __maybe_unused lazy_rcu_shrinker;
 
 #if defined(CONFIG_NO_HZ_FULL)
 	if (tick_nohz_full_running && !cpumask_empty(tick_nohz_full_mask))
@@ -1375,17 +1393,7 @@ void __init rcu_init_nohz(void)
 	if (!rcu_state.nocb_is_setup)
 		return;
 
-#ifdef CONFIG_RCU_LAZY
-	lazy_rcu_shrinker = shrinker_alloc(0, "rcu-lazy");
-	if (!lazy_rcu_shrinker) {
-		pr_err("Failed to allocate lazy_rcu shrinker!\n");
-	} else {
-		lazy_rcu_shrinker->count_objects = lazy_rcu_shrink_count;
-		lazy_rcu_shrinker->scan_objects = lazy_rcu_shrink_scan;
-
-		shrinker_register(lazy_rcu_shrinker);
-	}
-#endif // #ifdef CONFIG_RCU_LAZY
+	rcu_nocb_register_lazy_shrinker();
 
 	if (!cpumask_subset(rcu_nocb_mask, cpu_possible_mask)) {
 		pr_info("\tNote: kernel parameter 'rcu_nocbs=', 'nohz_full', or 'isolcpus=' contains nonexistent CPUs.\n");
@@ -1409,6 +1417,53 @@ void __init rcu_init_nohz(void)
 	rcu_organize_nocb_kthreads();
 }
 
+static DEFINE_MUTEX(rcu_nocb_lazy_mutex);
+
+/*
+ * Lazily initialize nocb infrastructure on the first call. Allocates
+ * rcu_nocb_mask and sets nocb_gp_rdp for every possible CPU so that
+ * rcu_nocb_cpu_isolate() can offload callbacks without rcu_nocbs= at boot.
+ */
+static noinline int rcu_nocb_lazy_init(void)
+{
+	if (rcu_state.nocb_is_setup)
+		return 0;
+
+	mutex_lock(&rcu_nocb_lazy_mutex);
+	if (!rcu_state.nocb_is_setup) {
+		if (!zalloc_cpumask_var(&rcu_nocb_mask, GFP_KERNEL)) {
+			mutex_unlock(&rcu_nocb_lazy_mutex);
+			return -ENOMEM;
+		}
+		rcu_organize_nocb_kthreads();
+		/*
+		 * Boot-time rcu_init_nohz() only registers the lazy-callback
+		 * shrinker when nocb_is_setup was already true at boot; this
+		 * is the first time it is set, so register it here instead.
+		 */
+		rcu_nocb_register_lazy_shrinker();
+		rcu_state.nocb_is_setup = true;
+	}
+	mutex_unlock(&rcu_nocb_lazy_mutex);
+	return 0;
+}
+
+/*
+ * Offload RCU callbacks for a CPU entering a kernel-noise isolated partition.
+ * @cpu must be offline. Lazily initializes nocb infrastructure on first use.
+ */
+int rcu_nocb_cpu_isolate(int cpu)
+{
+	int ret;
+
+	ret = rcu_nocb_lazy_init();
+	if (ret)
+		return ret;
+	rcu_spawn_cpu_nocb_kthread(cpu);
+	return rcu_nocb_cpu_offload(cpu);
+}
+EXPORT_SYMBOL_GPL(rcu_nocb_cpu_isolate);
+
 /* Initialize per-rcu_data variables for no-CBs CPUs. */
 static void __init rcu_boot_init_nocb_percpu_data(struct rcu_data *rdp)
 {
@@ -1501,7 +1556,7 @@ module_param(rcu_nocb_gp_stride, int, 0444);
 /*
  * Initialize GP-CB relationships for all no-CBs CPU.
  */
-static void __init rcu_organize_nocb_kthreads(void)
+static void rcu_organize_nocb_kthreads(void)
 {
 	int cpu;
 	bool firsttime = true;

-- 
2.43.0


^ permalink raw reply	[flat|nested] 14+ messages in thread

* [PATCH v5 07/12] watchdog: Sync watchdog_cpumask with HK_TYPE_KERNEL_NOISE on isolation
  2026-10-02 13:10 [PATCH v5 00/12] Dynamic Housekeeping Management (DHM) via CPUSets Qiliang Yuan
                   ` (5 preceding siblings ...)
  2026-10-02 13:10 ` [PATCH v5 06/12] rcu/nocb: Support lazy init for runtime CPU isolation Qiliang Yuan
@ 2026-10-02 13:10 ` Qiliang Yuan
  2026-10-02 15:26   ` Bradley Morgan
  2026-10-02 13:10 ` [PATCH v5 08/12] tick/nohz: Derive full-dynticks state from HK_TYPE_KERNEL_NOISE Qiliang Yuan
                   ` (4 subsequent siblings)
  11 siblings, 1 reply; 14+ messages in thread
From: Qiliang Yuan @ 2026-10-02 13:10 UTC (permalink / raw)
  To: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot,
	Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, Paul E. McKenney, Frederic Weisbecker,
	Neeraj Upadhyay, Joel Fernandes, Josh Triplett, Boqun Feng,
	Uladzislau Rezki, Mathieu Desnoyers, Lai Jiangshan, Zqiang,
	Anna-Maria Behnsen, Tejun Heo, Jonathan Corbet, Shuah Khan,
	Shuah Khan, Thomas Gleixner
  Cc: Waiman Long, linux-kernel, rcu, cgroups, linux-doc,
	linux-kselftest, Qiliang Yuan

The watchdog is initialized at boot to run on all housekeeping CPUs
(HK_TYPE_KERNEL_NOISE). When a cpuset isolated partition removes CPUs
from that mask at runtime, watchdog continues running on those CPUs
because nothing updates watchdog_cpumask.

Save the boot-time watchdog_cpumask as watchdog_cpumask_boot, which
captures the user's intended coverage (possibly narrowed via kernel
parameter or sysctl) before any runtime isolation. Introduce
lockup_detector_hk_update() which intersects this boot snapshot with
the current HK_TYPE_KERNEL_NOISE mask and reconfigures the detector.
This ensures that isolated CPUs are excluded while honoring any
manual narrowing the admin applied at or after boot.

lockup_detector_hk_update() snapshots the RCU-protected housekeeping
mask under rcu_read_lock(), then updates watchdog_cpumask and calls
__lockup_detector_reconfigure() under watchdog_mutex, matching the
same locking discipline used by proc_watchdog_cpumask().

Signed-off-by: Qiliang Yuan <odys.yuan@gmail.com>
---
 include/linux/nmi.h |  2 ++
 kernel/watchdog.c   | 24 ++++++++++++++++++++++++
 2 files changed, 26 insertions(+)

diff --git a/include/linux/nmi.h b/include/linux/nmi.h
index af69712df5f48..409af7884d939 100644
--- a/include/linux/nmi.h
+++ b/include/linux/nmi.h
@@ -37,6 +37,7 @@ extern int sysctl_hardlockup_all_cpu_backtrace;
 static inline void lockup_detector_init(void) { }
 static inline void lockup_detector_retry_init(void) { }
 static inline void lockup_detector_soft_poweroff(void) { }
+static inline void lockup_detector_hk_update(void) { }
 #endif /* !CONFIG_LOCKUP_DETECTOR */
 
 #ifdef CONFIG_SOFTLOCKUP_DETECTOR
@@ -120,6 +121,7 @@ void watchdog_hardlockup_enable(unsigned int cpu);
 void watchdog_hardlockup_disable(unsigned int cpu);
 
 void lockup_detector_reconfigure(void);
+void lockup_detector_hk_update(void);
 
 #ifdef CONFIG_HARDLOCKUP_DETECTOR_BUDDY
 void watchdog_buddy_check_hardlockup(int hrtimer_interrupts);
diff --git a/kernel/watchdog.c b/kernel/watchdog.c
index e567fbb0d4692..d4eccb8e337f1 100644
--- a/kernel/watchdog.c
+++ b/kernel/watchdog.c
@@ -53,6 +53,8 @@ static int __read_mostly watchdog_hardlockup_available;
 
 struct cpumask watchdog_cpumask __read_mostly;
 unsigned long *watchdog_cpumask_bits = cpumask_bits(&watchdog_cpumask);
+/* Boot snapshot: user's intended watchdog mask before any runtime isolation. */
+static struct cpumask watchdog_cpumask_boot __ro_after_init;
 
 #ifdef CONFIG_HARDLOCKUP_DETECTOR
 
@@ -1348,6 +1350,27 @@ static void __init lockup_detector_delay_init(struct work_struct *work)
 	lockup_detector_setup();
 }
 
+void lockup_detector_hk_update(void)
+{
+	cpumask_var_t new_mask;
+
+	if (!alloc_cpumask_var(&new_mask, GFP_KERNEL))
+		return;
+
+	rcu_read_lock();
+	cpumask_and(new_mask, &watchdog_cpumask_boot,
+		    housekeeping_cpumask_rcu(HK_TYPE_KERNEL_NOISE));
+	rcu_read_unlock();
+
+	mutex_lock(&watchdog_mutex);
+	cpumask_copy(&watchdog_cpumask, new_mask);
+	__lockup_detector_reconfigure(false);
+	mutex_unlock(&watchdog_mutex);
+
+	free_cpumask_var(new_mask);
+}
+EXPORT_SYMBOL_GPL(lockup_detector_hk_update);
+
 /*
  * lockup_detector_retry_init - retry init lockup detector if possible.
  *
@@ -1390,6 +1413,7 @@ void __init lockup_detector_init(void)
 
 	cpumask_copy(&watchdog_cpumask,
 		     housekeeping_cpumask(HK_TYPE_KERNEL_NOISE));
+	cpumask_copy(&watchdog_cpumask_boot, &watchdog_cpumask);
 
 	if (!watchdog_hardlockup_probe())
 		watchdog_hardlockup_available = true;

-- 
2.43.0


^ permalink raw reply	[flat|nested] 14+ messages in thread

* [PATCH v5 08/12] tick/nohz: Derive full-dynticks state from HK_TYPE_KERNEL_NOISE
  2026-10-02 13:10 [PATCH v5 00/12] Dynamic Housekeeping Management (DHM) via CPUSets Qiliang Yuan
                   ` (6 preceding siblings ...)
  2026-10-02 13:10 ` [PATCH v5 07/12] watchdog: Sync watchdog_cpumask with HK_TYPE_KERNEL_NOISE on isolation Qiliang Yuan
@ 2026-10-02 13:10 ` Qiliang Yuan
  2026-10-02 13:10 ` [PATCH v5 09/12] cpuset: Add dhm_cycling_cpus mask to suppress transient invalidation Qiliang Yuan
                   ` (3 subsequent siblings)
  11 siblings, 0 replies; 14+ messages in thread
From: Qiliang Yuan @ 2026-10-02 13:10 UTC (permalink / raw)
  To: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot,
	Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, Paul E. McKenney, Frederic Weisbecker,
	Neeraj Upadhyay, Joel Fernandes, Josh Triplett, Boqun Feng,
	Uladzislau Rezki, Mathieu Desnoyers, Lai Jiangshan, Zqiang,
	Anna-Maria Behnsen, Tejun Heo, Jonathan Corbet, Shuah Khan,
	Shuah Khan, Thomas Gleixner
  Cc: Waiman Long, linux-kernel, rcu, cgroups, linux-doc,
	linux-kselftest, Qiliang Yuan

tick_nohz_full_mask and tick_nohz_full_running duplicate state already
tracked by housekeeping: HK_TYPE_KERNEL_NOISE's cpumask is the
complement of the former, and housekeeping_enabled(HK_TYPE_KERNEL_NOISE)
the same as the latter.  Keeping both in sync is itself a source of
bugs now that DHM makes HK_TYPE_KERNEL_NOISE runtime-mutable.

Turn tick_nohz_full_enabled()/tick_nohz_full_cpu() from header inlines
into real functions in tick-sched.c that query housekeeping directly,
avoiding a tick.h <-> sched/isolation.h include cycle.  Remove
tick_nohz_full_mask, tick_nohz_full_running and tick_nohz_full_setup():
boot setup already records the same information in
housekeeping.cpumasks[HK_TYPE_KERNEL_NOISE].  Add
housekeeping_disable_type() for the one caller (tick_nohz_init()'s
arch-capability fallback) that needs to fully turn a type back off
after boot parsing already enabled it.

tick_nohz_cpu_isolate()/tick_nohz_cpu_deisolate() no longer need their
own mutex or mask bookkeeping: housekeeping_update_types() has already
updated the mask by the time they run, so they reduce to the
ct_cpu_track_user()/ct_cpu_untrack_user() context-tracking toggle.

Update the other direct readers (RCU's rcu_init_nohz(), the
nohz_full/housekeeping sysfs files in drivers/base/cpu.c, and
resctrl's cpumask_any_housekeeping()) to compute the full-dynticks set
as the complement of housekeeping_cpumask_rcu(HK_TYPE_KERNEL_NOISE)
instead of reading the removed mask.

tick_do_timer_cpu's hotplug protection and the "duty never relinquishes"
assertion need the same housekeeping-derived treatment, but not the
same predicate: the assertion in tick_sched_do_timer() should only
fire while a full-dynticks CPU genuinely exists right now, whereas
the hotplug protection in tick_nohz_cpu_hotpluggable() must stay
active across DHM's single-CPU isolate/de-isolate cycle even though
the published mask briefly looks empty.

Add tick_nohz_full_live(), checking HK_TYPE_KERNEL_NOISE is both
enabled and currently isolating at least one CPU, and use it for the
tick_sched_do_timer() assertion: HK_TYPE_KERNEL_NOISE can stay
permanently enabled after DHM's first runtime isolation even once
every CPU has been de-isolated again, and an enabled type with an
empty mask is an ordinary NO_HZ_IDLE duty handover, not a violation.

Keep tick_nohz_cpu_hotpluggable() on the plain housekeeping_enabled()
check instead: DHM's cpuset_update_sd_hk_unlock() only calls
housekeeping_update_types() to publish a new CPU's isolation after
remove_cpu() on it has already succeeded, so at the exact moment that
remove_cpu() call reaches this hotplug check, the live mask still
reflects the state from before this isolation and tick_nohz_full_live()
would see it as empty for every single-CPU isolation, not just the
very first one, leaving the actual tick_do_timer_cpu holder
unprotected each time.

Boot-time nohz_full=/isolcpus=nohz reaches the hotplug protection via
tick_nohz_init(), which is never called for DHM's zero-boot-param
runtime path since nohz_full= was never set at boot.  Factor the
cpuhp_setup_state_nocalls() call out of tick_nohz_init() into a new
tick_nohz_full_hotplug_init(), and call it from
housekeeping_update_types()'s HK_TYPE_KERNEL_NOISE first-enable path
as well, alongside the existing sched_tick_offload_init() call there.

Signed-off-by: Qiliang Yuan <odys.yuan@gmail.com>
---
 drivers/base/cpu.c              |  22 ++++--
 fs/resctrl/internal.h           |   6 +-
 include/linux/sched/isolation.h |   2 +
 include/linux/tick.h            |  40 +++--------
 kernel/rcu/tree_nocb.h          |  24 +++++--
 kernel/sched/isolation.c        |  26 +++++--
 kernel/time/tick-sched.c        | 156 +++++++++++++++++++++++++++++++++-------
 7 files changed, 205 insertions(+), 71 deletions(-)

diff --git a/drivers/base/cpu.c b/drivers/base/cpu.c
index 1f85fcbba867d..c492abd69fa1b 100644
--- a/drivers/base/cpu.c
+++ b/drivers/base/cpu.c
@@ -328,10 +328,24 @@ static ssize_t nohz_full_show(struct device *dev,
 				    struct device_attribute *attr,
 				    char *buf)
 {
-	if (cpumask_available(tick_nohz_full_mask))
-		return sysfs_emit(buf, "%*pbl\n",
-				  cpumask_pr_args(tick_nohz_full_mask));
-	return sysfs_emit(buf, "\n");
+	cpumask_var_t full_mask;
+	ssize_t len;
+
+	if (!housekeeping_enabled(HK_TYPE_KERNEL_NOISE))
+		return sysfs_emit(buf, "\n");
+
+	if (!alloc_cpumask_var(&full_mask, GFP_KERNEL))
+		return sysfs_emit(buf, "\n");
+
+	/* Full-dynticks CPUs are the complement of the housekeeping set. */
+	rcu_read_lock();
+	cpumask_andnot(full_mask, cpu_possible_mask,
+		       housekeeping_cpumask_rcu(HK_TYPE_KERNEL_NOISE));
+	rcu_read_unlock();
+
+	len = sysfs_emit(buf, "%*pbl\n", cpumask_pr_args(full_mask));
+	free_cpumask_var(full_mask);
+	return len;
 }
 static DEVICE_ATTR_RO(nohz_full);
 #endif
diff --git a/fs/resctrl/internal.h b/fs/resctrl/internal.h
index e62a277dee850..99c28507c1366 100644
--- a/fs/resctrl/internal.h
+++ b/fs/resctrl/internal.h
@@ -6,6 +6,7 @@
 #include <linux/kernfs.h>
 #include <linux/fs_context.h>
 #include <linux/tick.h>
+#include <linux/sched/isolation.h>
 
 #define CQM_LIMBOCHECK_INTERVAL	1000
 
@@ -28,7 +29,10 @@ cpumask_any_housekeeping(const struct cpumask *mask, int exclude_cpu)
 
 	/* Try to find a CPU that isn't nohz_full to use in preference */
 	if (tick_nohz_full_enabled()) {
-		cpu = cpumask_any_andnot_but(mask, tick_nohz_full_mask, exclude_cpu);
+		rcu_read_lock();
+		cpu = cpumask_any_and_but(mask, housekeeping_cpumask_rcu(HK_TYPE_KERNEL_NOISE),
+					  exclude_cpu);
+		rcu_read_unlock();
 		if (cpu < nr_cpu_ids)
 			return cpu;
 	}
diff --git a/include/linux/sched/isolation.h b/include/linux/sched/isolation.h
index 70602a74c1410..327c7b71bafe1 100644
--- a/include/linux/sched/isolation.h
+++ b/include/linux/sched/isolation.h
@@ -64,6 +64,7 @@ extern int housekeeping_update(struct cpumask *isol_mask);
 extern int housekeeping_update_types(unsigned long type_mask,
 				     struct cpumask *isol_mask);
 extern void __init housekeeping_init(void);
+extern void __init housekeeping_disable_type(enum hk_type type);
 
 #else
 
@@ -99,6 +100,7 @@ static inline int housekeeping_update(struct cpumask *isol_mask) { return 0; }
 static inline int housekeeping_update_types(unsigned long type_mask,
 					    struct cpumask *isol_mask) { return 0; }
 static inline void housekeeping_init(void) { }
+static inline void housekeeping_disable_type(enum hk_type type) { }
 #endif /* CONFIG_CPU_ISOLATION */
 
 static inline bool housekeeping_cpu(int cpu, enum hk_type type)
diff --git a/include/linux/tick.h b/include/linux/tick.h
index b121c5d53e308..58752fac3ce39 100644
--- a/include/linux/tick.h
+++ b/include/linux/tick.h
@@ -161,36 +161,14 @@ static inline ktime_t tick_nohz_get_sleep_length(ktime_t *delta_next)
 }
 #endif /* !CONFIG_NO_HZ_COMMON */
 
-/*
- * Mask of CPUs that are nohz_full.
- *
- * Users should be guarded by CONFIG_NO_HZ_FULL or a tick_nohz_full_cpu()
- * check.
- */
-extern cpumask_var_t tick_nohz_full_mask;
-
 #ifdef CONFIG_NO_HZ_FULL
-extern bool tick_nohz_full_running;
-
-static inline bool tick_nohz_full_enabled(void)
-{
-	if (!context_tracking_enabled())
-		return false;
-
-	return tick_nohz_full_running;
-}
-
 /*
- * Check if a CPU is part of the nohz_full subset. Arrange for evaluating
- * the cpu expression (typically smp_processor_id()) _after_ the static
- * key.
+ * tick_nohz_full_enabled() / tick_nohz_full_cpu() report the
+ * HK_TYPE_KERNEL_NOISE housekeeping state; they are implemented in
+ * tick-sched.c to avoid a tick.h <-> sched/isolation.h include cycle.
  */
-#define tick_nohz_full_cpu(_cpu) ({					\
-	bool __ret = false;						\
-	if (tick_nohz_full_enabled())					\
-		__ret = cpumask_test_cpu((_cpu), tick_nohz_full_mask);	\
-	__ret;								\
-})
+extern bool tick_nohz_full_enabled(void);
+extern bool tick_nohz_full_cpu(int cpu);
 
 extern void tick_nohz_dep_set(enum tick_dep_bits bit);
 extern void tick_nohz_dep_clear(enum tick_dep_bits bit);
@@ -205,6 +183,9 @@ extern void tick_nohz_dep_set_signal(struct task_struct *tsk,
 extern void tick_nohz_dep_clear_signal(struct signal_struct *signal,
 				       enum tick_dep_bits bit);
 extern bool tick_nohz_cpu_hotpluggable(unsigned int cpu);
+extern int tick_nohz_cpu_isolate(int cpu);
+extern void tick_nohz_cpu_deisolate(int cpu);
+extern int tick_nohz_full_hotplug_init(void);
 
 /*
  * The below are tick_nohz_[set,clear]_dep() wrappers that optimize off-cases
@@ -268,7 +249,6 @@ static inline void tick_dep_clear_signal(struct signal_struct *signal,
 
 extern void tick_nohz_full_kick_cpu(int cpu);
 extern void __tick_nohz_task_switch(void);
-extern void __init tick_nohz_full_setup(cpumask_var_t cpumask);
 #else
 static inline bool tick_nohz_full_enabled(void) { return false; }
 static inline bool tick_nohz_full_cpu(int cpu) { return false; }
@@ -276,6 +256,9 @@ static inline bool tick_nohz_full_cpu(int cpu) { return false; }
 static inline void tick_nohz_dep_set_cpu(int cpu, enum tick_dep_bits bit) { }
 static inline void tick_nohz_dep_clear_cpu(int cpu, enum tick_dep_bits bit) { }
 static inline bool tick_nohz_cpu_hotpluggable(unsigned int cpu) { return true; }
+static inline int tick_nohz_cpu_isolate(int cpu) { return -EINVAL; }
+static inline void tick_nohz_cpu_deisolate(int cpu) { }
+static inline int tick_nohz_full_hotplug_init(void) { return -EINVAL; }
 
 static inline void tick_dep_set(enum tick_dep_bits bit) { }
 static inline void tick_dep_clear(enum tick_dep_bits bit) { }
@@ -293,7 +276,6 @@ static inline void tick_dep_clear_signal(struct signal_struct *signal,
 
 static inline void tick_nohz_full_kick_cpu(int cpu) { }
 static inline void __tick_nohz_task_switch(void) { }
-static inline void tick_nohz_full_setup(cpumask_var_t cpumask) { }
 #endif
 
 static inline void tick_nohz_task_switch(void)
diff --git a/kernel/rcu/tree_nocb.h b/kernel/rcu/tree_nocb.h
index 99f3e4cec34e3..8a95aaa42ff01 100644
--- a/kernel/rcu/tree_nocb.h
+++ b/kernel/rcu/tree_nocb.h
@@ -1368,11 +1368,17 @@ void __init rcu_init_nohz(void)
 	int cpu;
 	struct rcu_data *rdp;
 	const struct cpumask *cpumask = NULL;
-
-#if defined(CONFIG_NO_HZ_FULL)
-	if (tick_nohz_full_running && !cpumask_empty(tick_nohz_full_mask))
-		cpumask = tick_nohz_full_mask;
-#endif
+	cpumask_var_t nohz_full_mask;
+	bool have_nohz_full_mask = false;
+
+	if (housekeeping_enabled(HK_TYPE_KERNEL_NOISE) &&
+	    alloc_cpumask_var(&nohz_full_mask, GFP_KERNEL)) {
+		have_nohz_full_mask = true;
+		cpumask_andnot(nohz_full_mask, cpu_possible_mask,
+			       housekeeping_cpumask(HK_TYPE_KERNEL_NOISE));
+		if (!cpumask_empty(nohz_full_mask))
+			cpumask = nohz_full_mask;
+	}
 
 	if (IS_ENABLED(CONFIG_RCU_NOCB_CPU_DEFAULT_ALL) &&
 	    !rcu_state.nocb_is_setup && !cpumask)
@@ -1382,7 +1388,7 @@ void __init rcu_init_nohz(void)
 		if (!cpumask_available(rcu_nocb_mask)) {
 			if (!zalloc_cpumask_var(&rcu_nocb_mask, GFP_KERNEL)) {
 				pr_info("rcu_nocb_mask allocation failed, callback offloading disabled.\n");
-				return;
+				goto out_free;
 			}
 		}
 
@@ -1391,7 +1397,7 @@ void __init rcu_init_nohz(void)
 	}
 
 	if (!rcu_state.nocb_is_setup)
-		return;
+		goto out_free;
 
 	rcu_nocb_register_lazy_shrinker();
 
@@ -1415,6 +1421,10 @@ void __init rcu_init_nohz(void)
 		rcu_segcblist_set_flags(&rdp->cblist, SEGCBLIST_OFFLOADED);
 	}
 	rcu_organize_nocb_kthreads();
+
+out_free:
+	if (have_nohz_full_mask)
+		free_cpumask_var(nohz_full_mask);
 }
 
 static DEFINE_MUTEX(rcu_nocb_lazy_mutex);
diff --git a/kernel/sched/isolation.c b/kernel/sched/isolation.c
index 7725514ac290e..35c8d5302c991 100644
--- a/kernel/sched/isolation.c
+++ b/kernel/sched/isolation.c
@@ -38,6 +38,20 @@ bool housekeeping_enabled(enum hk_type type)
 }
 EXPORT_SYMBOL_GPL(housekeeping_enabled);
 
+/*
+ * housekeeping_disable_type - Fully disable a housekeeping type at boot
+ * @type: Housekeeping type to disable
+ *
+ * Used by the rare boot fallback where a type's setup must be undone
+ * because a required arch capability turned out to be missing.  Clears
+ * the type's flag bit so housekeeping_enabled() and housekeeping_cpumask()
+ * fall back to "not configured" for it.
+ */
+void __init housekeeping_disable_type(enum hk_type type)
+{
+	WRITE_ONCE(housekeeping.flags, housekeeping.flags & ~BIT(type));
+}
+
 /*
  * Types that can change at runtime via cpuset isolated partitions.
  * Boot-only types (DOMAIN_BOOT) are always safe to read without lockdep.
@@ -299,10 +313,15 @@ int housekeeping_update_types(unsigned long type_mask,
 			 * was never allocated at boot since nohz_full= was
 			 * absent.  Allocate it now before CPUs cycle through
 			 * hotplug and sched_tick_stop() dereferences
-			 * tick_work_cpu.
+			 * tick_work_cpu.  Likewise, tick_nohz_init() never
+			 * ran this path's cpuhp registration, so the CPU
+			 * currently holding tick_do_timer_cpu duty has no
+			 * hotplug protection yet; install it now.
 			 */
-			if (type == HK_TYPE_KERNEL_NOISE)
+			if (type == HK_TYPE_KERNEL_NOISE) {
 				WARN_ON_ONCE(sched_tick_offload_init());
+				WARN_ON_ONCE(tick_nohz_full_hotplug_init());
+			}
 		}
 		rcu_assign_pointer(housekeeping.cpumasks[type], trials[type]);
 		trials[type] = NULL;
@@ -490,9 +509,6 @@ static int __init housekeeping_setup(char *str, unsigned long flags)
 			housekeeping_setup_type(type, housekeeping_staging);
 	}
 
-	if ((flags & HK_FLAG_KERNEL_NOISE) && !(housekeeping.flags & HK_FLAG_KERNEL_NOISE))
-		tick_nohz_full_setup(non_housekeeping_mask);
-
 	housekeeping.flags |= flags;
 	err = 1;
 
diff --git a/kernel/time/tick-sched.c b/kernel/time/tick-sched.c
index 8c53754c4ac4a..2d86035994372 100644
--- a/kernel/time/tick-sched.c
+++ b/kernel/time/tick-sched.c
@@ -22,6 +22,7 @@
 #include <linux/sched/stat.h>
 #include <linux/sched/nohz.h>
 #include <linux/sched/loadavg.h>
+#include <linux/sched/isolation.h>
 #include <linux/module.h>
 #include <linux/irq_work.h>
 #include <linux/posix-timers.h>
@@ -224,6 +225,20 @@ static bool tick_limited_update_jiffies64(struct tick_sched *ts, ktime_t now)
 
 #define MAX_STALLED_JIFFIES 5
 
+#ifdef CONFIG_NO_HZ_FULL
+/*
+ * True when HK_TYPE_KERNEL_NOISE is enabled and currently isolates at
+ * least one CPU. DHM can leave the type permanently enabled with an
+ * empty mask after a full runtime de-isolation; treat that state like
+ * ordinary NO_HZ_IDLE rather than full dynticks.
+ */
+static bool tick_nohz_full_live(void)
+{
+	return housekeeping_enabled(HK_TYPE_KERNEL_NOISE) &&
+	       !cpumask_full(housekeeping_cpumask(HK_TYPE_KERNEL_NOISE));
+}
+#endif
+
 static void tick_sched_do_timer(struct tick_sched *ts, ktime_t now)
 {
 	int tick_cpu, cpu = smp_processor_id();
@@ -235,14 +250,14 @@ static void tick_sched_do_timer(struct tick_sched *ts, ktime_t now)
 	 * this duty, then the jiffies update is still serialized by
 	 * 'jiffies_lock'.
 	 *
-	 * If nohz_full is enabled, this should not happen because the
-	 * 'tick_do_timer_cpu' CPU never relinquishes.
+	 * If a full-dynticks CPU is currently isolated, this should not
+	 * happen because the 'tick_do_timer_cpu' CPU never relinquishes.
 	 */
 	tick_cpu = READ_ONCE(tick_do_timer_cpu);
 
 	if (IS_ENABLED(CONFIG_NO_HZ_COMMON) && unlikely(tick_cpu == TICK_DO_TIMER_NONE)) {
 #ifdef CONFIG_NO_HZ_FULL
-		WARN_ON_ONCE(tick_nohz_full_running);
+		WARN_ON_ONCE(tick_nohz_full_live());
 #endif
 		WRITE_ONCE(tick_do_timer_cpu, cpu);
 		tick_cpu = cpu;
@@ -332,10 +347,35 @@ static enum hrtimer_restart tick_nohz_handler(struct hrtimer *timer)
 }
 
 #ifdef CONFIG_NO_HZ_FULL
-cpumask_var_t tick_nohz_full_mask;
-EXPORT_SYMBOL_GPL(tick_nohz_full_mask);
-bool tick_nohz_full_running;
-EXPORT_SYMBOL_GPL(tick_nohz_full_running);
+bool tick_nohz_full_enabled(void)
+{
+	if (!context_tracking_enabled())
+		return false;
+
+	return housekeeping_enabled(HK_TYPE_KERNEL_NOISE);
+}
+EXPORT_SYMBOL_GPL(tick_nohz_full_enabled);
+
+/*
+ * Check if a CPU is part of the nohz_full subset. Arrange for evaluating
+ * the cpu expression (typically smp_processor_id()) _after_ the static
+ * key.
+ */
+bool tick_nohz_full_cpu(int cpu)
+{
+	bool ret;
+
+	if (!tick_nohz_full_enabled())
+		return false;
+
+	rcu_read_lock();
+	ret = !cpumask_test_cpu(cpu, housekeeping_cpumask_rcu(HK_TYPE_KERNEL_NOISE));
+	rcu_read_unlock();
+
+	return ret;
+}
+EXPORT_SYMBOL_GPL(tick_nohz_full_cpu);
+
 static atomic_t tick_dep_mask;
 
 static bool check_tick_dependency(atomic_t *dep)
@@ -488,12 +528,15 @@ static void tick_nohz_full_kick_all(void)
 {
 	int cpu;
 
-	if (!tick_nohz_full_running)
+	if (!housekeeping_enabled(HK_TYPE_KERNEL_NOISE))
 		return;
 
 	preempt_disable();
-	for_each_cpu_and(cpu, tick_nohz_full_mask, cpu_online_mask)
+	rcu_read_lock();
+	for_each_cpu_andnot(cpu, cpu_online_mask,
+			    housekeeping_cpumask_rcu(HK_TYPE_KERNEL_NOISE))
 		tick_nohz_full_kick_cpu(cpu);
+	rcu_read_unlock();
 	preempt_enable();
 }
 
@@ -619,13 +662,31 @@ void __tick_nohz_task_switch(void)
 	}
 }
 
-/* Get the boot-time nohz CPU list from the kernel parameters. */
-void __init tick_nohz_full_setup(cpumask_var_t cpumask)
+/*
+ * tick_nohz_cpu_isolate - Add a CPU to the full-dynticks set at runtime.
+ * @cpu: the CPU to isolate; must be offline.
+ *
+ * The caller has already excluded @cpu from the HK_TYPE_KERNEL_NOISE
+ * housekeeping mask via housekeeping_update_types(), which is what
+ * tick_nohz_full_cpu() consults.  Activate per-CPU context tracking so
+ * that kernel/user transitions suppress the scheduler tick.
+ */
+int tick_nohz_cpu_isolate(int cpu)
+{
+	ct_cpu_track_user(cpu);
+	return 0;
+}
+EXPORT_SYMBOL_GPL(tick_nohz_cpu_isolate);
+
+/*
+ * tick_nohz_cpu_deisolate - Remove a CPU from the full-dynticks set.
+ * @cpu: the CPU to de-isolate; must be offline.
+ */
+void tick_nohz_cpu_deisolate(int cpu)
 {
-	alloc_bootmem_cpumask_var(&tick_nohz_full_mask);
-	cpumask_copy(tick_nohz_full_mask, cpumask);
-	tick_nohz_full_running = true;
+	ct_cpu_untrack_user(cpu);
 }
+EXPORT_SYMBOL_GPL(tick_nohz_cpu_deisolate);
 
 bool tick_nohz_cpu_hotpluggable(unsigned int cpu)
 {
@@ -633,8 +694,18 @@ bool tick_nohz_cpu_hotpluggable(unsigned int cpu)
 	 * The 'tick_do_timer_cpu' CPU handles housekeeping duty (unbound
 	 * timers, workqueues, timekeeping, ...) on behalf of full dynticks
 	 * CPUs. It must remain online when nohz full is enabled.
+	 *
+	 * Deliberately check the permanently-sticky HK_TYPE_KERNEL_NOISE
+	 * flag here, not tick_nohz_full_live()'s mask-aware variant: DHM
+	 * isolates one CPU at a time and only publishes the updated mask
+	 * (housekeeping_update_types()) after remove_cpu() succeeds, so
+	 * the live mask still looks empty at the exact moment a brand
+	 * new isolation's remove_cpu() call reaches this check. Gating
+	 * on the mask here would leave the duty holder unprotected for
+	 * every single-CPU isolation, not just the very first one.
 	 */
-	if (tick_nohz_full_running && READ_ONCE(tick_do_timer_cpu) == cpu)
+	if (housekeeping_enabled(HK_TYPE_KERNEL_NOISE) &&
+	    READ_ONCE(tick_do_timer_cpu) == cpu)
 		return false;
 	return true;
 }
@@ -644,11 +715,34 @@ static int tick_nohz_cpu_down(unsigned int cpu)
 	return tick_nohz_cpu_hotpluggable(cpu) ? 0 : -EBUSY;
 }
 
+/*
+ * tick_nohz_full_hotplug_init - Install tick_do_timer_cpu hotplug protection.
+ *
+ * Boot-time nohz_full=/isolcpus=nohz reaches this via tick_nohz_init().
+ * DHM's runtime first-enable path (no nohz_full= at boot) calls this
+ * directly instead, since tick_nohz_init() has already run and returned
+ * early by the time housekeeping_update_types() first sets
+ * HK_TYPE_KERNEL_NOISE. Without it, the CPU holding tick_do_timer_cpu
+ * duty has no hotplug protection and can be pulled down by
+ * remove_cpu(), dropping timekeeping duty with no notice.
+ */
+int tick_nohz_full_hotplug_init(void)
+{
+	int ret;
+
+	ret = cpuhp_setup_state_nocalls(CPUHP_AP_ONLINE_DYN,
+					 "kernel/nohz:predown", NULL,
+					 tick_nohz_cpu_down);
+	return ret < 0 ? ret : 0;
+}
+EXPORT_SYMBOL_GPL(tick_nohz_full_hotplug_init);
+
 void __init tick_nohz_init(void)
 {
+	cpumask_var_t full_mask;
 	int cpu, ret;
 
-	if (!tick_nohz_full_running)
+	if (!housekeeping_enabled(HK_TYPE_KERNEL_NOISE))
 		return;
 
 	/*
@@ -658,8 +752,7 @@ void __init tick_nohz_init(void)
 	 */
 	if (!arch_irq_work_has_interrupt()) {
 		pr_warn("NO_HZ: Can't run full dynticks because arch doesn't support IRQ work self-IPIs\n");
-		cpumask_clear(tick_nohz_full_mask);
-		tick_nohz_full_running = false;
+		housekeeping_disable_type(HK_TYPE_KERNEL_NOISE);
 		return;
 	}
 
@@ -667,22 +760,35 @@ void __init tick_nohz_init(void)
 			!IS_ENABLED(CONFIG_PM_SLEEP_SMP_NONZERO_CPU)) {
 		cpu = smp_processor_id();
 
-		if (cpumask_test_cpu(cpu, tick_nohz_full_mask)) {
+		if (!cpumask_test_cpu(cpu, housekeeping_cpumask(HK_TYPE_KERNEL_NOISE))) {
+			cpumask_var_t isolated;
+
 			pr_warn("NO_HZ: Clearing %d from nohz_full range "
 				"for timekeeping\n", cpu);
-			cpumask_clear_cpu(cpu, tick_nohz_full_mask);
+			if (alloc_cpumask_var(&isolated, GFP_KERNEL)) {
+				cpumask_andnot(isolated, cpu_possible_mask,
+					       housekeeping_cpumask(HK_TYPE_KERNEL_NOISE));
+				cpumask_clear_cpu(cpu, isolated);
+				WARN_ON_ONCE(housekeeping_update_types(BIT(HK_TYPE_KERNEL_NOISE),
+								       isolated));
+				free_cpumask_var(isolated);
+			}
 		}
 	}
 
-	for_each_cpu(cpu, tick_nohz_full_mask)
+	if (!alloc_cpumask_var(&full_mask, GFP_KERNEL))
+		return;
+	cpumask_andnot(full_mask, cpu_possible_mask,
+		       housekeeping_cpumask(HK_TYPE_KERNEL_NOISE));
+
+	for_each_cpu(cpu, full_mask)
 		ct_cpu_track_user_init(cpu);
 
-	ret = cpuhp_setup_state_nocalls(CPUHP_AP_ONLINE_DYN,
-					"kernel/nohz:predown", NULL,
-					tick_nohz_cpu_down);
+	ret = tick_nohz_full_hotplug_init();
 	WARN_ON(ret < 0);
 	pr_info("NO_HZ: Full dynticks CPUs: %*pbl.\n",
-		cpumask_pr_args(tick_nohz_full_mask));
+		cpumask_pr_args(full_mask));
+	free_cpumask_var(full_mask);
 }
 #endif /* #ifdef CONFIG_NO_HZ_FULL */
 

-- 
2.43.0


^ permalink raw reply	[flat|nested] 14+ messages in thread

* [PATCH v5 09/12] cpuset: Add dhm_cycling_cpus mask to suppress transient invalidation
  2026-10-02 13:10 [PATCH v5 00/12] Dynamic Housekeeping Management (DHM) via CPUSets Qiliang Yuan
                   ` (7 preceding siblings ...)
  2026-10-02 13:10 ` [PATCH v5 08/12] tick/nohz: Derive full-dynticks state from HK_TYPE_KERNEL_NOISE Qiliang Yuan
@ 2026-10-02 13:10 ` Qiliang Yuan
  2026-10-02 13:10 ` [PATCH v5 10/12] cpuset: Drive kernel-noise isolation via per-CPU hotplug cycling Qiliang Yuan
                   ` (2 subsequent siblings)
  11 siblings, 0 replies; 14+ messages in thread
From: Qiliang Yuan @ 2026-10-02 13:10 UTC (permalink / raw)
  To: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot,
	Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, Paul E. McKenney, Frederic Weisbecker,
	Neeraj Upadhyay, Joel Fernandes, Josh Triplett, Boqun Feng,
	Uladzislau Rezki, Mathieu Desnoyers, Lai Jiangshan, Zqiang,
	Anna-Maria Behnsen, Tejun Heo, Jonathan Corbet, Shuah Khan,
	Shuah Khan, Thomas Gleixner
  Cc: Waiman Long, linux-kernel, rcu, cgroups, linux-doc,
	linux-kselftest, Qiliang Yuan

When the cpuset code cycles a CPU through hotplug to apply kernel-noise
isolation (remove_cpu followed by add_cpu), the CPU disappears from
cpu_active_mask temporarily. cpuset_hotplug_update_tasks() sees an
empty effective CPU set on the isolated partition and issues
partcmd_invalidate, tearing down the partition. The subsequent add_cpu
brings the CPU back online, but the partition has already been marked
invalid and requires manual user intervention to restore.

Add a global dhm_cycling_cpus cpumask protected by dhm_cycling_lock.
The isolation cycling path sets the bits for CPUs being cycled before
calling remove_cpu(), clears them after add_cpu() completes.
cpuset_hotplug_update_tasks() checks whether any of the cpuset's
effective exclusive CPUs are in dhm_cycling_cpus and skips the
invalidation command when they are, treating the transient empty-CPU
state as expected rather than an error.

A global cpumask avoids the need to walk the cpuset tree to find the
owning cpuset during the cycling loop which runs without cpuset locks.

Signed-off-by: Qiliang Yuan <odys.yuan@gmail.com>
---
 kernel/cgroup/cpuset.c | 40 ++++++++++++++++++++++++++++++++++++++--
 1 file changed, 38 insertions(+), 2 deletions(-)

diff --git a/kernel/cgroup/cpuset.c b/kernel/cgroup/cpuset.c
index 8b3bb034adf62..6213e63cf348d 100644
--- a/kernel/cgroup/cpuset.c
+++ b/kernel/cgroup/cpuset.c
@@ -157,6 +157,14 @@ static bool		update_housekeeping;	/* RWCS */
  */
 static cpumask_var_t	isolated_hk_cpus;	/* T */
 
+/*
+ * CPUs currently being cycled through hotplug for kernel-noise isolation.
+ * Protected by dhm_cycling_lock; read in cpuset_hotplug_update_tasks() to
+ * suppress transient partition invalidation during the offline step.
+ */
+static DEFINE_SPINLOCK(dhm_cycling_lock);
+static cpumask_var_t	dhm_cycling_cpus;
+
 /*
  * A flag to force sched domain rebuild at the end of an operation.
  * It can be set in
@@ -3906,6 +3914,7 @@ int __init cpuset_init(void)
 	BUG_ON(!zalloc_cpumask_var(&subpartitions_cpus, GFP_KERNEL));
 	BUG_ON(!zalloc_cpumask_var(&isolated_cpus, GFP_KERNEL));
 	BUG_ON(!zalloc_cpumask_var(&isolated_hk_cpus, GFP_KERNEL));
+	BUG_ON(!zalloc_cpumask_var(&dhm_cycling_cpus, GFP_KERNEL));
 
 	cpumask_setall(top_cpuset.cpus_allowed);
 	nodes_setall(top_cpuset.mems_allowed);
@@ -3991,6 +4000,20 @@ static void cpuset_hotplug_update_tasks(struct cpuset *cs, struct tmpmasks *tmp)
 	if (remote && (cpumask_empty(subpartitions_cpus) ||
 			(cpumask_empty(&new_cpus) &&
 			 partition_is_populated(cs, NULL)))) {
+		bool cycling;
+
+		/*
+		 * Suppress transient invalidation when the offline is part
+		 * of a hotplug cycling step for kernel-noise isolation.
+		 */
+		spin_lock(&dhm_cycling_lock);
+		cycling = cpumask_available(dhm_cycling_cpus) &&
+			  cpumask_intersects(cs->effective_xcpus,
+					     dhm_cycling_cpus);
+		spin_unlock(&dhm_cycling_lock);
+		if (cycling)
+			goto unlock;
+
 		WRITE_ONCE(cs->prs_err, PERR_HOTPLUG);
 		remote_partition_disable(cs, tmp);
 		compute_effective_cpumask(&new_cpus, cs, parent);
@@ -4008,8 +4031,21 @@ static void cpuset_hotplug_update_tasks(struct cpuset *cs, struct tmpmasks *tmp)
 	if (is_local_partition(cs) &&
 	    (!is_partition_valid(parent) ||
 	     tasks_nocpu_error(parent, cs, &new_cpus) ||
-	     cpumask_empty(subpartitions_cpus)))
-		partcmd = partcmd_invalidate;
+	     cpumask_empty(subpartitions_cpus))) {
+		bool cycling;
+
+		/*
+		 * Suppress transient invalidation when the offline is part
+		 * of a hotplug cycling step for kernel-noise isolation.
+		 */
+		spin_lock(&dhm_cycling_lock);
+		cycling = cpumask_available(dhm_cycling_cpus) &&
+			  cpumask_intersects(cs->effective_xcpus,
+					     dhm_cycling_cpus);
+		spin_unlock(&dhm_cycling_lock);
+		if (!cycling)
+			partcmd = partcmd_invalidate;
+	}
 	/*
 	 * On the other hand, an invalid partition root may be transitioned
 	 * back to a regular one with a non-empty effective xcpus.

-- 
2.43.0


^ permalink raw reply	[flat|nested] 14+ messages in thread

* [PATCH v5 10/12] cpuset: Drive kernel-noise isolation via per-CPU hotplug cycling
  2026-10-02 13:10 [PATCH v5 00/12] Dynamic Housekeeping Management (DHM) via CPUSets Qiliang Yuan
                   ` (8 preceding siblings ...)
  2026-10-02 13:10 ` [PATCH v5 09/12] cpuset: Add dhm_cycling_cpus mask to suppress transient invalidation Qiliang Yuan
@ 2026-10-02 13:10 ` Qiliang Yuan
  2026-10-02 13:10 ` [PATCH v5 11/12] docs: cgroup-v2: Document kernel-noise isolation via isolated partitions Qiliang Yuan
  2026-10-02 13:10 ` [PATCH v5 12/12] selftests/cgroup: Add kernel-noise isolation test to cpuset selftest Qiliang Yuan
  11 siblings, 0 replies; 14+ messages in thread
From: Qiliang Yuan @ 2026-10-02 13:10 UTC (permalink / raw)
  To: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot,
	Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, Paul E. McKenney, Frederic Weisbecker,
	Neeraj Upadhyay, Joel Fernandes, Josh Triplett, Boqun Feng,
	Uladzislau Rezki, Mathieu Desnoyers, Lai Jiangshan, Zqiang,
	Anna-Maria Behnsen, Tejun Heo, Jonathan Corbet, Shuah Khan,
	Shuah Khan, Thomas Gleixner
  Cc: Waiman Long, linux-kernel, rcu, cgroups, linux-doc,
	linux-kselftest, Qiliang Yuan

Track A (cpuset_update_sd_hk_unlock) updates the HK_TYPE_KERNEL_NOISE
and HK_TYPE_MANAGED_IRQ cpumasks but performs no per-CPU reconfiguration.
Tick suppression, RCU callback offloading and managed-IRQ remapping only
take effect when the affected CPUs pass through the CPU hotplug machinery.

Implement dhm_cycle_isolated_cpus() and call it from
cpuset_update_sd_hk_unlock() with cpuset_top_mutex still held: per the
locking convention at the top of this file, cpuset_top_mutex is the
outermost lock, so remove_cpu()/add_cpu() may freely acquire
cpus_write_lock() underneath it.  Keeping the mutex held across the
whole cycle also matches dhm_prev_isolated's existing "protected by
cpuset_top_mutex" comment and prevents a second isolated-partition
update from entering cpuset_update_sd_hk_unlock() while a cycle for
the previous one is still in flight.

On isolation, for each newly-isolated CPU:
  1. remove_cpu()              - offline; dying callbacks migrate IRQs
  2. housekeeping_update_types() - publish this CPU as kernel-noise
                                 isolated now that it is actually offline
  3. tick_nohz_cpu_isolate()  - enable context tracking so the tick
                                 is suppressed for this CPU (B0/B3)
  4. rcu_nocb_cpu_isolate()   - lazy nocb init, spawn kthreads, offload
                                 callbacks (B1)
  5. add_cpu()                 - online; tick and IRQ online callbacks
                                 reconfigure against the now-published
                                 HK masks

On de-isolation, the reverse order is applied, publishing the CPU as
no longer kernel-noise isolated right after its own remove_cpu()
succeeds instead of before any CPU in the batch is touched.

The managed-IRQ remapping requires no explicit call:
irq_migrate_all_off_this_cpu() (dying callback) and
irq_affinity_online_cpu() (online callback) already consult the
updated HK_TYPE_MANAGED_IRQ mask.

dhm_prev_isolated tracks the previous isolation set so that only CPUs
whose state changed are cycled rather than the full isolation set.
lockup_detector_hk_update() (B2) is called once after all CPUs are
cycled to update the watchdog mask.

Some CPUs cannot be taken offline: cpu_is_hotpluggable() rejects a
CPU with hotplug disabled in the architecture (e.g. the x86-64 boot
CPU), and remove_cpu() itself can still fail for a CPU that passed
that filter (e.g. it turns out to be the last CPU in its sched
domain, or it is the current tick_do_timer_cpu holder and
tick_nohz_cpu_hotpluggable() rejects it).  Publishing
housekeeping_update_types() per CPU, strictly after that CPU's own
remove_cpu() has already succeeded, makes both cases handle
themselves: a CPU this function ends up skipping is simply never
added to the published mask, so HK_TYPE_KERNEL_NOISE never claims a
CPU is isolated while it keeps ticking and running RCU callbacks
normally, without needing a separate pre-filter pass or a
republish-on-failure step afterwards.  Also free
newly_isolated/newly_deisolated/cur_isolated on the early return for
a no-op update, which previously leaked the allocations.

Signed-off-by: Qiliang Yuan <odys.yuan@gmail.com>
---
 kernel/cgroup/cpuset.c | 155 ++++++++++++++++++++++++++++++++++++++++++++-----
 1 file changed, 140 insertions(+), 15 deletions(-)

diff --git a/kernel/cgroup/cpuset.c b/kernel/cgroup/cpuset.c
index 6213e63cf348d..9fa4bf504c3d9 100644
--- a/kernel/cgroup/cpuset.c
+++ b/kernel/cgroup/cpuset.c
@@ -20,6 +20,8 @@
  */
 #include "cpuset-internal.h"
 
+#include <linux/cpu.h>
+#include <linux/cpuhplock.h>
 #include <linux/init.h>
 #include <linux/interrupt.h>
 #include <linux/kernel.h>
@@ -33,7 +35,9 @@
 #include <linux/sched/task.h>
 #include <linux/security.h>
 #include <linux/oom.h>
+#include <linux/nmi.h>
 #include <linux/sched/isolation.h>
+#include <linux/tick.h>
 #include <linux/wait.h>
 #include <linux/workqueue.h>
 #include <linux/task_work.h>
@@ -165,6 +169,14 @@ static cpumask_var_t	isolated_hk_cpus;	/* T */
 static DEFINE_SPINLOCK(dhm_cycling_lock);
 static cpumask_var_t	dhm_cycling_cpus;
 
+/*
+ * Snapshot of the isolated CPUs from the previous housekeeping update.
+ * Used to compute the delta (newly isolated / newly de-isolated) so that
+ * only the changed CPUs are cycled rather than the full isolation set.
+ * Protected by cpuset_top_mutex.
+ */
+static cpumask_var_t	dhm_prev_isolated;
+
 /*
  * A flag to force sched domain rebuild at the end of an operation.
  * It can be set in
@@ -1423,6 +1435,108 @@ static bool prstate_housekeeping_conflict(int prstate, struct cpumask *new_cpus)
 	return false;
 }
 
+/*
+ * dhm_cycle_isolated_cpus - Apply kernel-noise isolation via hotplug cycling
+ *
+ * For each CPU newly entering isolation: cycle it offline, configure tick
+ * suppression and RCU callback offloading while it is offline, then bring
+ * it back online.  The managed-IRQ state is handled automatically by the
+ * existing irq_migrate_all_off_this_cpu() dying callback and the
+ * irq_affinity_online_cpu() online callback which both consult the
+ * already-updated HK_TYPE_MANAGED_IRQ mask.
+ *
+ * For each CPU leaving isolation: cycle it offline, de-offload RCU and
+ * restore the tick, then bring it back online.
+ *
+ * Called with cpuset_top_mutex held and no other cpuset or hotplug locks
+ * held: cpuset_top_mutex is the outermost lock, so remove_cpu()/add_cpu()
+ * may freely take cpus_write_lock() underneath it.
+ */
+static void dhm_cycle_isolated_cpus(const struct cpumask *new_isolated)
+{
+	static const unsigned long noise_types =
+		BIT(HK_TYPE_KERNEL_NOISE) | BIT(HK_TYPE_MANAGED_IRQ);
+	cpumask_var_t newly_isolated, newly_deisolated, cur_isolated;
+	int cpu;
+
+	if (!alloc_cpumask_var(&newly_isolated, GFP_KERNEL) ||
+	    !alloc_cpumask_var(&newly_deisolated, GFP_KERNEL) ||
+	    !alloc_cpumask_var(&cur_isolated, GFP_KERNEL)) {
+		free_cpumask_var(newly_isolated);
+		free_cpumask_var(newly_deisolated);
+		return;
+	}
+
+	cpumask_andnot(newly_isolated, new_isolated, dhm_prev_isolated);
+	cpumask_andnot(newly_deisolated, dhm_prev_isolated, new_isolated);
+	/* cur_isolated tracks the mask actually published so far. */
+	cpumask_copy(cur_isolated, dhm_prev_isolated);
+	cpumask_copy(dhm_prev_isolated, new_isolated);
+
+	if (cpumask_empty(newly_isolated) && cpumask_empty(newly_deisolated))
+		goto out_free;
+
+	/* Mark cycling CPUs so cpuset_hotplug_update_tasks skips invalidation */
+	spin_lock(&dhm_cycling_lock);
+	cpumask_or(dhm_cycling_cpus, newly_isolated, newly_deisolated);
+	spin_unlock(&dhm_cycling_lock);
+
+	/*
+	 * Publish each CPU's kernel-noise/managed-IRQ housekeeping state
+	 * strictly between its own remove_cpu() succeeding and add_cpu()
+	 * bringing it back, never as a batch before the whole cycle.
+	 * housekeeping_update_types() then always describes exactly which
+	 * CPUs are actually ticking and running RCU callbacks normally at
+	 * that instant, including the current tick_do_timer_cpu holder,
+	 * which tick_nohz_cpu_hotpluggable() must see as still a
+	 * housekeeping CPU for as long as it is still online and un-cycled.
+	 */
+	for_each_cpu(cpu, newly_isolated) {
+		if (!cpu_is_hotpluggable(cpu)) {
+			pr_warn_once("cpuset: CPU%d cannot be isolated (hotplug disabled)\n",
+				     cpu);
+			cpumask_clear_cpu(cpu, dhm_prev_isolated);
+			continue;
+		}
+		if (remove_cpu(cpu)) {
+			pr_warn_once("cpuset: failed to offline CPU%d for isolation\n",
+				     cpu);
+			cpumask_clear_cpu(cpu, dhm_prev_isolated);
+			continue;
+		}
+		cpumask_set_cpu(cpu, cur_isolated);
+		WARN_ON_ONCE(housekeeping_update_types(noise_types, cur_isolated));
+		WARN_ON_ONCE(tick_nohz_cpu_isolate(cpu));
+		WARN_ON_ONCE(rcu_nocb_cpu_isolate(cpu));
+		WARN_ON_ONCE(add_cpu(cpu));
+	}
+
+	for_each_cpu(cpu, newly_deisolated) {
+		if (remove_cpu(cpu)) {
+			pr_warn_once("cpuset: failed to offline CPU%d for de-isolation\n",
+				     cpu);
+			cpumask_set_cpu(cpu, dhm_prev_isolated);
+			continue;
+		}
+		cpumask_clear_cpu(cpu, cur_isolated);
+		WARN_ON_ONCE(housekeeping_update_types(noise_types, cur_isolated));
+		WARN_ON_ONCE(rcu_nocb_cpu_deoffload(cpu));
+		tick_nohz_cpu_deisolate(cpu);
+		WARN_ON_ONCE(add_cpu(cpu));
+	}
+
+	spin_lock(&dhm_cycling_lock);
+	cpumask_clear(dhm_cycling_cpus);
+	spin_unlock(&dhm_cycling_lock);
+
+	lockup_detector_hk_update();
+
+out_free:
+	free_cpumask_var(newly_isolated);
+	free_cpumask_var(newly_deisolated);
+	free_cpumask_var(cur_isolated);
+}
+
 /*
  * cpuset_update_sd_hk_unlock - Rebuild sched domains, update HK & unlock
  *
@@ -1439,8 +1553,6 @@ static void cpuset_update_sd_hk_unlock(void)
 		rebuild_sched_domains_locked();
 
 	if (update_housekeeping) {
-		static const unsigned long noise_types =
-			BIT(HK_TYPE_KERNEL_NOISE) | BIT(HK_TYPE_MANAGED_IRQ);
 		int ret;
 
 		update_housekeeping = false;
@@ -1450,9 +1562,16 @@ static void cpuset_update_sd_hk_unlock(void)
 		cpus_read_unlock();
 
 		/*
-		 * housekeeping_update() is now called without holding
-		 * cpus_read_lock and cpuset_mutex. Only cpuset_top_mutex
-		 * is still being held for mutual exclusion.
+		 * housekeeping_update() and housekeeping_update_types() are
+		 * now called without holding cpus_read_lock and cpuset_mutex.
+		 * cpuset_top_mutex stays held all the way through
+		 * dhm_cycle_isolated_cpus() below: per the locking convention
+		 * at the top of this file it is the outermost lock, so it may
+		 * legally nest around the cpus_write_lock() that remove_cpu()/
+		 * add_cpu() take. Holding it here is also what makes
+		 * dhm_prev_isolated's "protected by cpuset_top_mutex" comment
+		 * true, and keeps a second isolated-partition update from
+		 * entering this function while a cycle is still in flight.
 		 */
 
 		/*
@@ -1464,18 +1583,23 @@ static void cpuset_update_sd_hk_unlock(void)
 		WARN_ON_ONCE(ret);
 
 		/*
-		 * Only touch the kernel-noise housekeeping masks
-		 * (HK_TYPE_KERNEL_NOISE and HK_TYPE_MANAGED_IRQ) once the
-		 * sched domain update above actually succeeded: HK_TYPE_
-		 * KERNEL_NOISE must stay a subset of HK_TYPE_DOMAIN, so a CPU
-		 * can never end up tick-suppressed while still scheduled as
-		 * part of the normal (non-isolated) sched domain.  The tick,
-		 * RCU and managed-interrupt state is reconfigured as the
-		 * affected CPUs are cycled through the CPU hotplug machinery.
+		 * Only cycle CPUs through hotplug, applying the kernel-noise
+		 * types along the way, when the sched domain update above
+		 * actually succeeded.  HK_TYPE_KERNEL_NOISE must stay a
+		 * subset of HK_TYPE_DOMAIN, so a CPU can never end up
+		 * tick-suppressed while still scheduled as part of the
+		 * normal (non-isolated) sched domain.  The tick, RCU and
+		 * managed-interrupt state is reconfigured as the affected
+		 * CPUs are cycled through the CPU hotplug machinery;
+		 * dhm_cycle_isolated_cpus() publishes HK_TYPE_KERNEL_NOISE
+		 * and HK_TYPE_MANAGED_IRQ itself, one CPU at a time, strictly
+		 * between that CPU's own remove_cpu() and add_cpu(), so a
+		 * CPU that cpu_is_hotpluggable() rejects (e.g. the current
+		 * tick_do_timer_cpu) is simply never published as isolated.
 		 */
 		if (!ret)
-			WARN_ON_ONCE(housekeeping_update_types(noise_types,
-							       isolated_hk_cpus));
+			dhm_cycle_isolated_cpus(isolated_hk_cpus);
+
 		mutex_unlock(&cpuset_top_mutex);
 	} else {
 		cpuset_full_unlock();
@@ -3915,6 +4039,7 @@ int __init cpuset_init(void)
 	BUG_ON(!zalloc_cpumask_var(&isolated_cpus, GFP_KERNEL));
 	BUG_ON(!zalloc_cpumask_var(&isolated_hk_cpus, GFP_KERNEL));
 	BUG_ON(!zalloc_cpumask_var(&dhm_cycling_cpus, GFP_KERNEL));
+	BUG_ON(!zalloc_cpumask_var(&dhm_prev_isolated, GFP_KERNEL));
 
 	cpumask_setall(top_cpuset.cpus_allowed);
 	nodes_setall(top_cpuset.mems_allowed);

-- 
2.43.0


^ permalink raw reply	[flat|nested] 14+ messages in thread

* [PATCH v5 11/12] docs: cgroup-v2: Document kernel-noise isolation via isolated partitions
  2026-10-02 13:10 [PATCH v5 00/12] Dynamic Housekeeping Management (DHM) via CPUSets Qiliang Yuan
                   ` (9 preceding siblings ...)
  2026-10-02 13:10 ` [PATCH v5 10/12] cpuset: Drive kernel-noise isolation via per-CPU hotplug cycling Qiliang Yuan
@ 2026-10-02 13:10 ` Qiliang Yuan
  2026-10-02 13:10 ` [PATCH v5 12/12] selftests/cgroup: Add kernel-noise isolation test to cpuset selftest Qiliang Yuan
  11 siblings, 0 replies; 14+ messages in thread
From: Qiliang Yuan @ 2026-10-02 13:10 UTC (permalink / raw)
  To: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot,
	Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, Paul E. McKenney, Frederic Weisbecker,
	Neeraj Upadhyay, Joel Fernandes, Josh Triplett, Boqun Feng,
	Uladzislau Rezki, Mathieu Desnoyers, Lai Jiangshan, Zqiang,
	Anna-Maria Behnsen, Tejun Heo, Jonathan Corbet, Shuah Khan,
	Shuah Khan, Thomas Gleixner
  Cc: Waiman Long, linux-kernel, rcu, cgroups, linux-doc,
	linux-kselftest, Qiliang Yuan

Document that creating a cpuset isolated partition updates the
kernel-noise housekeeping masks (HK_TYPE_KERNEL_NOISE and
HK_TYPE_MANAGED_IRQ) in addition to the sched-domain mask, and
that destroying it restores the boot configuration.

No boot-time kernel parameters such as nohz_full= or rcu_nocbs=
are required; writing "isolated" to cpuset.cpus.partition is the
only mechanism needed.

Signed-off-by: Qiliang Yuan <odys.yuan@gmail.com>
---
 Documentation/admin-guide/cgroup-v2.rst | 17 +++++++++++++++++
 1 file changed, 17 insertions(+)

diff --git a/Documentation/admin-guide/cgroup-v2.rst b/Documentation/admin-guide/cgroup-v2.rst
index 86a2a0099178e..611c6f7ca4efe 100644
--- a/Documentation/admin-guide/cgroup-v2.rst
+++ b/Documentation/admin-guide/cgroup-v2.rst
@@ -2784,6 +2784,23 @@ Cpuset Interface Files
 	kernel boot command line option.  If those CPUs are to be put
 	into a partition, they have to be used in an isolated partition.
 
+	When an isolated partition is created or destroyed, the kernel
+	automatically drives runtime updates of the housekeeping masks
+	for kernel-noise types (nohz_full, RCU NOCB, managed IRQ
+	interrupts).  This extends isolation beyond scheduler domains:
+	the tick is stopped on isolated CPUs, RCU callbacks are
+	offloaded to housekeeping cores, and managed interrupts are
+	migrated away.  No boot-time kernel parameters such as
+	``nohz_full=`` or ``rcu_nocbs=`` are required; writing
+	``isolated`` to ``cpuset.cpus.partition`` is the only mechanism
+	needed.  No additional cgroupfs files are required.
+
+	CPUs with hotplug disabled (typically the boot CPU, CPU 0, on
+	x86-64) cannot be cycled offline for kernel-noise isolation.
+	The kernel emits a one-time warning and keeps those CPUs in
+	the tick and RCU-NOCB housekeeping set, even when they appear
+	in an isolated partition.
+
 
 Device controller
 -----------------

-- 
2.43.0


^ permalink raw reply	[flat|nested] 14+ messages in thread

* [PATCH v5 12/12] selftests/cgroup: Add kernel-noise isolation test to cpuset selftest
  2026-10-02 13:10 [PATCH v5 00/12] Dynamic Housekeeping Management (DHM) via CPUSets Qiliang Yuan
                   ` (10 preceding siblings ...)
  2026-10-02 13:10 ` [PATCH v5 11/12] docs: cgroup-v2: Document kernel-noise isolation via isolated partitions Qiliang Yuan
@ 2026-10-02 13:10 ` Qiliang Yuan
  11 siblings, 0 replies; 14+ messages in thread
From: Qiliang Yuan @ 2026-10-02 13:10 UTC (permalink / raw)
  To: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot,
	Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, Paul E. McKenney, Frederic Weisbecker,
	Neeraj Upadhyay, Joel Fernandes, Josh Triplett, Boqun Feng,
	Uladzislau Rezki, Mathieu Desnoyers, Lai Jiangshan, Zqiang,
	Anna-Maria Behnsen, Tejun Heo, Jonathan Corbet, Shuah Khan,
	Shuah Khan, Thomas Gleixner
  Cc: Waiman Long, linux-kernel, rcu, cgroups, linux-doc,
	linux-kselftest, Qiliang Yuan

Add test_hk_noise_isolated() to test_cpuset_prs.sh to verify that
creating and destroying an isolated partition updates the kernel-noise
housekeeping state, including the /sys/devices/system/cpu/nohz_full
attribute.  Add the cpu_in_cpulist() helper to correctly test membership
against a cpulist that may contain ranges.

Also detect and report whether the test is running in zero-boot-param
mode (no nohz_full= in /proc/cmdline).  When in zero-boot-param mode
the test confirms that nohz_full is activated by the DHM runtime path,
verifying that dhm_cycle_isolated_cpus() correctly enables tick
isolation without any boot-time setup.

Signed-off-by: Qiliang Yuan <odys.yuan@gmail.com>
---
 tools/testing/selftests/cgroup/test_cpuset_prs.sh | 578 ++++++++++++++++++++++
 1 file changed, 578 insertions(+)

diff --git a/tools/testing/selftests/cgroup/test_cpuset_prs.sh b/tools/testing/selftests/cgroup/test_cpuset_prs.sh
index 7efd5e645767c..d4ae35dd71563 100755
--- a/tools/testing/selftests/cgroup/test_cpuset_prs.sh
+++ b/tools/testing/selftests/cgroup/test_cpuset_prs.sh
@@ -1281,10 +1281,588 @@ test_inotify()
 	echo "" > cpuset.cpus
 }
 
+#
+# cpu_in_cpulist <cpu> <cpulist>
+#
+# Return 0 if <cpu> appears in <cpulist> (a kernel cpumask list such as
+# "0-3,8-31"), non-zero otherwise.  The kernel cpulist format uses ranges
+# ("lo-hi") and comma-separated items; a simple grep cannot detect that a
+# number falls in the middle of a range, so walk each element explicitly.
+#
+cpu_in_cpulist()
+{
+	local cpu=$1 list=$2 range lo hi
+	for range in $(echo "$list" | tr ',' ' '); do
+		if [[ "$range" == *-* ]]; then
+			lo=${range%-*}
+			hi=${range#*-}
+			[[ $cpu -ge $lo && $cpu -le $hi ]] && return 0
+		else
+			[[ $cpu -eq $range ]] && return 0
+		fi
+	done
+	return 1
+}
+
+#
+# verify_nohz_exact <baseline> <active_cpus> <current>
+#
+# Verify that <current> nohz_full equals <baseline> ∪ <active_cpus>.
+# <active_cpus> may be a cpulist range ("4-7") or empty string ("").
+# Catches both missing CPUs and unexpected extra CPUs.
+#
+verify_nohz_exact()
+{
+	local baseline=$1 active=$2 current=$3 cpu exp got
+	for cpu in $(seq 0 $((NR_CPUS - 1))); do
+		exp=0; got=0
+		cpu_in_cpulist $cpu "$baseline" && exp=1
+		[[ -n "$active" ]] && cpu_in_cpulist $cpu "$active" && exp=1
+		cpu_in_cpulist $cpu "$current" && got=1
+		[[ $exp -eq $got ]] || {
+			if [[ $got -eq 0 ]]; then
+				echo "FAIL: cpu${cpu} expected in nohz_full but absent" \
+				     "(baseline='$baseline' active='$active'" \
+				     "current='$current')"
+			else
+				echo "FAIL: cpu${cpu} unexpectedly in nohz_full" \
+				     "(baseline='$baseline' active='$active'" \
+				     "current='$current')"
+			fi
+			return 1
+		}
+	done
+	return 0
+}
+
+#
+# Test that isolated partition creation/destruction drives kernel-noise
+# housekeeping mask updates and remains correct under pressure.
+#
+# Requires: >=8 CPUs, no isolcpus= boot conflict, root
+#
+
+#
+# hk_noise_check_nocb_affinity <cpulist> <expect_isolated>
+#
+# When expect_isolated=1: verify rcuop/N kthreads for CPUs in cpulist do NOT
+# include those CPUs in their scheduler affinity (RCU NOCB active — callbacks
+# for the isolated CPU are offloaded to a different CPU).
+# When expect_isolated=0: verify CPUs are back in affinity (NOCB restored).
+# Silently skips CPUs whose rcuop/N thread is absent (no NOCB support).
+#
+hk_noise_check_nocb_affinity()
+{
+	local cpulist=$1 expect_isolated=$2
+	local cpu lo hi range pid aff_hex rev_hex nibble_pos nibble_char
+	local nibble_val bit_in_nibble bit failed=0
+
+	command -v taskset > /dev/null 2>&1 || return 0
+
+	for range in $(echo "$cpulist" | tr ',' ' '); do
+		if [[ "$range" == *-* ]]; then
+			lo=${range%-*}; hi=${range#*-}
+		else
+			lo=$range; hi=$range
+		fi
+		for cpu in $(seq "$lo" "$hi"); do
+			pid=$(ps -eo pid,comm | awk -v c="rcuop/$cpu" '$2==c{print $1}')
+			[[ -n "$pid" ]] || continue
+			aff_hex=$(taskset -p "$pid" 2>/dev/null | awk '{print $NF}')
+			[[ -n "$aff_hex" ]] || continue
+
+			# Extract bit <cpu> from the hex affinity mask.
+			# Each hex digit covers 4 CPUs; reverse the string to
+			# work from the LSB side.
+			rev_hex=$(echo "$aff_hex" | rev)
+			nibble_pos=$((cpu / 4))
+			nibble_char=${rev_hex:$nibble_pos:1}
+			if [[ -z "$nibble_char" ]]; then
+				nibble_val=0
+			else
+				nibble_val=$((16#$nibble_char))
+			fi
+			bit_in_nibble=$((cpu % 4))
+			bit=$(( (nibble_val >> bit_in_nibble) & 1 ))
+
+			if [[ $expect_isolated -eq 1 && $bit -eq 1 ]]; then
+				echo "FAIL: rcuop/$cpu affinity still includes" \
+				     "CPU$cpu after isolation (mask=0x$aff_hex)"
+				failed=1
+			elif [[ $expect_isolated -eq 0 && $bit -eq 0 ]]; then
+				echo "FAIL: rcuop/$cpu affinity still excludes" \
+				     "CPU$cpu after de-isolation (mask=0x$aff_hex)"
+				failed=1
+			fi
+		done
+	done
+	return $failed
+}
+
+test_hk_noise_isolated()
+{
+	local ISOL_BEFORE TEST_CPUS i PART ISOL_AFTER ISOL_RESTORE
+	local NOHZ_FILE NOHZ_BEFORE NOHZ_AFTER NOHZ_RESTORE
+	local HK_NOHZ_CHECK=0
+	local LOOPS=100
+	local CMDLINE HAS_BOOT_NOHZ=0
+	local DMESG_LINES_START
+	DMESG_LINES_START=$(dmesg | wc -l)
+
+	[[ $NR_CPUS -ge 8 ]] || {
+		echo "HK-noise test skipped: need >=8 CPUs, have $NR_CPUS"
+		return 0
+	}
+
+	# Detect whether CONFIG_NO_HZ_FULL is active: the sysfs attribute
+	# /sys/devices/system/cpu/nohz_full exposes the current nohz_full
+	# cpumask and is only present when NO_HZ_FULL is enabled.
+	NOHZ_FILE=/sys/devices/system/cpu/nohz_full
+	[[ -r "$NOHZ_FILE" ]] && HK_NOHZ_CHECK=1
+
+	# Determine if running in zero-boot-param mode.  DHM activates tick
+	# and RCU-NOCB isolation at runtime; no nohz_full= or rcu_nocbs=
+	# kernel boot parameters are required.
+	{ read -r CMDLINE < /proc/cmdline; } 2>/dev/null || CMDLINE=""
+	[[ $CMDLINE = *nohz_full=* ]] && HAS_BOOT_NOHZ=1
+	if [[ $HAS_BOOT_NOHZ -eq 0 ]]; then
+		console_msg "HK-noise: zero-boot-param mode" \
+		            "(no nohz_full= in /proc/cmdline -- testing DHM runtime path)"
+	else
+		console_msg "HK-noise: boot-param mode (nohz_full= present at boot)"
+	fi
+
+	cd $CGROUP2/test
+	echo member > cpuset.cpus.partition 2>/dev/null
+	echo "" > cpuset.cpus 2>/dev/null
+
+	ISOL_BEFORE=$(cat $CGROUP2/cpuset.cpus.isolated)
+	[[ $HK_NOHZ_CHECK -eq 1 ]] && NOHZ_BEFORE=$(cat $NOHZ_FILE)
+	TEST_CPUS="4-7"
+	echo $TEST_CPUS > cpuset.cpus
+
+	#
+	# Basic create/destroy cycle — verify domain isolation and
+	# kernel-noise (nohz_full) changes together.
+	#
+	console_msg "HK-noise: basic create/destroy cycle"
+	echo isolated > cpuset.cpus.partition
+
+	ISOL_AFTER=$(cat $CGROUP2/cpuset.cpus.isolated)
+	[[ $ISOL_AFTER != "$ISOL_BEFORE" ]] || {
+		echo "FAIL: isolated set unchanged after partition create"
+		exit 1
+	}
+
+	if [[ $HK_NOHZ_CHECK -eq 1 ]]; then
+		NOHZ_AFTER=$(cat $NOHZ_FILE)
+		verify_nohz_exact "$NOHZ_BEFORE" "$TEST_CPUS" "$NOHZ_AFTER" || exit 1
+		console_msg "HK-noise: nohz_full after isolation: $NOHZ_AFTER"
+	fi
+
+	# Verify RCU NOCB: rcuop/N kthreads for isolated CPUs must have those
+	# CPUs removed from their scheduler affinity mask.
+	hk_noise_check_nocb_affinity "$TEST_CPUS" 1 || exit 1
+
+	echo member > cpuset.cpus.partition
+
+	ISOL_RESTORE=$(cat $CGROUP2/cpuset.cpus.isolated)
+	[[ $ISOL_RESTORE = "$ISOL_BEFORE" ]] || {
+		echo "FAIL: expected '$ISOL_BEFORE' after destroy, got '$ISOL_RESTORE'"
+		exit 1
+	}
+
+	if [[ $HK_NOHZ_CHECK -eq 1 ]]; then
+		NOHZ_RESTORE=$(cat $NOHZ_FILE)
+		verify_nohz_exact "$NOHZ_BEFORE" "" "$NOHZ_RESTORE" || exit 1
+	fi
+
+	# Verify RCU NOCB restored: isolated CPUs must reappear in rcuop/N affinity.
+	hk_noise_check_nocb_affinity "$TEST_CPUS" 0 || exit 1
+
+	#
+	# Reject all-CPU isolation (must leave at least one housekeeping CPU)
+	#
+	console_msg "HK-noise: reject all-CPU isolation"
+	echo 0-$((NR_CPUS - 1)) > cpuset.cpus
+	echo isolated > cpuset.cpus.partition
+	PART=$(cat cpuset.cpus.partition)
+	[[ $PART = *invalid* || $PART = member ]] || {
+		echo "FAIL: all-CPU isolation was not rejected, got '$PART'"
+		exit 1
+	}
+
+	#
+	# SMT safety: partial sibling isolation
+	#
+	console_msg "HK-noise: SMT sibling constraint"
+	echo $TEST_CPUS > cpuset.cpus
+	echo isolated > cpuset.cpus.partition
+	PART=$(cat cpuset.cpus.partition)
+	[[ $PART = isolated ]] || {
+		echo "FAIL: could not create isolated partition, got '$PART'"
+		exit 1
+	}
+	echo member > cpuset.cpus.partition
+
+	#
+	# Non-hotpluggable CPU: must be skipped with a kernel warning without
+	# rejecting the partition; hotpluggable peers must still be isolated.
+	#
+	# A CPU whose online file is absent (e.g. CPU 0 on x86-64) has hotplug
+	# disabled.  DHM emits pr_warn_once and keeps it in the tick/RCU-NOCB
+	# housekeeping set; it must not appear in nohz_full after isolation.
+	# The remaining hotpluggable CPUs in the partition must still be isolated.
+	#
+	local FIXED_CPU="" c NOHZ_NOW
+	for c in $(seq 0 $((NR_CPUS - 1))); do
+		[[ -f /sys/devices/system/cpu/cpu${c}/online ]] || {
+			FIXED_CPU=$c
+			break
+		}
+	done
+	if [[ -n "$FIXED_CPU" ]]; then
+		console_msg "HK-noise: non-hotpluggable CPU${FIXED_CPU} skip"
+		echo "${FIXED_CPU},${TEST_CPUS}" > cpuset.cpus
+		echo isolated > cpuset.cpus.partition
+		PART=$(cat cpuset.cpus.partition)
+		[[ $PART = isolated ]] || {
+			echo "FAIL: partition rejected when including non-hotpluggable" \
+			     "CPU${FIXED_CPU}: got '$PART'"
+			exit 1
+		}
+		if [[ $HK_NOHZ_CHECK -eq 1 ]]; then
+			NOHZ_NOW=$(cat $NOHZ_FILE)
+			if cpu_in_cpulist $FIXED_CPU "$NOHZ_NOW"; then
+				echo "FAIL: non-hotpluggable CPU${FIXED_CPU} appeared" \
+				     "in nohz_full (should be skipped)"
+				exit 1
+			fi
+			local lo hi
+			lo=${TEST_CPUS%%-*}
+			hi=${TEST_CPUS##*-}
+			for cpu in $(seq "$lo" "$hi"); do
+				if ! cpu_in_cpulist $cpu "$NOHZ_NOW"; then
+					echo "FAIL: hotpluggable cpu${cpu} missing from" \
+					     "nohz_full in mixed partition (got: '$NOHZ_NOW')"
+					exit 1
+				fi
+			done
+			console_msg "HK-noise: CPU${FIXED_CPU} absent, ${TEST_CPUS} present" \
+			            "in nohz_full: $NOHZ_NOW"
+		fi
+		echo member > cpuset.cpus.partition
+		echo $TEST_CPUS > cpuset.cpus
+	else
+		console_msg "HK-noise: all CPUs hotpluggable; skip non-hotpluggable subtest"
+	fi
+
+	#
+	# Delta isolation: modify cpuset.cpus while the partition is isolated.
+	# dhm_prev_isolated must track the delta and update nohz_full in step.
+	#
+	local lo hi mid lower_cpus upper_cpu
+	lo=${TEST_CPUS%%-*}
+	hi=${TEST_CPUS##*-}
+	mid=$(( lo + (hi - lo) / 2 ))
+	lower_cpus="${lo}-${mid}"
+	upper_cpu=$(( mid + 1 ))
+	console_msg "HK-noise: delta isolation (shrink ${TEST_CPUS} → ${lower_cpus})"
+	echo $TEST_CPUS > cpuset.cpus
+	echo isolated > cpuset.cpus.partition
+	PART=$(cat cpuset.cpus.partition)
+	[[ $PART = isolated ]] || {
+		echo "FAIL: delta test: initial isolation failed, got '$PART'"
+		exit 1
+	}
+	echo $lower_cpus > cpuset.cpus
+	PART=$(cat cpuset.cpus.partition)
+	if [[ $PART = isolated ]]; then
+		if [[ $HK_NOHZ_CHECK -eq 1 ]]; then
+			NOHZ_NOW=$(cat $NOHZ_FILE)
+			verify_nohz_exact "$NOHZ_BEFORE" "$lower_cpus" "$NOHZ_NOW" || exit 1
+		fi
+		# Expand back to full TEST_CPUS and re-verify
+		echo $TEST_CPUS > cpuset.cpus
+		if [[ $HK_NOHZ_CHECK -eq 1 ]]; then
+			NOHZ_NOW=$(cat $NOHZ_FILE)
+			verify_nohz_exact "$NOHZ_BEFORE" "$TEST_CPUS" "$NOHZ_NOW" || exit 1
+		fi
+		echo member > cpuset.cpus.partition
+		if [[ $HK_NOHZ_CHECK -eq 1 ]]; then
+			NOHZ_NOW=$(cat $NOHZ_FILE)
+			verify_nohz_exact "$NOHZ_BEFORE" "" "$NOHZ_NOW" || exit 1
+		fi
+	else
+		console_msg "HK-noise: delta test: partition invalidated on shrink" \
+		            "('$PART') -- cpuset constraint, not a DHM bug; skipping"
+		echo member > cpuset.cpus.partition 2>/dev/null || true
+	fi
+	echo $TEST_CPUS > cpuset.cpus
+
+	#
+	# Nested partition: parent root → child isolated
+	#
+	console_msg "HK-noise: nested partition inheritance"
+	echo $TEST_CPUS > cpuset.cpus
+	test_partition root
+	mkdir -p HK_SUB
+	cd HK_SUB
+	echo "${lo}-$((lo + 1))" > cpuset.cpus
+	echo isolated > cpuset.cpus.partition
+	ISOL_AFTER=$(cat $CGROUP2/cpuset.cpus.isolated)
+	[[ -n $ISOL_AFTER ]] || {
+		echo "FAIL: nested isolated partition not reflected in cpuset.cpus.isolated"
+		exit 1
+	}
+	echo member > cpuset.cpus.partition
+	cd $CGROUP2/test
+	echo member > cpuset.cpus.partition
+	rmdir HK_SUB 2>/dev/null
+
+	#
+	# Pressure test: 100 create/destroy cycles with nohz_full verified
+	# on every cycle to catch mid-run state corruption.
+	#
+	console_msg "HK-noise: pressure test ($LOOPS cycles)"
+	echo $TEST_CPUS > cpuset.cpus
+	for i in $(seq 1 $LOOPS); do
+		echo isolated > cpuset.cpus.partition
+		PART=$(cat cpuset.cpus.partition)
+		[[ $PART = isolated ]] || {
+			echo "FAIL: cycle $i create failed, got '$PART'"
+			exit 1
+		}
+		if [[ $HK_NOHZ_CHECK -eq 1 ]]; then
+			NOHZ_NOW=$(cat $NOHZ_FILE)
+			verify_nohz_exact "$NOHZ_BEFORE" "$TEST_CPUS" "$NOHZ_NOW" || {
+				echo "FAIL: nohz_full wrong at cycle $i (isolated)"
+				exit 1
+			}
+		fi
+		echo member > cpuset.cpus.partition
+		PART=$(cat cpuset.cpus.partition)
+		[[ $PART = member ]] || {
+			echo "FAIL: cycle $i destroy failed, got '$PART'"
+			exit 1
+		}
+		if [[ $HK_NOHZ_CHECK -eq 1 ]]; then
+			NOHZ_NOW=$(cat $NOHZ_FILE)
+			verify_nohz_exact "$NOHZ_BEFORE" "" "$NOHZ_NOW" || {
+				echo "FAIL: nohz_full wrong at cycle $i (member)"
+				exit 1
+			}
+		fi
+	done
+
+	#
+	# Stability: after pressure test, verify final state
+	#
+	console_msg "HK-noise: post-pressure cleanup"
+	echo isolated > cpuset.cpus.partition
+	ISOL_AFTER=$(cat $CGROUP2/cpuset.cpus.isolated)
+	[[ -n $ISOL_AFTER ]] || {
+		echo "FAIL: isolated set empty after pressure test"
+		exit 1
+	}
+	echo member > cpuset.cpus.partition
+	echo "" > cpuset.cpus
+	ISOL_RESTORE=$(cat $CGROUP2/cpuset.cpus.isolated)
+	[[ $ISOL_RESTORE = "$ISOL_BEFORE" ]] || {
+		echo "FAIL: final isolated '$ISOL_RESTORE' != '$ISOL_BEFORE'"
+		exit 1
+	}
+
+	if [[ $HK_NOHZ_CHECK -eq 1 ]]; then
+		NOHZ_RESTORE=$(cat $NOHZ_FILE)
+		[[ "$NOHZ_RESTORE" = "$NOHZ_BEFORE" ]] || {
+			echo "FAIL: nohz_full not restored after pressure test:" \
+			     "expected '$NOHZ_BEFORE', got '$NOHZ_RESTORE'"
+			exit 1
+		}
+	fi
+
+	#
+	# Pressure with resident task: create/destroy cycles while a sleeping
+	# task occupies the isolated partition.  Exercises the dhm_cycling_cpus
+	# suppression path that prevents false partition invalidation when a
+	# task is present during hotplug cycling steps.
+	#
+	console_msg "HK-noise: pressure with resident task ($LOOPS cycles)"
+	echo $TEST_CPUS > cpuset.cpus
+	sleep 600 &
+	local TASK_PID=$!
+	echo $TASK_PID > cgroup.procs
+	for i in $(seq 1 $LOOPS); do
+		echo isolated > cpuset.cpus.partition
+		PART=$(cat cpuset.cpus.partition)
+		[[ $PART = isolated ]] || {
+			echo "FAIL: task-occupied cycle $i create failed, got '$PART'"
+			kill $TASK_PID 2>/dev/null; wait $TASK_PID 2>/dev/null
+			exit 1
+		}
+		if [[ $HK_NOHZ_CHECK -eq 1 ]]; then
+			NOHZ_NOW=$(cat $NOHZ_FILE)
+			verify_nohz_exact "$NOHZ_BEFORE" "$TEST_CPUS" "$NOHZ_NOW" || {
+				kill $TASK_PID 2>/dev/null; wait $TASK_PID 2>/dev/null
+				exit 1
+			}
+		fi
+		echo member > cpuset.cpus.partition
+		PART=$(cat cpuset.cpus.partition)
+		[[ $PART = member ]] || {
+			echo "FAIL: task-occupied cycle $i destroy failed, got '$PART'"
+			kill $TASK_PID 2>/dev/null; wait $TASK_PID 2>/dev/null
+			exit 1
+		}
+	done
+	kill $TASK_PID 2>/dev/null
+	wait $TASK_PID 2>/dev/null
+	echo "" > cpuset.cpus
+
+	#
+	# Concurrent partitions: two sibling cgroups each holding half of
+	# TEST_CPUS simultaneously isolated.  Verifies independent isolation
+	# and correct union in cpuset.cpus.isolated / nohz_full.
+	#
+	local lo hi mid lower_half upper_half NOHZ_BOTH
+	lo=${TEST_CPUS%%-*}; hi=${TEST_CPUS##*-}
+	mid=$(( lo + (hi - lo) / 2 ))
+	lower_half="${lo}-${mid}"
+	upper_half="$((mid + 1))-${hi}"
+	console_msg "HK-noise: concurrent partitions ($lower_half and $upper_half)"
+	echo $TEST_CPUS > cpuset.cpus
+	echo root > cpuset.cpus.partition
+	mkdir -p HK_A HK_B
+	echo "$lower_half" > HK_A/cpuset.cpus
+	echo isolated > HK_A/cpuset.cpus.partition
+	echo "$upper_half" > HK_B/cpuset.cpus
+	echo isolated > HK_B/cpuset.cpus.partition
+	local PA PB
+	PA=$(cat HK_A/cpuset.cpus.partition)
+	PB=$(cat HK_B/cpuset.cpus.partition)
+	[[ $PA = isolated ]] || {
+		echo "FAIL: concurrent partition HK_A not isolated, got '$PA'"
+		echo member > HK_A/cpuset.cpus.partition 2>/dev/null
+		echo member > HK_B/cpuset.cpus.partition 2>/dev/null
+		rmdir HK_A HK_B 2>/dev/null
+		echo member > cpuset.cpus.partition 2>/dev/null
+		exit 1
+	}
+	[[ $PB = isolated ]] || {
+		echo "FAIL: concurrent partition HK_B not isolated, got '$PB'"
+		echo member > HK_A/cpuset.cpus.partition 2>/dev/null
+		echo member > HK_B/cpuset.cpus.partition 2>/dev/null
+		rmdir HK_A HK_B 2>/dev/null
+		echo member > cpuset.cpus.partition 2>/dev/null
+		exit 1
+	}
+	if [[ $HK_NOHZ_CHECK -eq 1 ]]; then
+		NOHZ_BOTH=$(cat $NOHZ_FILE)
+		verify_nohz_exact "$NOHZ_BEFORE" "$TEST_CPUS" "$NOHZ_BOTH" || {
+			echo member > HK_A/cpuset.cpus.partition 2>/dev/null
+			echo member > HK_B/cpuset.cpus.partition 2>/dev/null
+			rmdir HK_A HK_B 2>/dev/null
+			echo member > cpuset.cpus.partition 2>/dev/null
+			exit 1
+		}
+	fi
+	hk_noise_check_nocb_affinity "$TEST_CPUS" 1 || {
+		echo member > HK_A/cpuset.cpus.partition 2>/dev/null
+		echo member > HK_B/cpuset.cpus.partition 2>/dev/null
+		rmdir HK_A HK_B 2>/dev/null
+		echo member > cpuset.cpus.partition 2>/dev/null
+		exit 1
+	}
+	echo member > HK_A/cpuset.cpus.partition
+	echo member > HK_B/cpuset.cpus.partition
+	rmdir HK_A HK_B
+	echo member > cpuset.cpus.partition
+	if [[ $HK_NOHZ_CHECK -eq 1 ]]; then
+		NOHZ_NOW=$(cat $NOHZ_FILE)
+		verify_nohz_exact "$NOHZ_BEFORE" "" "$NOHZ_NOW" || exit 1
+	fi
+	hk_noise_check_nocb_affinity "$TEST_CPUS" 0 || exit 1
+	echo "" > cpuset.cpus
+
+	#
+	# Concurrent create/destroy: two background subshells race on the same
+	# partition simultaneously.  Verifies that concurrent cpuset writes
+	# do not corrupt kernel state or trigger warnings.
+	#
+	console_msg "HK-noise: concurrent create/destroy race"
+	echo $TEST_CPUS > cpuset.cpus
+	local RACE_LOOPS=30 RACE_PID1 RACE_PID2
+	(for i in $(seq 1 $RACE_LOOPS); do
+		echo isolated > cpuset.cpus.partition 2>/dev/null
+		echo member   > cpuset.cpus.partition 2>/dev/null
+	done) &
+	RACE_PID1=$!
+	(for i in $(seq 1 $RACE_LOOPS); do
+		echo member   > cpuset.cpus.partition 2>/dev/null
+		echo isolated > cpuset.cpus.partition 2>/dev/null
+	done) &
+	RACE_PID2=$!
+	wait $RACE_PID1 $RACE_PID2
+	# Drive to a known-good state regardless of who won the last write.
+	echo member > cpuset.cpus.partition 2>/dev/null || true
+	PART=$(cat cpuset.cpus.partition)
+	[[ $PART = member ]] || {
+		echo "FAIL: concurrent race left partition in bad state: '$PART'"
+		exit 1
+	}
+	if [[ $HK_NOHZ_CHECK -eq 1 ]]; then
+		NOHZ_NOW=$(cat $NOHZ_FILE)
+		verify_nohz_exact "$NOHZ_BEFORE" "" "$NOHZ_NOW" || exit 1
+	fi
+	echo "" > cpuset.cpus
+
+	#
+	# Kernel hard-error check: none of the above scenarios must have
+	# triggered a BUG, OOPS, panic, or RCU stall in the kernel log.
+	# Also catch WARNINGs that implicate our subsystems (cpuset / rcu /
+	# nohz / housekeeping / irq_affinity); ignore unrelated WARNINGs from
+	# other kernel subsystems or user-space processes.
+	#
+	console_msg "HK-noise: checking kernel log for errors"
+	local new_errors
+	new_errors=$(dmesg | tail -n "+$((DMESG_LINES_START + 1))" | \
+		grep -c -E \
+		  'kernel BUG at|OOPS|Kernel panic|RCU Stall|scheduling while atomic' \
+		|| true)
+	local new_subsys_warns
+	new_subsys_warns=$(dmesg | tail -n "+$((DMESG_LINES_START + 1))" | \
+		grep 'WARNING:' | \
+		grep -c -E 'cpuset|rcu|nohz|housekeeping|irq_affinity|dhm' \
+		|| true)
+	local total_errors=$(( new_errors + new_subsys_warns ))
+	[[ $total_errors -eq 0 ]] || {
+		echo "FAIL: $total_errors kernel error(s)/warning(s) during HK-noise test:"
+		dmesg | tail -n "+$((DMESG_LINES_START + 1))" | \
+			grep -E 'kernel BUG at|OOPS|Kernel panic|RCU Stall|scheduling while atomic' | head -10
+		dmesg | tail -n "+$((DMESG_LINES_START + 1))" | \
+			grep 'WARNING:' | grep -E 'cpuset|rcu|nohz|housekeeping|irq_affinity|dhm' | head -10
+		exit 1
+	}
+
+	cd $CGROUP2
+	if [[ $HK_NOHZ_CHECK -eq 1 ]]; then
+		if [[ $HAS_BOOT_NOHZ -eq 0 ]]; then
+			console_msg "HK-noise: PASSED" \
+			            "(zero-boot-param: nohz_full verified via DHM runtime path)"
+		else
+			console_msg "HK-noise: PASSED (with nohz_full verification)"
+		fi
+	else
+		console_msg "HK-noise: PASSED (nohz_full skipped: CONFIG_NO_HZ_FULL not active)"
+	fi
+}
+
 trap cleanup 0 2 3 6
 run_state_test TEST_MATRIX
 run_remote_state_test REMOTE_TEST_MATRIX
 test_isolated
 test_boot_isolated
 test_inotify
+test_hk_noise_isolated
 echo "All tests PASSED."

-- 
2.43.0


^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH v5 07/12] watchdog: Sync watchdog_cpumask with HK_TYPE_KERNEL_NOISE on isolation
  2026-10-02 13:10 ` [PATCH v5 07/12] watchdog: Sync watchdog_cpumask with HK_TYPE_KERNEL_NOISE on isolation Qiliang Yuan
@ 2026-10-02 15:26   ` Bradley Morgan
  0 siblings, 0 replies; 14+ messages in thread
From: Bradley Morgan @ 2026-10-02 15:26 UTC (permalink / raw)
  To: odys.yuan
  Cc: anna-maria, boqun, bsegall, cgroups, corbet, dietmar.eggemann,
	frederic, jiangshanlai, joelagnelf, josh, juri.lelli, linux-doc,
	linux-kernel, linux-kselftest, longman, mathieu.desnoyers,
	mgorman, mingo, neeraj.upadhyay, paulmck, peterz, qiang.zhang,
	rcu, rostedt, shuah, skhan, tglx, tj, urezki, vincent.guittot,
	vschneid, pmladek

On 2 October 2026 14:10:27 BST, Qiliang Yuan <odys.yuan@gmail.com> wrote:
>The watchdog is initialized at boot to run on all housekeeping CPUs
>(HK_TYPE_KERNEL_NOISE). When a cpuset isolated partition removes CPUs
>from that mask at runtime, watchdog continues running on those CPUs
>because nothing updates watchdog_cpumask.
>
>Save the boot-time watchdog_cpumask as watchdog_cpumask_boot, which
>captures the user's intended coverage (possibly narrowed via kernel
>parameter or sysctl) before any runtime isolation. Introduce
>lockup_detector_hk_update() which intersects this boot snapshot with
>the current HK_TYPE_KERNEL_NOISE mask and reconfigures the detector.
>This ensures that isolated CPUs are excluded while honoring any
>manual narrowing the admin applied at or after boot.
>
>lockup_detector_hk_update() snapshots the RCU-protected housekeeping
>mask under rcu_read_lock(), then updates watchdog_cpumask and calls
>__lockup_detector_reconfigure() under watchdog_mutex, matching the
>same locking discipline used by proc_watchdog_cpumask().
>

Added Petr mladek, maybe he may be interested? See comments:

>Signed-off-by: Qiliang Yuan <odys.yuan@gmail.com>
>---
> include/linux/nmi.h |  2 ++
> kernel/watchdog.c   | 24 ++++++++++++++++++++++++
> 2 files changed, 26 insertions(+)
>
>diff --git a/include/linux/nmi.h b/include/linux/nmi.h
>index af69712df5f48..409af7884d939 100644
>--- a/include/linux/nmi.h
>+++ b/include/linux/nmi.h
>@@ -37,6 +37,7 @@ extern int sysctl_hardlockup_all_cpu_backtrace;
> static inline void lockup_detector_init(void) { }
> static inline void lockup_detector_retry_init(void) { }
> static inline void lockup_detector_soft_poweroff(void) { }
>+static inline void lockup_detector_hk_update(void) { }
> #endif /* !CONFIG_LOCKUP_DETECTOR */
> 
> #ifdef CONFIG_SOFTLOCKUP_DETECTOR
>@@ -120,6 +121,7 @@ void watchdog_hardlockup_enable(unsigned int cpu);
> void watchdog_hardlockup_disable(unsigned int cpu);
> 
> void lockup_detector_reconfigure(void);
>+void lockup_detector_hk_update(void);
> 
> #ifdef CONFIG_HARDLOCKUP_DETECTOR_BUDDY
> void watchdog_buddy_check_hardlockup(int hrtimer_interrupts);
>diff --git a/kernel/watchdog.c b/kernel/watchdog.c
>index e567fbb0d4692..d4eccb8e337f1 100644
>--- a/kernel/watchdog.c
>+++ b/kernel/watchdog.c
>@@ -53,6 +53,8 @@ static int __read_mostly watchdog_hardlockup_available;
> 
> struct cpumask watchdog_cpumask __read_mostly;
> unsigned long *watchdog_cpumask_bits = cpumask_bits(&watchdog_cpumask);
>+/* Boot snapshot: user's intended watchdog mask before any runtime isolation. */
>+static struct cpumask watchdog_cpumask_boot __ro_after_init;
> 
> #ifdef CONFIG_HARDLOCKUP_DETECTOR
> 
>@@ -1348,6 +1350,27 @@ static void __init lockup_detector_delay_init(struct work_struct *work)
> 	lockup_detector_setup();
> }
> 

Why no comment?, other than that,

Looks good to me, cheerss

Reviewed-by: Bradley Morgan <brads@mainlining.org>

>+void lockup_detector_hk_update(void)
>+{
>+	cpumask_var_t new_mask;
>+
>+	if (!alloc_cpumask_var(&new_mask, GFP_KERNEL))
>+		return;
>+
>+	rcu_read_lock();
>+	cpumask_and(new_mask, &watchdog_cpumask_boot,
>+		    housekeeping_cpumask_rcu(HK_TYPE_KERNEL_NOISE));
>+	rcu_read_unlock();
>+
>+	mutex_lock(&watchdog_mutex);
>+	cpumask_copy(&watchdog_cpumask, new_mask);
>+	__lockup_detector_reconfigure(false);
>+	mutex_unlock(&watchdog_mutex);
>+
>+	free_cpumask_var(new_mask);
>+}
>+EXPORT_SYMBOL_GPL(lockup_detector_hk_update);
>+
> /*
>  * lockup_detector_retry_init - retry init lockup detector if possible.
>  *
>@@ -1390,6 +1413,7 @@ void __init lockup_detector_init(void)
> 
> 	cpumask_copy(&watchdog_cpumask,
> 		     housekeeping_cpumask(HK_TYPE_KERNEL_NOISE));
>+	cpumask_copy(&watchdog_cpumask_boot, &watchdog_cpumask);
> 
> 	if (!watchdog_hardlockup_probe())
> 		watchdog_hardlockup_available = true;
>
>


--- Thanks!
"I'm not a very positive person" - Linus torvalds

^ permalink raw reply	[flat|nested] 14+ messages in thread

end of thread, other threads:[~2026-10-02 15:27 UTC | newest]

Thread overview: 14+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-02 13:10 [PATCH v5 00/12] Dynamic Housekeeping Management (DHM) via CPUSets Qiliang Yuan
2026-10-02 13:10 ` [PATCH v5 01/12] sched/isolation: Enforce nohz_full as a subset of isolcpus=domain at boot Qiliang Yuan
2026-10-02 13:10 ` [PATCH v5 02/12] sched/isolation: Add runtime housekeeping mask updates with boot snapshots Qiliang Yuan
2026-10-02 13:10 ` [PATCH v5 03/12] sched/isolation: RCU-protect runtime-mutable housekeeping cpumask readers Qiliang Yuan
2026-10-02 13:10 ` [PATCH v5 04/12] cpuset: Drive kernel-noise housekeeping from isolated partitions Qiliang Yuan
2026-10-02 13:10 ` [PATCH v5 05/12] context_tracking: Allow runtime per-CPU user tracking enable/disable Qiliang Yuan
2026-10-02 13:10 ` [PATCH v5 06/12] rcu/nocb: Support lazy init for runtime CPU isolation Qiliang Yuan
2026-10-02 13:10 ` [PATCH v5 07/12] watchdog: Sync watchdog_cpumask with HK_TYPE_KERNEL_NOISE on isolation Qiliang Yuan
2026-10-02 15:26   ` Bradley Morgan
2026-10-02 13:10 ` [PATCH v5 08/12] tick/nohz: Derive full-dynticks state from HK_TYPE_KERNEL_NOISE Qiliang Yuan
2026-10-02 13:10 ` [PATCH v5 09/12] cpuset: Add dhm_cycling_cpus mask to suppress transient invalidation Qiliang Yuan
2026-10-02 13:10 ` [PATCH v5 10/12] cpuset: Drive kernel-noise isolation via per-CPU hotplug cycling Qiliang Yuan
2026-10-02 13:10 ` [PATCH v5 11/12] docs: cgroup-v2: Document kernel-noise isolation via isolated partitions Qiliang Yuan
2026-10-02 13:10 ` [PATCH v5 12/12] selftests/cgroup: Add kernel-noise isolation test to cpuset selftest Qiliang Yuan

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®