mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v5 00/12] Dynamic Housekeeping Management (DHM) via CPUSets
@ 2026-10-02 13:10 Qiliang Yuan
  2026-10-02 13:10 ` [PATCH v5 01/12] sched/isolation: Enforce nohz_full as a subset of isolcpus=domain at boot Qiliang Yuan
                   ` (11 more replies)
  0 siblings, 12 replies; 14+ messages in thread
From: Qiliang Yuan @ 2026-10-02 13:10 UTC (permalink / raw)
  To: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot,
	Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, Paul E. McKenney, Frederic Weisbecker,
	Neeraj Upadhyay, Joel Fernandes, Josh Triplett, Boqun Feng,
	Uladzislau Rezki, Mathieu Desnoyers, Lai Jiangshan, Zqiang,
	Anna-Maria Behnsen, Tejun Heo, Jonathan Corbet, Shuah Khan,
	Shuah Khan, Thomas Gleixner
  Cc: Waiman Long, linux-kernel, rcu, cgroups, linux-doc,
	linux-kselftest, Qiliang Yuan

This series introduces Dynamic Housekeeping Management (DHM) to the Linux
kernel, enabling runtime reconfiguration of kernel-noise housekeeping
(nohz_full tick suppression, RCU NOCB offloading, and managed IRQ
migration) through the existing cgroup v2 cpuset isolated partition
mechanism — no new kernel ABI required.

When a cpuset partition is set to isolated mode, DHM cycles each CPU in
that partition offline, reconfigures the housekeeping masks (removing the
CPU from HK_TYPE_KERNEL_NOISE and related types), and brings it back
online.  The subsystems (tick/nohz, RCU NOCB, genirq) pick up the new
masks through their existing CPU hotplug callbacks.  Destroying the
partition reverses the cycle: each CPU is taken offline, restored to all
housekeeping masks, and brought back online.

Housekeeping cpumask pointers are RCU-protected to allow lock-free readers
during updates.  A global dhm_cycling_cpus mask suppresses transient
cpuset partition invalidation while CPUs are being cycled.

This work is related to Waiman Long's [PATCH-next 00/23] series
(20260421030351.281436-1-longman@redhat.com), which also targets runtime
cpuset housekeeping control.  DHM's distinguishing feature is zero-boot-
parameter activation: nohz_full+nocb isolation can be achieved at runtime
without any boot-time nohz_full= or rcu_nocbs= parameters.

Signed-off-by: Qiliang Yuan <odys.yuan@gmail.com>
---
V4 -> V5:
- Rebase onto current mainline (was based on a v7.1-era snapshot).
- Gate the HK_TYPE_KERNEL_NOISE/HK_TYPE_MANAGED_IRQ mask update on the
  HK_TYPE_DOMAIN update succeeding, so a CPU can never end up
  tick-suppressed while still part of the normal sched domain.
  Frederic asked for HK_TYPE_KERNEL_NOISE to stay a subset of
  HK_TYPE_DOMAIN.
- Keep cpuset_top_mutex held across dhm_cycle_isolated_cpus() instead
  of releasing it before the hotplug cycle, so dhm_prev_isolated's
  "protected by cpuset_top_mutex" invariant actually holds and a second
  isolated-partition update cannot race with one still in flight.
  Frederic asked whether cpus_read_lock should stay held here.
- Remove tick_nohz_full_mask and tick_nohz_full_running; derive
  tick_nohz_full_enabled()/tick_nohz_full_cpu() from
  housekeeping_enabled(HK_TYPE_KERNEL_NOISE) and
  housekeeping_cpumask_rcu(HK_TYPE_KERNEL_NOISE) instead, removing the
  now-redundant lazy-allocation and locking in
  tick_nohz_cpu_isolate()/deisolate().  Frederic pointed out the
  duplication.
- Publish each CPU's HK_TYPE_KERNEL_NOISE/HK_TYPE_MANAGED_IRQ mask
  update from dhm_cycle_isolated_cpus() itself, strictly between that
  CPU's own remove_cpu() succeeding and add_cpu() bringing it back,
  instead of batching the whole isolation set before or after the
  cycle.  A CPU the cycling loop skips (not hotpluggable, or
  remove_cpu() fails) is then simply never published as isolated,
  with no separate pre-filter or republish-on-failure step needed.
- Install tick_do_timer_cpu's hotplug protection from
  housekeeping_update_types()'s HK_TYPE_KERNEL_NOISE first-enable
  path too, via a new tick_nohz_full_hotplug_init(): tick_nohz_init()
  never runs on DHM's zero-boot-param path, so the current duty
  holder had no protection against being pulled down by remove_cpu().
- Keep tick_nohz_cpu_hotpluggable()'s duty-holder check on the plain
  housekeeping_enabled() flag rather than a mask-aware variant:
  housekeeping_update_types() only publishes a CPU's isolation after
  its own remove_cpu() already succeeded, so a mask-aware check would
  see every single-CPU isolation's own live mask as still empty and
  leave the duty holder unprotected each time, not just the first.
  Use the mask-aware variant only for tick_sched_do_timer()'s
  diagnostic WARN, since HK_TYPE_KERNEL_NOISE can stay enabled with an
  empty mask after a full runtime de-isolation, an ordinary
  NO_HZ_IDLE handover rather than a violation.
- Close a context_tracking_key activation race: a CPU whose
  kernel<->user transition lands in the window before every CPU
  observes static_branch_inc()'s patched code sees
  context_tracking_enabled() as still false and silently drops the
  transition, tripping CT_WARN_ON() on its next kernel entry.  Add
  context_tracking_activating, a plain per-CPU bool with no
  code-patching delay of its own, set via IPI before
  static_branch_inc() and cleared via IPI only after it returns, and
  route user_enter_irqoff()/user_exit_irqoff()/CT_WARN_ON()/ct_state()
  through it alongside the static key.
- Fix the boot-time nohz_full=/isolcpus=domain invariant check
  itself: it ran before static_branch_enable(&housekeeping_overridden),
  so housekeeping_cpumask() always returned cpu_possible_mask for
  both sides of the subset comparison and the check never actually
  rejected anything.  Enable housekeeping_overridden first.
- Fix a cpumask_var_t leak in dhm_cycle_isolated_cpus() on the no-op
  path (a partition update that nets to no CPU delta).
- Let HK_TYPE_MANAGED_IRQ use the same runtime first-enable path as
  HK_TYPE_KERNEL_NOISE; it was unconditionally skipped without a
  boot-time isolcpus=managed_irq, so managed-IRQ migration silently
  never activated on a zero-boot-parameter system.
- Finish converting HK_TYPE_MANAGED_IRQ/HK_TYPE_KERNEL_NOISE readers
  to housekeeping_cpumask_rcu(): vmbus_channel_set_cpu() in
  drivers/hv/vmbus_drv.c, rps_cpumask_housekeeping() in
  net/core/net-sysfs.c, and tmigr_isolated_exclude_cpumask() in
  kernel/time/timer_migration.c were still reading the mask
  without RCU protection.
- Register the CONFIG_RCU_LAZY shrinker from the runtime nocb
  first-enable path too; previously only the boot path did, so a
  system with no rcu_nocbs=/nohz_full= at boot never got lazy-callback
  reclaim once DHM enabled NOCB offload at runtime.
- Add a standalone patch enforcing nohz_full=/isolcpus=nohz as a
  subset of isolcpus=domain at boot: disable nohz_full entirely
  (including when it is passed alone, with no isolcpus=domain at
  all) if the invariant does not hold, checked in housekeeping_init()
  so the outcome does not depend on cmdline argument order.
  Frederic's superset invariant above only covers the cpuset-driven
  runtime path; this closes the same gap for the two boot
  parameters, which he separately said he'd enforce "if he could do
  it again".

V3 -> V4:
- Drop apply() callback architecture entirely (struct housekeeping_cbs,
  pre_validate/apply hooks, housekeeping_update_types() notification).
  Thomas Gleixner identified concurrent in-kernel mask apply as
  fundamentally unsafe ("broken beyond repair") and endorsed CPU-by-CPU
  hotplug cycling as the correct approach.
- Track A prerequisites (housekeeping boot type isolation, RCU reader
  protection, cpuset trigger) submitted upstream separately; v4 builds
  on top of those three committed patches.
- Fix rcu/nocb lazy_init: remove __init from rcu_organize_nocb_kthreads()
  forward declaration (tree.h) to prevent GCC from placing the function in
  .init.text, which caused an NX-protected page fault when
  rcu_nocb_cpu_isolate() was called at runtime.  Add noinline to
  rcu_nocb_lazy_init() to preserve the GCC IPA call chain.
- Fix cpuset remote partition cycling suppression: extend dhm_cycling_cpus
  guard to the remote-partition disable path in cpuset_hotplug_update_tasks(),
  preventing false partition invalidation during hotplug cycling steps.
  Fixes selftest TEST_MATRIX[69] (A2: expected 1-2, got 1-3).

V2 -> V3:
- Replace notifier chain with explicit per-type callback interface
  (struct housekeeping_cbs with .name, .pre_validate, .apply fields).
- RCU-protect all housekeeping cpumask pointers; callers must hold
  rcu_read_lock() or use housekeeping_cpumask_rcu() in apply() callbacks.
- Drop 5 patches from v2: HK_TYPE enum separation (upstream aliases are
  already correct), no-op timer/hrtimer patches, kthread dead code, and
  workqueue double-update.
- Fix deadlock in rcu_hk_workfn(): remove cpus_read_lock() wrapper around
  remove_cpu()/add_cpu() which take cpu_hotplug_lock write side.
- Fix UAF in rcu_hk_apply(): snapshot the housekeeping cpumask inside the
  work function under rcu_read_lock(), not at apply() time where the old
  pointer may be freed by synchronize_rcu() before the work runs.
- Fix tick apply(): snapshot housekeeping_cpumask_rcu() under
  rcu_read_lock() as required by lockdep for runtime-mutable types.
- Activate context_tracking dynamically via ct_cpu_track_user() /
  ct_cpu_untrack_user() in tick apply(), eliminating the dependency on
  CONFIG_CONTEXT_TRACKING_USER_FORCE flagged by tglx.
- Fix genirq apply(): snapshot HK_TYPE_MANAGED_IRQ mask under
  rcu_read_lock() before the IRQ iteration loop.
- Simplify cpuset noise_types to BIT(HK_TYPE_KERNEL_NOISE) |
  BIT(HK_TYPE_MANAGED_IRQ), replacing the redundant per-alias bitmask.
- housekeeping_update_types(): always use cpu_possible_mask as base
  for HK_TYPE_KERNEL_NOISE, so de-isolation restores the mask to all
  possible CPUs rather than leaving it at its last non-trivial value.
- Initialize watchdog_cpumask from HK_TYPE_KERNEL_NOISE (not
  HK_TYPE_TIMER) at boot; keep it in sync at runtime via a new
  housekeeping_cbs callback.
- Add kernel-noise selftest to test_cpuset_prs.sh, including
  cpu_in_cpulist() for correct cpulist range membership detection and
  nohz_full sysfs verification when CONFIG_NO_HZ_FULL is active.
- Add RCU caller fixes: sched/core (HK_TYPE_KERNEL_NOISE) and
  drivers/hv (HK_TYPE_MANAGED_IRQ) are required because those types
  are updated at runtime; hrtimer (HK_TYPE_TIMER) and arm64/topology
  (HK_TYPE_TICK) are defensive fixes.
- Reorder patches so all subsystem callbacks are registered before the
  cpuset patch that triggers housekeeping_update_types().

V1 -> V2:
- Rebrand series from DHEI to DHM (Dynamic Housekeeping Management).
- Drop custom sysfs interface entirely.
- Integrate housekeeping control into cgroup v2 cpuset isolated partition
  mechanism.
- Add SMT-aware isolation constraints to prevent splitting SMT siblings.
- Add comprehensive documentation and cgroup functional selftests.
- Refactor mask transition logic to use RCU-safe handover.

v4: https://lore.kernel.org/r/20260710-wujing-dhm-v4-0-2e912e5d9645@gmail.com
v3: https://lore.kernel.org/r/20260618-wujing-dhm-v3-0-28f1a4d83b68@gmail.com
v2: https://lore.kernel.org/r/20260413-wujing-dhm-v2-0-06df21caba5d@gmail.com
v1: https://lore.kernel.org/all/20260325-dhei-v12-final-v1-0-919cca23cadf@gmail.com

---
Qiliang Yuan (12):
      sched/isolation: Enforce nohz_full as a subset of isolcpus=domain at boot
      sched/isolation: Add runtime housekeeping mask updates with boot snapshots
      sched/isolation: RCU-protect runtime-mutable housekeeping cpumask readers
      cpuset: Drive kernel-noise housekeeping from isolated partitions
      context_tracking: Allow runtime per-CPU user tracking enable/disable
      rcu/nocb: Support lazy init for runtime CPU isolation
      watchdog: Sync watchdog_cpumask with HK_TYPE_KERNEL_NOISE on isolation
      tick/nohz: Derive full-dynticks state from HK_TYPE_KERNEL_NOISE
      cpuset: Add dhm_cycling_cpus mask to suppress transient invalidation
      cpuset: Drive kernel-noise isolation via per-CPU hotplug cycling
      docs: cgroup-v2: Document kernel-noise isolation via isolated partitions
      selftests/cgroup: Add kernel-noise isolation test to cpuset selftest

 Documentation/admin-guide/cgroup-v2.rst           |  17 +
 Documentation/admin-guide/kernel-parameters.txt   |  12 +
 arch/arm64/kernel/topology.c                      |   9 +-
 drivers/base/cpu.c                                |  42 +-
 drivers/hv/channel_mgmt.c                         |  50 +-
 drivers/hv/vmbus_drv.c                            |  13 +-
 fs/resctrl/internal.h                             |   6 +-
 include/linux/context_tracking.h                  |  16 +-
 include/linux/context_tracking_state.h            |  25 +-
 include/linux/nmi.h                               |   2 +
 include/linux/rcupdate.h                          |   2 +
 include/linux/sched/isolation.h                   |  34 +-
 include/linux/tick.h                              |  40 +-
 kernel/cgroup/cpuset.c                            | 203 +++++++-
 kernel/context_tracking.c                         |  91 +++-
 kernel/rcu/tree.h                                 |   2 +-
 kernel/rcu/tree_nocb.h                            | 105 +++-
 kernel/sched/core.c                               |   7 +-
 kernel/sched/isolation.c                          | 253 +++++++++-
 kernel/sched/sched.h                              |   2 +-
 kernel/time/hrtimer.c                             |   5 +-
 kernel/time/tick-sched.c                          | 158 +++++-
 kernel/time/timer_migration.c                     |   9 +-
 kernel/watchdog.c                                 |  26 +-
 net/core/net-sysfs.c                              |  10 +-
 tools/testing/selftests/cgroup/test_cpuset_prs.sh | 578 ++++++++++++++++++++++
 26 files changed, 1558 insertions(+), 159 deletions(-)
---
base-commit: 551c722f40809618230001baccf219193e22fc5a
change-id: 20260408-wujing-dhm-8f43e2d49cd8

Best regards,
-- 
Qiliang Yuan <odys.yuan@gmail.com>


^ permalink raw reply	[flat|nested] 14+ messages in thread

end of thread, other threads:[~2026-10-02 15:27 UTC | newest]

Thread overview: 14+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-02 13:10 [PATCH v5 00/12] Dynamic Housekeeping Management (DHM) via CPUSets Qiliang Yuan
2026-10-02 13:10 ` [PATCH v5 01/12] sched/isolation: Enforce nohz_full as a subset of isolcpus=domain at boot Qiliang Yuan
2026-10-02 13:10 ` [PATCH v5 02/12] sched/isolation: Add runtime housekeeping mask updates with boot snapshots Qiliang Yuan
2026-10-02 13:10 ` [PATCH v5 03/12] sched/isolation: RCU-protect runtime-mutable housekeeping cpumask readers Qiliang Yuan
2026-10-02 13:10 ` [PATCH v5 04/12] cpuset: Drive kernel-noise housekeeping from isolated partitions Qiliang Yuan
2026-10-02 13:10 ` [PATCH v5 05/12] context_tracking: Allow runtime per-CPU user tracking enable/disable Qiliang Yuan
2026-10-02 13:10 ` [PATCH v5 06/12] rcu/nocb: Support lazy init for runtime CPU isolation Qiliang Yuan
2026-10-02 13:10 ` [PATCH v5 07/12] watchdog: Sync watchdog_cpumask with HK_TYPE_KERNEL_NOISE on isolation Qiliang Yuan
2026-10-02 15:26   ` Bradley Morgan
2026-10-02 13:10 ` [PATCH v5 08/12] tick/nohz: Derive full-dynticks state from HK_TYPE_KERNEL_NOISE Qiliang Yuan
2026-10-02 13:10 ` [PATCH v5 09/12] cpuset: Add dhm_cycling_cpus mask to suppress transient invalidation Qiliang Yuan
2026-10-02 13:10 ` [PATCH v5 10/12] cpuset: Drive kernel-noise isolation via per-CPU hotplug cycling Qiliang Yuan
2026-10-02 13:10 ` [PATCH v5 11/12] docs: cgroup-v2: Document kernel-noise isolation via isolated partitions Qiliang Yuan
2026-10-02 13:10 ` [PATCH v5 12/12] selftests/cgroup: Add kernel-noise isolation test to cpuset selftest Qiliang Yuan

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®