From: Shrikanth Hegde <sshegde@linux.ibm.com>
To: linux-kernel@vger.kernel.org, mingo@kernel.org,
peterz@infradead.org, juri.lelli@redhat.com,
vincent.guittot@linaro.org, yury.norov@gmail.com,
kprateek.nayak@amd.com, iii@linux.ibm.com, corbet@lwn.net,
meted@linux.ibm.com, ynorov@nvidia.com
Cc: sshegde@linux.ibm.com, tglx@kernel.org,
gregkh@linuxfoundation.org, pbonzini@redhat.com,
seanjc@google.com, vschneid@redhat.com, huschle@linux.ibm.com,
rostedt@goodmis.org, dietmar.eggemann@arm.com,
maddy@linux.ibm.com, srikar@linux.ibm.com, hdanton@sina.com,
chleroy@kernel.org, vineeth@bitbyteword.org, frederic@kernel.org,
arighi@nvidia.com, pauld@redhat.com, christian.loehle@arm.com,
tj@kernel.org, tommaso.cucinotta@gmail.com, maz@kernel.org,
rafael@kernel.org, rdunlap@infradead.org, kernellwp@gmail.com,
linux-doc@vger.kernel.org, jgross@suse.com,
virtualization@lists.linux.dev, sunlightlinux@gmail.com
Subject: [PATCH v14 08/13] sched/core: Push current task from non preferred CPU
Date: Mon, 28 Sep 2026 11:07:23 +0530 [thread overview]
Message-ID: <20260928053728.797539-9-sshegde@linux.ibm.com> (raw)
In-Reply-To: <20260928053728.797539-1-sshegde@linux.ibm.com>
Actively push out the current running task on a non-preferred CPU (NPC).
Since the task is currently running, a stopper thread must be queued
to push the task out. However, if the task is pinned only to
non-preferred CPUs, it will continue running there.
This helps to maintain userspace affinities, unlike CPU hotplug
or isolated cpusets.
The implementation follows the structure of __balance_push_cpu_stop(),
but is kept separate to handle the preferred-CPU specific conditions and
pending-work state under CONFIG_PREFERRED_CPU.
Add the npc_push_work_pending flag to protect the work buffer.
For now, only the currently running task is pushed out. This keeps the code
simpler. In the future, an optimization may be added to move all queued
tasks on the runqueue.
This works only for the FAIR scheduling class.
Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com>
---
kernel/sched/core.c | 83 ++++++++++++++++++++++++++++++++++++++++++++
kernel/sched/sched.h | 9 +++++
2 files changed, 92 insertions(+)
diff --git a/kernel/sched/core.c b/kernel/sched/core.c
index 04400f934cc7..5049eff58fb7 100644
--- a/kernel/sched/core.c
+++ b/kernel/sched/core.c
@@ -5809,6 +5809,9 @@ void sched_tick(void)
unsigned long hw_pressure;
u64 resched_latency;
+ if (!cpu_preferred(cpu))
+ sched_push_current_non_preferred_cpu(rq);
+
if (housekeeping_cpu(cpu, HK_TYPE_KERNEL_NOISE))
arch_scale_freq_tick();
@@ -11201,3 +11204,83 @@ void sched_change_end(struct sched_change_ctx *ctx)
p->sched_class->prio_changed(rq, p, ctx->prio);
}
}
+
+#ifdef CONFIG_PREFERRED_CPU
+static DEFINE_PER_CPU(struct cpu_stop_work, npc_push_task_work);
+
+static int sched_non_preferred_cpu_push_stop(void *arg)
+{
+ struct task_struct *p = arg;
+ struct rq *rq = this_rq();
+ struct rq_flags rf;
+ int cpu;
+
+ if (cpu_preferred(rq->cpu)) {
+ scoped_guard(rq_lock_irqsave, rq)
+ rq->npc_push_work_pending = false;
+ put_task_struct(p);
+ return 0;
+ }
+
+ scoped_guard (raw_spinlock_irq, &p->pi_lock) {
+ /*
+ * select_fallback_rq() may acquire the rq lock in case of
+ * fallback. So call it before grabbing rq lock. If the task
+ * migrates to another CPU before the rq lock is acquired,
+ * subsequent validation of task's current rq will help to
+ * safely bail out.
+ */
+ cpu = select_fallback_rq(rq->cpu, p);
+ rq_lock(rq, &rf);
+ rq->npc_push_work_pending = false;
+ update_rq_clock(rq);
+ context_unsafe_alias(rq);
+
+ if (task_rq(p) == rq && task_on_rq_queued(p))
+ rq = __migrate_task(rq, &rf, p, cpu);
+ rq_unlock(rq, &rf);
+ }
+
+ put_task_struct(p);
+ return 0;
+}
+
+/*
+ * Push the current task running on non-preferred CPU(npc).
+ * Using this non preferred CPU will lead to more contention
+ * in the host. So it is better not to use this CPU.
+ *
+ * Since task is running, call a stopper to push the task out. This is
+ * similar to how task moves during hotplug. In select_fallback_rq() a
+ * preferred CPU will be chosen and henceforth task shouldn't come back to
+ * this CPU again.
+ *
+ * Works for FAIR class only.
+ *
+ * If task is affined only on non-preferred CPUs, no point in moving it out.
+ */
+void sched_push_current_non_preferred_cpu(struct rq *rq)
+{
+ struct task_struct *push_task = rq->curr;
+
+ scoped_guard(rq_lock, rq) {
+ /* Push the task if its explicit affinity allows */
+ if (!task_can_migrate_to_preferred(push_task, rq->cpu))
+ return;
+
+ /* There is already a stopper thread. Don't race with it. */
+ if (rq->npc_push_work_pending)
+ return;
+
+ if (is_migration_disabled(push_task))
+ return;
+
+ rq->npc_push_work_pending = true;
+ }
+
+ /* sched_tick runs with interrupts disabled. */
+ get_task_struct(push_task);
+ stop_one_cpu_nowait(rq->cpu, sched_non_preferred_cpu_push_stop,
+ push_task, this_cpu_ptr(&npc_push_task_work));
+}
+#endif
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index b98084e1f5b0..ee482bb12a66 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -1326,6 +1326,9 @@ struct rq {
#ifdef CONFIG_PARAVIRT_TIME_ACCOUNTING
u64 prev_steal_time_rq;
#endif
+#ifdef CONFIG_PREFERRED_CPU
+ bool npc_push_work_pending;
+#endif
/* calc_load related fields */
unsigned long calc_load_update;
@@ -4292,4 +4295,10 @@ DEFINE_CLASS_IS_UNCONDITIONAL(sched_change)
#include "ext/ext.h"
+#ifdef CONFIG_PREFERRED_CPU
+void sched_push_current_non_preferred_cpu(struct rq *rq);
+#else /* !CONFIG_PREFERRED_CPU */
+static inline void sched_push_current_non_preferred_cpu(struct rq *rq) { }
+#endif
+
#endif /* _KERNEL_SCHED_SCHED_H */
--
2.52.0
next prev parent reply other threads:[~2026-09-28 5:39 UTC|newest]
Thread overview: 30+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-28 5:37 [PATCH v14 00/13] sched, steal_governor: Introduce preferred CPUs and steal-driven vCPU backoff Shrikanth Hegde
2026-09-28 5:37 ` [PATCH v14 01/13] sched/cputime: Add kcpustat_field_total helper Shrikanth Hegde
2026-09-28 7:01 ` [tip: sched/core] " tip-bot2 for Shrikanth Hegde
2026-09-28 5:37 ` [PATCH v14 02/13] cpumask: Introduce cpumask_intersects_and Shrikanth Hegde
2026-09-28 7:01 ` [tip: sched/core] " tip-bot2 for Shrikanth Hegde
2026-09-28 5:37 ` [PATCH v14 03/13] sched/docs: Document cpu_preferred_mask and Preferred CPU concept Shrikanth Hegde
2026-09-28 7:01 ` [tip: sched/core] " tip-bot2 for Shrikanth Hegde
2026-09-28 5:37 ` [PATCH v14 04/13] cpumask: Introduce cpu_preferred_mask Shrikanth Hegde
2026-09-28 7:01 ` [tip: sched/core] " tip-bot2 for Shrikanth Hegde
2026-09-28 5:37 ` [PATCH v14 05/13] sysfs: Add preferred CPU file Shrikanth Hegde
2026-09-28 7:01 ` [tip: sched/core] " tip-bot2 for Shrikanth Hegde
2026-09-28 5:37 ` [PATCH v14 06/13] sched/core: Try to use a preferred CPU in is_cpu_allowed Shrikanth Hegde
2026-09-28 7:01 ` [tip: sched/core] " tip-bot2 for Shrikanth Hegde
2026-09-28 5:37 ` [PATCH v14 07/13] sched/fair: Load balance only among preferred CPUs Shrikanth Hegde
2026-09-28 7:01 ` [tip: sched/core] " tip-bot2 for Shrikanth Hegde
2026-09-28 5:37 ` Shrikanth Hegde [this message]
2026-09-28 7:01 ` [tip: sched/core] sched/core: Push current task from non preferred CPU tip-bot2 for Shrikanth Hegde
2026-09-28 5:37 ` [PATCH v14 09/13] sched/debug: Add migration stats due to non preferred CPUs Shrikanth Hegde
2026-09-28 7:01 ` [tip: sched/core] " tip-bot2 for Shrikanth Hegde
2026-09-29 12:18 ` [PATCH v14 09/13] " Nathan Chancellor
2026-09-29 12:43 ` Shrikanth Hegde
2026-09-29 14:57 ` Shrikanth Hegde
2026-09-28 5:37 ` [PATCH v14 10/13] virt: Introduce steal governor driver Shrikanth Hegde
2026-09-28 7:01 ` [tip: sched/core] " tip-bot2 for Shrikanth Hegde
2026-09-28 5:37 ` [PATCH v14 11/13] virt/steal_governor: Add control knobs for handling steal values Shrikanth Hegde
2026-09-28 7:01 ` [tip: sched/core] " tip-bot2 for Shrikanth Hegde
2026-09-28 5:37 ` [PATCH v14 12/13] virt/steal_governor: Implement steal_governor policy loop Shrikanth Hegde
2026-09-28 7:01 ` [tip: sched/core] " tip-bot2 for Shrikanth Hegde
2026-09-28 5:37 ` [PATCH v14 13/13] virt/steal_governor: Enable the driver Shrikanth Hegde
2026-09-28 7:01 ` [tip: sched/core] " tip-bot2 for Shrikanth Hegde
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260928053728.797539-9-sshegde@linux.ibm.com \
--to=sshegde@linux.ibm.com \
--cc=arighi@nvidia.com \
--cc=chleroy@kernel.org \
--cc=christian.loehle@arm.com \
--cc=corbet@lwn.net \
--cc=dietmar.eggemann@arm.com \
--cc=frederic@kernel.org \
--cc=gregkh@linuxfoundation.org \
--cc=hdanton@sina.com \
--cc=huschle@linux.ibm.com \
--cc=iii@linux.ibm.com \
--cc=jgross@suse.com \
--cc=juri.lelli@redhat.com \
--cc=kernellwp@gmail.com \
--cc=kprateek.nayak@amd.com \
--cc=linux-doc@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=maddy@linux.ibm.com \
--cc=maz@kernel.org \
--cc=meted@linux.ibm.com \
--cc=mingo@kernel.org \
--cc=pauld@redhat.com \
--cc=pbonzini@redhat.com \
--cc=peterz@infradead.org \
--cc=rafael@kernel.org \
--cc=rdunlap@infradead.org \
--cc=rostedt@goodmis.org \
--cc=seanjc@google.com \
--cc=srikar@linux.ibm.com \
--cc=sunlightlinux@gmail.com \
--cc=tglx@kernel.org \
--cc=tj@kernel.org \
--cc=tommaso.cucinotta@gmail.com \
--cc=vincent.guittot@linaro.org \
--cc=vineeth@bitbyteword.org \
--cc=virtualization@lists.linux.dev \
--cc=vschneid@redhat.com \
--cc=ynorov@nvidia.com \
--cc=yury.norov@gmail.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®