mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Peter Zijlstra <peterz@infradead.org>
To: quzicheng315@gmail.com
Cc: arighi@nvidia.com, brho@google.com, bsegall@google.com,
	changwoo@igalia.com, dietmar.eggemann@arm.com, haoluo@google.com,
	joshdon@google.com, juri.lelli@redhat.co, kprateek.nayak@amd.com,
	linux-kernel@vger.kernel.org, mgorman@suse.de, mingo@redhat.com,
	quzicheng@huawei.com, rostedt@goodmis.org,
	sched-ext@lists.linux.dev, tanghui20@huawei.com, tj@kernel.org,
	vincent.guittot@linaro.org, void@manifault.com,
	vschneid@redhat.com, zhangqiao22@huawei.com
Subject: Re: [PATCH v2] sched_ext: Rebuild fair weight on ext to fair switches
Date: Wed, 27 May 2026 13:26:24 +0200	[thread overview]
Message-ID: <20260527112624.GT3126523@noisy.programming.kicks-ass.net> (raw)
In-Reply-To: <20260527094037.3494671-1-quzicheng315@gmail.com>

On Wed, May 27, 2026 at 05:40:37PM +0800, quzicheng315@gmail.com wrote:
> From: Zicheng Qu <quzicheng315@gmail.com>
> 
> Tasks running on sched_ext do not use p->se.load as their active
> scheduling weight. Their nice-derived weight is maintained as
> p->scx.weight instead.
> 
> When such a task switches back to fair, CFS expects p->se.load to match
> the task's current policy/static_prio before the task is enqueued.
> However, not all ext to fair transitions rebuild p->se.load. For
> example, scx_root_disable() switches tasks back to fair directly, and
> partial mode can move a task from SCHED_EXT to SCHED_NORMAL through
> sched_setscheduler(). In the latter case, set_load_weight(p, true) runs
> while p->sched_class is still ext_sched_class, so reweight_task_scx()
> updates p->scx.weight but leaves p->se.load stale.
> 
> Rebuild the fair load weight in sched_change_end() when the class switch
> is from ext_sched_class to fair_sched_class. This is after the class has
> been changed and before the task is enqueued on fair, so CFS sees a
> native load_weight derived from the task's current policy/static_prio.
> 
> Fixes: f0e1a0643a59 ("sched_ext: Implement BPF extensible scheduler class")
> Signed-off-by: Zicheng Qu <quzicheng@huawei.com>
> ---
> Changes in v2:
> - Move the fix from scx_root_disable() to sched_change_end() so the same
>   ext-to-fair rebuild also covers partial mode SCHED_EXT to SCHED_NORMAL
>   transitions through sched_setscheduler(), as Andrea pointed out.
> 
>  kernel/sched/core.c |  2 ++
>  kernel/sched/ext.h  | 11 +++++++++++
>  2 files changed, 13 insertions(+)
> 
> diff --git a/kernel/sched/core.c b/kernel/sched/core.c
> index b8871449d3c6..c694aabc451a 100644
> --- a/kernel/sched/core.c
> +++ b/kernel/sched/core.c
> @@ -11200,6 +11200,8 @@ void sched_change_end(struct sched_change_ctx *ctx)
>  	 */
>  	WARN_ON_ONCE(p->sched_class != ctx->class && !(ctx->flags & ENQUEUE_CLASS));
>  
> +	scx_rebuild_fair_weight_on_class_switch(p, ctx->class, p->sched_class);
> +
>  	if ((ctx->flags & ENQUEUE_CLASS) && p->sched_class->switching_to)
>  		p->sched_class->switching_to(rq, p);
>  
> diff --git a/kernel/sched/ext.h b/kernel/sched/ext.h
> index 0b7fc46aee08..1f8248c897af 100644
> --- a/kernel/sched/ext.h
> +++ b/kernel/sched/ext.h
> @@ -35,6 +35,14 @@ static inline bool task_on_scx(const struct task_struct *p)
>  	return scx_enabled() && p->sched_class == &ext_sched_class;
>  }
>  
> +static inline void scx_rebuild_fair_weight_on_class_switch(struct task_struct *p,
> +							   const struct sched_class *old_class,
> +							   const struct sched_class *new_class)
> +{
> +	if (old_class == &ext_sched_class && new_class == &fair_sched_class)
> +		set_load_weight(p, false);
> +}
> +
>  #ifdef CONFIG_SCHED_CORE
>  bool scx_prio_less(const struct task_struct *a, const struct task_struct *b,
>  		   bool in_fi);
> @@ -55,6 +63,9 @@ static inline int scx_check_setscheduler(struct task_struct *p, int policy) { re
>  static inline bool task_on_scx(const struct task_struct *p) { return false; }
>  static inline bool scx_allow_ttwu_queue(const struct task_struct *p) { return true; }
>  static inline void init_sched_ext_class(void) {}
> +static inline void scx_rebuild_fair_weight_on_class_switch(struct task_struct *p,
> +							   const struct sched_class *old_class,
> +							   const struct sched_class *new_class) {}
>  
>  #endif	/* CONFIG_SCHED_CLASS_EXT */

This is truly horrible. We have 4 class methods involved with switching
classes and you stick in a random call in a place that is called when no
class is changed.

Would not something like this work?

diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 62a2dcb0d03e..a2eb43bd73b9 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -14957,6 +14957,11 @@ static void switched_from_fair(struct rq *rq, struct task_struct *p)
 	detach_task_cfs_rq(p);
 }
 
+static void switching_to_fair(struct rq *rq, struct task_struct *p)
+{
+	set_load_weight(p, false);
+}
+
 static void switched_to_fair(struct rq *rq, struct task_struct *p)
 {
 	WARN_ON_ONCE(p->se.sched_delayed);
@@ -15351,6 +15356,7 @@ DEFINE_SCHED_CLASS(fair) = {
 	.prio_changed		= prio_changed_fair,
 	.switching_from		= switching_from_fair,
 	.switched_from		= switched_from_fair,
+	.switching_to		= switching_to_fair,
 	.switched_to		= switched_to_fair,
 
 	.get_rr_interval	= get_rr_interval_fair,

  reply	other threads:[~2026-05-27 11:26 UTC|newest]

Thread overview: 8+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-05-26 13:52 [PATCH] sched_ext: Rebuild fair weight when disabling BPF scheduler Zicheng Qu
2026-05-26 17:20 ` Andrea Righi
2026-05-27  9:40   ` [PATCH v2] sched_ext: Rebuild fair weight on ext to fair switches quzicheng315
2026-05-27 11:26     ` Peter Zijlstra [this message]
2026-05-28  2:53       ` Zicheng Qu
2026-05-28  9:25         ` Peter Zijlstra
2026-05-28 13:12           ` [PATCH v3] sched/fair: Rebuild load weight when switching to fair quzicheng315
2026-05-28 14:27             ` Tejun Heo

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260527112624.GT3126523@noisy.programming.kicks-ass.net \
    --to=peterz@infradead.org \
    --cc=arighi@nvidia.com \
    --cc=brho@google.com \
    --cc=bsegall@google.com \
    --cc=changwoo@igalia.com \
    --cc=dietmar.eggemann@arm.com \
    --cc=haoluo@google.com \
    --cc=joshdon@google.com \
    --cc=juri.lelli@redhat.co \
    --cc=kprateek.nayak@amd.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=mgorman@suse.de \
    --cc=mingo@redhat.com \
    --cc=quzicheng315@gmail.com \
    --cc=quzicheng@huawei.com \
    --cc=rostedt@goodmis.org \
    --cc=sched-ext@lists.linux.dev \
    --cc=tanghui20@huawei.com \
    --cc=tj@kernel.org \
    --cc=vincent.guittot@linaro.org \
    --cc=void@manifault.com \
    --cc=vschneid@redhat.com \
    --cc=zhangqiao22@huawei.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®