From: Juri Lelli <juri.lelli@redhat.com>
To: Yuri Andriaccio <yurand2000@gmail.com>
Cc: Ingo Molnar <mingo@redhat.com>,
Peter Zijlstra <peterz@infradead.org>,
Vincent Guittot <vincent.guittot@linaro.org>,
Dietmar Eggemann <dietmar.eggemann@arm.com>,
Steven Rostedt <rostedt@goodmis.org>,
Ben Segall <bsegall@google.com>, Mel Gorman <mgorman@suse.de>,
Valentin Schneider <vschneid@redhat.com>,
linux-kernel@vger.kernel.org,
Luca Abeni <luca.abeni@santannapisa.it>,
Yuri Andriaccio <yuri.andriaccio@santannapisa.it>
Subject: Re: [RFC PATCH v3 12/24] sched/rt: Update task event callbacks for HCBS scheduling
Date: Thu, 9 Oct 2025 08:54:25 +0200 [thread overview]
Message-ID: <aOdcIbsqwhJTdGjL@jlelli-thinkpadt14gen4.remote.csb> (raw)
In-Reply-To: <20250929092221.10947-13-yurand2000@gmail.com>
Hello,
On 29/09/25 11:22, Yuri Andriaccio wrote:
> Update wakeup_preempt_rt, switched_{from/to}_rt and prio_changed_rt with
> rt-cgroup's specific preemption rules.
> Add checks whether a rt-task can be attached or not to a rt-cgroup.
> Update task_is_throttled_rt for SCHED_CORE.
A little dry. :) This is telling the what, can you please add also the
why and how?
> Co-developed-by: Alessio Balsini <a.balsini@sssup.it>
> Signed-off-by: Alessio Balsini <a.balsini@sssup.it>
> Co-developed-by: Andrea Parri <parri.andrea@gmail.com>
> Signed-off-by: Andrea Parri <parri.andrea@gmail.com>
> Co-developed-by: luca abeni <luca.abeni@santannapisa.it>
> Signed-off-by: luca abeni <luca.abeni@santannapisa.it>
> Signed-off-by: Yuri Andriaccio <yurand2000@gmail.com>
> ---
> kernel/sched/core.c | 2 +-
> kernel/sched/rt.c | 88 ++++++++++++++++++++++++++++++++++++++---
> kernel/sched/syscalls.c | 13 ++++++
> 3 files changed, 96 insertions(+), 7 deletions(-)
>
> diff --git a/kernel/sched/core.c b/kernel/sched/core.c
> index e5b4facee24..2cfbe3b7b17 100644
> --- a/kernel/sched/core.c
> +++ b/kernel/sched/core.c
> @@ -9346,7 +9346,7 @@ static int cpu_cgroup_can_attach(struct cgroup_taskset *tset)
> goto scx_check;
>
> cgroup_taskset_for_each(task, css, tset) {
> - if (!sched_rt_can_attach(css_tg(css), task))
> + if (rt_task(task) && !sched_rt_can_attach(css_tg(css), task))
> return -EINVAL;
> }
> scx_check:
> diff --git a/kernel/sched/rt.c b/kernel/sched/rt.c
> index d9442f64c6b..ce114823fe7 100644
> --- a/kernel/sched/rt.c
> +++ b/kernel/sched/rt.c
> @@ -946,6 +946,50 @@ static void wakeup_preempt_rt(struct rq *rq, struct task_struct *p, int flags)
> {
> struct task_struct *donor = rq->donor;
>
> + if (!rt_group_sched_enabled())
Should we actually be using rt_group_sched_enabled() (instead of
ifdeffery) everywhere to check if group scheduling is enabled and
differentiate behavior?
> + goto no_group_sched;
> +
> + /*
> + * Preemption checks are different if the waking task and the current
> + * task are running on the global runqueue or in a cgroup.
> + * The following rules apply:
> + * - dl-tasks (and equally dl_servers) always preempt FIFO/RR tasks.
We already touched upon this and maybe we can leave it for later, but
the fact that dl_servers always preempt FIFO/RR is going to be a change
of behavior wrt legacy RT group scheduling. Legacy users might have
expectations that setting an high static priority means those tasks will
always be scheduled first and this will break that. Maybe we want to
convince them to change their assumptions, but we need to at least
cleary document the new behavior (and why it is right/better - if it is
indeed :). Not sure if we can also find a way to be "back-compatible"
and if we want to do that.
> + * - if curr is inside a cgroup (i.e. run by a dl_server) and
> + * waking is not, do nothing.
> + * - if waking is inside a cgroup but not curr, always reschedule.
> + * - if they are both on the global runqueue, run the standard code.
> + * - if they are both in the same cgroup, check for tasks priorities.
> + * - if they are both in a cgroup, but not the same one, check whether
> + * the woken task's dl_server preempts the current's dl_server.
> + */
> + if (is_dl_group(rt_rq_of_se(&p->rt)) &&
> + is_dl_group(rt_rq_of_se(&rq->curr->rt))) {
> + struct sched_dl_entity *woken_dl_se, *curr_dl_se;
> +
> + woken_dl_se = dl_group_of(rt_rq_of_se(&p->rt));
> + curr_dl_se = dl_group_of(rt_rq_of_se(&rq->curr->rt));
> +
> + if (rt_rq_of_se(&p->rt)->tg == rt_rq_of_se(&rq->curr->rt)->tg) {
What about checking if woken_dl_se and curr_dl_se are the same dl_se?
> + if (p->prio < rq->curr->prio)
> + resched_curr(rq);
> +
> + return;
> + }
> +
> + if (dl_entity_preempt(woken_dl_se, curr_dl_se))
> + resched_curr(rq);
> +
> + return;
> +
> + } else if (is_dl_group(rt_rq_of_se(&p->rt))) {
> + resched_curr(rq);
> + return;
> +
> + } else if (is_dl_group(rt_rq_of_se(&rq->curr->rt))) {
> + return;
> + }
> +
> +no_group_sched:
> if (p->prio < donor->prio) {
> resched_curr(rq);
> return;
...
> @@ -1750,8 +1797,17 @@ static void switched_to_rt(struct rq *rq, struct task_struct *p)
> * then see if we can move to another run queue.
> */
> if (task_on_rq_queued(p)) {
> +
> +#ifndef CONFIG_RT_GROUP_SCHED
> if (p->nr_cpus_allowed > 1 && rq->rt.overloaded)
> rt_queue_push_tasks(rt_rq_of_se(&p->rt));
> +#else
> + if (rt_rq_of_se(&p->rt)->overloaded) {
Is this empty because intra-group migration is coming with some future
change? Even if that's the case, I believe this wants a comment
explaining why this branch is empty.
> + } else {
> + if (p->prio < rq->curr->prio)
> + resched_curr(rq);
> + }
> +#endif
> if (p->prio < rq->donor->prio && cpu_online(cpu_of(rq)))
> resched_curr(rq);
> }
...
> @@ -1876,7 +1943,16 @@ static unsigned int get_rr_interval_rt(struct rq *rq, struct task_struct *task)
> #ifdef CONFIG_SCHED_CORE
> static int task_is_throttled_rt(struct task_struct *p, int cpu)
> {
> +#ifdef CONFIG_RT_GROUP_SCHED
> + struct rt_rq *rt_rq;
> +
> + rt_rq = task_group(p)->rt_rq[cpu];
> + WARN_ON(!rt_group_sched_enabled() && rt_rq->tg != &root_task_group);
> +
> + return dl_group_of(rt_rq)->dl_throttled;
> +#else
> return 0;
> +#endif
> }
> #endif /* CONFIG_SCHED_CORE */
>
> @@ -2131,7 +2207,7 @@ static int sched_rt_global_constraints(void)
> int sched_rt_can_attach(struct task_group *tg, struct task_struct *tsk)
tsk argument is not going to be used anymore with this, please remove
it. Also maybe rename to something closer to what the function actually
checks for, e.g. [sched_]rt_group_has_runtime.
> {
> /* Don't accept real-time tasks when there is no way for them to run */
> - if (rt_group_sched_enabled() && rt_task(tsk) && tg->rt_bandwidth.rt_runtime == 0)
> + if (rt_group_sched_enabled() && tg->dl_bandwidth.dl_runtime == 0)
> return 0;
Thanks,
Juri
next prev parent reply other threads:[~2025-10-09 6:54 UTC|newest]
Thread overview: 47+ messages / expand[flat|nested] mbox.gz Atom feed top
2025-09-29 9:21 [RFC PATCH v3 00/24] Hierarchical Constant Bandwidth Server Yuri Andriaccio
2025-09-29 9:21 ` [RFC PATCH v3 01/24] sched/deadline: Do not access dl_se->rq directly Yuri Andriaccio
2025-10-02 13:10 ` Juri Lelli
2025-09-29 9:21 ` [RFC PATCH v3 02/24] sched/deadline: Distinct between dl_rq and my_q Yuri Andriaccio
2025-10-02 13:29 ` Juri Lelli
2025-09-29 9:22 ` [RFC PATCH v3 03/24] sched/rt: Pass an rt_rq instead of an rq where needed Yuri Andriaccio
2025-10-02 14:01 ` Juri Lelli
2025-09-29 9:22 ` [RFC PATCH v3 04/24] sched/rt: Move some functions from rt.c to sched.h Yuri Andriaccio
2025-10-02 14:12 ` Juri Lelli
2025-09-29 9:22 ` [RFC PATCH v3 05/24] sched/rt: Disable RT_GROUP_SCHED Yuri Andriaccio
2025-10-02 15:35 ` Juri Lelli
2025-09-29 9:22 ` [RFC PATCH v3 06/24] sched/rt: Introduce HCBS specific structs in task_group Yuri Andriaccio
2025-10-02 15:58 ` Juri Lelli
2025-09-29 9:22 ` [RFC PATCH v3 07/24] sched/core: Initialize root_task_group Yuri Andriaccio
2025-10-02 16:33 ` Juri Lelli
2025-09-29 9:22 ` [RFC PATCH v3 08/24] sched/deadline: Add dl_init_tg Yuri Andriaccio
2025-10-08 6:25 ` Juri Lelli
2025-09-29 9:22 ` [RFC PATCH v3 09/24] sched/rt: Add {alloc/free}_rt_sched_group Yuri Andriaccio
2025-10-08 7:28 ` Juri Lelli
2025-09-29 9:22 ` [RFC PATCH v3 10/24] sched/deadline: Account rt-cgroups bandwidth in deadline tasks schedulability tests Yuri Andriaccio
2025-09-29 9:22 ` [RFC PATCH v3 11/24] sched/rt: Add rt-cgroups' dl-servers operations Yuri Andriaccio
2025-10-08 10:26 ` Juri Lelli
2025-09-29 9:22 ` [RFC PATCH v3 12/24] sched/rt: Update task event callbacks for HCBS scheduling Yuri Andriaccio
2025-10-09 6:54 ` Juri Lelli [this message]
2025-09-29 9:22 ` [RFC PATCH v3 13/24] sched/rt: Update rt-cgroup schedulability checks Yuri Andriaccio
2025-09-29 11:03 ` Markus Elfring
2025-10-09 9:51 ` Juri Lelli
2025-09-29 9:22 ` [RFC PATCH v3 14/24] sched/rt: Allow zeroing the runtime of the root control group Yuri Andriaccio
2025-10-09 13:54 ` Juri Lelli
2025-09-29 9:22 ` [RFC PATCH v3 15/24] sched/rt: Remove old RT_GROUP_SCHED data structures Yuri Andriaccio
2025-09-29 9:22 ` [RFC PATCH v3 16/24] sched/core: Cgroup v2 support Yuri Andriaccio
2025-09-29 9:22 ` [RFC PATCH v3 17/24] sched/rt: Remove support for cgroups-v1 Yuri Andriaccio
2025-10-15 11:51 ` Juri Lelli
2025-09-29 9:22 ` [RFC PATCH v3 18/24] sched/deadline: Allow deeper hierarchies of RT cgroups Yuri Andriaccio
2025-10-15 14:24 ` Juri Lelli
2025-09-29 9:22 ` [RFC PATCH v3 19/24] sched/rt: Add rt-cgroup migration Yuri Andriaccio
2025-09-29 9:22 ` [RFC PATCH v3 20/24] sched/rt: Add HCBS migration related checks and function calls Yuri Andriaccio
2025-09-29 9:22 ` [RFC PATCH v3 21/24] sched/deadline: Make rt-cgroup's servers pull tasks on timer replenishment Yuri Andriaccio
2025-09-29 9:22 ` [RFC PATCH v3 22/24] sched/deadline: Fix HCBS migrations on server stop Yuri Andriaccio
2025-09-29 9:22 ` [RFC PATCH v3 23/24] sched/core: Execute enqueued balance callbacks when changing allowed CPUs Yuri Andriaccio
2025-09-29 9:22 ` [RFC PATCH v3 24/24] sched/core: Execute enqueued balance callbacks when migrating task betweeen cgroups Yuri Andriaccio
2025-10-02 9:00 ` [RFC PATCH v3 00/24] Hierarchical Constant Bandwidth Server Juri Lelli
2025-10-15 14:35 ` Juri Lelli
2025-10-15 15:17 ` Yuri Andriaccio
2025-10-20 9:40 ` Juri Lelli
2025-10-24 8:02 ` luca abeni
2025-11-03 10:32 ` Juri Lelli
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=aOdcIbsqwhJTdGjL@jlelli-thinkpadt14gen4.remote.csb \
--to=juri.lelli@redhat.com \
--cc=bsegall@google.com \
--cc=dietmar.eggemann@arm.com \
--cc=linux-kernel@vger.kernel.org \
--cc=luca.abeni@santannapisa.it \
--cc=mgorman@suse.de \
--cc=mingo@redhat.com \
--cc=peterz@infradead.org \
--cc=rostedt@goodmis.org \
--cc=vincent.guittot@linaro.org \
--cc=vschneid@redhat.com \
--cc=yurand2000@gmail.com \
--cc=yuri.andriaccio@santannapisa.it \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®