From: Gabriele Monaco <gmonaco@redhat.com>
To: Peter Zijlstra <peterz@infradead.org>
Cc: Juri Lelli <juri.lelli@redhat.com>,
linux-kernel@vger.kernel.org, Ingo Molnar <mingo@redhat.com>,
Clark Williams <williams@redhat.com>,
arighi@nvidia.com
Subject: Re: [RFC PATCH] sched/deadline: Avoid dl_server boosting with expired deadline
Date: Tue, 11 Nov 2025 10:58:51 +0100 [thread overview]
Message-ID: <b8454329ce061f0c113b25b9400e2b0771fa9e27.camel@redhat.com> (raw)
In-Reply-To: <20251101000057.GA2184199@noisy.programming.kicks-ass.net>
On Sat, 2025-11-01 at 01:00 +0100, Peter Zijlstra wrote:
> On Fri, Oct 31, 2025 at 04:44:55PM +0100, Peter Zijlstra wrote:
>
> > Anyway, back to noodling on how to make it stop on idle :-)
>
> This seems to behave (at least, it did before the cleanup).
>
> It has the fancy comment, ascii-art and everything. Hopefully we'll all
> get less confused when looking at this in the future.
>
I tested the 3 patches from your tree and I confirm the server behaves as
expected in my test case:
* it properly stops and restarts
* it doesn't boost after deadline or consumed runtime
I'm going to write a model like the one you drew to catch if there are issues
with boosting when not necessary.
From the traces I see the server stays in the state running until the end of my
test (30s) although the CPU should be busy only for the first 5s.
I see things like this, where the server starts as running, bootsts and stops
(while scheduling out, not in the timer):
<idle>-0 11.412790: sched_dl_server_start: comm=server pid=-9 runtime=34227068 deadline=12209337602 yielded=0
<idle>-0 11.412795: sched_wakeup: ksoftirqd/9:94 [120] CPU:009
<idle>-0 11.412802: bprint: pick_task_dl: Picked task 94 (running 1)
<idle>-0 11.412809: event_throttle: -9: armed x sched_switch_in -> running (final)
<idle>-0 11.412809: sched_switch: swapper/9:0 [120] R ==> ksoftirqd/9:94 [120]
ksoftirqd/9-94 11.412833: event_throttle: -9: running x sched_switch_out -> preempted
ksoftirqd/9-94 11.412834: sched_dl_server_stop: comm=server pid=-9 runtime=34190225 deadline=12209337602 yielded=0
ksoftirqd/9-94 11.412840: sched_switch: ksoftirqd/9:94 [120] S ==> irq/50-virtio1:213 [49]
Is there a reason why dl_server_stop() doesn't reset the running flag? In this
case if the CPU is only running fair tasks in bursts (like ksoftirqd), it will
continue boosting them and stopping the server when they go to sleep, wont' it?
Thanks,
Gabriele
> ---
> include/linux/sched.h | 15 +--
> kernel/sched/deadline.c | 233
> +++++++++++++++++++++++++++++++++++++++++++++++-
> 2 files changed, 240 insertions(+), 8 deletions(-)
>
> --- a/include/linux/sched.h
> +++ b/include/linux/sched.h
> @@ -685,20 +685,22 @@ struct sched_dl_entity {
> *
> * @dl_server tells if this is a server entity.
> *
> - * @dl_defer tells if this is a deferred or regular server. For
> - * now only defer server exists.
> - *
> - * @dl_defer_armed tells if the deferrable server is waiting
> - * for the replenishment timer to activate it.
> - *
> * @dl_server_active tells if the dlserver is active(started).
> * dlserver is started on first cfs enqueue on an idle runqueue
> * and is stopped when a dequeue results in 0 cfs tasks on the
> * runqueue. In other words, dlserver is active only when cpu's
> * runqueue has atleast one cfs task.
> *
> + * @dl_defer tells if this is a deferred or regular server. For
> + * now only defer server exists.
> + *
> + * @dl_defer_armed tells if the deferrable server is waiting
> + * for the replenishment timer to activate it.
> + *
> * @dl_defer_running tells if the deferrable server is actually
> * running, skipping the defer phase.
> + *
> + * @dl_defer_idle tracks idle state
> */
> unsigned int dl_throttled : 1;
> unsigned int dl_yielded : 1;
> @@ -709,6 +711,7 @@ struct sched_dl_entity {
> unsigned int dl_defer : 1;
> unsigned int dl_defer_armed : 1;
> unsigned int dl_defer_running : 1;
> + unsigned int dl_defer_idle : 1;
>
> /*
> * Bandwidth enforcement timer. Each -deadline task has its
> --- a/kernel/sched/deadline.c
> +++ b/kernel/sched/deadline.c
> @@ -1173,6 +1173,11 @@ static enum hrtimer_restart dl_server_ti
> */
> rq->curr->sched_class->update_curr(rq);
>
> + if (dl_se->dl_defer_idle) {
> + dl_server_stop(dl_se);
> + return HRTIMER_NORESTART;
> + }
> +
> if (dl_se->dl_defer_armed) {
> /*
> * First check if the server could consume runtime in
> background.
> @@ -1420,10 +1425,11 @@ s64 dl_scaled_delta_exec(struct rq *rq,
> }
>
> static inline void
> -update_stats_dequeue_dl(struct dl_rq *dl_rq, struct sched_dl_entity *dl_se,
> - int flags);
> +update_stats_dequeue_dl(struct dl_rq *dl_rq, struct sched_dl_entity *dl_se,
> int flags);
> +
> static void update_curr_dl_se(struct rq *rq, struct sched_dl_entity *dl_se,
> s64 delta_exec)
> {
> + bool idle = rq->curr == rq->idle;
> s64 scaled_delta_exec;
>
> if (unlikely(delta_exec <= 0)) {
> @@ -1444,6 +1450,9 @@ static void update_curr_dl_se(struct rq
>
> dl_se->runtime -= scaled_delta_exec;
>
> + if (dl_se->dl_defer_idle && !idle)
> + dl_se->dl_defer_idle = 0;
> +
> /*
> * The fair server can consume its runtime while throttled (not
> queued/
> * running as regular CFS).
> @@ -1454,6 +1463,29 @@ static void update_curr_dl_se(struct rq
> */
> if (dl_se->dl_defer && dl_se->dl_throttled &&
> dl_runtime_exceeded(dl_se)) {
> /*
> + * Non-servers would never get time accounted while
> throttled.
> + */
> + WARN_ON_ONCE(!dl_server(dl_se));
> +
> + /*
> + * While the server is marked idle, do not push out the
> + * activation further, instead wait for the period timer
> + * to lapse and stop the server.
> + */
> + if (dl_se->dl_defer_idle && idle) {
> + /*
> + * The timer is at the zero-laxity point, this means
> + * dl_server_stop() / dl_server_start() can happen
> + * while now < deadline. This means
> update_dl_entity()
> + * will not replenish. Additionally start_dl_timer()
> + * will be set for 'deadline - runtime'. Negative
> + * runtime will not do.
> + */
> + dl_se->runtime = 0;
> + return;
> + }
> +
> + /*
> * If the server was previously activated - the starving
> condition
> * took place, it this point it went away because the fair
> scheduler
> * was able to get runtime in background. So return to the
> initial
> @@ -1465,6 +1497,9 @@ static void update_curr_dl_se(struct rq
>
> replenish_dl_new_period(dl_se, dl_se->rq);
>
> + if (idle)
> + dl_se->dl_defer_idle = 1;
> +
> /*
> * Not being able to start the timer seems problematic. If it
> could not
> * be started for whatever reason, we need to "unthrottle"
> the DL server
> @@ -1560,6 +1595,199 @@ void dl_server_update(struct sched_dl_en
> update_curr_dl_se(dl_se->rq, dl_se, delta_exec);
> }
>
> +/*
> + * dl_server && dl_defer:
> + *
> + * 6
> + * +--------------------+
> + * v |
> + * +-------------+ 4 +-----------+ 5 +------------------+
> + * +-> | A:init | <--- | D:running | -----> | E:replenish-wait |
> + * | +-------------+ +-----------+ +------------------+
> + * | | | 1 ^ ^ |
> + * | | 1 +----------+ | 3 |
> + * | v | |
> + * | +--------------------------------+ 2 |
> + * | | | ----+ |
> + * | 8 | B:zero_laxity-wait | | |
> + * | | | <---+ |
> + * | +--------------------------------+ |
> + * | | ^ ^ 2 |
> + * | | 7 | 2 +--------------------+
> + * | v |
> + * | +-------------+ |
> + * +-- | C:idle-wait | -+
> + * +-------------+
> + * ^ 7 |
> + * +---------+
> + *
> + *
> + * [A] - init
> + * dl_server_active = 0
> + * dl_throttled = 0
> + * dl_defer_armed = 0
> + * dl_defer_running = 0/1
> + * dl_defer_idle = 0
> + *
> + * [B] - zero_laxity-wait
> + * dl_server_active = 1
> + * dl_throttled = 1
> + * dl_defer_armed = 1
> + * dl_defer_running = 0
> + * dl_defer_idle = 0
> + *
> + * [C] - idle-wait
> + * dl_server_active = 1
> + * dl_throttled = 1
> + * dl_defer_armed = 1
> + * dl_defer_running = 0
> + * dl_defer_idle = 1
> + *
> + * [D] - running
> + * dl_server_active = 1
> + * dl_throttled = 0
> + * dl_defer_armed = 0
> + * dl_defer_running = 1
> + * dl_defer_idle = 0
> + *
> + * [E] - replenish-wait
> + * dl_server_active = 1
> + * dl_throttled = 1
> + * dl_defer_armed = 0
> + * dl_defer_running = 1
> + * dl_defer_idle = 0
> + *
> + *
> + * [1] A->B, A->D
> + * dl_server_start()
> + * dl_server_active = 1;
> + * enqueue_dl_entity()
> + * update_dl_entity(WAKEUP)
> + * if (!dl_defer_running)
> + * dl_defer_armed = 1;
> + * dl_throttled = 1;
> + * if (dl_throttled && start_dl_timer())
> + * return;
> + * // start server into waiting for zero-laxity
> + * __enqueue_dl_entity();
> + * // start running
> + *
> + * // deplete server runtime from client-class
> + * [2] B->B, C->B, E->B
> + * dl_server_update()
> + * update_curr_dl_se()
> + * if (dl_defer_idle)
> + * dl_defer_idle = 0;
> + * if (dl_defer && dl_throttled && dl_runtime_exceeded())
> + * dl_defer_running = 0;
> + * hrtimer_try_to_cancel(); // stop timer
> + * replenish_dl_new_period()
> + * // fwd period
> + * dl_throttled = 1;
> + * dl_defer_armed = 1;
> + * start_dl_timer(); // restart timer
> + * // back into waiting for zero-laxity
> + *
> + * // timer actually fires means we have runtime
> + * [3] B->D
> + * dl_server_timer()
> + * if (dl_defer_armed)
> + * dl_defer_running = 1;
> + * enqueue_dl_entity(REPLENISH)
> + * replenish_dl_entity()
> + * // fwd period
> + * if (dl_throttled)
> + * dl_throttled = 0;
> + * if (dl_defer_armed)
> + * dl_defer_armed = 0;
> + * __enqueue_dl_entity();
> + * // goto [4]
> + *
> + * // schedule server
> + * [4] D->A
> + * pick_task_dl()
> + * p = server_pick_task();
> + * if (!p)
> + * dl_server_stop()
> + * dequeue_dl_entity();
> + * hrtimer_try_to_cancel();
> + * dl_defer_armed = 0;
> + * dl_throttled = 0;
> + * dl_server_active = 0;
> + * // goto [1]
> + * return p;
> + *
> + * // server running
> + * [5] D->E
> + * update_curr_dl_se()
> + * if (dl_runtime_exceeded())
> + * dl_throttled = 1;
> + * dequeue_dl_entity();
> + * start_dl_timer();
> + * // replenish-timer
> + *
> + * // server exchausted
> + * [6] E->D
> + * dl_server_timer()
> + * enqueue_dl_entity(REPLENISH)
> + * replenish_dl_entity()
> + * fwd-period
> + * if (dl_throttled)
> + * dl_throttled = 0;
> + * __enqueue_dl_entity();
> + * // goto [4]
> + *
> + * // deplete server runtime from idle
> + * [7] B->C, C->C
> + * dl_server_update_idle()
> + * update_curr_dl_se()
> + * if (dl_defer && dl_throttled && dl_runtime_exceeded())
> + * if (dl_defer_idle)
> + * return;
> + * dl_defer_running = 0;
> + * hrtimer_try_to_cancel();
> + * replenish_dl_new_period()
> + * // fwd period
> + * dl_throttled = 1;
> + * dl_defer_armed = 1;
> + * dl_defer_idle = 1;
> + * start_dl_timer(); // restart timer
> + *
> + * // stop idle server
> + * [8] C->A
> + * dl_server_timer()
> + * if (dl_defer_idle)
> + * dl_server_stop();
> + *
> + *
> + * digraph dl_server {
> + * "A:init" -> "B:zero_laxity-wait" [label="1:dl_server_start"]
> + * "A:init" -> "D:running" [label="1:dl_server_start"]
> + * "B:zero_laxity-wait" -> "B:zero_laxity-wait"
> [label="2:dl_server_update"]
> + * "B:zero_laxity-wait" -> "C:idle-wait"
> [label="7:dl_server_update_idle"]
> + * "B:zero_laxity-wait" -> "D:running" [label="3:dl_server_timer"]
> + * "C:idle-wait" -> "C:idle-wait"
> [label="7:dl_server_update_idle"]
> + * "C:idle-wait" -> "B:zero_laxity-wait"
> [label="2:dl_server_update"]
> + * "C:idle-wait" -> "A:init" [label="8:dl_server_timer"]
> + * "D:running" -> "A:init" [label="4:pick_task_dl"]
> + * "D:running" -> "E:replenish-wait"
> [label="5:update_curr_dl_se"]
> + * "E:replenish-wait" -> "B:zero_laxity-wait"
> [label="2:dl_server_update"]
> + * "E:replenish-wait" -> "D:running" [label="6:dl_server_timer"]
> + * }
> + *
> + *
> + * Notes:
> + *
> + * - When there are fair tasks running the most likely loop is [2]->[2].
> + * the dl_server never actually runs, the timer never fires.
> + *
> + * - When there is actual fair starvation; the timer fires and starts the
> + * dl_server. This will then throttle and replenish like a normal DL
> + * task. Notably it will not 'defer' again.
> + *
> + * - When idle it will push the activation forward once, and then wait
> + * for the timer to hit or a non-idle update to restart things.
> + */
> void dl_server_start(struct sched_dl_entity *dl_se)
> {
> struct rq *rq = dl_se->rq;
> @@ -1590,6 +1818,7 @@ void dl_server_stop(struct sched_dl_enti
> hrtimer_try_to_cancel(&dl_se->dl_timer);
> dl_se->dl_defer_armed = 0;
> dl_se->dl_throttled = 0;
> + dl_se->dl_defer_idle = 0;
> dl_se->dl_server_active = 0;
> }
>
next prev parent reply other threads:[~2025-11-11 9:58 UTC|newest]
Thread overview: 26+ messages / expand[flat|nested] mbox.gz Atom feed top
2025-10-07 12:29 Gabriele Monaco
2025-10-14 9:30 ` Gabriele Monaco
2025-10-14 9:54 ` Peter Zijlstra
2025-10-14 10:05 ` Gabriele Monaco
2025-10-14 10:25 ` Peter Zijlstra
2025-10-14 15:32 ` Gabriele Monaco
2025-10-14 16:01 ` Juri Lelli
2025-10-14 19:33 ` Peter Zijlstra
2025-10-15 5:40 ` Juri Lelli
2025-10-20 14:11 ` Peter Zijlstra
2025-10-22 10:11 ` Gabriele Monaco
2025-10-30 18:42 ` Peter Zijlstra
2025-10-31 13:05 ` Peter Zijlstra
2025-10-31 13:24 ` Gabriele Monaco
2025-10-31 15:20 ` Peter Zijlstra
2025-10-31 15:41 ` Gabriele Monaco
2025-10-31 15:44 ` Peter Zijlstra
2025-10-31 15:51 ` Gabriele Monaco
2025-11-01 0:00 ` Peter Zijlstra
2025-11-11 9:58 ` Gabriele Monaco [this message]
2025-11-11 11:17 ` Peter Zijlstra
2025-11-11 11:24 ` Peter Zijlstra
2025-11-11 11:37 ` [tip: sched/core] sched/deadline: Fix dl_server stop condition tip-bot2 for Peter Zijlstra
2025-11-01 0:08 ` [RFC PATCH] sched/deadline: Avoid dl_server boosting with expired deadline Peter Zijlstra
2025-11-01 8:43 ` Gabriele Monaco
2025-11-11 11:37 ` [tip: sched/core] sched/deadline: Fix dl_server time accounting tip-bot2 for Peter Zijlstra
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=b8454329ce061f0c113b25b9400e2b0771fa9e27.camel@redhat.com \
--to=gmonaco@redhat.com \
--cc=arighi@nvidia.com \
--cc=juri.lelli@redhat.com \
--cc=linux-kernel@vger.kernel.org \
--cc=mingo@redhat.com \
--cc=peterz@infradead.org \
--cc=williams@redhat.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®