From: Andrea Righi <arighi@nvidia.com>
To: Tejun Heo <tj@kernel.org>, David Vernet <void@manifault.com>,
Changwoo Min <changwoo@igalia.com>,
Ingo Molnar <mingo@redhat.com>,
Peter Zijlstra <peterz@infradead.org>,
Juri Lelli <juri.lelli@redhat.com>,
Vincent Guittot <vincent.guittot@linaro.org>
Cc: Dietmar Eggemann <dietmar.eggemann@arm.com>,
Steven Rostedt <rostedt@goodmis.org>,
Ben Segall <bsegall@google.com>, Mel Gorman <mgorman@suse.de>,
Valentin Schneider <vschneid@redhat.com>,
K Prateek Nayak <kprateek.nayak@amd.com>,
Christian Loehle <christian.loehle@arm.com>,
Phil Auld <pauld@redhat.com>, Koba Ko <kobak@nvidia.com>,
Joel Fernandes <joelagnelf@nvidia.com>,
Richard Cheng <icheng@nvidia.com>,
Cheng-Yang Chou <yphbchou0911@gmail.com>,
sched-ext@lists.linux.dev, linux-kernel@vger.kernel.org
Subject: [PATCHSET v2 sched_ext/for-7.2] sched_ext: Auto-manage ext/fair dl_server bandwidth
Date: Tue, 26 May 2026 10:27:54 +0200 [thread overview]
Message-ID: <20260526082954.550958-1-arighi@nvidia.com> (raw)
Currently, a fixed bandwidth is reserved at boot for both the fair and ext
deadline servers, and this reservation remains unchanged unless explicitly
modified via debugfs. As a result, both servers permanently contribute to global
bandwidth accounting, regardless of whether a BPF scheduler is active.
While unused bandwidth can still be reclaimed at runtime by other classes, this
static reservation prevents RT from fully utilizing available headroom in
situations where one of the sched_ext or fair class is guaranteed to be inactive
(for example, when no BPF scheduler is loaded, or when sched_ext runs in full
mode and replaces fair).
As discussed at the VIII OSPM summit in Cambridge [1], a better solution would
be to dynamically register and unregister deadline server bandwidth based on the
active sched_ext state. This allows the kernel to automatically enable bandwidth
accounting only for the scheduling class that is currently active, while
disabling it for inactive ones.
This patch series implements this automatic register/unregister logic. The
sched_ext total_bw kselftest is also modified to validate the correct behavior
across the different scheduling configurations and ensure that bandwidth
accounting follows the expected state transitions.
[1] https://retis.santannapisa.it/ospm-summit/
Git tree: git://git.kernel.org/pub/scm/linux/kernel/git/arighi/linux.git dl-server-bw-v2
Changes in v2:
- Rework the sched_ext enable path as suggested by Peter: attach ext_server
before committing the scheduler switch and fail the enable if admission
control rejects the reservation; detach fair_server only after a successful
full-mode switch.
- Added dl_server_swap_bw() for the disable/recovery path so ext_server detach
and fair_server reattach happen under the same dl_b->lock, closing the
window where concurrent SCHED_DEADLINE admission could steal the freed
bandwidth (reported by Sashiko).
- Fixed the attach/detach accounting issue reported by Sashiko by updating
rq->dl.this_bw together with root-domain total_bw, draining active or
non-contending servers before detach and preventing detached servers from
starting.
- Reuse dl_rq_change_utilization() to drain the server, so the detach path goes
through the same machinery as dl_server_apply_params()
- Made root-domain accounting honor the same cpu_active() conditions used by
root-domain rebuilds, while preserving runtime/period updates made while a
server is detached.
- Fixed the total_bw selftest issues reported by Sashiko: check fclose()
errors for debugfs writes, preserve per-CPU fair_server runtime values, and
restore all CPUs on cleanup even if one write fails.
- Link to v1: https://lore.kernel.org/all/20260521174509.1534623-1-arighi@nvidia.com/
Andrea Righi (2):
sched_ext: Auto-register/unregister dl_server reservations
selftests/sched_ext: Validate dl_server attach/detach in total_bw test
include/linux/sched.h | 6 +
kernel/sched/deadline.c | 207 +++++++++++++++++++++++++--
kernel/sched/ext.c | 71 +++++++++
kernel/sched/sched.h | 4 +
tools/testing/selftests/sched_ext/total_bw.c | 201 +++++++++++++++++++++++++-
5 files changed, 480 insertions(+), 9 deletions(-)
next reply other threads:[~2026-05-26 8:30 UTC|newest]
Thread overview: 3+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-05-26 8:27 Andrea Righi [this message]
2026-05-26 8:27 ` [PATCH 1/2] sched_ext: Auto-register/unregister dl_server reservations Andrea Righi
2026-05-26 8:27 ` [PATCH 2/2] selftests/sched_ext: Validate dl_server attach/detach in total_bw test Andrea Righi
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260526082954.550958-1-arighi@nvidia.com \
--to=arighi@nvidia.com \
--cc=bsegall@google.com \
--cc=changwoo@igalia.com \
--cc=christian.loehle@arm.com \
--cc=dietmar.eggemann@arm.com \
--cc=icheng@nvidia.com \
--cc=joelagnelf@nvidia.com \
--cc=juri.lelli@redhat.com \
--cc=kobak@nvidia.com \
--cc=kprateek.nayak@amd.com \
--cc=linux-kernel@vger.kernel.org \
--cc=mgorman@suse.de \
--cc=mingo@redhat.com \
--cc=pauld@redhat.com \
--cc=peterz@infradead.org \
--cc=rostedt@goodmis.org \
--cc=sched-ext@lists.linux.dev \
--cc=tj@kernel.org \
--cc=vincent.guittot@linaro.org \
--cc=void@manifault.com \
--cc=vschneid@redhat.com \
--cc=yphbchou0911@gmail.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®