mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Andrea Righi <arighi@nvidia.com>
To: Tejun Heo <tj@kernel.org>, David Vernet <void@manifault.com>,
	Changwoo Min <changwoo@igalia.com>,
	John Stultz <jstultz@google.com>
Cc: Ingo Molnar <mingo@redhat.com>,
	Peter Zijlstra <peterz@infradead.org>,
	Juri Lelli <juri.lelli@redhat.com>,
	Vincent Guittot <vincent.guittot@linaro.org>,
	Dietmar Eggemann <dietmar.eggemann@arm.com>,
	Steven Rostedt <rostedt@goodmis.org>,
	Ben Segall <bsegall@google.com>, Mel Gorman <mgorman@suse.de>,
	Valentin Schneider <vschneid@redhat.com>,
	K Prateek Nayak <kprateek.nayak@amd.com>,
	Christian Loehle <christian.loehle@arm.com>,
	David Dai <david.dai@linux.dev>, Koba Ko <kobak@nvidia.com>,
	Aiqun Yu <aiqun.yu@oss.qualcomm.com>,
	Shuah Khan <shuah@kernel.org>,
	sched-ext@lists.linux.dev, linux-kernel@vger.kernel.org
Subject: [PATCH 12/16] sched_ext: Track proxy execution for NOHZ_FULL
Date: Tue, 22 Sep 2026 18:51:51 +0200	[thread overview]
Message-ID: <20260922165445.943315-13-arighi@nvidia.com> (raw)
In-Reply-To: <20260922165445.943315-1-arighi@nvidia.com>

scx_can_stop_tick() currently treats any blocked rq->donor as an active
proxy session. However, rq->donor can temporarily still identify an
outgoing blocked task while the scheduler selects an ordinary execution
context. A dependency update in that window can keep TICK_DEP_BIT_SCHED
set after the context switch.

Track whether proxy resolution actually selected a different execution
context. Keep the tick active while that state is set instead of
deriving it from a potentially stale donor.

When proxy execution ends, clear the state and queue a balance callback
to reevaluate the tick dependency after context_switch() updates
rq->curr. This lets sched_can_stop_tick() observe the complete donor and
execution selection without adding work outside the existing
proxy-execution branch.

Conservatively re-enable the periodic scheduler tick for the duration of
proxy execution. Allowing proxy execution itself to run tickless would
require making the remote NOHZ scheduler tick aware of the split between
rq->curr and rq->donor, which is left as a future improvement.

Perform the tick update from scx_proxy_reenqueue_retry(), which already
runs after proxy resolution, so scheduler core needs no additional hook.
Compile the additional runqueue state and callback out when
CONFIG_NO_HZ_FULL is disabled.

Signed-off-by: Andrea Righi <arighi@nvidia.com>
---
 kernel/sched/ext/ext.c | 44 +++++++++++++++++++++++++++++++++++++++---
 kernel/sched/sched.h   |  4 ++++
 2 files changed, 45 insertions(+), 3 deletions(-)

diff --git a/kernel/sched/ext/ext.c b/kernel/sched/ext/ext.c
index 1f4016c04d527..247da76021044 100644
--- a/kernel/sched/ext/ext.c
+++ b/kernel/sched/ext/ext.c
@@ -1130,13 +1130,51 @@ static void schedule_deferred_locked(struct rq *rq)
 	schedule_deferred(rq);
 }
 
+#ifdef CONFIG_NO_HZ_FULL
+static void scx_proxy_tick_bal_cb(struct rq *rq)
+{
+	sched_update_tick_dependency(rq);
+}
+
+static void scx_proxy_update_tick(struct rq *rq, struct task_struct *next)
+{
+	bool proxy = next != rq->donor;
+	bool was_proxy = rq->scx.flags & SCX_RQ_PROXY_TICK;
+
+	if (likely(proxy == was_proxy))
+		return;
+
+	if (proxy) {
+		/* Keep proxy execution tick-driven for now. */
+		rq->scx.flags |= SCX_RQ_PROXY_TICK;
+		tick_nohz_dep_set_cpu(cpu_of(rq), TICK_DEP_BIT_SCHED);
+	} else {
+		rq->scx.flags &= ~SCX_RQ_PROXY_TICK;
+		/*
+		 * The selected donor is already visible, but rq->curr still
+		 * identifies the outgoing execution context. Reevaluate after
+		 * context_switch() updates rq->curr so sched_can_stop_tick() sees
+		 * the complete selection.
+		 */
+		queue_balance_callback(rq, &rq->scx.proxy_tick_bal_cb,
+				       scx_proxy_tick_bal_cb);
+	}
+}
+#endif
+
 /*
- * Retry proxy-rejected tasks which couldn't be reenqueued by an earlier drain.
+ * Complete sched_ext bookkeeping after proxy resolution and retry tasks which
+ * couldn't be reenqueued by an earlier reject DSQ drain.
  */
 void scx_proxy_reenqueue_retry(struct rq *rq, struct task_struct *next)
 {
 	lockdep_assert_rq_held(rq);
 
+#ifdef CONFIG_NO_HZ_FULL
+	if (scx_enabled() && tick_nohz_full_cpu(cpu_of(rq)))
+		scx_proxy_update_tick(rq, next);
+#endif
+
 	if (rq->scx.flags & SCX_RQ_PROXY_RETRY) {
 		rq->scx.flags &= ~SCX_RQ_PROXY_RETRY;
 		schedule_deferred_locked(rq);
@@ -4966,8 +5004,8 @@ bool scx_can_stop_tick(struct rq *rq)
 	struct task_struct *p = rq->donor;
 	struct scx_sched *sch = scx_task_sched(p);
 
-	/* Keep the tick running while a blocked proxy donor is selected. */
-	if (p->is_blocked)
+	/* Proxy execution is conservatively tick-driven for now. */
+	if (rq->scx.flags & SCX_RQ_PROXY_TICK)
 		return false;
 
 	if (p->sched_class != &ext_sched_class)
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index 52c60a884994f..083074edd0bab 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -789,6 +789,7 @@ enum scx_rq_flags {
 	SCX_RQ_SUB_IDLE_RENOTIFY	= 1 << 7, /* sub-scheds are owed update_idle() */
 	SCX_RQ_ROOT_IDLE_RENOTIFY	= 1 << 8, /* the root is owed update_idle() */
 	SCX_RQ_PROXY_RETRY	= 1 << 9, /* proxy-rejected tasks need retry */
+	SCX_RQ_PROXY_TICK	= 1 << 10, /* proxy execution requires the tick */
 
 	SCX_RQ_IN_WAKEUP	= 1 << 16,
 	SCX_RQ_IN_DISPATCH	= 1 << 17,
@@ -843,6 +844,9 @@ struct scx_rq {
 	struct list_head	deferred_reenq_users;	/* user DSQs requesting reenq */
 	struct balance_callback	deferred_bal_cb;
 	struct balance_callback	kick_sync_bal_cb;
+#ifdef CONFIG_NO_HZ_FULL
+	struct balance_callback	proxy_tick_bal_cb;
+#endif
 	struct irq_work		deferred_irq_work;
 	struct irq_work		kick_cpus_irq_work;
 };
-- 
2.55.0


  parent reply	other threads:[~2026-09-22 16:55 UTC|newest]

Thread overview: 17+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-22 16:51 [PATCHSET v14 sched_ext/for-7.4] sched: Make proxy execution compatible with sched_ext Andrea Righi
2026-09-22 16:51 ` [PATCH 01/16] sched/core: Drop mutex locks before proxy rescheduling Andrea Righi
2026-09-22 16:51 ` [PATCH 02/16] sched/core: Dequeue waking proxy donors before reset Andrea Righi
2026-09-22 16:51 ` [PATCH 03/16] sched/core: Mark wakeups completed through ttwu_runnable() Andrea Righi
2026-09-22 16:51 ` [PATCH 04/16] sched: Add helper to block retained proxy donors Andrea Righi
2026-09-22 16:51 ` [PATCH 05/16] sched: Add sched_ext hooks for proxy execution Andrea Righi
2026-09-22 16:51 ` [PATCH 06/16] sched_ext: Block proxy donors before taking control Andrea Righi
2026-09-22 16:51 ` [PATCH 07/16] sched_ext: Fix ops.running/stopping() pairing for proxy-exec donors Andrea Righi
2026-09-22 16:51 ` [PATCH 08/16] sched_ext: Move reject DSQ draining into core Andrea Righi
2026-09-22 16:51 ` [PATCH 09/16] sched_ext: Generalize the reject DSQ reenqueue path Andrea Righi
2026-09-22 16:51 ` [PATCH 10/16] sched_ext: Handle proxy-exec races in remote DSQ transfers Andrea Righi
2026-09-22 16:51 ` [PATCH 11/16] sched_ext: Split curr|donor references properly Andrea Righi
2026-09-22 16:51 ` Andrea Righi [this message]
2026-09-22 16:51 ` [PATCH 13/16] sched_ext: Delegate proxy donor admission to BPF schedulers Andrea Righi
2026-09-22 16:51 ` [PATCH 14/16] sched_ext: Add selftest for blocked donor admission Andrea Righi
2026-09-22 16:51 ` [PATCH 15/16] sched_ext: scx_qmap: Add proxy execution support Andrea Righi
2026-09-22 16:51 ` [PATCH 16/16] sched: Allow enabling proxy exec with sched_ext Andrea Righi

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260922165445.943315-13-arighi@nvidia.com \
    --to=arighi@nvidia.com \
    --cc=aiqun.yu@oss.qualcomm.com \
    --cc=bsegall@google.com \
    --cc=changwoo@igalia.com \
    --cc=christian.loehle@arm.com \
    --cc=david.dai@linux.dev \
    --cc=dietmar.eggemann@arm.com \
    --cc=jstultz@google.com \
    --cc=juri.lelli@redhat.com \
    --cc=kobak@nvidia.com \
    --cc=kprateek.nayak@amd.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=mgorman@suse.de \
    --cc=mingo@redhat.com \
    --cc=peterz@infradead.org \
    --cc=rostedt@goodmis.org \
    --cc=sched-ext@lists.linux.dev \
    --cc=shuah@kernel.org \
    --cc=tj@kernel.org \
    --cc=vincent.guittot@linaro.org \
    --cc=void@manifault.com \
    --cc=vschneid@redhat.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®