* [PATCH v3 0/2] sched: Fix execution-context tick handling under proxy execution
@ 2026-09-04 8:52 Hui Su
2026-09-04 8:52 ` [PATCH v3 1/2] sched/numa: Drive NUMA task tick from execution context Hui Su
2026-09-04 8:52 ` [PATCH v3 2/2] sched/cache: Drive cache " Hui Su
0 siblings, 2 replies; 5+ messages in thread
From: Hui Su @ 2026-09-04 8:52 UTC (permalink / raw)
To: peterz, mingo, tim.c.chen, yu.c.chen, kprateek.nayak
Cc: juri.lelli, vincent.guittot, dietmar.eggemann, rostedt, bsegall,
mgorman, vschneid, connoro, jstultz, linux-kernel
Proxy execution separates the scheduling context in rq->donor from the
execution context in rq->curr. Scheduler tick hooks which operate on
execution-context state must use rq->curr, while scheduling-context
state must continue to use the donor.
Move NUMA and cache task tick handling into a common execution-context
helper. This ensures that both hooks run when a fair task executes on
behalf of an RT or deadline donor, while the remaining fair-class tick
bookkeeping stays associated with the scheduling context.
Changes in v2 (since v1):
- Move NUMA and cache execution-context tick handling from
task_tick_fair() to sched_tick().
- Invoke the hooks when rq->curr is a fair task, allowing them to run
when the donor belongs to another scheduling class.
- Add the corresponding calls to sched_tick_remote() to preserve
full-dynticks behavior.
Changes in v3 (since v2):
- Factor NUMA and cache execution-context tick handling into
sched_tick_exec_ctx(), keeping sched_tick() and sched_tick_remote()
in sync.
- Keep sched_tick_remote()'s existing curr-based scheduling-class
dispatch unchanged and invoke the common helper afterward.
- Reword the NUMA rationale around execution-context state and the
task/mm associated with the execution context, with sum_exec_runtime as
supporting state rather than the sole reason for the change.
- Document why misfit, overutilized and core-scheduling tick handling
remains associated with the scheduling context.
- Keep the task_tick_core() consumed-slice accounting issue discussed
during review separate from this series.
Link: https://lore.kernel.org/r/20260903041154.2479761-1-sh_def@163.com
Tested with CONFIG_SCHED_PROXY_EXEC=y and all four NUMA/cache
configuration combinations, including W=1 scheduler object builds and
a full bzImage/modules build. The first commit was also built
independently, and a CONFIG_NO_HZ_FULL=y kernel was built and booted.
A temporary QEMU proxy-mutex reproducer exercised both a normal FAIR
path and a FAIR execution task running on behalf of an SCHED_FIFO donor.
Using the same workload with QEMU 11.1.0 exposing two L3 domains, the
unpatched kernel did not invoke task_tick_numa() or task_tick_cache() for
the FAIR execution task during proxy execution. With this series, both
hooks were repeatedly observed with rq->curr while rq->donor remained the
RT scheduling context. NUMA task work was queued for the execution task;
cache work was also queued and its mm scan epoch advanced. Three proxy
episodes completed, and the normal FAIR path was observed with
rq->curr == rq->donor. The full-dynticks remote tick path was observed
invoking sched_tick_exec_ctx() on CPU 1. No warnings, BUGs, oopses or
panics were observed.
The instrumentation and reproducer were kept outside this series.
Hui Su (2):
sched/numa: Drive NUMA task tick from execution context
sched/cache: Drive cache task tick from execution context
kernel/sched/core.c | 15 +++++++++++++++
kernel/sched/fair.c | 18 +++++++++---------
kernel/sched/sched.h | 2 ++
3 files changed, 26 insertions(+), 9 deletions(-)
^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH v3 1/2] sched/numa: Drive NUMA task tick from execution context
2026-09-04 8:52 [PATCH v3 0/2] sched: Fix execution-context tick handling under proxy execution Hui Su
@ 2026-09-04 8:52 ` Hui Su
2026-09-04 15:52 ` Chen Yu
2026-09-04 17:17 ` Tim Chen
2026-09-04 8:52 ` [PATCH v3 2/2] sched/cache: Drive cache " Hui Su
1 sibling, 2 replies; 5+ messages in thread
From: Hui Su @ 2026-09-04 8:52 UTC (permalink / raw)
To: peterz, mingo, tim.c.chen, yu.c.chen, kprateek.nayak
Cc: juri.lelli, vincent.guittot, dietmar.eggemann, rostedt, bsegall,
mgorman, vschneid, connoro, jstultz, linux-kernel
Proxy execution separates the scheduling context in rq->donor from the
execution context in rq->curr. sched_tick() invokes task_tick() for the
donor's scheduling class.
task_tick_numa() operates on state associated with the task actually
executing, including its mm and NUMA work state. With proxy execution,
rq->donor provides the scheduling context while rq->curr identifies the
execution context.
Task-level execution runtime is likewise accounted to rq->curr, and
task_tick_numa() uses that runtime to drive periodic NUMA scanning.
Keeping task_tick_numa() under task_tick_fair() also means that it is not
invoked when a fair task executes on behalf of an RT or deadline donor.
Move NUMA tick handling into a scheduler helper for the execution
context, and invoke it from both sched_tick() and sched_tick_remote().
Fixes: 7de9d4f94638 ("sched: Start blocked_on chain processing in find_proxy_task()")
Suggested-by: K Prateek Nayak <kprateek.nayak@amd.com>
Suggested-by: Tim Chen <tim.c.chen@linux.intel.com>
Signed-off-by: Hui Su <sh_def@163.com>
---
kernel/sched/core.c | 14 ++++++++++++++
kernel/sched/fair.c | 7 ++-----
kernel/sched/sched.h | 1 +
3 files changed, 17 insertions(+), 5 deletions(-)
diff --git a/kernel/sched/core.c b/kernel/sched/core.c
index f78275192036..4db55e4ace9e 100644
--- a/kernel/sched/core.c
+++ b/kernel/sched/core.c
@@ -5762,6 +5762,17 @@ static int __init setup_resched_latency_warn_ms(char *str)
}
__setup("resched_latency_warn_ms=", setup_resched_latency_warn_ms);
+static void sched_tick_exec_ctx(struct rq *rq)
+{
+ struct task_struct *curr = rq->curr;
+
+ if (curr->sched_class != &fair_sched_class)
+ return;
+
+ if (static_branch_unlikely(&sched_numa_balancing))
+ task_tick_numa(rq, curr);
+}
+
/*
* This function gets called by the timer code, with HZ frequency.
* We call it with interrupts disabled.
@@ -5794,6 +5805,8 @@ void sched_tick(void)
resched_curr(rq);
donor->sched_class->task_tick(rq, donor, 0);
+ sched_tick_exec_ctx(rq);
+
if (sched_feat(LATENCY_WARN))
resched_latency = cpu_resched_latency(rq);
calc_global_load_tick(rq);
@@ -5890,6 +5903,7 @@ static void sched_tick_remote(struct work_struct *work)
WARN_ON_ONCE(delta > (u64)NSEC_PER_SEC * 30);
}
curr->sched_class->task_tick(rq, curr, 0);
+ sched_tick_exec_ctx(rq);
calc_load_nohz_remote(rq);
}
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 8dff37059faf..55f0460e4ae3 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -4425,7 +4425,7 @@ void init_numa_balancing(u64 clone_flags, struct task_struct *p)
/*
* Drive the periodic memory faults..
*/
-static void task_tick_numa(struct rq *rq, struct task_struct *curr)
+void task_tick_numa(struct rq *rq, struct task_struct *curr)
{
struct callback_head *work = &curr->numa_work;
u64 period, now;
@@ -4491,7 +4491,7 @@ static void update_scan_period(struct task_struct *p, int new_cpu)
#else /* !CONFIG_NUMA_BALANCING: */
-static void task_tick_numa(struct rq *rq, struct task_struct *curr)
+void task_tick_numa(struct rq *rq, struct task_struct *curr)
{
}
@@ -15042,9 +15042,6 @@ static void task_tick_fair(struct rq *rq, struct task_struct *curr, int queued)
if (queued)
return;
- if (static_branch_unlikely(&sched_numa_balancing))
- task_tick_numa(rq, curr);
-
task_tick_cache(rq, curr);
update_misfit_status(curr, rq);
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index e656c7059bf8..4d619f272b15 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -4152,6 +4152,7 @@ extern void sched_cache_active_set(void);
void sched_domains_free_llc_id(int cpu);
extern void init_sched_mm(struct task_struct *p);
+void task_tick_numa(struct rq *rq, struct task_struct *p);
extern u64 avg_vruntime(struct cfs_rq *cfs_rq);
extern int entity_eligible(struct cfs_rq *cfs_rq, struct sched_entity *se);
^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH v3 2/2] sched/cache: Drive cache task tick from execution context
2026-09-04 8:52 [PATCH v3 0/2] sched: Fix execution-context tick handling under proxy execution Hui Su
2026-09-04 8:52 ` [PATCH v3 1/2] sched/numa: Drive NUMA task tick from execution context Hui Su
@ 2026-09-04 8:52 ` Hui Su
1 sibling, 0 replies; 5+ messages in thread
From: Hui Su @ 2026-09-04 8:52 UTC (permalink / raw)
To: peterz, mingo, tim.c.chen, yu.c.chen, kprateek.nayak
Cc: juri.lelli, vincent.guittot, dietmar.eggemann, rostedt, bsegall,
mgorman, vschneid, connoro, jstultz, linux-kernel
Cache-aware scheduling accounts CPU runtime to the mm of the task
actually executing. update_se() passes the execution task to
account_mm_sched() for this purpose.
With proxy execution, however, sched_tick() invokes task_tick() for the
scheduling context in rq->donor. task_tick_cache() is currently called
from task_tick_fair(), so it is skipped when a fair task executes on
behalf of an RT or deadline donor.
In that case account_mm_sched() continues to advance runtime accounting
for rq->curr, while task_tick_cache() does not advance the corresponding
mm scan epoch. Once the epoch becomes stale, account_mm_sched() can
invalidate the mm's preferred LLC.
Move cache tick handling into sched_tick_exec_ctx(), alongside NUMA tick
handling, and run it when the execution context is a fair task. Use the
same helper from sched_tick() and sched_tick_remote() so both tick paths
handle the execution context consistently.
Keep the remaining task_tick_fair() bookkeeping with its task argument,
since misfit, overutilized, and core scheduling state belong to the
scheduling context.
Fixes: df0d98475954 ("sched/cache: Introduce infrastructure for cache-aware load balancing")
Suggested-by: Tim Chen <tim.c.chen@linux.intel.com>
Signed-off-by: Hui Su <sh_def@163.com>
---
kernel/sched/core.c | 1 +
kernel/sched/fair.c | 11 +++++++----
kernel/sched/sched.h | 1 +
3 files changed, 9 insertions(+), 4 deletions(-)
diff --git a/kernel/sched/core.c b/kernel/sched/core.c
index 4db55e4ace9e..7b8d4b06207d 100644
--- a/kernel/sched/core.c
+++ b/kernel/sched/core.c
@@ -5771,6 +5771,7 @@ static void sched_tick_exec_ctx(struct rq *rq)
if (static_branch_unlikely(&sched_numa_balancing))
task_tick_numa(rq, curr);
+ task_tick_cache(rq, curr);
}
/*
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 55f0460e4ae3..c519bb193850 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -1774,7 +1774,7 @@ void account_mm_sched(struct rq *rq, struct task_struct *p, s64 delta_exec)
}
}
-static void task_tick_cache(struct rq *rq, struct task_struct *p)
+void task_tick_cache(struct rq *rq, struct task_struct *p)
{
struct callback_head *work = &p->cache_work;
struct mm_struct *mm = p->mm;
@@ -1996,7 +1996,7 @@ static inline void account_mm_sched(struct rq *rq, struct task_struct *p,
void init_sched_mm(struct task_struct *p) { }
-static void task_tick_cache(struct rq *rq, struct task_struct *p) { }
+void task_tick_cache(struct rq *rq, struct task_struct *p) { }
static inline int get_pref_llc(struct task_struct *p,
struct mm_struct *mm)
@@ -15042,8 +15042,11 @@ static void task_tick_fair(struct rq *rq, struct task_struct *curr, int queued)
if (queued)
return;
- task_tick_cache(rq, curr);
-
+ /*
+ * Misfit, overutilized and core scheduling state belong to the
+ * scheduling context, and therefore stay with @curr rather than
+ * rq->curr. See sched_tick_exec_ctx() for execution-context work.
+ */
update_misfit_status(curr, rq);
check_update_overutilized_status(task_rq(curr));
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index 4d619f272b15..5d1f5ee47bf1 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -4153,6 +4153,7 @@ void sched_domains_free_llc_id(int cpu);
extern void init_sched_mm(struct task_struct *p);
void task_tick_numa(struct rq *rq, struct task_struct *p);
+void task_tick_cache(struct rq *rq, struct task_struct *p);
extern u64 avg_vruntime(struct cfs_rq *cfs_rq);
extern int entity_eligible(struct cfs_rq *cfs_rq, struct sched_entity *se);
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH v3 1/2] sched/numa: Drive NUMA task tick from execution context
2026-09-04 8:52 ` [PATCH v3 1/2] sched/numa: Drive NUMA task tick from execution context Hui Su
@ 2026-09-04 15:52 ` Chen Yu
2026-09-04 17:17 ` Tim Chen
1 sibling, 0 replies; 5+ messages in thread
From: Chen Yu @ 2026-09-04 15:52 UTC (permalink / raw)
To: Hui Su
Cc: peterz, mingo, tim.c.chen, kprateek.nayak, juri.lelli,
vincent.guittot, dietmar.eggemann, rostedt, bsegall, mgorman,
vschneid, connoro, jstultz, linux-kernel
On Fri, Sep 04, 2026 at 04:52:43PM +0800, Hui Su wrote:
> Proxy execution separates the scheduling context in rq->donor from the
> execution context in rq->curr. sched_tick() invokes task_tick() for the
> donor's scheduling class.
>
> task_tick_numa() operates on state associated with the task actually
> executing, including its mm and NUMA work state. With proxy execution,
> rq->donor provides the scheduling context while rq->curr identifies the
> execution context.
>
> Task-level execution runtime is likewise accounted to rq->curr, and
> task_tick_numa() uses that runtime to drive periodic NUMA scanning.
> Keeping task_tick_numa() under task_tick_fair() also means that it is not
> invoked when a fair task executes on behalf of an RT or deadline donor.
>
> Move NUMA tick handling into a scheduler helper for the execution
> context, and invoke it from both sched_tick() and sched_tick_remote().
>
Both patches look good to me, let me launch a test and verify it works
as expected and report back later.
thanks,
Chenyu
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH v3 1/2] sched/numa: Drive NUMA task tick from execution context
2026-09-04 8:52 ` [PATCH v3 1/2] sched/numa: Drive NUMA task tick from execution context Hui Su
2026-09-04 15:52 ` Chen Yu
@ 2026-09-04 17:17 ` Tim Chen
1 sibling, 0 replies; 5+ messages in thread
From: Tim Chen @ 2026-09-04 17:17 UTC (permalink / raw)
To: Hui Su, peterz, mingo, yu.c.chen, kprateek.nayak
Cc: juri.lelli, vincent.guittot, dietmar.eggemann, rostedt, bsegall,
mgorman, vschneid, connoro, jstultz, linux-kernel
On Fri, 2026-09-04 at 16:52 +0800, Hui Su wrote:
> Proxy execution separates the scheduling context in rq->donor from the
> execution context in rq->curr. sched_tick() invokes task_tick() for the
> donor's scheduling class.
>
> task_tick_numa() operates on state associated with the task actually
> executing, including its mm and NUMA work state. With proxy execution,
> rq->donor provides the scheduling context while rq->curr identifies the
> execution context.
>
> Task-level execution runtime is likewise accounted to rq->curr, and
> task_tick_numa() uses that runtime to drive periodic NUMA scanning.
> Keeping task_tick_numa() under task_tick_fair() also means that it is not
> invoked when a fair task executes on behalf of an RT or deadline donor.
>
> Move NUMA tick handling into a scheduler helper for the execution
> context, and invoke it from both sched_tick() and sched_tick_remote().
>
> Fixes: 7de9d4f94638 ("sched: Start blocked_on chain processing in find_proxy_task()")
> Suggested-by: K Prateek Nayak <kprateek.nayak@amd.com>
> Suggested-by: Tim Chen <tim.c.chen@linux.intel.com>
> Signed-off-by: Hui Su <sh_def@163.com>
> ---
> kernel/sched/core.c | 14 ++++++++++++++
> kernel/sched/fair.c | 7 ++-----
> kernel/sched/sched.h | 1 +
> 3 files changed, 17 insertions(+), 5 deletions(-)
>
> diff --git a/kernel/sched/core.c b/kernel/sched/core.c
> index f78275192036..4db55e4ace9e 100644
> --- a/kernel/sched/core.c
> +++ b/kernel/sched/core.c
> @@ -5762,6 +5762,17 @@ static int __init setup_resched_latency_warn_ms(char *str)
> }
> __setup("resched_latency_warn_ms=", setup_resched_latency_warn_ms);
>
> +static void sched_tick_exec_ctx(struct rq *rq)
Just a minor nit. We could consider putting sched_tick_exec_ctx()
in fair.c and export it instead.
That allows task_tick_numa() and task_tick_cache() declaration to
remain static.
No big deal either way.
Otherwise the two patches in the series look good to me.
Reviewed-by: Tim Chen <tim.c.chen@linux.intel.com>
Tim
> +{
> + struct task_struct *curr = rq->curr;
> +
> + if (curr->sched_class != &fair_sched_class)
> + return;
> +
> + if (static_branch_unlikely(&sched_numa_balancing))
> + task_tick_numa(rq, curr);
> +}
> +
> /*
> * This function gets called by the timer code, with HZ frequency.
> * We call it with interrupts disabled.
> @@ -5794,6 +5805,8 @@ void sched_tick(void)
> resched_curr(rq);
>
> donor->sched_class->task_tick(rq, donor, 0);
> + sched_tick_exec_ctx(rq);
> +
> if (sched_feat(LATENCY_WARN))
> resched_latency = cpu_resched_latency(rq);
> calc_global_load_tick(rq);
> @@ -5890,6 +5903,7 @@ static void sched_tick_remote(struct work_struct *work)
> WARN_ON_ONCE(delta > (u64)NSEC_PER_SEC * 30);
> }
> curr->sched_class->task_tick(rq, curr, 0);
> + sched_tick_exec_ctx(rq);
>
> calc_load_nohz_remote(rq);
> }
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index 8dff37059faf..55f0460e4ae3 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -4425,7 +4425,7 @@ void init_numa_balancing(u64 clone_flags, struct task_struct *p)
> /*
> * Drive the periodic memory faults..
> */
> -static void task_tick_numa(struct rq *rq, struct task_struct *curr)
> +void task_tick_numa(struct rq *rq, struct task_struct *curr)
> {
> struct callback_head *work = &curr->numa_work;
> u64 period, now;
> @@ -4491,7 +4491,7 @@ static void update_scan_period(struct task_struct *p, int new_cpu)
>
> #else /* !CONFIG_NUMA_BALANCING: */
>
> -static void task_tick_numa(struct rq *rq, struct task_struct *curr)
> +void task_tick_numa(struct rq *rq, struct task_struct *curr)
> {
> }
>
> @@ -15042,9 +15042,6 @@ static void task_tick_fair(struct rq *rq, struct task_struct *curr, int queued)
> if (queued)
> return;
>
> - if (static_branch_unlikely(&sched_numa_balancing))
> - task_tick_numa(rq, curr);
> -
> task_tick_cache(rq, curr);
>
> update_misfit_status(curr, rq);
> diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
> index e656c7059bf8..4d619f272b15 100644
> --- a/kernel/sched/sched.h
> +++ b/kernel/sched/sched.h
> @@ -4152,6 +4152,7 @@ extern void sched_cache_active_set(void);
> void sched_domains_free_llc_id(int cpu);
>
> extern void init_sched_mm(struct task_struct *p);
> +void task_tick_numa(struct rq *rq, struct task_struct *p);
>
> extern u64 avg_vruntime(struct cfs_rq *cfs_rq);
> extern int entity_eligible(struct cfs_rq *cfs_rq, struct sched_entity *se);
>
^ permalink raw reply [flat|nested] 5+ messages in thread
end of thread, other threads:[~2026-09-04 17:17 UTC | newest]
Thread overview: 5+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-04 8:52 [PATCH v3 0/2] sched: Fix execution-context tick handling under proxy execution Hui Su
2026-09-04 8:52 ` [PATCH v3 1/2] sched/numa: Drive NUMA task tick from execution context Hui Su
2026-09-04 15:52 ` Chen Yu
2026-09-04 17:17 ` Tim Chen
2026-09-04 8:52 ` [PATCH v3 2/2] sched/cache: Drive cache " Hui Su
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®