* Re: [RFC PATCH 1/1] sched/eevdf: Skip eligibility check for current entity during wakeup preemption
2024-03-25 6:02 ` [RFC PATCH 1/1] sched/eevdf: Skip eligibility check for current entity during wakeup preemption K Prateek Nayak
@ 2024-03-25 15:13 ` Youssef Esmat
2024-03-26 3:06 ` K Prateek Nayak
2024-04-24 15:07 ` Peter Zijlstra
2026-07-12 11:31 ` [PATCH] [Question] sched/fair: Task starvation with RUN_TO_PARITY_WAKEUP under group topologies Chen Jinghuang
2 siblings, 1 reply; 11+ messages in thread
From: Youssef Esmat @ 2024-03-25 15:13 UTC (permalink / raw)
To: K Prateek Nayak
Cc: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot,
Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Daniel Bristot de Oliveira, Valentin Schneider, linux-kernel,
Tobias Huschle, Luis Machado, Chen Yu, Abel Wu, Tianchen Ding,
Xuewen Yan, Gautham R. Shenoy
On Mon, Mar 25, 2024 at 1:03 AM K Prateek Nayak <kprateek.nayak@amd.com> wrote:
>
> With the curr entity's eligibility check, a wakeup preemption is very
> likely when an entity with positive lag joins the runqueue pushing the
> avg_vruntime of the runqueue backwards, making the vruntime of the
> current entity ineligible. This leads to aggressive wakeup preemption
> which was previously guarded by wakeup_granularity_ns in legacy CFS.
> Below figure depicts one such aggressive preemption scenario with EEVDF
> in DeathStarBench [1]:
>
> deadline for Nginx
> |
> +-------+ | |
> /-- | Nginx | -|------------------> |
> | +-------+ | |
> | |
> | -----------|-------------------------------> vruntime timeline
> | \--> rq->avg_vruntime
> |
> | wakes service on the same runqueue since system is busy
> |
> | +---------+|
> \-->| Service || (service has +ve lag pushes avg_vruntime backwards)
> +---------+|
> | |
> wakeup | +--|-----+ |
> preempts \---->| N|ginx | --------------------> | {deadline for Nginx}
> +--|-----+ |
> (Nginx ineligible)
> -----------|-------------------------------> vruntime timeline
> \--> rq->avg_vruntime
>
> When NGINX server is involuntarily switched out, it cannot accept any
> incoming request, leading to longer turn around time for the clients and
> thus loss in DeathStarBench throughput.
>
> ==================================================================
> Test : DeathStarBench
> Units : Normalized latency
> Interpretation: Lower is better
> Statistic : Mean
> ==================================================================
> tip 1.00
> eevdf 1.14 (+14.61%)
>
> For current running task, skip eligibility check in pick_eevdf() if it
> has not exhausted the slice promised to it during selection despite the
> situation having changed since. The behavior is guarded by
> RUN_TO_PARITY_WAKEUP sched_feat to simplify testing. With
> RUN_TO_PARITY_WAKEUP enabled, performance loss seen with DeathStarBench
> since the merge of EEVDF disappears. Following are the results from
> testing on a Dual Socket 3rd Generation EPYC server (2 x 64C/128T):
>
> ==================================================================
> Test : DeathStarBench
> Units : Normalized throughput
> Interpretation: Higher is better
> Statistic : Mean
> ==================================================================
> Pinning scaling tip run-to-parity-wakeup(pct imp)
> 1CCD 1 1.00 1.16 (%diff: 16%)
> 2CCD 2 1.00 1.03 (%diff: 3%)
> 4CCD 4 1.00 1.12 (%diff: 12%)
> 8CCD 8 1.00 1.05 (%diff: 6%)
>
> With spec_rstack_overflow=off, the DeathStarBench performance with the
> proposed solution is same as the performance on v6.5 release before
> EEVDF was merged.
Thanks for sharing this Prateek.
We actually noticed we could also gain performance by disabling
eligibility checks (but disable it on all paths).
The following are a few threads we had on the topic:
Discussion around eligibility:
https://lore.kernel.org/lkml/CA+q576MS0-MV1Oy-eecvmYpvNT3tqxD8syzrpxQ-Zk310hvRbw@mail.gmail.com/
Some of our results:
https://lore.kernel.org/lkml/CA+q576Mov1jpdfZhPBoy_hiVh3xSWuJjXdP3nS4zfpqfOXtq7Q@mail.gmail.com/
Sched feature to disable eligibility:
https://lore.kernel.org/lkml/20231013030213.2472697-1-youssefesmat@chromium.org/
>
> This may lead to newly waking task waiting longer for its turn on the
> CPU, however, testing on the same system did not reveal any consistent
> regressions with the standard benchmarks.
>
> Link: https://github.com/delimitrou/DeathStarBench/ [1]
> Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com>
> ---
> kernel/sched/fair.c | 24 ++++++++++++++++++++----
> kernel/sched/features.h | 1 +
> 2 files changed, 21 insertions(+), 4 deletions(-)
>
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index 6a16129f9a5c..a9b145a4eab0 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -875,7 +875,7 @@ struct sched_entity *__pick_first_entity(struct cfs_rq *cfs_rq)
> *
> * Which allows tree pruning through eligibility.
> */
> -static struct sched_entity *pick_eevdf(struct cfs_rq *cfs_rq)
> +static struct sched_entity *pick_eevdf(struct cfs_rq *cfs_rq, bool wakeup_preempt)
> {
> struct rb_node *node = cfs_rq->tasks_timeline.rb_root.rb_node;
> struct sched_entity *se = __pick_first_entity(cfs_rq);
> @@ -889,7 +889,23 @@ static struct sched_entity *pick_eevdf(struct cfs_rq *cfs_rq)
> if (cfs_rq->nr_running == 1)
> return curr && curr->on_rq ? curr : se;
>
> - if (curr && (!curr->on_rq || !entity_eligible(cfs_rq, curr)))
> + if (curr && !curr->on_rq)
> + curr = NULL;
> +
> + /*
> + * When an entity with positive lag wakes up, it pushes the
> + * avg_vruntime of the runqueue backwards. This may causes the
> + * current entity to be ineligible soon into its run leading to
> + * wakeup preemption.
> + *
> + * To prevent such aggressive preemption of the current running
> + * entity during task wakeups, skip the eligibility check if the
> + * slice promised to the entity since its selection has not yet
> + * elapsed.
> + */
> + if (curr &&
> + !(sched_feat(RUN_TO_PARITY_WAKEUP) && wakeup_preempt && curr->vlag == curr->deadline) &&
> + !entity_eligible(cfs_rq, curr))
> curr = NULL;
>
> /*
> @@ -5460,7 +5476,7 @@ pick_next_entity(struct cfs_rq *cfs_rq)
> cfs_rq->next && entity_eligible(cfs_rq, cfs_rq->next))
> return cfs_rq->next;
>
> - return pick_eevdf(cfs_rq);
> + return pick_eevdf(cfs_rq, false);
> }
>
> static bool check_cfs_rq_runtime(struct cfs_rq *cfs_rq);
> @@ -8340,7 +8356,7 @@ static void check_preempt_wakeup_fair(struct rq *rq, struct task_struct *p, int
> /*
> * XXX pick_eevdf(cfs_rq) != se ?
> */
> - if (pick_eevdf(cfs_rq) == pse)
> + if (pick_eevdf(cfs_rq, true) == pse)
> goto preempt;
>
> return;
> diff --git a/kernel/sched/features.h b/kernel/sched/features.h
> index 143f55df890b..027bab5b4031 100644
> --- a/kernel/sched/features.h
> +++ b/kernel/sched/features.h
> @@ -7,6 +7,7 @@
> SCHED_FEAT(PLACE_LAG, true)
> SCHED_FEAT(PLACE_DEADLINE_INITIAL, true)
> SCHED_FEAT(RUN_TO_PARITY, true)
> +SCHED_FEAT(RUN_TO_PARITY_WAKEUP, true)
>
> /*
> * Prefer to schedule the task we woke last (assuming it failed
> --
> 2.34.1
>
^ permalink raw reply [flat|nested] 11+ messages in thread* Re: [RFC PATCH 1/1] sched/eevdf: Skip eligibility check for current entity during wakeup preemption
2024-03-25 15:13 ` Youssef Esmat
@ 2024-03-26 3:06 ` K Prateek Nayak
2024-04-17 6:08 ` K Prateek Nayak
0 siblings, 1 reply; 11+ messages in thread
From: K Prateek Nayak @ 2024-03-26 3:06 UTC (permalink / raw)
To: Youssef Esmat
Cc: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot,
Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Daniel Bristot de Oliveira, Valentin Schneider, linux-kernel,
Tobias Huschle, Luis Machado, Chen Yu, Abel Wu, Tianchen Ding,
Xuewen Yan, Gautham R. Shenoy
Hello Youssef,
On 3/25/2024 8:43 PM, Youssef Esmat wrote:
> On Mon, Mar 25, 2024 at 1:03 AM K Prateek Nayak <kprateek.nayak@amd.com> wrote:
>>
>> With the curr entity's eligibility check, a wakeup preemption is very
>> likely when an entity with positive lag joins the runqueue pushing the
>> avg_vruntime of the runqueue backwards, making the vruntime of the
>> current entity ineligible. This leads to aggressive wakeup preemption
>> which was previously guarded by wakeup_granularity_ns in legacy CFS.
>> Below figure depicts one such aggressive preemption scenario with EEVDF
>> in DeathStarBench [1]:
>>
>> deadline for Nginx
>> |
>> +-------+ | |
>> /-- | Nginx | -|------------------> |
>> | +-------+ | |
>> | |
>> | -----------|-------------------------------> vruntime timeline
>> | \--> rq->avg_vruntime
>> |
>> | wakes service on the same runqueue since system is busy
>> |
>> | +---------+|
>> \-->| Service || (service has +ve lag pushes avg_vruntime backwards)
>> +---------+|
>> | |
>> wakeup | +--|-----+ |
>> preempts \---->| N|ginx | --------------------> | {deadline for Nginx}
>> +--|-----+ |
>> (Nginx ineligible)
>> -----------|-------------------------------> vruntime timeline
>> \--> rq->avg_vruntime
>>
>> When NGINX server is involuntarily switched out, it cannot accept any
>> incoming request, leading to longer turn around time for the clients and
>> thus loss in DeathStarBench throughput.
>>
>> ==================================================================
>> Test : DeathStarBench
>> Units : Normalized latency
>> Interpretation: Lower is better
>> Statistic : Mean
>> ==================================================================
>> tip 1.00
>> eevdf 1.14 (+14.61%)
>>
>> For current running task, skip eligibility check in pick_eevdf() if it
>> has not exhausted the slice promised to it during selection despite the
>> situation having changed since. The behavior is guarded by
>> RUN_TO_PARITY_WAKEUP sched_feat to simplify testing. With
>> RUN_TO_PARITY_WAKEUP enabled, performance loss seen with DeathStarBench
>> since the merge of EEVDF disappears. Following are the results from
>> testing on a Dual Socket 3rd Generation EPYC server (2 x 64C/128T):
>>
>> ==================================================================
>> Test : DeathStarBench
>> Units : Normalized throughput
>> Interpretation: Higher is better
>> Statistic : Mean
>> ==================================================================
>> Pinning scaling tip run-to-parity-wakeup(pct imp)
>> 1CCD 1 1.00 1.16 (%diff: 16%)
>> 2CCD 2 1.00 1.03 (%diff: 3%)
>> 4CCD 4 1.00 1.12 (%diff: 12%)
>> 8CCD 8 1.00 1.05 (%diff: 6%)
>>
>> With spec_rstack_overflow=off, the DeathStarBench performance with the
>> proposed solution is same as the performance on v6.5 release before
>> EEVDF was merged.
>
> Thanks for sharing this Prateek.
> We actually noticed we could also gain performance by disabling
> eligibility checks (but disable it on all paths).
> The following are a few threads we had on the topic:
>
> Discussion around eligibility:
> https://lore.kernel.org/lkml/CA+q576MS0-MV1Oy-eecvmYpvNT3tqxD8syzrpxQ-Zk310hvRbw@mail.gmail.com/
> Some of our results:
> https://lore.kernel.org/lkml/CA+q576Mov1jpdfZhPBoy_hiVh3xSWuJjXdP3nS4zfpqfOXtq7Q@mail.gmail.com/
> Sched feature to disable eligibility:
> https://lore.kernel.org/lkml/20231013030213.2472697-1-youssefesmat@chromium.org/
>
Thank you for pointing me to the discussions. I'll give this a spin on
my machine and report back what I see. Hope some of it will help during
the OSPM discussion :)
>>
>> This may lead to newly waking task waiting longer for its turn on the
>> CPU, however, testing on the same system did not reveal any consistent
>> regressions with the standard benchmarks.
>>
>> Link: https://github.com/delimitrou/DeathStarBench/ [1]
>> Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com>
>> ---
>> [..snip..]
>>
--
Thanks and Regards,
Prateek
^ permalink raw reply [flat|nested] 11+ messages in thread* Re: [RFC PATCH 1/1] sched/eevdf: Skip eligibility check for current entity during wakeup preemption
2024-03-26 3:06 ` K Prateek Nayak
@ 2024-04-17 6:08 ` K Prateek Nayak
0 siblings, 0 replies; 11+ messages in thread
From: K Prateek Nayak @ 2024-04-17 6:08 UTC (permalink / raw)
To: Youssef Esmat
Cc: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot,
Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Daniel Bristot de Oliveira, Valentin Schneider, linux-kernel,
Tobias Huschle, Luis Machado, Chen Yu, Abel Wu, Tianchen Ding,
Xuewen Yan, Gautham R. Shenoy
Hello Youssef,
On 3/26/2024 8:36 AM, K Prateek Nayak wrote:
>> [..snip..]
>>
>> Thanks for sharing this Prateek.
>> We actually noticed we could also gain performance by disabling
>> eligibility checks (but disable it on all paths).
>> The following are a few threads we had on the topic:
>>
>> Discussion around eligibility:
>> https://lore.kernel.org/lkml/CA+q576MS0-MV1Oy-eecvmYpvNT3tqxD8syzrpxQ-Zk310hvRbw@mail.gmail.com/
>> Some of our results:
>> https://lore.kernel.org/lkml/CA+q576Mov1jpdfZhPBoy_hiVh3xSWuJjXdP3nS4zfpqfOXtq7Q@mail.gmail.com/
>> Sched feature to disable eligibility:
>> https://lore.kernel.org/lkml/20231013030213.2472697-1-youssefesmat@chromium.org/
>>
>
> Thank you for pointing me to the discussions. I'll give this a spin on
> my machine and report back what I see. Hope some of it will help during
> the OSPM discussion :)
Sorry about the delay but on a positive note, I do not see any
concerning regressions after dropping the eligibility criteria. I'll
leave the full results from my testing below.
o System Details
- 3rd Generation EPYC System
- 2 x 64C/128T
- NPS1 mode
o Kernels
tip: tip:sched/core at commit 4475cd8bfd9b
("sched/balancing: Simplify the sg_status
bitmask and use separate ->overloaded and
->overutilized flags")
eie: (everyone is eligible)
tip + vruntime_eligible() and entity_eligible()
always returns true.
o Results
==================================================================
Test : hackbench
Units : Normalized time in seconds
Interpretation: Lower is better
Statistic : AMean
==================================================================
Case: tip[pct imp](CV) eie[pct imp](CV)
1-groups 1.00 [ -0.00]( 1.94) 0.95 [ 5.11]( 2.56)
2-groups 1.00 [ -0.00]( 2.41) 0.97 [ 2.80]( 1.52)
4-groups 1.00 [ -0.00]( 1.16) 0.95 [ 5.01]( 1.04)
8-groups 1.00 [ -0.00]( 1.72) 0.96 [ 4.37]( 1.01)
16-groups 1.00 [ -0.00]( 2.16) 0.94 [ 5.88]( 2.30)
==================================================================
Test : tbench
Units : Normalized throughput
Interpretation: Higher is better
Statistic : AMean
==================================================================
Clients: tip[pct imp](CV) eie[pct imp](CV)
1 1.00 [ 0.00]( 0.69) 1.00 [ 0.05]( 0.61)
2 1.00 [ 0.00]( 0.25) 1.00 [ 0.06]( 0.51)
4 1.00 [ 0.00]( 1.04) 0.98 [ -1.69]( 1.21)
8 1.00 [ 0.00]( 0.72) 1.00 [ -0.13]( 0.56)
16 1.00 [ 0.00]( 2.40) 1.00 [ 0.43]( 0.63)
32 1.00 [ 0.00]( 0.62) 0.98 [ -1.80]( 2.18)
64 1.00 [ 0.00]( 1.19) 0.98 [ -2.13]( 1.26)
128 1.00 [ 0.00]( 0.91) 1.00 [ 0.37]( 0.50)
256 1.00 [ 0.00]( 0.52) 1.00 [ -0.11]( 0.21)
512 1.00 [ 0.00]( 0.36) 1.02 [ 1.54]( 0.58)
1024 1.00 [ 0.00]( 0.26) 1.01 [ 1.21]( 0.41)
==================================================================
Test : stream-10
Units : Normalized Bandwidth, MB/s
Interpretation: Higher is better
Statistic : HMean
==================================================================
Test: tip[pct imp](CV) eie[pct imp](CV)
Copy 1.00 [ 0.00]( 5.01) 1.01 [ 1.27]( 4.63)
Scale 1.00 [ 0.00]( 6.93) 1.03 [ 2.66]( 5.20)
Add 1.00 [ 0.00]( 5.94) 1.03 [ 3.41]( 4.99)
Triad 1.00 [ 0.00]( 6.40) 0.95 [ -4.69]( 8.29)
==================================================================
Test : stream-100
Units : Normalized Bandwidth, MB/s
Interpretation: Higher is better
Statistic : HMean
==================================================================
Test: tip[pct imp](CV) eie[pct imp](CV)
Copy 1.00 [ 0.00]( 2.84) 1.00 [ -0.37]( 2.44)
Scale 1.00 [ 0.00]( 5.26) 1.00 [ 0.21]( 3.88)
Add 1.00 [ 0.00]( 4.98) 1.00 [ 0.11]( 1.15)
Triad 1.00 [ 0.00]( 1.60) 0.96 [ -3.72]( 5.26)
==================================================================
Test : netperf
Units : Normalized Througput
Interpretation: Higher is better
Statistic : AMean
==================================================================
Clients: tip[pct imp](CV) eie[pct imp](CV)
1-clients 1.00 [ 0.00]( 0.90) 1.00 [ -0.09]( 0.16)
2-clients 1.00 [ 0.00]( 0.77) 0.99 [ -0.89]( 0.97)
4-clients 1.00 [ 0.00]( 0.63) 0.99 [ -1.03]( 1.53)
8-clients 1.00 [ 0.00]( 0.52) 0.99 [ -0.86]( 1.66)
16-clients 1.00 [ 0.00]( 0.43) 0.99 [ -0.91]( 0.79)
32-clients 1.00 [ 0.00]( 0.88) 0.98 [ -2.37]( 1.42)
64-clients 1.00 [ 0.00]( 1.63) 0.96 [ -4.07]( 0.91) *
128-clients 1.00 [ 0.00]( 0.94) 1.00 [ -0.30]( 0.94)
256-clients 1.00 [ 0.00]( 5.08) 0.95 [ -4.95]( 3.36)
512-clients 1.00 [ 0.00](51.89) 0.99 [ -0.93](51.00)
* This seems to be the only point of regression with low CV. I'll
rerun this and report back if I see a consistent dip but for the
time being I'm not worried.
==================================================================
Test : schbench
Units : Normalized 99th percentile latency in us
Interpretation: Lower is better
Statistic : Median
==================================================================
#workers: tip[pct imp](CV) eie[pct imp](CV)
1 1.00 [ -0.00](30.01) 0.97 [ 3.12](14.32)
2 1.00 [ -0.00](26.14) 1.23 [-22.58](13.48)
4 1.00 [ -0.00](13.22) 1.00 [ -0.00]( 6.04)
8 1.00 [ -0.00]( 6.23) 1.00 [ -0.00](13.09)
16 1.00 [ -0.00]( 3.49) 1.02 [ -1.69]( 3.43)
32 1.00 [ -0.00]( 2.20) 0.98 [ 2.13]( 2.47)
64 1.00 [ -0.00]( 7.17) 0.88 [ 12.50]( 3.18)
128 1.00 [ -0.00]( 2.79) 1.02 [ -2.46]( 8.29)
256 1.00 [ -0.00](13.02) 1.01 [ -1.34](37.58)
512 1.00 [ -0.00]( 4.27) 0.79 [ 21.49]( 2.41)
==================================================================
Test : DeathStarBench
Units : Normalized throughput
Interpretation: Higher is better
Statistic : Mean
==================================================================
Pinning scaling tip eie (pct imp)
1CCD 1 1.00 1.15 (%diff: 15.68%)
2CCD 2 1.00 0.99 (%diff: -1.12%)
4CCD 4 1.00 1.11 (%diff: 11.65%)
8CCD 8 1.00 1.05 (%diff: 4.98%)
--
>
> [..snip..]
>
I'll try to get data from more workloads, will update the thread with
when it arrives.
--
Thanks and Regards,
Prateek
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [RFC PATCH 1/1] sched/eevdf: Skip eligibility check for current entity during wakeup preemption
2024-03-25 6:02 ` [RFC PATCH 1/1] sched/eevdf: Skip eligibility check for current entity during wakeup preemption K Prateek Nayak
2024-03-25 15:13 ` Youssef Esmat
@ 2024-04-24 15:07 ` Peter Zijlstra
2024-04-25 3:34 ` K Prateek Nayak
2026-07-12 11:31 ` [PATCH] [Question] sched/fair: Task starvation with RUN_TO_PARITY_WAKEUP under group topologies Chen Jinghuang
2 siblings, 1 reply; 11+ messages in thread
From: Peter Zijlstra @ 2024-04-24 15:07 UTC (permalink / raw)
To: K Prateek Nayak
Cc: Ingo Molnar, Juri Lelli, Vincent Guittot, Dietmar Eggemann,
Steven Rostedt, Ben Segall, Mel Gorman,
Daniel Bristot de Oliveira, Valentin Schneider, linux-kernel,
Tobias Huschle, Luis Machado, Chen Yu, Abel Wu, Tianchen Ding,
Youssef Esmat, Xuewen Yan, Gautham R. Shenoy
On Mon, Mar 25, 2024 at 11:32:26AM +0530, K Prateek Nayak wrote:
> With the curr entity's eligibility check, a wakeup preemption is very
> likely when an entity with positive lag joins the runqueue pushing the
> avg_vruntime of the runqueue backwards, making the vruntime of the
> current entity ineligible. This leads to aggressive wakeup preemption
> which was previously guarded by wakeup_granularity_ns in legacy CFS.
> Below figure depicts one such aggressive preemption scenario with EEVDF
> in DeathStarBench [1]:
>
> deadline for Nginx
> |
> +-------+ | |
> /-- | Nginx | -|------------------> |
> | +-------+ | |
> | |
> | -----------|-------------------------------> vruntime timeline
> | \--> rq->avg_vruntime
> |
> | wakes service on the same runqueue since system is busy
> |
> | +---------+|
> \-->| Service || (service has +ve lag pushes avg_vruntime backwards)
> +---------+|
> | |
> wakeup | +--|-----+ |
> preempts \---->| N|ginx | --------------------> | {deadline for Nginx}
> +--|-----+ |
> (Nginx ineligible)
> -----------|-------------------------------> vruntime timeline
> \--> rq->avg_vruntime
This graph is really hard to interpret. If you want to illustrate
avg_vruntime moves back, you should not align it. That's really
disorienting.
In both (upper and lower) nginx has the same vruntime thus *that* should
be aligned. The lower will have service placed left and with that
avg_vruntime also moves left, rendering nginx in-eligible.
> Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com>
> ---
> kernel/sched/fair.c | 24 ++++++++++++++++++++----
> kernel/sched/features.h | 1 +
> 2 files changed, 21 insertions(+), 4 deletions(-)
>
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index 6a16129f9a5c..a9b145a4eab0 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -875,7 +875,7 @@ struct sched_entity *__pick_first_entity(struct cfs_rq *cfs_rq)
> *
> * Which allows tree pruning through eligibility.
> */
> -static struct sched_entity *pick_eevdf(struct cfs_rq *cfs_rq)
> +static struct sched_entity *pick_eevdf(struct cfs_rq *cfs_rq, bool wakeup_preempt)
> {
> struct rb_node *node = cfs_rq->tasks_timeline.rb_root.rb_node;
> struct sched_entity *se = __pick_first_entity(cfs_rq);
> @@ -889,7 +889,23 @@ static struct sched_entity *pick_eevdf(struct cfs_rq *cfs_rq)
> if (cfs_rq->nr_running == 1)
> return curr && curr->on_rq ? curr : se;
>
> - if (curr && (!curr->on_rq || !entity_eligible(cfs_rq, curr)))
> + if (curr && !curr->on_rq)
> + curr = NULL;
> +
> + /*
> + * When an entity with positive lag wakes up, it pushes the
> + * avg_vruntime of the runqueue backwards. This may causes the
> + * current entity to be ineligible soon into its run leading to
> + * wakeup preemption.
> + *
> + * To prevent such aggressive preemption of the current running
> + * entity during task wakeups, skip the eligibility check if the
> + * slice promised to the entity since its selection has not yet
> + * elapsed.
> + */
> + if (curr &&
> + !(sched_feat(RUN_TO_PARITY_WAKEUP) && wakeup_preempt && curr->vlag == curr->deadline) &&
> + !entity_eligible(cfs_rq, curr))
> curr = NULL;
>
> /*
So I see what you want to do, but this is highly unreadable.
I'll try something like the below on top of queue/sched/eevdf, but I
should probably first look at fixing those reported crashes on that tree
:/
---
kernel/sched/fair.c | 60 ++++++++++++++++++++++++++++++++++---------------
kernel/sched/features.h | 11 +++++----
2 files changed, 49 insertions(+), 22 deletions(-)
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 8a9206d8532f..23977ed1cb2c 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -855,6 +855,39 @@ struct sched_entity *__pick_first_entity(struct cfs_rq *cfs_rq)
return __node_2_se(left);
}
+static inline bool pick_curr(struct cfs_rq *cfs_rq,
+ struct sched_entity *curr, struct sched_entity *wakee)
+{
+ /*
+ * Nothing to preserve...
+ */
+ if (!curr || !sched_feat(RESPECT_SLICE))
+ return false;
+
+ /*
+ * Allow preemption at the 0-lag point -- even if not all of the slice
+ * is consumed. Note: placement of positive lag can push V left and render
+ * @curr instantly ineligible irrespective the time on-cpu.
+ */
+ if (sched_feat(RUN_TO_PARITY) && !entity_eligible(cfs_rq, curr))
+ return false;
+
+ /*
+ * Don't preserve @curr when the @wakee has a shorter slice and earlier
+ * deadline. IOW, explicitly allow preemption.
+ */
+ if (sched_feat(PREEMPT_SHORT) && wakee &&
+ wakee->slice < curr->slice &&
+ (s64)(wakee->deadline - curr->deadline) < 0)
+ return false;
+
+ /*
+ * Preserve @curr to allow it to finish its first slice.
+ * See the HACK in set_next_entity().
+ */
+ return curr->vlag == curr->deadline;
+}
+
/*
* Earliest Eligible Virtual Deadline First
*
@@ -874,28 +907,27 @@ struct sched_entity *__pick_first_entity(struct cfs_rq *cfs_rq)
*
* Which allows tree pruning through eligibility.
*/
-static struct sched_entity *pick_eevdf(struct cfs_rq *cfs_rq)
+static struct sched_entity *pick_eevdf(struct cfs_rq *cfs_rq, struct sched_entity *wakee)
{
struct rb_node *node = cfs_rq->tasks_timeline.rb_root.rb_node;
struct sched_entity *se = __pick_first_entity(cfs_rq);
struct sched_entity *curr = cfs_rq->curr;
struct sched_entity *best = NULL;
+ if (curr && !curr->on_rq)
+ curr = NULL;
+
/*
* We can safely skip eligibility check if there is only one entity
* in this cfs_rq, saving some cycles.
*/
if (cfs_rq->nr_running == 1)
- return curr && curr->on_rq ? curr : se;
-
- if (curr && (!curr->on_rq || !entity_eligible(cfs_rq, curr)))
- curr = NULL;
+ return curr ?: se;
/*
- * Once selected, run a task until it either becomes non-eligible or
- * until it gets a new slice. See the HACK in set_next_entity().
+ * Preserve @curr to let it finish its slice.
*/
- if (sched_feat(RUN_TO_PARITY) && curr && curr->vlag == curr->deadline)
+ if (pick_curr(cfs_rq, curr, wakee))
return curr;
/* Pick the leftmost entity if it's eligible */
@@ -5507,7 +5539,7 @@ pick_next_entity(struct rq *rq, struct cfs_rq *cfs_rq)
return cfs_rq->next;
}
- struct sched_entity *se = pick_eevdf(cfs_rq);
+ struct sched_entity *se = pick_eevdf(cfs_rq, NULL);
if (se->sched_delayed) {
dequeue_entities(rq, se, DEQUEUE_SLEEP | DEQUEUE_DELAYED);
SCHED_WARN_ON(se->sched_delayed);
@@ -8548,15 +8580,7 @@ static void check_preempt_wakeup_fair(struct rq *rq, struct task_struct *p, int
cfs_rq = cfs_rq_of(se);
update_curr(cfs_rq);
- if (sched_feat(PREEMPT_SHORT) && pse->slice < se->slice &&
- entity_eligible(cfs_rq, pse) &&
- (s64)(pse->deadline - se->deadline) < 0 &&
- se->vlag == se->deadline) {
- /* negate RUN_TO_PARITY */
- se->vlag = se->deadline - 1;
- }
-
- if (pick_eevdf(cfs_rq) == pse)
+ if (pick_eevdf(cfs_rq, pse) == pse)
goto preempt;
return;
diff --git a/kernel/sched/features.h b/kernel/sched/features.h
index 64ce99cf04ec..2285dc30294c 100644
--- a/kernel/sched/features.h
+++ b/kernel/sched/features.h
@@ -10,12 +10,15 @@ SCHED_FEAT(PLACE_LAG, true)
*/
SCHED_FEAT(PLACE_DEADLINE_INITIAL, true)
/*
- * Inhibit (wakeup) preemption until the current task has either matched the
- * 0-lag point or until is has exhausted it's slice.
+ * Inhibit (wakeup) preemption until the current task has exhausted its slice.
*/
-SCHED_FEAT(RUN_TO_PARITY, true)
+SCHED_FEAT(RESPECT_SLICE, true)
/*
- * Allow tasks with a shorter slice to disregard RUN_TO_PARITY
+ * Relax RESPECT_SLICE to allow preemption once current has reached 0-lag.
+ */
+SCHED_FEAT(RUN_TO_PARITY, false)
+/*
+ * Allow tasks with a shorter slice to disregard RESPECT_SLICE
*/
SCHED_FEAT(PREEMPT_SHORT, true)
^ permalink raw reply [flat|nested] 11+ messages in thread* Re: [RFC PATCH 1/1] sched/eevdf: Skip eligibility check for current entity during wakeup preemption
2024-04-24 15:07 ` Peter Zijlstra
@ 2024-04-25 3:34 ` K Prateek Nayak
0 siblings, 0 replies; 11+ messages in thread
From: K Prateek Nayak @ 2024-04-25 3:34 UTC (permalink / raw)
To: Peter Zijlstra
Cc: Ingo Molnar, Juri Lelli, Vincent Guittot, Dietmar Eggemann,
Steven Rostedt, Ben Segall, Mel Gorman,
Daniel Bristot de Oliveira, Valentin Schneider, linux-kernel,
Tobias Huschle, Luis Machado, Chen Yu, Abel Wu, Tianchen Ding,
Youssef Esmat, Xuewen Yan, Gautham R. Shenoy
Hello Peter,
On 4/24/2024 8:37 PM, Peter Zijlstra wrote:
> On Mon, Mar 25, 2024 at 11:32:26AM +0530, K Prateek Nayak wrote:
>> With the curr entity's eligibility check, a wakeup preemption is very
>> likely when an entity with positive lag joins the runqueue pushing the
>> avg_vruntime of the runqueue backwards, making the vruntime of the
>> current entity ineligible. This leads to aggressive wakeup preemption
>> which was previously guarded by wakeup_granularity_ns in legacy CFS.
>> Below figure depicts one such aggressive preemption scenario with EEVDF
>> in DeathStarBench [1]:
>>
>> deadline for Nginx
>> |
>> +-------+ | |
>> /-- | Nginx | -|------------------> |
>> | +-------+ | |
>> | |
>> | -----------|-------------------------------> vruntime timeline
>> | \--> rq->avg_vruntime
>> |
>> | wakes service on the same runqueue since system is busy
>> |
>> | +---------+|
>> \-->| Service || (service has +ve lag pushes avg_vruntime backwards)
>> +---------+|
>> | |
>> wakeup | +--|-----+ |
>> preempts \---->| N|ginx | --------------------> | {deadline for Nginx}
>> +--|-----+ |
>> (Nginx ineligible)
>> -----------|-------------------------------> vruntime timeline
>> \--> rq->avg_vruntime
>
> This graph is really hard to interpret. If you want to illustrate
> avg_vruntime moves back, you should not align it. That's really
> disorienting.
>
> In both (upper and lower) nginx has the same vruntime thus *that* should
> be aligned. The lower will have service placed left and with that
> avg_vruntime also moves left, rendering nginx in-eligible.
Sorry about that. I'll keep this in mind for for future illustrations.
>
>
>> Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com>
>> ---
>> kernel/sched/fair.c | 24 ++++++++++++++++++++----
>> kernel/sched/features.h | 1 +
>> 2 files changed, 21 insertions(+), 4 deletions(-)
>>
>> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
>> index 6a16129f9a5c..a9b145a4eab0 100644
>> --- a/kernel/sched/fair.c
>> +++ b/kernel/sched/fair.c
>> @@ -875,7 +875,7 @@ struct sched_entity *__pick_first_entity(struct cfs_rq *cfs_rq)
>> *
>> * Which allows tree pruning through eligibility.
>> */
>> -static struct sched_entity *pick_eevdf(struct cfs_rq *cfs_rq)
>> +static struct sched_entity *pick_eevdf(struct cfs_rq *cfs_rq, bool wakeup_preempt)
>> {
>> struct rb_node *node = cfs_rq->tasks_timeline.rb_root.rb_node;
>> struct sched_entity *se = __pick_first_entity(cfs_rq);
>> @@ -889,7 +889,23 @@ static struct sched_entity *pick_eevdf(struct cfs_rq *cfs_rq)
>> if (cfs_rq->nr_running == 1)
>> return curr && curr->on_rq ? curr : se;
>>
>> - if (curr && (!curr->on_rq || !entity_eligible(cfs_rq, curr)))
>> + if (curr && !curr->on_rq)
>> + curr = NULL;
>> +
>> + /*
>> + * When an entity with positive lag wakes up, it pushes the
>> + * avg_vruntime of the runqueue backwards. This may causes the
>> + * current entity to be ineligible soon into its run leading to
>> + * wakeup preemption.
>> + *
>> + * To prevent such aggressive preemption of the current running
>> + * entity during task wakeups, skip the eligibility check if the
>> + * slice promised to the entity since its selection has not yet
>> + * elapsed.
>> + */
>> + if (curr &&
>> + !(sched_feat(RUN_TO_PARITY_WAKEUP) && wakeup_preempt && curr->vlag == curr->deadline) &&
>> + !entity_eligible(cfs_rq, curr))
>> curr = NULL;
>>
>> /*
>
> So I see what you want to do, but this is highly unreadable.
>
> I'll try something like the below on top of queue/sched/eevdf, but I
> should probably first look at fixing those reported crashes on that tree
> :/
I'll give this a try since Mike's suggestion seems to have fixed the
crash I was observing :) Thank you for suggesting this alternative.
--
Thanks and Regards,
Prateek
>
> ---
> kernel/sched/fair.c | 60 ++++++++++++++++++++++++++++++++++---------------
> kernel/sched/features.h | 11 +++++----
> 2 files changed, 49 insertions(+), 22 deletions(-)
>
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index 8a9206d8532f..23977ed1cb2c 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -855,6 +855,39 @@ struct sched_entity *__pick_first_entity(struct cfs_rq *cfs_rq)
> return __node_2_se(left);
> }
>
> +static inline bool pick_curr(struct cfs_rq *cfs_rq,
> + struct sched_entity *curr, struct sched_entity *wakee)
> +{
> + /*
> + * Nothing to preserve...
> + */
> + if (!curr || !sched_feat(RESPECT_SLICE))
> + return false;
> +
> + /*
> + * Allow preemption at the 0-lag point -- even if not all of the slice
> + * is consumed. Note: placement of positive lag can push V left and render
> + * @curr instantly ineligible irrespective the time on-cpu.
> + */
> + if (sched_feat(RUN_TO_PARITY) && !entity_eligible(cfs_rq, curr))
> + return false;
> +
> + /*
> + * Don't preserve @curr when the @wakee has a shorter slice and earlier
> + * deadline. IOW, explicitly allow preemption.
> + */
> + if (sched_feat(PREEMPT_SHORT) && wakee &&
> + wakee->slice < curr->slice &&
> + (s64)(wakee->deadline - curr->deadline) < 0)
> + return false;
> +
> + /*
> + * Preserve @curr to allow it to finish its first slice.
> + * See the HACK in set_next_entity().
> + */
> + return curr->vlag == curr->deadline;
> +}
> +
> /*
> * Earliest Eligible Virtual Deadline First
> *
> @@ -874,28 +907,27 @@ struct sched_entity *__pick_first_entity(struct cfs_rq *cfs_rq)
> *
> * Which allows tree pruning through eligibility.
> */
> -static struct sched_entity *pick_eevdf(struct cfs_rq *cfs_rq)
> +static struct sched_entity *pick_eevdf(struct cfs_rq *cfs_rq, struct sched_entity *wakee)
> {
> struct rb_node *node = cfs_rq->tasks_timeline.rb_root.rb_node;
> struct sched_entity *se = __pick_first_entity(cfs_rq);
> struct sched_entity *curr = cfs_rq->curr;
> struct sched_entity *best = NULL;
>
> + if (curr && !curr->on_rq)
> + curr = NULL;
> +
> /*
> * We can safely skip eligibility check if there is only one entity
> * in this cfs_rq, saving some cycles.
> */
> if (cfs_rq->nr_running == 1)
> - return curr && curr->on_rq ? curr : se;
> -
> - if (curr && (!curr->on_rq || !entity_eligible(cfs_rq, curr)))
> - curr = NULL;
> + return curr ?: se;
>
> /*
> - * Once selected, run a task until it either becomes non-eligible or
> - * until it gets a new slice. See the HACK in set_next_entity().
> + * Preserve @curr to let it finish its slice.
> */
> - if (sched_feat(RUN_TO_PARITY) && curr && curr->vlag == curr->deadline)
> + if (pick_curr(cfs_rq, curr, wakee))
> return curr;
>
> /* Pick the leftmost entity if it's eligible */
> @@ -5507,7 +5539,7 @@ pick_next_entity(struct rq *rq, struct cfs_rq *cfs_rq)
> return cfs_rq->next;
> }
>
> - struct sched_entity *se = pick_eevdf(cfs_rq);
> + struct sched_entity *se = pick_eevdf(cfs_rq, NULL);
> if (se->sched_delayed) {
> dequeue_entities(rq, se, DEQUEUE_SLEEP | DEQUEUE_DELAYED);
> SCHED_WARN_ON(se->sched_delayed);
> @@ -8548,15 +8580,7 @@ static void check_preempt_wakeup_fair(struct rq *rq, struct task_struct *p, int
> cfs_rq = cfs_rq_of(se);
> update_curr(cfs_rq);
>
> - if (sched_feat(PREEMPT_SHORT) && pse->slice < se->slice &&
> - entity_eligible(cfs_rq, pse) &&
> - (s64)(pse->deadline - se->deadline) < 0 &&
> - se->vlag == se->deadline) {
> - /* negate RUN_TO_PARITY */
> - se->vlag = se->deadline - 1;
> - }
> -
> - if (pick_eevdf(cfs_rq) == pse)
> + if (pick_eevdf(cfs_rq, pse) == pse)
> goto preempt;
>
> return;
> diff --git a/kernel/sched/features.h b/kernel/sched/features.h
> index 64ce99cf04ec..2285dc30294c 100644
> --- a/kernel/sched/features.h
> +++ b/kernel/sched/features.h
> @@ -10,12 +10,15 @@ SCHED_FEAT(PLACE_LAG, true)
> */
> SCHED_FEAT(PLACE_DEADLINE_INITIAL, true)
> /*
> - * Inhibit (wakeup) preemption until the current task has either matched the
> - * 0-lag point or until is has exhausted it's slice.
> + * Inhibit (wakeup) preemption until the current task has exhausted its slice.
> */
> -SCHED_FEAT(RUN_TO_PARITY, true)
> +SCHED_FEAT(RESPECT_SLICE, true)
> /*
> - * Allow tasks with a shorter slice to disregard RUN_TO_PARITY
> + * Relax RESPECT_SLICE to allow preemption once current has reached 0-lag.
> + */
> +SCHED_FEAT(RUN_TO_PARITY, false)
> +/*
> + * Allow tasks with a shorter slice to disregard RESPECT_SLICE
> */
> SCHED_FEAT(PREEMPT_SHORT, true)
>
^ permalink raw reply [flat|nested] 11+ messages in thread
* [PATCH] [Question] sched/fair: Task starvation with RUN_TO_PARITY_WAKEUP under group topologies
2024-03-25 6:02 ` [RFC PATCH 1/1] sched/eevdf: Skip eligibility check for current entity during wakeup preemption K Prateek Nayak
2024-03-25 15:13 ` Youssef Esmat
2024-04-24 15:07 ` Peter Zijlstra
@ 2026-07-12 11:31 ` Chen Jinghuang
2026-07-13 3:25 ` K Prateek Nayak
2 siblings, 1 reply; 11+ messages in thread
From: Chen Jinghuang @ 2026-07-12 11:31 UTC (permalink / raw)
To: kprateek.nayak
Cc: bristot, bsegall, dietmar.eggemann, dtcccc, gautham.shenoy,
huschle, juri.lelli, linux-kernel, luis.machado, mgorman, mingo,
peterz, rostedt, vincent.guittot, vschneid, wuyun.abel,
xuewen.yan94, youssefesmat, yu.c.chen
Hi all,
While adapting the RUN_TO_PARITY_WAKEUP feature on top of mainline v7.2-rc2,
I encountered a severe task starvation issue under specific cgroup topologies.
Specifically, a running task completely hogs the CPU and prevents any
preemption.
Test Topology and Observed Behavior:
- CPU 0: Bound with two tasks: A1 (non-sleeping task) from cgroup A, and
B1 (frequent sleep/wake task) from cgroup B.
- CPU 20: Bound with other tasks from cgroup A (A2, A3... which sleep/wake
frequently) and task C1 (non-sleeping task) from cgroup C.
- All cgroups maintain default shares.
CPU 0 CPU 20 / Other CPUs
+----------------------------+ +----------------------------+
| cgroup A cgroup B | | cgroup A cgroup C |
| |- A1 (Run) |- B1 | | |- A2, A3... |- C1 |
| (Non-sleeping) (50us S) | | (Freq S/W) (Non-sleeping)|
| (10us W) | | |
+----------------------------+ +----------------------------+
|
v
A2/A3 frequent sleep/wake -> Frequent reweight of gse(A) on CPU 0
-> Triggers vlag clamping -> Unidirectional avruntime drift
-> gse->vruntime drops abnormally (while gse weight remains stable)
-> protect_slice() returns true -> pick_eevdf always returns curr
-> Task A1 constantly hogs the CPU
According to function_graph traces, the frequent sleep/wake cycles of
sibling tasks (A2/A3) on CPU 20 cause gse(A) on CPU 0 to undergo continuous
reweighting via update_cfs_cgroup() -> reweight_entity(). This triggers the
following chain reaction:
1. gse->vlag triggers the lag clamping logic in entity_lag().
2. Due to the clamping limit, A1->vruntime drifts unidirectionally against
avg_vruntime (i.e., gse->vruntime drops abnormally even though its weight
remains stable).
3. protect_slice() constantly returns true, forcing pick_eevdf() to always
return curr(A1) during wakeup preemption checks.
4. As a result, A1 indefinitely hogs CPU 0. Even in entity_tick(),
update_curr() fails to trigger resched_curr(). Both wakeup and periodic
preemptions end up picking curr.
Analysis:
Experimentally, either disabling RUN_TO_PARITY_WAKEUP or reverting the vlag
clamping patches completely resolves this starvation issue.
This indicates a problematic interaction between the extended runtime
introduced by RUN_TO_PARITY_WAKEUP and the vlag clamping mechanism. While
A1 runs beyond its expected share, frequent sibling-induced reweighting
pushes vlag to its limit, pulling A1->vruntime artificially closer to
avg_vruntime. However, A1->deadline and A1->vprot are still rescaled in
rescale_entity(). This mathematical asymmetry effectively ensures that
A1->vruntime < A1->vprot`always holds true. Consequently, protect_slice()
permanently returns true, trapping the scheduler into believing A1 has not
yet exhausted its promised slice, thereby completely blocking legitimate
context switches.
Discussion:
I'd love to hear your thoughts on how we can fix this issue without losing
the throughput gains of RUN_TO_PARITY_WAKEUP.
Thanks all.
Signed-off-by: Chen Jinghuang <chenjinghuang2@huawei.com>
---
kernel/sched/fair.c | 18 +++++++++++++++++-
1 file changed, 17 insertions(+), 1 deletion(-)
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index d78467ec6ee1..00869624d914 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -1157,7 +1157,23 @@ static struct sched_entity *pick_eevdf(struct cfs_rq *cfs_rq, bool protect)
return cfs_rq->next;
}
- if (curr && (!curr->on_rq || !entity_eligible(cfs_rq, curr)))
+ if (curr && !curr->on_rq)
+ curr = NULL;
+
+ /*
+ * When an entity with positive lag wakes up, it pushes the
+ * avg_vruntime of the runqueue backwards. This may causes the
+ * current entity to be ineligible soon into its run leading to
+ * wakeup preemption.
+ *
+ * To prevent such aggressive preemption of the current running
+ * entity during task wakeups, skip the eligibility check if the
+ * slice promised to the entity since its selection has not yet
+ * elapsed.
+ */
+ if (curr &&
+ !(sched_feat(RUN_TO_PARITY_WAKEUP) && protect && protect_slice(curr)) &&
+ !entity_eligible(cfs_rq, curr))
curr = NULL;
if (curr && protect && protect_slice(curr))
--
2.34.1
^ permalink raw reply [flat|nested] 11+ messages in thread* Re: [PATCH] [Question] sched/fair: Task starvation with RUN_TO_PARITY_WAKEUP under group topologies
2026-07-12 11:31 ` [PATCH] [Question] sched/fair: Task starvation with RUN_TO_PARITY_WAKEUP under group topologies Chen Jinghuang
@ 2026-07-13 3:25 ` K Prateek Nayak
0 siblings, 0 replies; 11+ messages in thread
From: K Prateek Nayak @ 2026-07-13 3:25 UTC (permalink / raw)
To: Chen Jinghuang
Cc: bristot, bsegall, dietmar.eggemann, dtcccc, gautham.shenoy,
huschle, juri.lelli, linux-kernel, luis.machado, mgorman, mingo,
peterz, rostedt, vincent.guittot, vschneid, wuyun.abel,
xuewen.yan94, youssefesmat, yu.c.chen
Hello Chen,
On 7/12/2026 5:01 PM, Chen Jinghuang wrote:
> Hi all,
>
> While adapting the RUN_TO_PARITY_WAKEUP feature on top of mainline v7.2-rc2,
> I encountered a severe task starvation issue under specific cgroup topologies.
> Specifically, a running task completely hogs the CPU and prevents any
> preemption.
I'm assuming you are referring to
https://lore.kernel.org/lkml/20240325060226.1540-2-kprateek.nayak@amd.com/
There was a reason it was not mainlined for this exact issue.
...
> Discussion:
> I'd love to hear your thoughts on how we can fix this issue without losing
> the throughput gains of RUN_TO_PARITY_WAKEUP.
Run your tasks as SCHED_BATCH and they'll forego wakeup preemption.
This is the exact recipe that had helped a bunch of workloads that
suffered throughput degradation from wakeup preemption after EEVDF
was introduced.
If you have well behaved workloads, you can also play around with
custom slice (sched_attr.sched_runtime for fair tasks) which gets
propagated down the cgroup hierarchy and should be better behaved.
>
> Thanks all.
>
> Signed-off-by: Chen Jinghuang <chenjinghuang2@huawei.com>
> ---
> kernel/sched/fair.c | 18 +++++++++++++++++-
> 1 file changed, 17 insertions(+), 1 deletion(-)
>
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index d78467ec6ee1..00869624d914 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -1157,7 +1157,23 @@ static struct sched_entity *pick_eevdf(struct cfs_rq *cfs_rq, bool protect)
> return cfs_rq->next;
> }
>
> - if (curr && (!curr->on_rq || !entity_eligible(cfs_rq, curr)))
> + if (curr && !curr->on_rq)
> + curr = NULL;
> +
> + /*
> + * When an entity with positive lag wakes up, it pushes the
> + * avg_vruntime of the runqueue backwards. This may causes the
> + * current entity to be ineligible soon into its run leading to
> + * wakeup preemption.
> + *
> + * To prevent such aggressive preemption of the current running
> + * entity during task wakeups, skip the eligibility check if the
> + * slice promised to the entity since its selection has not yet
> + * elapsed.
> + */
> + if (curr &&
> + !(sched_feat(RUN_TO_PARITY_WAKEUP) && protect && protect_slice(curr)) &&
> + !entity_eligible(cfs_rq, curr))
> curr = NULL;
We also have a bunch of changes in tip:sched/core that improves wakeup
preemption with custom slices. You may want to check them out too:
https://git.kernel.org/pub/scm/linux/kernel/git/tip/tip.git/log/?h=sched/core
>
> if (curr && protect && protect_slice(curr))
--
Thanks and Regards,
Prateek
^ permalink raw reply [flat|nested] 11+ messages in thread