* [PATCH 0/2] sched/eevdf: Fix slice protection across state changes
@ 2026-09-30 8:47 Christian Loehle
2026-09-30 8:47 ` [PATCH 1/2] sched/eevdf: Cap protection when current has the shortest slice Christian Loehle
2026-09-30 8:47 ` [PATCH 2/2] sched/eevdf: Keep expired protection expired across reweighting Christian Loehle
0 siblings, 2 replies; 7+ messages in thread
From: Christian Loehle @ 2026-09-30 8:47 UTC (permalink / raw)
To: linux-kernel, mingo, peterz, vincent.guittot
Cc: dietmar.eggemann, kayracizmeci, Christian Loehle
These two fixes address slice-protection boundaries that become stale
after a request-size or weight change, found when testing a few edge
cases for Vincent's series:
https://lore.kernel.org/lkml/20260921152238.3804392-1-vincent.guittot@linaro.org/
Patch 1 caps every fresh protection grant at the selected minimum slice.
A preserved deadline can outlast a task's new, shorter request, including
when current itself owns the minimum slice. In that case the existing
slice != se->slice condition skips the cap. This completes the minimum-
slice bound from aae2a33ea662 ("sched/eevdf: Ensure that vprot will never go
above a min slice").
Patch 2 keeps expired protection expired when reweighting moves vruntime.
An old absolute vprot can otherwise become live again and cause current
to be selected over an eligible entity with an earlier deadline. Anchor
the expired boundary at the new vruntime while retaining the existing
rescaling of live protection.
For the first case a task reducing its request from 100 ms to 100 us can
receive over 26 ms of fresh protection while competing with a 1 ms task.
For the second case directed cgroup-weight changes show expired protection
becoming live and affecting selection.
Based on v7.3-rc5 plus:
c72945693b90 sched: Restart fair hrtick after same-task repicks
aae2a33ea662 sched/eevdf: Ensure that vprot will never go above a min slice
4bf32ec3327d sched/eevdf: Align update_protect_slice to set_protect_slice
d2e010082757 sched/eevdf: Handle more short slice waking cases
(Patch 2 is independent, patch 1 follows Vincent's minimum-slice change)
Thanks,
Christian
Christian Loehle (2):
sched/eevdf: Cap protection when current has the shortest slice
sched/eevdf: Keep expired protection expired across reweighting
kernel/sched/fair.c | 12 +++++++-----
1 file changed, 7 insertions(+), 5 deletions(-)
^ permalink raw reply [flat|nested] 7+ messages in thread
* [PATCH 1/2] sched/eevdf: Cap protection when current has the shortest slice
2026-09-30 8:47 [PATCH 0/2] sched/eevdf: Fix slice protection across state changes Christian Loehle
@ 2026-09-30 8:47 ` Christian Loehle
2026-09-30 13:15 ` Christian Loehle
2026-09-30 8:47 ` [PATCH 2/2] sched/eevdf: Keep expired protection expired across reweighting Christian Loehle
1 sibling, 1 reply; 7+ messages in thread
From: Christian Loehle @ 2026-09-30 8:47 UTC (permalink / raw)
To: linux-kernel, mingo, peterz, vincent.guittot
Cc: dietmar.eggemann, kayracizmeci, Christian Loehle
Changing a task's slice does not discard its remaining request. With
PLACE_REL_DEADLINE, the preserved deadline can therefore extend beyond
the new slice after sched_setattr() reduces it.
set_protect_slice() starts from that deadline and only applies the slice
cap when another entity has a shorter slice. If current itself has the
shortest slice, a later pick can protect it for the remainder of the old
request. update_curr() then need not request another selection when a
competitor becomes eligible.
For a preserved deadline d beyond the new virtual slice vslice:
v v + vslice d
|---------------------|----------------------|
old: |<---------------- protection -------------->|
new: |<---- protection --->|
Commit aae2a33ea662 ("sched/eevdf: Ensure that vprot will never go above a
min slice") caps protection even when the ineligibility boundary is later,
but leaves this slice == se->slice case uncovered.
Apply the minimum-slice cap unconditionally, including when it is
current's own slice. Keep the earlier ineligibility boundary for the
PREEMPT_SHORT case when a shorter slice is competing. This preserves the
outstanding request while limiting protection at each fresh pick.
Fixes: 82e9d0456e06 ("sched/fair: Avoid re-setting virtual deadline on 'migrations'")
Signed-off-by: Christian Loehle <christian.loehle@arm.com>
---
kernel/sched/fair.c | 9 ++++-----
1 file changed, 4 insertions(+), 5 deletions(-)
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 85bf02570473..868c3911337a 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -1128,13 +1128,12 @@ static inline void set_protect_slice(struct cfs_rq *cfs_rq, struct sched_entity
slice = cfs_rq_min_slice(cfs_rq);
slice = min(slice, se->slice);
+ /* A preserved deadline can extend beyond the current slice. */
+ vprot = min_vruntime(vprot, se->vruntime + calc_delta_fair(slice, se));
/* If there are shorter slices than se's one */
- if (slice != se->slice) {
- vprot = min_vruntime(vprot, se->vruntime + calc_delta_fair(slice, se));
- if (sched_feat(PREEMPT_SHORT))
- vprot = min_vruntime(vprot, ineligible_vruntime(cfs_rq));
- }
+ if (sched_feat(PREEMPT_SHORT) && slice != se->slice)
+ vprot = min_vruntime(vprot, ineligible_vruntime(cfs_rq));
se->vprot = vprot;
}
^ permalink raw reply [flat|nested] 7+ messages in thread
* [PATCH 2/2] sched/eevdf: Keep expired protection expired across reweighting
2026-09-30 8:47 [PATCH 0/2] sched/eevdf: Fix slice protection across state changes Christian Loehle
2026-09-30 8:47 ` [PATCH 1/2] sched/eevdf: Cap protection when current has the shortest slice Christian Loehle
@ 2026-09-30 8:47 ` Christian Loehle
2026-09-30 13:37 ` Kayra Cizmeci
1 sibling, 1 reply; 7+ messages in thread
From: Christian Loehle @ 2026-09-30 8:47 UTC (permalink / raw)
To: linux-kernel, mingo, peterz, vincent.guittot
Cc: dietmar.eggemann, kayracizmeci, Christian Loehle
reweight_eevdf() rescales live slice protection, but leaves an expired
vprot unchanged when moving vruntime. A reweight can move vruntime behind
the old vprot. For example, reducing the weight of an entity with positive
lag can do so:
before reweight: vprot <= vruntime
after reweight: vruntime < vprot
protect_slice() consequently becomes true again, even though no fresh
protection was granted.
This also happens from task_tick_fair() after a queued HRTICK has
requested a reschedule. A four-task rt-app workload with 100 us, 1 ms,
10 ms and 100 ms requests can then repick current despite a runnable,
eligible entity having an earlier deadline. The stale protection takes
precedence in pick_eevdf(). Cgroup weight changes also expose this with
HRTICK disabled.
Separating vprot from vlag allowed the expired absolute boundary to survive
the lag update and rescaling. Previously, those writes to vlag overwrote
the shared storage. Commit ff38424030f9 ("sched/eevdf: Update se->vprot in
reweight_entity()") subsequently handled live protection, but left the
expired case unchanged.
Keep an expired current entity's protection at its new vruntime:
after reweight: vprot = vruntime
This keeps protect_slice() false. Retain the existing rescaling for
protection that was still live and leave non-current entities alone.
Fixes: 80390ead2080 ("sched/fair: Separate se->vlag from se->vprot")
Signed-off-by: Christian Loehle <christian.loehle@arm.com>
---
kernel/sched/fair.c | 3 +++
1 file changed, 3 insertions(+)
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 868c3911337a..d10eeea75f13 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -4951,6 +4951,9 @@ static void reweight_eevdf(struct cfs_rq *cfs_rq, struct sched_entity *se,
se->deadline += avruntime;
se->rel_deadline = 0;
se->vruntime = avruntime - se->vlag;
+ /* Reweighting must not revive expired slice protection. */
+ if (curr && !rel_vprot)
+ se->vprot = se->vruntime;
if (!curr)
__enqueue_entity(cfs_rq, se);
^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: [PATCH 1/2] sched/eevdf: Cap protection when current has the shortest slice
2026-09-30 8:47 ` [PATCH 1/2] sched/eevdf: Cap protection when current has the shortest slice Christian Loehle
@ 2026-09-30 13:15 ` Christian Loehle
0 siblings, 0 replies; 7+ messages in thread
From: Christian Loehle @ 2026-09-30 13:15 UTC (permalink / raw)
To: linux-kernel, mingo, peterz, vincent.guittot
Cc: dietmar.eggemann, kayracizmeci, Elif Topuz
On 9/30/26 09:47, Christian Loehle wrote:
> Changing a task's slice does not discard its remaining request. With
> PLACE_REL_DEADLINE, the preserved deadline can therefore extend beyond
> the new slice after sched_setattr() reduces it.
>
> set_protect_slice() starts from that deadline and only applies the slice
> cap when another entity has a shorter slice. If current itself has the
> shortest slice, a later pick can protect it for the remainder of the old
> request. update_curr() then need not request another selection when a
> competitor becomes eligible.
>
> For a preserved deadline d beyond the new virtual slice vslice:
>
> v v + vslice d
> |---------------------|----------------------|
> old: |<---------------- protection -------------->|
> new: |<---- protection --->|
>
> Commit aae2a33ea662 ("sched/eevdf: Ensure that vprot will never go above a
> min slice") caps protection even when the ineligibility boundary is later,
> but leaves this slice == se->slice case uncovered.
>
> Apply the minimum-slice cap unconditionally, including when it is
> current's own slice. Keep the earlier ineligibility boundary for the
> PREEMPT_SHORT case when a shorter slice is competing. This preserves the
> outstanding request while limiting protection at each fresh pick.
>
> Fixes: 82e9d0456e06 ("sched/fair: Avoid re-setting virtual deadline on 'migrations'")
> Signed-off-by: Christian Loehle <christian.loehle@arm.com>
> ---
> kernel/sched/fair.c | 9 ++++-----
> 1 file changed, 4 insertions(+), 5 deletions(-)
>
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index 85bf02570473..868c3911337a 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -1128,13 +1128,12 @@ static inline void set_protect_slice(struct cfs_rq *cfs_rq, struct sched_entity
> slice = cfs_rq_min_slice(cfs_rq);
>
> slice = min(slice, se->slice);
> + /* A preserved deadline can extend beyond the current slice. */
> + vprot = min_vruntime(vprot, se->vruntime + calc_delta_fair(slice, se));
>
> /* If there are shorter slices than se's one */
> - if (slice != se->slice) {
> - vprot = min_vruntime(vprot, se->vruntime + calc_delta_fair(slice, se));
> - if (sched_feat(PREEMPT_SHORT))
> - vprot = min_vruntime(vprot, ineligible_vruntime(cfs_rq));
> - }
> + if (sched_feat(PREEMPT_SHORT) && slice != se->slice)
> + vprot = min_vruntime(vprot, ineligible_vruntime(cfs_rq));
>
> se->vprot = vprot;
> }
Elif (+CC) made me aware that apart from the slice protection boundary and vruntime, which is:
set_protect_slice(): deadline, se->vruntime
update_protect_slice(): vprot, min(se->vruntime, avg_vruntime())
this now mirrors update_protect_slice() so should probably be merged into one, I'll resend.
^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: [PATCH 2/2] sched/eevdf: Keep expired protection expired across reweighting
2026-09-30 8:47 ` [PATCH 2/2] sched/eevdf: Keep expired protection expired across reweighting Christian Loehle
@ 2026-09-30 13:37 ` Kayra Cizmeci
2026-09-30 15:28 ` Christian Loehle
0 siblings, 1 reply; 7+ messages in thread
From: Kayra Cizmeci @ 2026-09-30 13:37 UTC (permalink / raw)
To: christian.loehle
Cc: dietmar.eggemann, kayracizmeci, linux-kernel, mingo, peterz,
vincent.guittot
Hi Christian,
> reweight_eevdf() rescales live slice protection, but leaves an expired
> vprot unchanged when moving vruntime. A reweight can move vruntime behind
> the old vprot. For example, reducing the weight of an entity with positive
> lag can do so:
>
> before reweight: vprot <= vruntime
> after reweight: vruntime < vprot
>
> protect_slice() consequently becomes true again, even though no fresh
> protection was granted.
> This also happens from task_tick_fair() after a queued HRTICK has
> requested a reschedule. A four-task rt-app workload with 100 us, 1 ms,
> 10 ms and 100 ms requests can then repick current despite a runnable,
> eligible entity having an earlier deadline. The stale protection takes
> precedence in pick_eevdf(). Cgroup weight changes also expose this with
> HRTICK disabled.
> Separating vprot from vlag allowed the expired absolute boundary to survive
> the lag update and rescaling. Previously, those writes to vlag overwrote
> the shared storage. Commit ff38424030f9 ("sched/eevdf: Update se->vprot in
> reweight_entity()") subsequently handled live protection, but left the
> expired case unchanged.
> Keep an expired current entity's protection at its new vruntime:
>
> after reweight: vprot = vruntime
>
> This keeps protect_slice() false. Retain the existing rescaling for
> protection that was still live and leave non-current entities alone.
>
> Fixes: 80390ead2080 ("sched/fair: Separate se->vlag from se->vprot")
> Signed-off-by: Christian Loehle <christian.loehle@arm.com>
> ---
> kernel/sched/fair.c | 3 +++
> 1 file changed, 3 insertions(+)
>
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index 868c3911337a..d10eeea75f13 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -4951,6 +4951,9 @@ static void reweight_eevdf(struct cfs_rq *cfs_rq, struct sched_entity *se,
> se->deadline += avruntime;
> se->rel_deadline = 0;
> se->vruntime = avruntime - se->vlag;
> + /* Reweighting must not revive expired slice protection. */
> + if (curr && !rel_vprot)
> + se->vprot = se->vruntime;
>
> if (!curr)
> __enqueue_entity(cfs_rq, se);
Okay, so scene:
we have a rq like this:
+----+
|root|
+----+
/\
/ \
/ \
+------+ +------+
|task_a| |task_b|
+------+ +------+
When, sched_change_end() activates for task_a, (Assuming it's Fair Class and weight is different than h_load.weight)
enqueue_task_fair() calls reweight_eevdf() with on_rq = false so the block never works. Then we call place_entity() and vruntime goes back.
On normal enqueue, this will be tolerated with vprot getting written over. But, sched_change_end() calls set_next_task()
that calls set_next_task_fair() with SNT_NORMAL. On that case set_protect_slice() is not called. reweight_eevdf() is called
again on that path, but because the first call did the job this one just returns early.
I could be missing something tho. If I'm not, should we fold this into 2/2 or should I send a patch about this?
Thanks,
Kayra
^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: [PATCH 2/2] sched/eevdf: Keep expired protection expired across reweighting
2026-09-30 13:37 ` Kayra Cizmeci
@ 2026-09-30 15:28 ` Christian Loehle
2026-09-30 15:34 ` Kayra Cizmeci
0 siblings, 1 reply; 7+ messages in thread
From: Christian Loehle @ 2026-09-30 15:28 UTC (permalink / raw)
To: Kayra Cizmeci
Cc: dietmar.eggemann, linux-kernel, mingo, peterz, vincent.guittot
On 9/30/26 14:37, Kayra Cizmeci wrote:
> Hi Christian,
>
>> reweight_eevdf() rescales live slice protection, but leaves an expired
>> vprot unchanged when moving vruntime. A reweight can move vruntime behind
>> the old vprot. For example, reducing the weight of an entity with positive
>> lag can do so:
>>
>> before reweight: vprot <= vruntime
>> after reweight: vruntime < vprot
>>
>> protect_slice() consequently becomes true again, even though no fresh
>> protection was granted.
>
>> This also happens from task_tick_fair() after a queued HRTICK has
>> requested a reschedule. A four-task rt-app workload with 100 us, 1 ms,
>> 10 ms and 100 ms requests can then repick current despite a runnable,
>> eligible entity having an earlier deadline. The stale protection takes
>> precedence in pick_eevdf(). Cgroup weight changes also expose this with
>> HRTICK disabled.
>
>> Separating vprot from vlag allowed the expired absolute boundary to survive
>> the lag update and rescaling. Previously, those writes to vlag overwrote
>> the shared storage. Commit ff38424030f9 ("sched/eevdf: Update se->vprot in
>> reweight_entity()") subsequently handled live protection, but left the
>> expired case unchanged.
>
>> Keep an expired current entity's protection at its new vruntime:
>>
>> after reweight: vprot = vruntime
>>
>> This keeps protect_slice() false. Retain the existing rescaling for
>> protection that was still live and leave non-current entities alone.
>>
>> Fixes: 80390ead2080 ("sched/fair: Separate se->vlag from se->vprot")
>> Signed-off-by: Christian Loehle <christian.loehle@arm.com>
>> ---
>> kernel/sched/fair.c | 3 +++
>> 1 file changed, 3 insertions(+)
>>
>> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
>> index 868c3911337a..d10eeea75f13 100644
>> --- a/kernel/sched/fair.c
>> +++ b/kernel/sched/fair.c
>> @@ -4951,6 +4951,9 @@ static void reweight_eevdf(struct cfs_rq *cfs_rq, struct sched_entity *se,
>> se->deadline += avruntime;
>> se->rel_deadline = 0;
>> se->vruntime = avruntime - se->vlag;
>> + /* Reweighting must not revive expired slice protection. */
>> + if (curr && !rel_vprot)
>> + se->vprot = se->vruntime;
>>
>> if (!curr)
>> __enqueue_entity(cfs_rq, se);
>
>
> Okay, so scene:
>
> we have a rq like this:
> +----+
> |root|
> +----+
> /\
> / \
> / \
> +------+ +------+
> |task_a| |task_b|
> +------+ +------+
>
> When, sched_change_end() activates for task_a, (Assuming it's Fair Class and weight is different than h_load.weight)
> enqueue_task_fair() calls reweight_eevdf() with on_rq = false so the block never works. Then we call place_entity() and vruntime goes back.
> On normal enqueue, this will be tolerated with vprot getting written over. But, sched_change_end() calls set_next_task()
> that calls set_next_task_fair() with SNT_NORMAL. On that case set_protect_slice() is not called. reweight_eevdf() is called
> again on that path, but because the first call did the job this one just returns early.
>
> I could be missing something tho. If I'm not, should we fold this into 2/2 or should I send a patch about this?
>
Thanks, looks legit to me and I did a quick rt-app + trace analysis to confirm.
I'll fold that in with you as a reporter if you don't mind.
I might start including the rt-app workloads in the cover-letter before I
lose track of them myself...
^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: [PATCH 2/2] sched/eevdf: Keep expired protection expired across reweighting
2026-09-30 15:28 ` Christian Loehle
@ 2026-09-30 15:34 ` Kayra Cizmeci
0 siblings, 0 replies; 7+ messages in thread
From: Kayra Cizmeci @ 2026-09-30 15:34 UTC (permalink / raw)
To: christian.loehle
Cc: dietmar.eggemann, kayracizmeci, linux-kernel, mingo, peterz,
vincent.guittot
> Thanks, looks legit to me and I did a quick rt-app + trace analysis to confirm.
> I'll fold that in with you as a reporter if you don't mind.
Yeah, I'll won't :-).
^ permalink raw reply [flat|nested] 7+ messages in thread
end of thread, other threads:[~2026-09-30 15:34 UTC | newest]
Thread overview: 7+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-30 8:47 [PATCH 0/2] sched/eevdf: Fix slice protection across state changes Christian Loehle
2026-09-30 8:47 ` [PATCH 1/2] sched/eevdf: Cap protection when current has the shortest slice Christian Loehle
2026-09-30 13:15 ` Christian Loehle
2026-09-30 8:47 ` [PATCH 2/2] sched/eevdf: Keep expired protection expired across reweighting Christian Loehle
2026-09-30 13:37 ` Kayra Cizmeci
2026-09-30 15:28 ` Christian Loehle
2026-09-30 15:34 ` Kayra Cizmeci
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®