mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v2 0/5] sched/eevdf: Keep slice protection boundaries consistent
@ 2026-10-01 13:40 Christian Loehle
  2026-10-01 13:40 ` [PATCH v2 1/5] sched/fair: Take slice protection into account when arming HRTICK Christian Loehle
                   ` (4 more replies)
  0 siblings, 5 replies; 12+ messages in thread
From: Christian Loehle @ 2026-10-01 13:40 UTC (permalink / raw)
  To: linux-kernel, mingo, peterz, vincent.guittot
  Cc: dietmar.eggemann, kayracizmeci, kprateek.nayak, elif.topuz, sh,
	Christian Loehle

EEVDF slice protection is an absolute virtual-runtime boundary. Several
paths currently treat that boundary inconsistently with the task's current
slice, or with the timer that makes the scheduler reconsider current.

Patch 1 makes fair HRTICK expire at live protection rather than only at the
later virtual deadline. Patch 2 consolidates the fresh and update
protection calculations. The remaining patches cap protection after a
slice change, prevent reweighting from reviving expired protection, and
update protection when sched_setattr() restores the running task.

All workloads below pin every task to CPU1, use SCHED_OTHER, and keep the
controller on another physical core. rt-app's dl-runtime values are in us;
sched_getattr() verified the resulting sched_runtime values. The tests
enable PREEMPT_SHORT, RUN_TO_PARITY and PLACE_REL_DEADLINE, and run with
both HRTICK and NO_HRTICK (except for patch 1).

The HRTICK/deadline mismatch in patch 1 is exposed by two continuously
runnable equal-weight tasks:

  "short": { "cpus": [1], "loop": -1,
             "dl-runtime": 100,  "run": 1000000 },
  "long":  { "cpus": [1], "loop": -1,
             "dl-runtime": 1000, "run": 1000000 }

On the HZ=250 control, the short task's median sched_switch run interval is
1005 us even with HRTICK enabled. With patch 1 it is 101 us. Both tasks
retain approximately half of the CPU in either case. (The error is therefore
the reconsideration time rather than CPU share.)

With RUN_TO_PARITY disabled, vprot instead marks current's minimum progress
quantum. HRTICK must reconsider current at that boundary as well. After an
expired-protection repick it falls back to the deadline without renewing
the quantum.

The oversized fresh protection in patch 3 is exposed by retaining a 100 ms
relative deadline while reducing the request to 100 us:

  "shrink": {
    "cpus": [1], "loop": -1, "dl-runtime": 100000,
    "delay": 700000,
    "phases": {
      "warmup": { "loop": 1,  "dl-runtime": 100000, "run": 20000 },
      "hold":   { "loop": -1, "dl-runtime": 100,    "run": 1000 }
    }
  },
  "peer": { "cpus": [1], "loop": -1,
            "dl-runtime": 1000, "run": 10000 }

After the change, inspect the next fresh SNT_PICK of shrink. The control
grants between 13 and 28 ms of protection to the 100 us task because its
own slice is the runqueue minimum. Patch 3 caps that fresh boundary to the
current request. In the trace-free companion test, the control's complete
post-change run intervals reached 69--83 ms with HRTICK enabled; the
isolated fix reduced the maximum to 3--4 ms. Wall time includes interrupts
and legitimate repicks, so the vprot observation is the correctness
criterion.

The expired-protection reweighting bug in patch 4 is exposed by three
continuously runnable tasks in separate cgroup-v2 CPU groups:

  group a, weight 10000:
    "long":  { "cpus": [1], "loop": -1,
               "dl-runtime": 100000, "run": 10000 }
  group b, weight 100:
    "short": { "cpus": [1], "loop": -1,
               "dl-runtime": 100, "run": 10000 }
  group c, weight 100:
    "mid":   { "cpus": [1], "loop": -1,
               "dl-runtime": 1000, "run": 10000 }

While rt-app runs, change group a's cpu.weight every 400 ms through:

  10000, 1, 100, 1, 10000, 100, 1, 10000

On the control this produced 32 expired-to-live vprot transitions in 16
captures and seven selections of current over a runnable, eligible peer
with an earlier deadline. With patch 4, 181 captures containing 441920
queued current weight changes produced no revival and no such wrong
selection.

Kayra Cizmeci reported the same invariant failure when
sched_change_begin() temporarily dequeues current. This rt-app phase loop
makes the weight change after running beyond a 100 us request:

  "change": {
    "cpus": [1], "loop": -1, "priority": -10, "dl-runtime": 100,
    "phases": {
      "heavy": { "loop": 1, "priority": -10,
                 "dl-runtime": 100, "run": 117 },
      "light": { "loop": 1, "priority": 10,
                 "dl-runtime": 100, "run": 173 }
    }
  },
  "peer": { "cpus": [1], "loop": -1,
            "dl-runtime": 1000, "run": 10000 }

The diagnostic trace records vprot <= vruntime at __setparam_fair() entry
and checks that the following SNT_NORMAL restore does not make vprot live
again.

The restore bug in patch 5 is exposed at the phase boundary itself. Use a
long-slice peer so the task being changed owns the minimum after the
change:

  "change": {
    "cpus": [1], "loop": -1, "dl-runtime": 100000,
    "delay": 700000,
    "phases": {
      "warmup": { "loop": 1,  "dl-runtime": 100000, "run": 20000 },
      "hold":   { "loop": -1, "dl-runtime": 100,    "run": 1000 }
    }
  },
  "peer": { "cpus": [1], "loop": -1,
            "dl-runtime": 100000, "run": 10000 }

---
Changes since v1:
- Expand the standalone HRTICK fix into a series covering related
  protection-boundary inconsistencies found during testing.
- Share the fresh and update protection calculation.
- Fix protection after slice changes, reweighting and SNT_NORMAL restores.
- Request lazy rescheduling when an SNT_NORMAL update expires protection.
- Add HZ=250/HZ=1000 and HRTICK/NO_HRTICK regression testing.

Link: https://lore.kernel.org/r/20260929194711.2689811-1-christian.loehle@arm.com/

Christian Loehle (5):
  sched/fair: Take slice protection into account when arming HRTICK
  sched/eevdf: Share the slice protection calculation
  sched/eevdf: Cap protection when current has the shortest slice
  sched/eevdf: Keep expired protection expired across reweighting
  sched/eevdf: Update protection after restoring current

 kernel/sched/fair.c | 72 +++++++++++++++++++++++++++------------------
 1 file changed, 44 insertions(+), 28 deletions(-)


base-commit: 01d1f30564aac8d53c4ba2de2bfc81fcef206721
-- 
2.34.1

^ permalink raw reply	[flat|nested] 12+ messages in thread

* [PATCH v2 1/5] sched/fair: Take slice protection into account when arming HRTICK
  2026-10-01 13:40 [PATCH v2 0/5] sched/eevdf: Keep slice protection boundaries consistent Christian Loehle
@ 2026-10-01 13:40 ` Christian Loehle
  2026-10-01 13:43   ` Peter Zijlstra
  2026-10-01 13:40 ` [PATCH v2 2/5] sched/eevdf: Share the slice protection calculation Christian Loehle
                   ` (3 subsequent siblings)
  4 siblings, 1 reply; 12+ messages in thread
From: Christian Loehle @ 2026-10-01 13:40 UTC (permalink / raw)
  To: linux-kernel, mingo, peterz, vincent.guittot
  Cc: dietmar.eggemann, kayracizmeci, kprateek.nayak, elif.topuz, sh,
	Christian Loehle

EEVDF can end a task's slice protection before its virtual deadline when
shorter requests compete on the runqueue. update_curr() requests a new
selection once protection expires, but hrtick_start_fair() still arms the
timer for the virtual deadline. Without an intervening scheduling event,
the next opportunity to reconsider the task can therefore arrive much
later than the protection boundary.

Arm the fair hrtick for the earlier of the virtual deadline and live
slice protection. Keep the deadline fallback once protection expires.
With RUN_TO_PARITY disabled, vprot remains the minimum progress quantum,
so the timer must reconsider current at that boundary too. A same-task
repick restarts the timer without renewing protection, so newly eligible
tasks can compete without waiting for another protected quantum.

A wakeup can shorten protection without preempting current. The enqueue
path calls hrtick_update() before wakeup_preempt_fair() updates protection,
and hrtick_update() leaves an active timer alone. Recompute the fair
hrtick after update_protect_slice() changes vprot, including when a timer
is already active. If the new protection boundary has already passed,
request lazy rescheduling, as update_curr() does on protection expiry,
instead of falling back to the later deadline.

Fixes: 74eec63661d4 ("sched/fair: Fix NO_RUN_TO_PARITY case")
Reported-by: Vincent Guittot <vincent.guittot@linaro.org>
Link: https://lore.kernel.org/r/CAKfTPtD_aJeAh1biW6gKCLK6BnYcNtbz-FJvxUzZ1tD3V=43mw@mail.gmail.com/
Link: https://lore.kernel.org/r/CAKfTPtApFGcYA22Qtg3OyYEt6EJw_FW9_VV=_vgE_HtNY3V9og@mail.gmail.com/
Signed-off-by: Christian Loehle <christian.loehle@arm.com>
---
 kernel/sched/fair.c | 18 ++++++++++++++++--
 1 file changed, 16 insertions(+), 2 deletions(-)

diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index a39c48704d40..85bf02570473 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -8018,9 +8018,13 @@ static void hrtick_start_fair(struct rq *rq, struct task_struct *p)
 		return;
 
 	/*
-	 * Compute time until virtual deadline
+	 * Reconsider the pick when live slice protection expires, which can be
+	 * earlier than the virtual deadline when shorter requests compete.
+	 * Without live protection, use the virtual deadline instead.
 	 */
 	vdelta = se->deadline - se->vruntime;
+	if (protect_slice(se))
+		vdelta = min_vruntime(se->deadline, se->vprot) - se->vruntime;
 	if ((s64)vdelta < 0) {
 		if (task_current_donor(rq, p))
 			resched_curr(rq);
@@ -10261,8 +10265,18 @@ static void wakeup_preempt_fair(struct rq *rq, struct task_struct *p, int wake_f
 	if (preempt_action == PREEMPT_WAKEUP_SHORT && entity_eligible(cfs_rq, pse))
 		goto preempt;
 update:
-	if (sched_feat(RUN_TO_PARITY))
+	if (sched_feat(RUN_TO_PARITY)) {
+		u64 old_vprot = se->vprot;
+
 		update_protect_slice(cfs_rq, se);
+		if (hrtick_enabled_fair(rq) && se->vprot != old_vprot) {
+			/* The enqueue may have armed the timer before this update. */
+			if (protect_slice(se))
+				hrtick_start_fair(rq, donor);
+			else
+				resched_curr_lazy(rq);
+		}
+	}
 
 	return;
 
-- 
2.34.1

^ permalink raw reply	[flat|nested] 12+ messages in thread

* [PATCH v2 2/5] sched/eevdf: Share the slice protection calculation
  2026-10-01 13:40 [PATCH v2 0/5] sched/eevdf: Keep slice protection boundaries consistent Christian Loehle
  2026-10-01 13:40 ` [PATCH v2 1/5] sched/fair: Take slice protection into account when arming HRTICK Christian Loehle
@ 2026-10-01 13:40 ` Christian Loehle
  2026-10-01 13:40 ` [PATCH v2 3/5] sched/eevdf: Cap protection when current has the shortest slice Christian Loehle
                   ` (2 subsequent siblings)
  4 siblings, 0 replies; 12+ messages in thread
From: Christian Loehle @ 2026-10-01 13:40 UTC (permalink / raw)
  To: linux-kernel, mingo, peterz, vincent.guittot
  Cc: dietmar.eggemann, kayracizmeci, kprateek.nayak, elif.topuz, sh,
	Christian Loehle

set_protect_slice() and update_protect_slice() duplicate the slice
selection and vprot limiting rules. Merge the calculation into
set_protect_slice().

No functional change intended.

Suggested-by: Elif Topuz <elif.topuz@arm.com>
Signed-off-by: Christian Loehle <christian.loehle@arm.com>
---
 kernel/sched/fair.c | 39 +++++++++++++++------------------------
 1 file changed, 15 insertions(+), 24 deletions(-)

diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 85bf02570473..3f881ded5457 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -1119,43 +1119,34 @@ struct sched_entity *__pick_first_entity(struct cfs_rq *cfs_rq)
  * When run to parity is disabled, we give a minimum quantum to the running
  * entity to ensure progress.
  */
-static inline void set_protect_slice(struct cfs_rq *cfs_rq, struct sched_entity *se)
+static inline void set_protect_slice(struct cfs_rq *cfs_rq, struct sched_entity *se,
+				     bool update)
 {
 	u64 slice = normalized_sysctl_sched_base_slice;
+	u64 vruntime = se->vruntime;
 	u64 vprot = se->deadline;
 
+	if (update) {
+		vruntime = min_vruntime(se->vruntime, avg_vruntime(cfs_rq));
+		vprot = se->vprot;
+	}
+
 	if (sched_feat(RUN_TO_PARITY))
 		slice = cfs_rq_min_slice(cfs_rq);
 
-	slice = min(slice, se->slice);
+	if (!update)
+		slice = min(slice, se->slice);
 
 	/* If there are shorter slices than se's one */
-	if (slice != se->slice) {
-		vprot = min_vruntime(vprot, se->vruntime + calc_delta_fair(slice, se));
-		if (sched_feat(PREEMPT_SHORT))
+	if (update || slice != se->slice) {
+		vprot = min_vruntime(vprot, vruntime + calc_delta_fair(slice, se));
+		if (sched_feat(PREEMPT_SHORT) && slice != se->slice)
 			vprot = min_vruntime(vprot, ineligible_vruntime(cfs_rq));
 	}
 
 	se->vprot = vprot;
 }
 
-static inline void update_protect_slice(struct cfs_rq *cfs_rq, struct sched_entity *se)
-{
-	u64 vruntime = min_vruntime(se->vruntime, avg_vruntime(cfs_rq));
-	u64 slice = normalized_sysctl_sched_base_slice;
-	u64 vprot;
-
-	if (sched_feat(RUN_TO_PARITY))
-		slice = cfs_rq_min_slice(cfs_rq);
-
-	vprot = min_vruntime(se->vprot, vruntime + calc_delta_fair(slice, se));
-
-	if (sched_feat(PREEMPT_SHORT) && slice != se->slice)
-		vprot = min_vruntime(vprot, ineligible_vruntime(cfs_rq));
-
-	se->vprot = vprot;
-}
-
 static inline bool protect_slice(struct sched_entity *se)
 {
 	return vruntime_cmp(se->vruntime, "<", se->vprot);
@@ -10268,7 +10259,7 @@ static void wakeup_preempt_fair(struct rq *rq, struct task_struct *p, int wake_f
 	if (sched_feat(RUN_TO_PARITY)) {
 		u64 old_vprot = se->vprot;
 
-		update_protect_slice(cfs_rq, se);
+		set_protect_slice(cfs_rq, se, true);
 		if (hrtick_enabled_fair(rq) && se->vprot != old_vprot) {
 			/* The enqueue may have armed the timer before this update. */
 			if (protect_slice(se))
@@ -15576,7 +15567,7 @@ static void set_next_task_fair(struct rq *rq, struct task_struct *p, enum snt_e
 	if (on_rq) {
 		reweight_eevdf(cfs_rq, se, weight, se->on_rq);
 		if (first)
-			set_protect_slice(cfs_rq, se);
+			set_protect_slice(cfs_rq, se, false);
 	}
 
 	if (task_on_rq_queued(p)) {
-- 
2.34.1

^ permalink raw reply	[flat|nested] 12+ messages in thread

* [PATCH v2 3/5] sched/eevdf: Cap protection when current has the shortest slice
  2026-10-01 13:40 [PATCH v2 0/5] sched/eevdf: Keep slice protection boundaries consistent Christian Loehle
  2026-10-01 13:40 ` [PATCH v2 1/5] sched/fair: Take slice protection into account when arming HRTICK Christian Loehle
  2026-10-01 13:40 ` [PATCH v2 2/5] sched/eevdf: Share the slice protection calculation Christian Loehle
@ 2026-10-01 13:40 ` Christian Loehle
  2026-10-01 13:55   ` Peter Zijlstra
  2026-10-01 13:40 ` [PATCH v2 4/5] sched/eevdf: Keep expired protection expired across reweighting Christian Loehle
  2026-10-01 13:40 ` [PATCH v2 5/5] sched/eevdf: Update protection after restoring current Christian Loehle
  4 siblings, 1 reply; 12+ messages in thread
From: Christian Loehle @ 2026-10-01 13:40 UTC (permalink / raw)
  To: linux-kernel, mingo, peterz, vincent.guittot
  Cc: dietmar.eggemann, kayracizmeci, kprateek.nayak, elif.topuz, sh,
	Christian Loehle

Changing a task slice preserves its remaining request. With
PLACE_REL_DEADLINE, the deadline can therefore extend beyond the new
slice after sched_setattr() reduces it.

The protection calculation starts from that deadline for a fresh grant,
or from the existing boundary for an update. It only applies the slice
cap when another entity has a shorter slice. If current itself has the
shortest slice, protection can remain live for the remainder of the old
request.

Apply the slice bound for both fresh grants and updates, including when it
is current's own slice. Keep the earlier ineligibility boundary for the
PREEMPT_SHORT case when a shorter slice is competing.

Fixes: 82e9d0456e06 ("sched/fair: Avoid re-setting virtual deadline on 'migrations'")
Signed-off-by: Christian Loehle <christian.loehle@arm.com>
---
 kernel/sched/fair.c | 12 +++++-------
 1 file changed, 5 insertions(+), 7 deletions(-)

diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 3f881ded5457..15fa967273f6 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -1134,15 +1134,13 @@ static inline void set_protect_slice(struct cfs_rq *cfs_rq, struct sched_entity
 	if (sched_feat(RUN_TO_PARITY))
 		slice = cfs_rq_min_slice(cfs_rq);
 
-	if (!update)
-		slice = min(slice, se->slice);
+	slice = min(slice, se->slice);
+	/* A preserved deadline can extend beyond the current slice. */
+	vprot = min_vruntime(vprot, vruntime + calc_delta_fair(slice, se));
 
 	/* If there are shorter slices than se's one */
-	if (update || slice != se->slice) {
-		vprot = min_vruntime(vprot, vruntime + calc_delta_fair(slice, se));
-		if (sched_feat(PREEMPT_SHORT) && slice != se->slice)
-			vprot = min_vruntime(vprot, ineligible_vruntime(cfs_rq));
-	}
+	if (sched_feat(PREEMPT_SHORT) && slice != se->slice)
+		vprot = min_vruntime(vprot, ineligible_vruntime(cfs_rq));
 
 	se->vprot = vprot;
 }
-- 
2.34.1

^ permalink raw reply	[flat|nested] 12+ messages in thread

* [PATCH v2 4/5] sched/eevdf: Keep expired protection expired across reweighting
  2026-10-01 13:40 [PATCH v2 0/5] sched/eevdf: Keep slice protection boundaries consistent Christian Loehle
                   ` (2 preceding siblings ...)
  2026-10-01 13:40 ` [PATCH v2 3/5] sched/eevdf: Cap protection when current has the shortest slice Christian Loehle
@ 2026-10-01 13:40 ` Christian Loehle
  2026-10-01 16:46   ` Kayra Cizmeci
  2026-10-01 13:40 ` [PATCH v2 5/5] sched/eevdf: Update protection after restoring current Christian Loehle
  4 siblings, 1 reply; 12+ messages in thread
From: Christian Loehle @ 2026-10-01 13:40 UTC (permalink / raw)
  To: linux-kernel, mingo, peterz, vincent.guittot
  Cc: dietmar.eggemann, kayracizmeci, kprateek.nayak, elif.topuz, sh,
	Christian Loehle

reweight_eevdf() rescales live slice protection, but leaves an expired
vprot unchanged when moving vruntime. A reweight can move vruntime behind
the old vprot. For example, reducing the weight of an entity with positive
lag can do so:

    before reweight:       vprot <= vruntime
    after reweight:        vruntime < vprot

protect_slice() consequently becomes true again, even though no fresh
protection was granted.

This can also happen in the HRTICK callback. entity_tick() can expire
protection and request a reschedule, after which task_tick_fair() calls
reweight_eevdf() before schedule() gets a chance to select another entity.
The reweight can move vruntime behind the stale vprot, making
protect_slice() true again before the pending selection. A four-task rt-app
workload with 100 us, 1 ms, 10 ms and 100 ms requests can then repick
current despite a runnable, eligible entity having an earlier deadline.
Cgroup weight changes also expose this with HRTICK disabled.

sched_change_begin() exposes another path by temporarily dequeuing current.
enqueue_task_fair() then reweights it with on_rq clear before
place_entity() moves vruntime. Remember whether protection was expired
before placement so the later SNT_NORMAL restore cannot revive it.

Separating vprot from vlag allowed the expired absolute boundary to
survive the lag update and rescaling. Previously, those writes to vlag
overwrote the shared storage. Commit ff38424030f9 ("sched/eevdf: Update
se->vprot in reweight_entity()") subsequently handled live protection,
but left the expired case unchanged.

Keep an expired current entity's protection at its new vruntime:

    after reweight:        vprot = vruntime

This keeps protect_slice() false. Retain the existing rescaling for
protection that was still live and leave non-current entities alone.

Fixes: 80390ead2080 ("sched/fair: Separate se->vlag from se->vprot")
Reported-by: Kayra Cizmeci <kayracizmeci@gmail.com>
Link: https://lore.kernel.org/lkml/20260930133716.214471-1-kayracizmeci@gmail.com/
Signed-off-by: Christian Loehle <christian.loehle@arm.com>
---
 kernel/sched/fair.c | 9 ++++++++-
 1 file changed, 8 insertions(+), 1 deletion(-)

diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 15fa967273f6..bffc560bc7b8 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -4941,6 +4941,9 @@ static void reweight_eevdf(struct cfs_rq *cfs_rq, struct sched_entity *se,
 		se->deadline += avruntime;
 		se->rel_deadline = 0;
 		se->vruntime = avruntime - se->vlag;
+		/* Reweighting must not revive expired slice protection. */
+		if (curr && !rel_vprot)
+			se->vprot = se->vruntime;
 
 		if (!curr)
 			__enqueue_entity(cfs_rq, se);
@@ -8204,7 +8207,7 @@ enqueue_task_fair(struct rq *rq, struct task_struct *p, int flags)
 	struct sched_entity *se = &p->se;
 	struct cfs_rq *cfs_rq = &rq->cfs;
 	unsigned long weight;
-	bool curr;
+	bool curr, expired = false;
 
 	if (task_is_throttled(p) && enqueue_throttled_task(p))
 		return;
@@ -8237,6 +8240,8 @@ enqueue_task_fair(struct rq *rq, struct task_struct *p, int flags)
 	 * XXX comment on the curr thing
 	 */
 	curr = (cfs_rq->curr == se);
+	if (!curr && task_current_donor(rq, p))
+		expired = !protect_slice(se);
 	if (curr)
 		place_entity(cfs_rq, se, flags);
 
@@ -8248,6 +8253,8 @@ enqueue_task_fair(struct rq *rq, struct task_struct *p, int flags)
 	if (!curr) {
 		reweight_eevdf(cfs_rq, se, weight, false);
 		place_entity(cfs_rq, se, flags | ENQUEUE_QUEUED);
+		if (expired)
+			se->vprot = se->vruntime;
 		__enqueue_entity(cfs_rq, se);
 	}
 
-- 
2.34.1

^ permalink raw reply	[flat|nested] 12+ messages in thread

* [PATCH v2 5/5] sched/eevdf: Update protection after restoring current
  2026-10-01 13:40 [PATCH v2 0/5] sched/eevdf: Keep slice protection boundaries consistent Christian Loehle
                   ` (3 preceding siblings ...)
  2026-10-01 13:40 ` [PATCH v2 4/5] sched/eevdf: Keep expired protection expired across reweighting Christian Loehle
@ 2026-10-01 13:40 ` Christian Loehle
  4 siblings, 0 replies; 12+ messages in thread
From: Christian Loehle @ 2026-10-01 13:40 UTC (permalink / raw)
  To: linux-kernel, mingo, peterz, vincent.guittot
  Cc: dietmar.eggemann, kayracizmeci, kprateek.nayak, elif.topuz, sh,
	Christian Loehle

A running task changed through the sched_change pattern is restored
with SNT_NORMAL. Only fresh picks set slice protection, so the restore
currently leaves vprot unchanged.

If sched_setattr() shortens the task slice, the old boundary can remain
live far beyond the new request. Lower the existing boundary when
restoring current. If protection remains live, rearm HRTICK because its
previous expiry may be later than the new boundary. The update can also
move vprot behind current's vruntime. In that case, request lazy-resched
when another fair entity is queued instead of letting hrtick_start_fair()
fall back to the later deadline.

Fixes: bcd74b2ffdd0 ("sched/fair: Only set slice protection at pick time")
Signed-off-by: Christian Loehle <christian.loehle@arm.com>
---
 kernel/sched/fair.c | 12 +++++++++---
 1 file changed, 9 insertions(+), 3 deletions(-)

diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index bffc560bc7b8..59f9c08f968a 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -15571,8 +15571,7 @@ static void set_next_task_fair(struct rq *rq, struct task_struct *p, enum snt_e
 
 	if (on_rq) {
 		reweight_eevdf(cfs_rq, se, weight, se->on_rq);
-		if (first)
-			set_protect_slice(cfs_rq, se, false);
+		set_protect_slice(cfs_rq, se, !first);
 	}
 
 	if (task_on_rq_queued(p)) {
@@ -15582,8 +15581,15 @@ static void set_next_task_fair(struct rq *rq, struct task_struct *p, enum snt_e
 		 */
 		list_move(&se->group_node, &rq->cfs_tasks);
 	}
-	if (!first)
+	if (!first) {
+		if (!protect_slice(se)) {
+			if (rq->cfs.h_nr_queued > 1)
+				resched_curr_lazy(rq);
+		} else if (hrtick_enabled_fair(rq)) {
+			hrtick_start_fair(rq, p);
+		}
 		return;
+	}
 
 	WARN_ON_ONCE(se->sched_delayed);
 
-- 
2.34.1

^ permalink raw reply	[flat|nested] 12+ messages in thread

* Re: [PATCH v2 1/5] sched/fair: Take slice protection into account when arming HRTICK
  2026-10-01 13:40 ` [PATCH v2 1/5] sched/fair: Take slice protection into account when arming HRTICK Christian Loehle
@ 2026-10-01 13:43   ` Peter Zijlstra
  2026-10-01 14:01     ` Christian Loehle
  0 siblings, 1 reply; 12+ messages in thread
From: Peter Zijlstra @ 2026-10-01 13:43 UTC (permalink / raw)
  To: Christian Loehle
  Cc: linux-kernel, mingo, vincent.guittot, dietmar.eggemann,
	kayracizmeci, kprateek.nayak, elif.topuz, sh

On Thu, Oct 01, 2026 at 02:40:12PM +0100, Christian Loehle wrote:
> EEVDF can end a task's slice protection before its virtual deadline when
> shorter requests compete on the runqueue. update_curr() requests a new
> selection once protection expires, but hrtick_start_fair() still arms the
> timer for the virtual deadline. Without an intervening scheduling event,
> the next opportunity to reconsider the task can therefore arrive much
> later than the protection boundary.
> 
> Arm the fair hrtick for the earlier of the virtual deadline and live
> slice protection. Keep the deadline fallback once protection expires.
> With RUN_TO_PARITY disabled, vprot remains the minimum progress quantum,
> so the timer must reconsider current at that boundary too. A same-task
> repick restarts the timer without renewing protection, so newly eligible
> tasks can compete without waiting for another protected quantum.
> 
> A wakeup can shorten protection without preempting current. The enqueue
> path calls hrtick_update() before wakeup_preempt_fair() updates protection,
> and hrtick_update() leaves an active timer alone. Recompute the fair
> hrtick after update_protect_slice() changes vprot, including when a timer
> is already active. If the new protection boundary has already passed,
> request lazy rescheduling, as update_curr() does on protection expiry,
> instead of falling back to the later deadline.
> 

Right, so I don't much like this. This wrecks the steady state behaviour
in favour of our 'dodgy' wakeup heuristics.

I would much rather we stick to the paper for steady state, and let
wakeups be wakeups.

^ permalink raw reply	[flat|nested] 12+ messages in thread

* Re: [PATCH v2 3/5] sched/eevdf: Cap protection when current has the shortest slice
  2026-10-01 13:40 ` [PATCH v2 3/5] sched/eevdf: Cap protection when current has the shortest slice Christian Loehle
@ 2026-10-01 13:55   ` Peter Zijlstra
  2026-10-01 13:57     ` Christian Loehle
  0 siblings, 1 reply; 12+ messages in thread
From: Peter Zijlstra @ 2026-10-01 13:55 UTC (permalink / raw)
  To: Christian Loehle
  Cc: linux-kernel, mingo, vincent.guittot, dietmar.eggemann,
	kayracizmeci, kprateek.nayak, elif.topuz, sh

On Thu, Oct 01, 2026 at 02:40:14PM +0100, Christian Loehle wrote:
> Changing a task slice preserves its remaining request. With
> PLACE_REL_DEADLINE, the deadline can therefore extend beyond the new
> slice after sched_setattr() reduces it.

Or, the new slice is only effective after the current expires.

Is there a reason we care about this? Changelog didn't mention.

^ permalink raw reply	[flat|nested] 12+ messages in thread

* Re: [PATCH v2 3/5] sched/eevdf: Cap protection when current has the shortest slice
  2026-10-01 13:55   ` Peter Zijlstra
@ 2026-10-01 13:57     ` Christian Loehle
  0 siblings, 0 replies; 12+ messages in thread
From: Christian Loehle @ 2026-10-01 13:57 UTC (permalink / raw)
  To: Peter Zijlstra
  Cc: linux-kernel, mingo, vincent.guittot, dietmar.eggemann,
	kayracizmeci, kprateek.nayak, elif.topuz, sh

On 10/1/26 14:55, Peter Zijlstra wrote:
> On Thu, Oct 01, 2026 at 02:40:14PM +0100, Christian Loehle wrote:
>> Changing a task slice preserves its remaining request. With
>> PLACE_REL_DEADLINE, the deadline can therefore extend beyond the new
>> slice after sched_setattr() reduces it.
> 
> Or, the new slice is only effective after the current expires.
> 
> Is there a reason we care about this? Changelog didn't mention.

Yeah I wasn't sure about this, it does make the tests a little more
hairy as this obviously races...

^ permalink raw reply	[flat|nested] 12+ messages in thread

* Re: [PATCH v2 1/5] sched/fair: Take slice protection into account when arming HRTICK
  2026-10-01 13:43   ` Peter Zijlstra
@ 2026-10-01 14:01     ` Christian Loehle
  0 siblings, 0 replies; 12+ messages in thread
From: Christian Loehle @ 2026-10-01 14:01 UTC (permalink / raw)
  To: Peter Zijlstra
  Cc: linux-kernel, mingo, vincent.guittot, dietmar.eggemann,
	kayracizmeci, kprateek.nayak, elif.topuz, sh

On 10/1/26 14:43, Peter Zijlstra wrote:
> On Thu, Oct 01, 2026 at 02:40:12PM +0100, Christian Loehle wrote:
>> EEVDF can end a task's slice protection before its virtual deadline when
>> shorter requests compete on the runqueue. update_curr() requests a new
>> selection once protection expires, but hrtick_start_fair() still arms the
>> timer for the virtual deadline. Without an intervening scheduling event,
>> the next opportunity to reconsider the task can therefore arrive much
>> later than the protection boundary.
>>
>> Arm the fair hrtick for the earlier of the virtual deadline and live
>> slice protection. Keep the deadline fallback once protection expires.
>> With RUN_TO_PARITY disabled, vprot remains the minimum progress quantum,
>> so the timer must reconsider current at that boundary too. A same-task
>> repick restarts the timer without renewing protection, so newly eligible
>> tasks can compete without waiting for another protected quantum.
>>
>> A wakeup can shorten protection without preempting current. The enqueue
>> path calls hrtick_update() before wakeup_preempt_fair() updates protection,
>> and hrtick_update() leaves an active timer alone. Recompute the fair
>> hrtick after update_protect_slice() changes vprot, including when a timer
>> is already active. If the new protection boundary has already passed,
>> request lazy rescheduling, as update_curr() does on protection expiry,
>> instead of falling back to the later deadline.
>>
> 
> Right, so I don't much like this. This wrecks the steady state behaviour
> in favour of our 'dodgy' wakeup heuristics.
> 
> I would much rather we stick to the paper for steady state, and let
> wakeups be wakeups.

Alright, I guess the same reasoning then for 5/5?
Again these are all just from throwing a bunch of artificial rt-app tests
against Vincent's patchset and seeing what sticks...

^ permalink raw reply	[flat|nested] 12+ messages in thread

* Re: [PATCH v2 4/5] sched/eevdf: Keep expired protection expired across reweighting
  2026-10-01 13:40 ` [PATCH v2 4/5] sched/eevdf: Keep expired protection expired across reweighting Christian Loehle
@ 2026-10-01 16:46   ` Kayra Cizmeci
  2026-10-01 21:51     ` Christian Loehle
  0 siblings, 1 reply; 12+ messages in thread
From: Kayra Cizmeci @ 2026-10-01 16:46 UTC (permalink / raw)
  To: christian.loehle
  Cc: dietmar.eggemann, elif.topuz, kayracizmeci, kprateek.nayak,
	linux-kernel, mingo, peterz, sh, vincent.guittot

Sorry for being a bit late :-(

> @@ -4941,6 +4941,9 @@ static void reweight_eevdf(struct cfs_rq *cfs_rq, struct sched_entity *se,
>  		se->deadline += avruntime;
>  		se->rel_deadline = 0;
>  		se->vruntime = avruntime - se->vlag;
> +		/* Reweighting must not revive expired slice protection. */
> +		if (curr && !rel_vprot)
> +			se->vprot = se->vruntime;
>  
>  		if (!curr)
 			__enqueue_entity(cfs_rq, se);
> @@ -8204,7 +8207,7 @@ enqueue_task_fair(struct rq *rq, struct task_struct *p, int flags)
>  	struct sched_entity *se = &p->se;
>  	struct cfs_rq *cfs_rq = &rq->cfs;
>  	unsigned long weight;
> -	bool curr;
> +	bool curr, expired = false;
>  
>  	if (task_is_throttled(p) && enqueue_throttled_task(p))
>  		return;
> @@ -8237,6 +8240,8 @@ enqueue_task_fair(struct rq *rq, struct task_struct *p, int flags)
>  	 * XXX comment on the curr thing
>  	 */
>  	curr = (cfs_rq->curr == se);
> +	if (!curr && task_current_donor(rq, p))
> +		expired = !protect_slice(se);
>   	if (curr)
>  		place_entity(cfs_rq, se, flags);
> 

> @@ -8248,6 +8253,8 @@ enqueue_task_fair(struct rq *rq, struct task_struct *p, int flags)
>  	if (!curr) {
>  		reweight_eevdf(cfs_rq, se, weight, false);
>  		place_entity(cfs_rq, se, flags | ENQUEUE_QUEUED);
> +		if (expired)
> +			se->vprot = se->vruntime;
>  		__enqueue_entity(cfs_rq, se);
>  	}
>  
> -- 

There are a few thing here:

This (!curr) branch has been removed from tip sched/core, here: https://patch.msgid.link/20260911155449.1249726-1-kayracizmeci@gmail.com
So, heads-up.

Also, when we have alive (or live, as y'all call it, I think it's better this way tho :>) 
protection, and (!expired) reweight_eevdf() doesn't enter the on_rq path and
alive protection is not reweighted. Not tested tho. And I may be getting something wrong.
This is not about this patch tho.

Your commit says "Commit ff38424030f9 ("sched/eevdf: Update
se->vprot in reweight_entity()") subsequently handled live protection,
but left the expired case unchanged." too. So it's gotta change if I'm not missing something.

If I'm not tho that can mean a patch. :-). 

But IDK if that's intentional or not :-(. I think it's not but I'm not sure.

Anyway,

Thanks,

Regards,

Kayra :>



^ permalink raw reply	[flat|nested] 12+ messages in thread

* Re: [PATCH v2 4/5] sched/eevdf: Keep expired protection expired across reweighting
  2026-10-01 16:46   ` Kayra Cizmeci
@ 2026-10-01 21:51     ` Christian Loehle
  0 siblings, 0 replies; 12+ messages in thread
From: Christian Loehle @ 2026-10-01 21:51 UTC (permalink / raw)
  To: Kayra Cizmeci
  Cc: dietmar.eggemann, elif.topuz, kprateek.nayak, linux-kernel,
	mingo, peterz, sh, vincent.guittot

On 10/1/26 17:46, Kayra Cizmeci wrote:
> Sorry for being a bit late :-(
> 
>> @@ -4941,6 +4941,9 @@ static void reweight_eevdf(struct cfs_rq *cfs_rq, struct sched_entity *se,
>>  		se->deadline += avruntime;
>>  		se->rel_deadline = 0;
>>  		se->vruntime = avruntime - se->vlag;
>> +		/* Reweighting must not revive expired slice protection. */
>> +		if (curr && !rel_vprot)
>> +			se->vprot = se->vruntime;
>>  
>>  		if (!curr)
>  			__enqueue_entity(cfs_rq, se);
>> @@ -8204,7 +8207,7 @@ enqueue_task_fair(struct rq *rq, struct task_struct *p, int flags)
>>  	struct sched_entity *se = &p->se;
>>  	struct cfs_rq *cfs_rq = &rq->cfs;
>>  	unsigned long weight;
>> -	bool curr;
>> +	bool curr, expired = false;
>>  
>>  	if (task_is_throttled(p) && enqueue_throttled_task(p))
>>  		return;
>> @@ -8237,6 +8240,8 @@ enqueue_task_fair(struct rq *rq, struct task_struct *p, int flags)
>>  	 * XXX comment on the curr thing
>>  	 */
>>  	curr = (cfs_rq->curr == se);
>> +	if (!curr && task_current_donor(rq, p))
>> +		expired = !protect_slice(se);
>>   	if (curr)
>>  		place_entity(cfs_rq, se, flags);
>>
> 
>> @@ -8248,6 +8253,8 @@ enqueue_task_fair(struct rq *rq, struct task_struct *p, int flags)
>>  	if (!curr) {
>>  		reweight_eevdf(cfs_rq, se, weight, false);
>>  		place_entity(cfs_rq, se, flags | ENQUEUE_QUEUED);
>> +		if (expired)
>> +			se->vprot = se->vruntime;
>>  		__enqueue_entity(cfs_rq, se);
>>  	}
>>  
>> -- 
> 
> There are a few thing here:
> 
> This (!curr) branch has been removed from tip sched/core, here: https://patch.msgid.link/20260911155449.1249726-1-kayracizmeci@gmail.com
> So, heads-up.

Duh yeah, and the fixes I needed are also on sched/core by now, so
no more reason for this awkward base.

> 
> Also, when we have alive (or live, as y'all call it, I think it's better this way tho :>) 
> protection, and (!expired) reweight_eevdf() doesn't enter the on_rq path and
> alive protection is not reweighted. Not tested tho. And I may be getting something wrong.
> This is not about this patch tho.

No, you're correct, do you wanna fold mine into your fix and send that out?
(I think it looks better squashed but feel free to just pick mine up as your 1/2
if you disagree)
With 1/5 and 3+5/5 dropped there's only the refactor remaining and I might
as well send that as standalone.

> 
> Your commit says "Commit ff38424030f9 ("sched/eevdf: Update
> se->vprot in reweight_entity()") subsequently handled live protection,
> but left the expired case unchanged." too. So it's gotta change if I'm not missing something.

Correct!
Thanks!

^ permalink raw reply	[flat|nested] 12+ messages in thread

end of thread, other threads:[~2026-10-01 21:52 UTC | newest]

Thread overview: 12+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-01 13:40 [PATCH v2 0/5] sched/eevdf: Keep slice protection boundaries consistent Christian Loehle
2026-10-01 13:40 ` [PATCH v2 1/5] sched/fair: Take slice protection into account when arming HRTICK Christian Loehle
2026-10-01 13:43   ` Peter Zijlstra
2026-10-01 14:01     ` Christian Loehle
2026-10-01 13:40 ` [PATCH v2 2/5] sched/eevdf: Share the slice protection calculation Christian Loehle
2026-10-01 13:40 ` [PATCH v2 3/5] sched/eevdf: Cap protection when current has the shortest slice Christian Loehle
2026-10-01 13:55   ` Peter Zijlstra
2026-10-01 13:57     ` Christian Loehle
2026-10-01 13:40 ` [PATCH v2 4/5] sched/eevdf: Keep expired protection expired across reweighting Christian Loehle
2026-10-01 16:46   ` Kayra Cizmeci
2026-10-01 21:51     ` Christian Loehle
2026-10-01 13:40 ` [PATCH v2 5/5] sched/eevdf: Update protection after restoring current Christian Loehle

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®