* Re: [PATCH v2 2/6] sched/rt: Optimize select_task_rq_rt() for non-RT curr task
@ 2014-11-04 12:21 Hillf Danton
0 siblings, 0 replies; 6+ messages in thread
From: Hillf Danton @ 2014-11-04 12:21 UTC (permalink / raw)
To: 'pang.xunlei'
Cc: linux-kernel, 'Ingo Molnar', 'Peter Zijlstra',
'Steven Rostedt', 'Juri Lelli'
>
> When selecting the cpu for a waking RT task, if curr is a non-RT
> task which is bound only on this cpu, then we can give it a chance
> to select a different cpu(definitely an idle cpu if existing) for
> the RT task to avoid curr starving.
>
> Signed-off-by: pang.xunlei <pang.xunlei@linaro.org>
> ---
> kernel/sched/rt.c | 10 +++++++---
> 1 file changed, 7 insertions(+), 3 deletions(-)
>
> diff --git a/kernel/sched/rt.c b/kernel/sched/rt.c
> index da6922e..dc1f7f0 100644
> --- a/kernel/sched/rt.c
> +++ b/kernel/sched/rt.c
> @@ -1340,6 +1340,11 @@ select_task_rq_rt(struct task_struct *p, int cpu, int sd_flag, int flags)
> * runqueue. Otherwise simply start this RT task
> * on its current runqueue.
> *
> + * If the current task on @p's runqueue is a non-RT task,
> + * and this task is bound on current runqueue, then try to
> + * see if we can wake this RT task up on a different runqueue,
> + * we will definitely find an idle cpu if there is any.
> + *
> * We want to avoid overloading runqueues. If the woken
> * task is a higher priority, then it will stay on this CPU
> * and the lower prio task should be moved to another CPU.
> @@ -1356,9 +1361,8 @@ select_task_rq_rt(struct task_struct *p, int cpu, int sd_flag, int flags)
> * This test is optimistic, if we get it wrong the load-balancer
> * will have to sort it out.
> */
> - if (curr && unlikely(rt_task(curr)) &&
> - (curr->nr_cpus_allowed < 2 ||
> - curr->prio <= p->prio)) {
> + if (curr && unlikely(curr->nr_cpus_allowed < 2 ||
> + curr->prio <= p->prio)) {
Nack, it is no meaning to compare apple against orange.
Hillf
> int target = find_lowest_rq(p);
>
> if (target != -1)
> --
> 1.7.9.5
^ permalink raw reply [flat|nested] 6+ messages in thread* [PATCH v2 1/6] sched/cpupri: Deal with cpupri.pri_to_cpu[CPUPRI_IDLE] for idle cases
@ 2014-11-04 11:13 pang.xunlei
2014-11-04 11:13 ` [PATCH v2 2/6] sched/rt: Optimize select_task_rq_rt() for non-RT curr task pang.xunlei
0 siblings, 1 reply; 6+ messages in thread
From: pang.xunlei @ 2014-11-04 11:13 UTC (permalink / raw)
To: linux-kernel
Cc: Ingo Molnar, Peter Zijlstra, Steven Rostedt, Juri Lelli, pang.xunlei
When a runqueue runs out of RT tasks, it may have non-RT tasks or
none tasks(idle). Currently, RT balance treats the two cases equally
and manipulates cpupri.pri_to_cpu[CPUPRI_NORMAL] only which may cause
problems.
For instance, 4 cpus system, non-RT task1 is running on cpu0, RT
task2 is running on cpu3, cpu1/cpu2 both are idle. Then RT task3
(usually CPU-intensive) is waken up or created on cpu3, it will
be placed to cpu0 (see find_lowest_rq()) causing task1 starving
until cfs load balance places task1 to another cpu, or even worse
if task1 is bound on cpu0. So, it would be reasonable to put task3
to cpu1 or cpu2 which is idle(even though doing this may break the
energy-saving idle state).
This patch tackles the problem by operating pri_to_cpu[CPUPRI_IDLE]
of cpupri according to the stages of idle task, so that when pushing
or selecting RT tasks through find_lowest_rq(), it will try to find
one idle cpu as the goal.
Signed-off-by: pang.xunlei <pang.xunlei@linaro.org>
---
kernel/sched/idle_task.c | 3 +++
kernel/sched/rt.c | 21 +++++++++++++++++++++
kernel/sched/sched.h | 6 ++++++
3 files changed, 30 insertions(+)
diff --git a/kernel/sched/idle_task.c b/kernel/sched/idle_task.c
index 67ad4e7..e053347 100644
--- a/kernel/sched/idle_task.c
+++ b/kernel/sched/idle_task.c
@@ -26,6 +26,8 @@ static void check_preempt_curr_idle(struct rq *rq, struct task_struct *p, int fl
static struct task_struct *
pick_next_task_idle(struct rq *rq, struct task_struct *prev)
{
+ idle_enter_rt(rq);
+
put_prev_task(rq, prev);
schedstat_inc(rq, sched_goidle);
@@ -47,6 +49,7 @@ dequeue_task_idle(struct rq *rq, struct task_struct *p, int flags)
static void put_prev_task_idle(struct rq *rq, struct task_struct *prev)
{
+ idle_exit_rt(rq);
idle_exit_fair(rq);
rq_last_tick_reset(rq);
}
diff --git a/kernel/sched/rt.c b/kernel/sched/rt.c
index d024e6c..da6922e 100644
--- a/kernel/sched/rt.c
+++ b/kernel/sched/rt.c
@@ -992,6 +992,27 @@ enqueue_top_rt_rq(struct rt_rq *rt_rq)
#if defined CONFIG_SMP
+/* Set CPUPRI_IDLE bitmap for this cpu when entering idle. */
+void idle_enter_rt(struct rq *this_rq)
+{
+ struct cpupri *cp = &this_rq->rd->cpupri;
+ int currpri = cp->cpu_to_pri[this_rq->cpu];
+
+ BUG_ON(currpri != CPUPRI_NORMAL);
+ cpupri_set(cp, this_rq->cpu, MAX_PRIO);
+}
+
+/* Set CPUPRI_NORMAL bitmap for this cpu when exiting from idle. */
+void idle_exit_rt(struct rq *this_rq)
+{
+ struct cpupri *cp = &this_rq->rd->cpupri;
+ int currpri = cp->cpu_to_pri[this_rq->cpu];
+
+ /* RT tasks may be queued before, this judgement is needed. */
+ if (currpri == CPUPRI_IDLE)
+ cpupri_set(cp, this_rq->cpu, MAX_RT_PRIO);
+}
+
static void
inc_rt_prio_smp(struct rt_rq *rt_rq, int prio, int prev_prio)
{
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index 24156c84..cc603fa 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -1162,11 +1162,17 @@ extern void update_group_capacity(struct sched_domain *sd, int cpu);
extern void trigger_load_balance(struct rq *rq);
+extern void idle_enter_rt(struct rq *this_rq);
+extern void idle_exit_rt(struct rq *this_rq);
+
extern void idle_enter_fair(struct rq *this_rq);
extern void idle_exit_fair(struct rq *this_rq);
#else
+static inline void idle_enter_rt(struct rq *rq) { }
+static inline void idle_exit_rt(struct rq *rq) { }
+
static inline void idle_enter_fair(struct rq *rq) { }
static inline void idle_exit_fair(struct rq *rq) { }
--
1.7.9.5
^ permalink raw reply [flat|nested] 6+ messages in thread* [PATCH v2 2/6] sched/rt: Optimize select_task_rq_rt() for non-RT curr task
2014-11-04 11:13 [PATCH v2 1/6] sched/cpupri: Deal with cpupri.pri_to_cpu[CPUPRI_IDLE] for idle cases pang.xunlei
@ 2014-11-04 11:13 ` pang.xunlei
2014-11-04 12:52 ` Steven Rostedt
0 siblings, 1 reply; 6+ messages in thread
From: pang.xunlei @ 2014-11-04 11:13 UTC (permalink / raw)
To: linux-kernel
Cc: Ingo Molnar, Peter Zijlstra, Steven Rostedt, Juri Lelli, pang.xunlei
When selecting the cpu for a waking RT task, if curr is a non-RT
task which is bound only on this cpu, then we can give it a chance
to select a different cpu(definitely an idle cpu if existing) for
the RT task to avoid curr starving.
Signed-off-by: pang.xunlei <pang.xunlei@linaro.org>
---
kernel/sched/rt.c | 10 +++++++---
1 file changed, 7 insertions(+), 3 deletions(-)
diff --git a/kernel/sched/rt.c b/kernel/sched/rt.c
index da6922e..dc1f7f0 100644
--- a/kernel/sched/rt.c
+++ b/kernel/sched/rt.c
@@ -1340,6 +1340,11 @@ select_task_rq_rt(struct task_struct *p, int cpu, int sd_flag, int flags)
* runqueue. Otherwise simply start this RT task
* on its current runqueue.
*
+ * If the current task on @p's runqueue is a non-RT task,
+ * and this task is bound on current runqueue, then try to
+ * see if we can wake this RT task up on a different runqueue,
+ * we will definitely find an idle cpu if there is any.
+ *
* We want to avoid overloading runqueues. If the woken
* task is a higher priority, then it will stay on this CPU
* and the lower prio task should be moved to another CPU.
@@ -1356,9 +1361,8 @@ select_task_rq_rt(struct task_struct *p, int cpu, int sd_flag, int flags)
* This test is optimistic, if we get it wrong the load-balancer
* will have to sort it out.
*/
- if (curr && unlikely(rt_task(curr)) &&
- (curr->nr_cpus_allowed < 2 ||
- curr->prio <= p->prio)) {
+ if (curr && unlikely(curr->nr_cpus_allowed < 2 ||
+ curr->prio <= p->prio)) {
int target = find_lowest_rq(p);
if (target != -1)
--
1.7.9.5
^ permalink raw reply [flat|nested] 6+ messages in thread* Re: [PATCH v2 2/6] sched/rt: Optimize select_task_rq_rt() for non-RT curr task
2014-11-04 11:13 ` [PATCH v2 2/6] sched/rt: Optimize select_task_rq_rt() for non-RT curr task pang.xunlei
@ 2014-11-04 12:52 ` Steven Rostedt
2014-11-04 14:29 ` pang.xunlei
0 siblings, 1 reply; 6+ messages in thread
From: Steven Rostedt @ 2014-11-04 12:52 UTC (permalink / raw)
To: pang.xunlei; +Cc: linux-kernel, Ingo Molnar, Peter Zijlstra, Juri Lelli
On Tue, 4 Nov 2014 19:13:01 +0800
"pang.xunlei" <pang.xunlei@linaro.org> wrote:
> When selecting the cpu for a waking RT task, if curr is a non-RT
> task which is bound only on this cpu, then we can give it a chance
> to select a different cpu(definitely an idle cpu if existing) for
> the RT task to avoid curr starving.
Absolutely not! An RT task doesn't give a crap if a non RT task is
bound to a CPU or not. We are not going to migrate an RT task to be
nice to a bounded non-RT task.
Migration is not cheap. It causes cache misses and TLB flushes. This is
not something that should be taken lightly.
Nack
-- Steve
^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [PATCH v2 2/6] sched/rt: Optimize select_task_rq_rt() for non-RT curr task
2014-11-04 12:52 ` Steven Rostedt
@ 2014-11-04 14:29 ` pang.xunlei
2014-11-04 14:47 ` Steven Rostedt
0 siblings, 1 reply; 6+ messages in thread
From: pang.xunlei @ 2014-11-04 14:29 UTC (permalink / raw)
To: Steven Rostedt; +Cc: lkml, Ingo Molnar, Peter Zijlstra, Juri Lelli
On 4 November 2014 20:52, Steven Rostedt <rostedt@goodmis.org> wrote:
> On Tue, 4 Nov 2014 19:13:01 +0800
> "pang.xunlei" <pang.xunlei@linaro.org> wrote:
>
>> When selecting the cpu for a waking RT task, if curr is a non-RT
>> task which is bound only on this cpu, then we can give it a chance
>> to select a different cpu(definitely an idle cpu if existing) for
>> the RT task to avoid curr starving.
>
> Absolutely not! An RT task doesn't give a crap if a non RT task is
> bound to a CPU or not. We are not going to migrate an RT task to be
> nice to a bounded non-RT task.
>
> Migration is not cheap. It causes cache misses and TLB flushes. This is
> not something that should be taken lightly.
Ok, thanks!
But I think the PUSH operation optimized by the former patch is reasonable,
since PUSH itselft does involve the Migration. Do I miss something?
>
> Nack
>
> -- Steve
>
^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [PATCH v2 2/6] sched/rt: Optimize select_task_rq_rt() for non-RT curr task
2014-11-04 14:29 ` pang.xunlei
@ 2014-11-04 14:47 ` Steven Rostedt
2014-11-04 15:09 ` pang.xunlei
0 siblings, 1 reply; 6+ messages in thread
From: Steven Rostedt @ 2014-11-04 14:47 UTC (permalink / raw)
To: pang.xunlei; +Cc: lkml, Ingo Molnar, Peter Zijlstra, Juri Lelli
On Tue, 4 Nov 2014 22:29:24 +0800
"pang.xunlei" <pang.xunlei@linaro.org> wrote:
> > Migration is not cheap. It causes cache misses and TLB flushes. This is
> > not something that should be taken lightly.
> Ok, thanks!
> But I think the PUSH operation optimized by the former patch is reasonable,
> since PUSH itselft does involve the Migration. Do I miss something?
For the first patch you may be right, but I want to think about it some
more. I want to make sure we are not adding any other type of overhead
with the extra calls.
-- Steve
^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [PATCH v2 2/6] sched/rt: Optimize select_task_rq_rt() for non-RT curr task
2014-11-04 14:47 ` Steven Rostedt
@ 2014-11-04 15:09 ` pang.xunlei
0 siblings, 0 replies; 6+ messages in thread
From: pang.xunlei @ 2014-11-04 15:09 UTC (permalink / raw)
To: Steven Rostedt; +Cc: lkml, Ingo Molnar, Peter Zijlstra, Juri Lelli
On 4 November 2014 22:47, Steven Rostedt <rostedt@goodmis.org> wrote:
> On Tue, 4 Nov 2014 22:29:24 +0800
> "pang.xunlei" <pang.xunlei@linaro.org> wrote:
>
>
>> > Migration is not cheap. It causes cache misses and TLB flushes. This is
>> > not something that should be taken lightly.
>> Ok, thanks!
>> But I think the PUSH operation optimized by the former patch is reasonable,
>> since PUSH itselft does involve the Migration. Do I miss something?
>
> For the first patch you may be right, but I want to think about it some
> more. I want to make sure we are not adding any other type of overhead
> with the extra calls.
Yes, this may cause some overhead/latency in idle especially its exit
stage, if that can't be accepted, I think it can also be done just in
find_lowest_rq() after cpupri_find(), we can modify cpupri_find() for
example to return a pri_to_cpu[] index plus one instead of 1, then if
the return index equals CPUPRI_NORMAL+1, then iterate the
"lowest_mask" with something like cpu_idle() judgement to select the
idle cpu.
>
> -- Steve
^ permalink raw reply [flat|nested] 6+ messages in thread
end of thread, other threads:[~2014-11-04 15:09 UTC | newest]
Thread overview: 6+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2014-11-04 12:21 [PATCH v2 2/6] sched/rt: Optimize select_task_rq_rt() for non-RT curr task Hillf Danton
-- strict thread matches above, loose matches on Subject: below --
2014-11-04 11:13 [PATCH v2 1/6] sched/cpupri: Deal with cpupri.pri_to_cpu[CPUPRI_IDLE] for idle cases pang.xunlei
2014-11-04 11:13 ` [PATCH v2 2/6] sched/rt: Optimize select_task_rq_rt() for non-RT curr task pang.xunlei
2014-11-04 12:52 ` Steven Rostedt
2014-11-04 14:29 ` pang.xunlei
2014-11-04 14:47 ` Steven Rostedt
2014-11-04 15:09 ` pang.xunlei
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®