From: Alex Shi <alex.shi@intel.com>
To: rob@landley.net, mingo@redhat.com, peterz@infradead.org
Cc: gregkh@linuxfoundation.org, andre.przywara@amd.com, rjw@sisk.pl,
paul.gortmaker@windriver.com, akpm@linux-foundation.org,
paulmck@linux.vnet.ibm.com, linux-kernel@vger.kernel.org,
pjt@google.com, vincent.guittot@linaro.org
Subject: [PATCH 17/18] sched: power aware load balance,
Date: Mon, 10 Dec 2012 16:22:33 +0800 [thread overview]
Message-ID: <1355127754-8444-18-git-send-email-alex.shi@intel.com> (raw)
In-Reply-To: <1355127754-8444-1-git-send-email-alex.shi@intel.com>
This patch enabled the power aware consideration in load balance.
As mentioned in the power aware scheduler proposal, Power aware
scheduling has 2 assumptions:
1, race to idle is helpful for power saving
2, shrink tasks on less sched_groups will reduce power consumption
The first assumption make performance policy take over scheduling when
system busy.
The second assumption make power aware scheduling try to move
disperse tasks into fewer groups until that groups are full of tasks.
This patch reuse some of Suresh's power saving load balance code.
The enabling logical summary here:
1, Collect power aware scheduler statistics during performance load
balance statistics collection.
2, if the balance cpu is eligible for power load balance, just do it
and forget performance load balance. but if the domain is suitable for
power balance, while the cpu is not appropriate, stop both
power/performance balance, else do performance load balance.
A test can show the effort on different policy:
for ((i = 0; i < I; i++)) ; do while true; do :; done & done
On my SNB laptop with 4core* HT: the data is Watts
powersaving balance performance
i = 2 40 54 54
i = 4 57 64* 68
i = 8 68 68 68
Note:
When i = 4 with balance policy, the power may change in 57~68Watt,
since the HT capacity and core capacity are both 1.
on SNB EP machine with 2 sockets * 8 cores * HT:
powersaving balance performance
i = 4 190 201 238
i = 8 205 241 268
i = 16 271 348 376
If system has few continued tasks, use power policy can get
the performance/power gain. Like sysbench fileio randrw test with 16
thread on the SNB EP box,
Signed-off-by: Alex Shi <alex.shi@intel.com>
---
kernel/sched/fair.c | 128 +++++++++++++++++++++++++++++++++++++++++++++++++-
1 files changed, 125 insertions(+), 3 deletions(-)
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 27630ae..e2ba22f 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -3124,6 +3124,7 @@ struct sd_lb_stats {
unsigned int sd_utils; /* sum utilizations of this domain */
unsigned long sd_capacity; /* capacity of this domain */
struct sched_group *group_leader; /* Group which relieves group_min */
+ struct sched_group *group_min; /* Least loaded group in sd */
unsigned long min_load_per_task; /* load_per_task in group_min */
unsigned int leader_util; /* sum utilizations of group_leader */
unsigned int min_util; /* sum utilizations of group_min */
@@ -4125,6 +4126,111 @@ static unsigned long task_h_load(struct task_struct *p)
#endif
/********** Helpers for find_busiest_group ************************/
+
+/**
+ * init_sd_lb_power_stats - Initialize power savings statistics for
+ * the given sched_domain, during load balancing.
+ *
+ * @env: The load balancing environment.
+ * @sds: Variable containing the statistics for sd.
+ */
+static inline void init_sd_lb_power_stats(struct lb_env *env,
+ struct sd_lb_stats *sds)
+{
+ if (sched_policy == SCHED_POLICY_PERFORMANCE ||
+ env->idle == CPU_NOT_IDLE) {
+ env->power_lb = 0;
+ env->perf_lb = 1;
+ return;
+ }
+ env->perf_lb = 0;
+ env->power_lb = 1;
+ sds->min_util= UINT_MAX;
+ sds->leader_util= 0;
+}
+
+/**
+ * update_sd_lb_power_stats - Update the power saving stats for a
+ * sched_domain while performing load balancing.
+ *
+ * @env: The load balancing environment.
+ * @group: sched_group belonging to the sched_domain under consideration.
+ * @sds: Variable containing the statistics of the sched_domain
+ * @local_group: Does group contain the CPU for which we're performing
+ * load balancing?
+ * @sgs: Variable containing the statistics of the group.
+ */
+static inline void update_sd_lb_power_stats(struct lb_env *env,
+ struct sched_group *group, struct sd_lb_stats *sds,
+ int local_group, struct sg_lb_stats *sgs)
+{
+ unsigned long threshold, threshold_util;
+
+ if (!env->power_lb)
+ return;
+
+ if (sched_policy == SCHED_POLICY_POWERSAVING)
+ threshold = sgs->group_weight;
+ else
+ threshold = sgs->group_capacity;
+ threshold_util = threshold * FULL_UTIL;
+
+ /*
+ * If the local group is idle or full loaded
+ * no need to do power savings balance at this domain
+ */
+ if (local_group && (!sgs->sum_nr_running ||
+ (sgs->sum_nr_running == threshold &&
+ sgs->group_utils >= threshold_util)))
+ env->power_lb = 0;
+
+ /*
+ * Do performance load balance if any group overload or maybe
+ * potentially overload.
+ */
+ if (sgs->group_utils > threshold * 100 ||
+ sgs->sum_nr_running > threshold) {
+ env->perf_lb = 1;
+ env->power_lb = 0;
+ }
+
+ /*
+ * If a group is idle,
+ * don't include that group in power savings calculations
+ */
+ if (!env->power_lb || !sgs->sum_nr_running)
+ return;
+
+ /*
+ * Calculate the group which has the least non-idle load.
+ * This is the group from where we need to pick up the load
+ * for saving power
+ */
+ if ((sgs->group_utils < sds->min_util) ||
+ (sgs->group_utils == sds->min_util &&
+ group_first_cpu(group) > group_first_cpu(sds->group_min))) {
+ sds->group_min = group;
+ sds->min_util = sgs->group_utils;
+ sds->min_load_per_task = sgs->sum_weighted_load /
+ sgs->sum_nr_running;
+ }
+
+ /*
+ * Calculate the group which is almost near its
+ * capacity but still has some space to pick up some load
+ * from other group and save more power
+ */
+ if (sgs->group_utils + FULL_UTIL > threshold * 100)
+ return;
+
+ if (sgs->group_utils > sds->leader_util ||
+ (sgs->group_utils == sds->leader_util && sds->group_leader &&
+ group_first_cpu(group) < group_first_cpu(sds->group_leader))) {
+ sds->group_leader = group;
+ sds->leader_util = sgs->group_utils;
+ }
+}
+
/**
* get_sd_load_idx - Obtain the load index for a given sched domain.
* @sd: The sched_domain whose load_idx is to be obtained.
@@ -4364,6 +4470,8 @@ static inline void update_sg_lb_stats(struct lb_env *env,
sgs->group_load += load;
sgs->sum_nr_running += nr_running;
sgs->sum_weighted_load += weighted_cpuload(i);
+ sgs->group_utils += rq->util;
+
if (idle_cpu(i))
sgs->idle_cpus++;
}
@@ -4472,6 +4580,7 @@ static inline void update_sd_lb_stats(struct lb_env *env,
if (child && child->flags & SD_PREFER_SIBLING)
prefer_sibling = 1;
+ init_sd_lb_power_stats(env, sds);
load_idx = get_sd_load_idx(env->sd, env->idle);
do {
@@ -4523,6 +4632,7 @@ static inline void update_sd_lb_stats(struct lb_env *env,
sds->group_imb = sgs.group_imb;
}
+ update_sd_lb_power_stats(env, sg, sds, local_group, &sgs);
sg = sg->next;
} while (sg != env->sd->groups);
}
@@ -4740,6 +4850,19 @@ find_busiest_group(struct lb_env *env, int *balance)
*/
update_sd_lb_stats(env, balance, &sds);
+ if (!env->perf_lb && !env->power_lb)
+ return NULL;
+
+ if (env->power_lb) {
+ if (sds.this == sds.group_leader &&
+ sds.group_leader != sds.group_min) {
+ env->imbalance = sds.min_load_per_task;
+ return sds.group_min;
+ }
+ env->power_lb = 0;
+ return NULL;
+ }
+
/*
* this_cpu is not the appropriate cpu to perform load balancing at
* this level.
@@ -4917,8 +5040,8 @@ static int load_balance(int this_cpu, struct rq *this_rq,
.idle = idle,
.loop_break = sched_nr_migrate_break,
.cpus = cpus,
- .power_lb = 0,
- .perf_lb = 1,
+ .power_lb = 1,
+ .perf_lb = 0,
};
cpumask_copy(cpus, cpu_active_mask);
@@ -5996,7 +6119,6 @@ void unregister_fair_sched_group(struct task_group *tg, int cpu) { }
#endif /* CONFIG_FAIR_GROUP_SCHED */
-
static unsigned int get_rr_interval_fair(struct rq *rq, struct task_struct *task)
{
struct sched_entity *se = &task->se;
--
1.7.5.1
next prev parent reply other threads:[~2012-12-10 8:26 UTC|newest]
Thread overview: 56+ messages / expand[flat|nested] mbox.gz Atom feed top
2012-12-10 8:22 [PATCH 0/18] sched: simplified fork, enable load average into LB and power awareness scheduling Alex Shi
2012-12-10 8:22 ` [PATCH 01/18] sched: select_task_rq_fair clean up Alex Shi
2012-12-11 4:23 ` Preeti U Murthy
2012-12-11 5:28 ` Alex Shi
2012-12-11 6:30 ` Preeti U Murthy
2012-12-11 11:53 ` Alex Shi
2012-12-12 5:26 ` Preeti U Murthy
2012-12-21 4:28 ` Namhyung Kim
2012-12-23 12:17 ` Alex Shi
2012-12-10 8:22 ` [PATCH 02/18] sched: fix find_idlest_group mess logical Alex Shi
2012-12-11 5:08 ` Preeti U Murthy
2012-12-11 5:29 ` Alex Shi
2012-12-11 5:50 ` Preeti U Murthy
2012-12-11 11:55 ` Alex Shi
2012-12-10 8:22 ` [PATCH 03/18] sched: don't need go to smaller sched domain Alex Shi
2012-12-10 8:22 ` [PATCH 04/18] sched: remove domain iterations in fork/exec/wake Alex Shi
2012-12-10 8:22 ` [PATCH 05/18] sched: load tracking bug fix Alex Shi
2012-12-10 8:22 ` [PATCH 06/18] sched: set initial load avg of new forked task as its load weight Alex Shi
2012-12-21 4:33 ` Namhyung Kim
2012-12-23 12:00 ` Alex Shi
2012-12-10 8:22 ` [PATCH 07/18] sched: compute runnable load avg in cpu_load and cpu_avg_load_per_task Alex Shi
2012-12-12 3:57 ` Preeti U Murthy
2012-12-12 5:52 ` Alex Shi
2012-12-13 8:45 ` Alex Shi
2012-12-21 4:35 ` Namhyung Kim
2012-12-23 11:42 ` Alex Shi
2012-12-10 8:22 ` [PATCH 08/18] sched: consider runnable load average in move_tasks Alex Shi
2012-12-12 4:41 ` Preeti U Murthy
2012-12-12 6:26 ` Alex Shi
2012-12-21 4:43 ` Namhyung Kim
2012-12-23 12:29 ` Alex Shi
2012-12-10 8:22 ` [PATCH 09/18] Revert "sched: Introduce temporary FAIR_GROUP_SCHED dependency for load-tracking" Alex Shi
2012-12-10 8:22 ` [PATCH 10/18] sched: add sched_policy in kernel Alex Shi
2012-12-10 8:22 ` [PATCH 11/18] sched: add sched_policy and it's sysfs interface Alex Shi
2012-12-10 8:22 ` [PATCH 12/18] sched: log the cpu utilization at rq Alex Shi
2012-12-10 8:22 ` [PATCH 13/18] sched: add power aware scheduling in fork/exec/wake Alex Shi
2012-12-10 8:22 ` [PATCH 14/18] sched: add power/performance balance allowed flag Alex Shi
2012-12-10 8:22 ` [PATCH 15/18] sched: don't care if the local group has capacity Alex Shi
2012-12-10 8:22 ` [PATCH 16/18] sched: pull all tasks from source group Alex Shi
2012-12-10 8:22 ` Alex Shi [this message]
2012-12-10 8:22 ` [PATCH 18/18] sched: lazy powersaving balance Alex Shi
2012-12-11 0:51 ` [PATCH 0/18] sched: simplified fork, enable load average into LB and power awareness scheduling Alex Shi
2012-12-11 12:10 ` Alex Shi
2012-12-11 15:48 ` Borislav Petkov
2012-12-11 16:03 ` Arjan van de Ven
2012-12-11 16:13 ` Borislav Petkov
2012-12-11 16:40 ` Arjan van de Ven
2012-12-12 9:52 ` Amit Kucheria
2012-12-12 13:55 ` Alex Shi
2012-12-12 14:21 ` Vincent Guittot
2012-12-13 2:51 ` Alex Shi
2012-12-12 14:41 ` Borislav Petkov
2012-12-13 3:07 ` Alex Shi
2012-12-13 11:35 ` Borislav Petkov
2012-12-14 1:56 ` Alex Shi
2012-12-12 1:14 ` Alex Shi
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=1355127754-8444-18-git-send-email-alex.shi@intel.com \
--to=alex.shi@intel.com \
--cc=akpm@linux-foundation.org \
--cc=andre.przywara@amd.com \
--cc=gregkh@linuxfoundation.org \
--cc=linux-kernel@vger.kernel.org \
--cc=mingo@redhat.com \
--cc=paul.gortmaker@windriver.com \
--cc=paulmck@linux.vnet.ibm.com \
--cc=peterz@infradead.org \
--cc=pjt@google.com \
--cc=rjw@sisk.pl \
--cc=rob@landley.net \
--cc=vincent.guittot@linaro.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®