From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org X-Spam-Level: X-Spam-Status: No, score=-15.3 required=3.0 tests=BAYES_00, HEADER_FROM_DIFFERENT_DOMAINS,INCLUDES_CR_TRAILER,INCLUDES_PATCH, MAILING_LIST_MULTI,NICE_REPLY_A,SPF_HELO_NONE,SPF_PASS,URIBL_BLOCKED, USER_AGENT_SANE_1 autolearn=unavailable autolearn_force=no version=3.4.0 Received: from mail.kernel.org (mail.kernel.org [198.145.29.99]) by smtp.lore.kernel.org (Postfix) with ESMTP id BDD89C433DB for ; Wed, 10 Mar 2021 11:24:11 +0000 (UTC) Received: from vger.kernel.org (vger.kernel.org [23.128.96.18]) by mail.kernel.org (Postfix) with ESMTP id 7E18A64FD7 for ; Wed, 10 Mar 2021 11:24:11 +0000 (UTC) Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S231958AbhCJLWh (ORCPT ); Wed, 10 Mar 2021 06:22:37 -0500 Received: from foss.arm.com ([217.140.110.172]:44402 "EHLO foss.arm.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S232598AbhCJLWT (ORCPT ); Wed, 10 Mar 2021 06:22:19 -0500 Received: from usa-sjc-imap-foss1.foss.arm.com (unknown [10.121.207.14]) by usa-sjc-mx-foss1.foss.arm.com (Postfix) with ESMTP id 3AAA41FB; Wed, 10 Mar 2021 03:22:19 -0800 (PST) Received: from [10.57.15.210] (unknown [10.57.15.210]) by usa-sjc-imap-foss1.foss.arm.com (Postfix) with ESMTPSA id 577B83F85F; Wed, 10 Mar 2021 03:22:15 -0800 (PST) Subject: Re: [PATCH v3 5/5] powercap/drivers/dtpm: Scale the power with the load To: Daniel Lezcano Cc: rafael@kernel.org, linux-kernel@vger.kernel.org, linux-pm@vger.kernel.org References: <20210310110212.26512-1-daniel.lezcano@linaro.org> <20210310110212.26512-5-daniel.lezcano@linaro.org> From: Lukasz Luba Message-ID: Date: Wed, 10 Mar 2021 11:22:10 +0000 User-Agent: Mozilla/5.0 (X11; Linux x86_64; rv:60.0) Gecko/20100101 Thunderbird/60.9.0 MIME-Version: 1.0 In-Reply-To: <20210310110212.26512-5-daniel.lezcano@linaro.org> Content-Type: text/plain; charset=utf-8; format=flowed Content-Language: en-US Content-Transfer-Encoding: 7bit Precedence: bulk List-ID: X-Mailing-List: linux-kernel@vger.kernel.org On 3/10/21 11:02 AM, Daniel Lezcano wrote: > Currently the power consumption is based on the current OPP power > assuming the entire performance domain is fully loaded. > > That gives very gross power estimation and we can do much better by > using the load to scale the power consumption. > > Use the utilization to normalize and scale the power usage over the > max possible power. > > Tested on a rock960 with 2 big CPUS, the power consumption estimation > conforms with the expected one. > > Before this change: > > ~$ ~/dhrystone -t 1 -l 10000& > ~$ cat /sys/devices/virtual/powercap/dtpm/dtpm:0/dtpm:0:1/constraint_0_max_power_uw > 2260000 > > After this change: > > ~$ ~/dhrystone -t 1 -l 10000& > ~$ cat /sys/devices/virtual/powercap/dtpm/dtpm:0/dtpm:0:1/constraint_0_max_power_uw > 1130000 > > ~$ ~/dhrystone -t 2 -l 10000& > ~$ cat /sys/devices/virtual/powercap/dtpm/dtpm:0/dtpm:0:1/constraint_0_max_power_uw > 2260000 > > Signed-off-by: Daniel Lezcano > --- > > V3: > - Fixed uninitialized 'cpu' in scaled_power_uw() > V2: > - Replaced cpumask by em_span_cpus > - Changed 'util' metrics variable types > - Optimized utilization scaling power computation > - Renamed parameter name for scale_pd_power_uw() > --- > drivers/powercap/dtpm_cpu.c | 46 +++++++++++++++++++++++++++++++------ > 1 file changed, 39 insertions(+), 7 deletions(-) > > diff --git a/drivers/powercap/dtpm_cpu.c b/drivers/powercap/dtpm_cpu.c > index ac7f2e7e262f..47854923d958 100644 > --- a/drivers/powercap/dtpm_cpu.c > +++ b/drivers/powercap/dtpm_cpu.c > @@ -68,27 +68,59 @@ static u64 set_pd_power_limit(struct dtpm *dtpm, u64 power_limit) > return power_limit; > } > > +static u64 scale_pd_power_uw(struct cpumask *pd_mask, u64 power) > +{ > + unsigned long max = 0, sum_util = 0; > + int cpu; > + > + for_each_cpu_and(cpu, pd_mask, cpu_online_mask) { > + > + /* > + * The capacity is the same for all CPUs belonging to > + * the same perf domain, so a single call to > + * arch_scale_cpu_capacity() is enough. However, we > + * need the CPU parameter to be initialized by the > + * loop, so the call ends up in this block. > + * > + * We can initialize 'max' with a cpumask_first() call > + * before the loop but the bits computation is not > + * worth given the arch_scale_cpu_capacity() just > + * returns a value where the resulting assembly code > + * will be optimized by the compiler. > + */ > + max = arch_scale_cpu_capacity(cpu); > + sum_util += sched_cpu_util(cpu, max); > + } > + > + /* > + * In the improbable case where all the CPUs of the perf > + * domain are offline, 'max' will be zero and will lead to an > + * illegal operation with a zero division. > + */ > + return max ? (power * ((sum_util << 10) / max)) >> 10 : 0; > +} > + > static u64 get_pd_power_uw(struct dtpm *dtpm) > { > struct dtpm_cpu *dtpm_cpu = to_dtpm_cpu(dtpm); > struct em_perf_domain *pd; > - struct cpumask cpus; > + struct cpumask *pd_mask; > unsigned long freq; > - int i, nr_cpus; > + int i; > > pd = em_cpu_get(dtpm_cpu->cpu); > - freq = cpufreq_quick_get(dtpm_cpu->cpu); > > - cpumask_and(&cpus, cpu_online_mask, to_cpumask(pd->cpus)); > - nr_cpus = cpumask_weight(&cpus); > + pd_mask = em_span_cpus(pd); > + > + freq = cpufreq_quick_get(dtpm_cpu->cpu); > > for (i = 0; i < pd->nr_perf_states; i++) { > > if (pd->table[i].frequency < freq) > continue; > > - return pd->table[i].power * > - MICROWATT_PER_MILLIWATT * nr_cpus; > + return scale_pd_power_uw(pd_mask, pd->table[i].power * > + MICROWATT_PER_MILLIWATT); > } > > return 0; > LGTM Reviewed-by: Lukasz Luba