From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1754055AbcBVJeH (ORCPT ); Mon, 22 Feb 2016 04:34:07 -0500 Received: from mga01.intel.com ([192.55.52.88]:10087 "EHLO mga01.intel.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1753463AbcBVJeB (ORCPT ); Mon, 22 Feb 2016 04:34:01 -0500 X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="5.22,483,1449561600"; d="scan'208";a="750815971" Message-ID: <1456133633.3659.4.camel@linux.intel.com> Subject: Re: [PATCH] intel-pstate: Update frequencies of policy->cpus only from ->set_policy() From: Joonas Lahtinen To: Viresh Kumar , Rafael Wysocki , Srinivas Pandruvada , Len Brown , Daniel Vetter Cc: linaro-kernel@lists.linaro.org, linux-pm@vger.kernel.org, linux-kernel@vger.kernel.org, intel-gfx@lists.freedesktop.org Date: Mon, 22 Feb 2016 11:33:53 +0200 In-Reply-To: References: Organization: Intel Finland Oy - BIC 0357606-4 - Westendinkatu 7, 02160 Espoo Content-Type: text/plain; charset="UTF-8" X-Mailer: Evolution 3.18.4 (3.18.4-1.fc23) Mime-Version: 1.0 Content-Transfer-Encoding: 8bit Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org Hi, This fixes the issue for my machine, we'll try in our CI system, too. CC'd Daniel for that. By R-b and T-b below. On ma, 2016-02-22 at 10:27 +0530, Viresh Kumar wrote: > The intel-pstate driver is using intel_pstate_hwp_set() from two > separate paths, i.e. ->set_policy() callback and sysfs update path for > the files present in /sys/devices/system/cpu/intel_pstate/ directory. > > While an update to the sysfs path applies to all the CPUs being managed > by the driver (which essentially means all the online CPUs), the update > via the ->set_policy() callback applies to a smaller group of CPUs > managed by the policy for which ->set_policy() is called. > > And so, intel_pstate_hwp_set() should update frequencies of only the > CPUs that are part of policy->cpus mask, while it is called from > ->set_policy() callback. > > In order to do that, add a parameter (cpumask) to intel_pstate_hwp_set() > and apply the frequency changes only to the concerned CPUs. > > For ->set_policy() path, we are only concerned about policy->cpus, and > so policy->rwsem lock taken by the core prior to calling ->set_policy() > is enough to take care of any races. The larger lock acquired by > get_online_cpus() is required only for the updates to sysfs files. > > Add another routine, intel_pstate_hwp_set_online_cpus(), and call it > from the sysfs update paths. > > This also fixes a lockdep reported recently, where policy->rwsem and > get_online_cpus() could have been acquired in any order causing an ABBA > deadlock. The sequence of events leading to that was: > > intel_pstate_init(...) > ...cpufreq_online(...) > down_write(&policy->rwsem); // Locks policy->rwsem > ... > cpufreq_init_policy(policy); > ...intel_pstate_hwp_set(); > get_online_cpus(); // Temporarily locks cpu_hotplug.lock > ... > up_write(&policy->rwsem); > > pm_suspend(...) > ...disable_nonboot_cpus() > _cpu_down() > cpu_hotplug_begin(); // Locks cpu_hotplug.lock > __cpu_notify(CPU_DOWN_PREPARE, ...); > ...cpufreq_offline_prepare(); > down_write(&policy->rwsem); // Locks policy->rwsem > > Reported-by: Joonas Lahtinen Tested-by: Joonas Lahtinen Reviewed-by: Joonas Lahtinen > Signed-off-by: Viresh Kumar > --- >  drivers/cpufreq/intel_pstate.c | 21 ++++++++++++--------- >  1 file changed, 12 insertions(+), 9 deletions(-) > > diff --git a/drivers/cpufreq/intel_pstate.c b/drivers/cpufreq/intel_pstate.c > index f4d85c2ae7b1..2e7058a2479d 100644 > --- a/drivers/cpufreq/intel_pstate.c > +++ b/drivers/cpufreq/intel_pstate.c > @@ -287,7 +287,7 @@ static inline void update_turbo_state(void) >    cpu->pstate.max_pstate == cpu->pstate.turbo_pstate); >  } >   > -static void intel_pstate_hwp_set(void) > +static void intel_pstate_hwp_set(const struct cpumask *cpumask) >  { >   int min, hw_min, max, hw_max, cpu, range, adj_range; >   u64 value, cap; > @@ -297,9 +297,7 @@ static void intel_pstate_hwp_set(void) >   hw_max = HWP_HIGHEST_PERF(cap); >   range = hw_max - hw_min; >   > - get_online_cpus(); > - > - for_each_online_cpu(cpu) { > + for_each_cpu(cpu, cpumask) { >   rdmsrl_on_cpu(cpu, MSR_HWP_REQUEST, &value); >   adj_range = limits->min_perf_pct * range / 100; >   min = hw_min + adj_range; > @@ -318,7 +316,12 @@ static void intel_pstate_hwp_set(void) >   value |= HWP_MAX_PERF(max); >   wrmsrl_on_cpu(cpu, MSR_HWP_REQUEST, value); >   } > +} >   > +static void intel_pstate_hwp_set_online_cpus(void) > +{ > + get_online_cpus(); > + intel_pstate_hwp_set(cpu_online_mask); >   put_online_cpus(); >  } >   > @@ -440,7 +443,7 @@ static ssize_t store_no_turbo(struct kobject *a, struct attribute *b, >   limits->no_turbo = clamp_t(int, input, 0, 1); >   >   if (hwp_active) > - intel_pstate_hwp_set(); > + intel_pstate_hwp_set_online_cpus(); >   >   return count; >  } > @@ -466,7 +469,7 @@ static ssize_t store_max_perf_pct(struct kobject *a, struct attribute *b, >     int_tofp(100)); >   >   if (hwp_active) > - intel_pstate_hwp_set(); > + intel_pstate_hwp_set_online_cpus(); >   return count; >  } >   > @@ -491,7 +494,7 @@ static ssize_t store_min_perf_pct(struct kobject *a, struct attribute *b, >     int_tofp(100)); >   >   if (hwp_active) > - intel_pstate_hwp_set(); > + intel_pstate_hwp_set_online_cpus(); >   return count; >  } >   > @@ -1112,7 +1115,7 @@ static int intel_pstate_set_policy(struct cpufreq_policy *policy) >   pr_debug("intel_pstate: set performance\n"); >   limits = &performance_limits; >   if (hwp_active) > - intel_pstate_hwp_set(); > + intel_pstate_hwp_set(policy->cpus); >   return 0; >   } >   > @@ -1144,7 +1147,7 @@ static int intel_pstate_set_policy(struct cpufreq_policy *policy) >     int_tofp(100)); >   >   if (hwp_active) > - intel_pstate_hwp_set(); > + intel_pstate_hwp_set(policy->cpus); >   >   return 0; >  } -- Joonas Lahtinen Open Source Technology Center Intel Corporation