* [PATCH 0/3] cpufreq: Resolve CPPC frequencies to performance levels
@ 2026-09-29 10:29 Christian Loehle
2026-09-29 10:29 ` [PATCH 1/3] cpufreq: Add a driver frequency resolution callback Christian Loehle
` (4 more replies)
0 siblings, 5 replies; 6+ messages in thread
From: Christian Loehle @ 2026-09-29 10:29 UTC (permalink / raw)
To: rafael
Cc: zhenglifeng1, zhongqiu.han, viresh.kumar, linux-kernel,
linux-acpi, linux-arm-kernel, linux-doc, Mario Limonciello,
K Prateek Nayak, Huang Rui, Perry Yuan, Gautham R . Shenoy,
Vanshidhar Konda, Shubhang Kaushik, Pierre Gondois,
Beata Michalska, Dietmar Eggemann, Ionela Voinescu, Sudeep Holla,
Lukasz Luba, Jeremy Linton, Peter Zijlstra, jonathanh, zhanjie9,
Vincent Guittot, Jonathan Corbet, Shuah Khan, Randy Dunlap,
Christian Loehle
For table-based cpufreq drivers, the core resolves requests to supported
frequency-table entries before governors such as schedutil compare them
with their cached target. Different kHz requests selecting the same entry
therefore need not reach the driver again.
Table-less cppc-cpufreq lacks this resolution: the core returns the requested
kHz value even if firmware exposes only a few CPPC performance levels.
Different requests can miss the governor's cache yet program the same
Desired Performance value. The tested ARM AGI CPU exposes only 61 levels,
with bounds visible at:
/sys/devices/system/cpu/cpuX/acpi_cppc/{lowest,highest}_perf
Add ->resolve_freq() so table-less ->target() drivers can canonicalize
requests without constructing a frequency table. CPPC resolves requests
before sugov_update_next_freq(), allowing its existing cache to skip
equivalent requests. Raw limit notifications still trigger
reconsideration, but only changed resolved limits force a driver update.
Failed limit writes remain pending for retry.
Testing on an ARM AGI CPU (61 distinct CPPC levels) with schedutil gives
the following end-to-end results across 16 iterations:
schbench -m 2 -t 31 -F 256 -n 5 -R 18000 -r 60 -w 20 -i 60
metric baseline resolve-freq change
median p99 3364 us 3280 us -2.5%
mean of run p99s 3664.2 us 3299.0 us -10.0%
worst run p99 4360 us 3500 us -19.7%
median throughput 16431.08 RPS 16460.69 RPS +0.2%
In an instrumented run of the same workload, cppc_set_perf() calls fell
from 1,136,868 to 866,703, a 23.8% reduction.
Christian Loehle (3):
cpufreq: Add a driver frequency resolution callback
cpufreq: CPPC: Resolve frequencies to performance levels
cpufreq: Skip updates for unchanged resolved limits
Documentation/admin-guide/pm/cpufreq.rst | 4 +
Documentation/cpu-freq/cpu-drivers.rst | 19 ++
drivers/cpufreq/cppc_cpufreq.c | 257 +++++++++++++++++++++--
drivers/cpufreq/cpufreq.c | 99 +++++++--
include/linux/cpufreq.h | 11 +
kernel/sched/cpufreq_schedutil.c | 19 +-
6 files changed, 359 insertions(+), 50 deletions(-)
base-commit: 6375e61c01e93e35ee7acd336a689ac1fae4b509
^ permalink raw reply [flat|nested] 6+ messages in thread
* [PATCH 1/3] cpufreq: Add a driver frequency resolution callback
2026-09-29 10:29 [PATCH 0/3] cpufreq: Resolve CPPC frequencies to performance levels Christian Loehle
@ 2026-09-29 10:29 ` Christian Loehle
2026-09-29 10:29 ` [PATCH 2/3] cpufreq: CPPC: Resolve frequencies to performance levels Christian Loehle
` (3 subsequent siblings)
4 siblings, 0 replies; 6+ messages in thread
From: Christian Loehle @ 2026-09-29 10:29 UTC (permalink / raw)
To: rafael
Cc: zhenglifeng1, zhongqiu.han, viresh.kumar, linux-kernel,
linux-acpi, linux-arm-kernel, linux-doc, Mario Limonciello,
K Prateek Nayak, Huang Rui, Perry Yuan, Gautham R . Shenoy,
Vanshidhar Konda, Shubhang Kaushik, Pierre Gondois,
Beata Michalska, Dietmar Eggemann, Ionela Voinescu, Sudeep Holla,
Lukasz Luba, Jeremy Linton, Peter Zijlstra, jonathanh, zhanjie9,
Vincent Guittot, Jonathan Corbet, Shuah Khan, Randy Dunlap,
Christian Loehle
Without a frequency table, cpufreq treats the policy range as continuous
even when the driver selects discrete performance levels. Governors can
then issue different kHz requests for the same driver setting.
Add ->resolve_freq() for table-less ->target() drivers to canonicalize
requests within the supplied limits using CPUFREQ_RELATION_{L,H,C}.
Document the callback and verification contracts so callers can safely
cache resolved requests.
Signed-off-by: Christian Loehle <christian.loehle@arm.com>
---
Documentation/admin-guide/pm/cpufreq.rst | 4 ++++
Documentation/cpu-freq/cpu-drivers.rst | 19 +++++++++++++++++++
drivers/cpufreq/cpufreq.c | 17 +++++++++++++++--
include/linux/cpufreq.h | 10 ++++++++++
4 files changed, 48 insertions(+), 2 deletions(-)
diff --git a/Documentation/admin-guide/pm/cpufreq.rst b/Documentation/admin-guide/pm/cpufreq.rst
index 34baf20cc202..e634b87a62a8 100644
--- a/Documentation/admin-guide/pm/cpufreq.rst
+++ b/Documentation/admin-guide/pm/cpufreq.rst
@@ -144,6 +144,10 @@ that belong to the same policy (including both online and offline CPUs). That
mask is then used by the core to populate the policy pointers for all of the
CPUs in it.
+A table-less driver whose discrete frequencies are derived at runtime can
+instead provide a ``->resolve_freq()`` callback to map arbitrary requests to
+deterministic, supported frequencies.
+
The next major initialization step for a new policy object is to attach a
scaling governor to it (to begin with, that is the default scaling governor
determined by the kernel command line or configuration, but it may be changed
diff --git a/Documentation/cpu-freq/cpu-drivers.rst b/Documentation/cpu-freq/cpu-drivers.rst
index 17c69f83691e..4f327760bc04 100644
--- a/Documentation/cpu-freq/cpu-drivers.rst
+++ b/Documentation/cpu-freq/cpu-drivers.rst
@@ -86,6 +86,9 @@ And optionally
.set_boost - A pointer to a per-policy function to enable/disable boost
frequencies.
+ .resolve_freq - A pointer to a frequency-resolution function for table-less
+ drivers with discrete, runtime-derived frequencies. See below.
+
1.2 Per-CPU Initialization
--------------------------
@@ -170,6 +173,22 @@ limits on their own. These shall use the ->setpolicy() callback.
1.5. target/target_index
------------------------
+Table-less ``->target()`` drivers may provide ``->resolve_freq()`` to map a
+clamped target to a supported frequency within the supplied limits.
+``CPUFREQ_RELATION_L`` selects the lowest frequency at or above the target,
+H the highest at or below it, and C the closest, choosing higher on ties.
+If no supported frequency is at or above the target, L returns the highest
+supported frequency in the interval. If none is at or below the target,
+H returns the lowest supported frequency in the interval.
+``CPUFREQ_RELATION_E`` is stripped before the call.
+
+Equivalent requests must resolve to the same frequency, which the target
+callbacks must map to one canonical driver request even if performance levels
+share a kHz value. This lets callers cache resolved requests. ``->verify()``
+must leave a supported frequency in every accepted limit interval.
+The callback must not sleep: it may run in scheduler context.
+Policies providing both a frequency table and this callback are rejected.
+
The target_index call has two arguments: ``struct cpufreq_policy *policy``,
and ``unsigned int`` index (into the exposed frequency table).
diff --git a/drivers/cpufreq/cpufreq.c b/drivers/cpufreq/cpufreq.c
index 54dde8419bdc..44bda2f32fcf 100644
--- a/drivers/cpufreq/cpufreq.c
+++ b/drivers/cpufreq/cpufreq.c
@@ -480,8 +480,12 @@ static unsigned int __resolve_freq(struct cpufreq_policy *policy,
target_freq = clamp_val(target_freq, min, max);
- if (!policy->freq_table)
+ if (!policy->freq_table) {
+ if (cpufreq_driver->resolve_freq)
+ return cpufreq_driver->resolve_freq(policy, target_freq, min, max,
+ relation & ~CPUFREQ_RELATION_E);
return target_freq;
+ }
idx = cpufreq_frequency_table_target(policy, target_freq, min, max, relation);
policy->cached_resolved_idx = idx;
@@ -495,7 +499,10 @@ static unsigned int __resolve_freq(struct cpufreq_policy *policy,
* @policy: associated policy to interrogate
* @target_freq: target frequency to resolve.
*
- * The target to driver frequency mapping is cached in the policy.
+ * The frequency-table resolution path caches the mapping in the policy.
+ *
+ * Keep the policy active and exclude driver teardown; a policy reference
+ * alone does not protect driver-private data.
*
* Return: Lowest driver-supported frequency greater than or equal to the
* given target_freq, subject to policy (min/max) and driver limitations.
@@ -1445,6 +1452,11 @@ static int cpufreq_policy_online(struct cpufreq_policy *policy,
* If there is a problem with its frequency table, take it
* offline and drop it.
*/
+ if (policy->freq_table && cpufreq_driver->resolve_freq) {
+ ret = -EINVAL;
+ goto out_offline_policy;
+ }
+
ret = cpufreq_table_validate_and_sort(policy);
if (ret)
goto out_offline_policy;
@@ -2925,6 +2937,7 @@ int cpufreq_register_driver(struct cpufreq_driver *driver_data)
if (!driver_data || !driver_data->verify || !driver_data->init ||
(driver_data->target_index && driver_data->target) ||
+ (driver_data->resolve_freq && !driver_data->target) ||
(!!driver_data->setpolicy == (driver_data->target_index || driver_data->target)) ||
(!driver_data->get_intermediate != !driver_data->target_intermediate) ||
(!driver_data->online != !driver_data->offline) ||
diff --git a/include/linux/cpufreq.h b/include/linux/cpufreq.h
index d3d0d9d02aa4..a0a7619d11fd 100644
--- a/include/linux/cpufreq.h
+++ b/include/linux/cpufreq.h
@@ -364,6 +364,16 @@ struct cpufreq_driver {
int (*target)(struct cpufreq_policy *policy,
unsigned int target_freq,
unsigned int relation); /* Deprecated */
+ /*
+ * Optional for table-less ->target() drivers. Resolve a clamped request
+ * within the supplied limits using CPUFREQ_RELATION_{L,H,C}. Must not
+ * sleep. See Documentation/cpu-freq/cpu-drivers.rst for the contract.
+ */
+ unsigned int (*resolve_freq)(struct cpufreq_policy *policy,
+ unsigned int target_freq,
+ unsigned int min_freq,
+ unsigned int max_freq,
+ unsigned int relation);
int (*target_index)(struct cpufreq_policy *policy,
unsigned int index);
unsigned int (*fast_switch)(struct cpufreq_policy *policy,
^ permalink raw reply [flat|nested] 6+ messages in thread
* [PATCH 2/3] cpufreq: CPPC: Resolve frequencies to performance levels
2026-09-29 10:29 [PATCH 0/3] cpufreq: Resolve CPPC frequencies to performance levels Christian Loehle
2026-09-29 10:29 ` [PATCH 1/3] cpufreq: Add a driver frequency resolution callback Christian Loehle
@ 2026-09-29 10:29 ` Christian Loehle
2026-09-29 10:29 ` [PATCH 3/3] cpufreq: Skip updates for unchanged resolved limits Christian Loehle
` (2 subsequent siblings)
4 siblings, 0 replies; 6+ messages in thread
From: Christian Loehle @ 2026-09-29 10:29 UTC (permalink / raw)
To: rafael
Cc: zhenglifeng1, zhongqiu.han, viresh.kumar, linux-kernel,
linux-acpi, linux-arm-kernel, linux-doc, Mario Limonciello,
K Prateek Nayak, Huang Rui, Perry Yuan, Gautham R . Shenoy,
Vanshidhar Konda, Shubhang Kaushik, Pierre Gondois,
Beata Michalska, Dietmar Eggemann, Ionela Voinescu, Sudeep Holla,
Lukasz Luba, Jeremy Linton, Peter Zijlstra, jonathanh, zhanjie9,
Vincent Guittot, Jonathan Corbet, Shuah Khan, Randy Dunlap,
Christian Loehle
Different kHz requests can select the same CPPC performance level.
Implement ->resolve_freq() so governors such as schedutil can skip
redundant writes.
Precompute the affine conversion and invert its integer rounding directly.
Share bounded conversions with ->target() and ->fast_switch(), choosing the
first performance level when several share a kHz value. Recompute limits
from each policy snapshot to avoid caching sysfs-shared mutable controls.
Cap intervals at or below nominal kHz at Nominal Performance, excluding
boosted levels that alias it. Unchanged frequency limits then imply
unchanged performance limits across boost toggles.
After clamping to CPU limits, snap the limits to supported frequencies.
If none lies in the interval, collapse both limits to the highest supported
frequency not above its maximum, as frequency-table verification does.
Signed-off-by: Christian Loehle <christian.loehle@arm.com>
---
drivers/cpufreq/cppc_cpufreq.c | 257 ++++++++++++++++++++++++++++++---
1 file changed, 239 insertions(+), 18 deletions(-)
diff --git a/drivers/cpufreq/cppc_cpufreq.c b/drivers/cpufreq/cppc_cpufreq.c
index 4ea444ff889e..fa85efb9f701 100644
--- a/drivers/cpufreq/cppc_cpufreq.c
+++ b/drivers/cpufreq/cppc_cpufreq.c
@@ -18,8 +18,10 @@
#include <linux/cpufreq.h>
#include <linux/irq_work.h>
#include <linux/kthread.h>
+#include <linux/math64.h>
#include <linux/mutex.h>
#include <linux/time.h>
+#include <linux/units.h>
#include <linux/vmalloc.h>
#include <uapi/linux/sched/types.h>
@@ -29,6 +31,24 @@
static struct cpufreq_driver cppc_cpufreq_driver;
+struct cppc_perf_freq_map {
+ s64 offset;
+ u64 multiplier;
+ u32 divisor;
+};
+
+struct cppc_cpufreq_data {
+ struct cppc_cpudata cpu_data;
+ struct cppc_perf_freq_map map;
+ unsigned int nominal_khz;
+};
+
+static struct cppc_cpufreq_data *
+cppc_cpufreq_data(struct cppc_cpudata *cpu_data)
+{
+ return container_of(cpu_data, struct cppc_cpufreq_data, cpu_data);
+}
+
#ifdef CONFIG_ACPI_CPPC_CPUFREQ_FIE
static enum {
FIE_UNSET = -1,
@@ -302,11 +322,159 @@ static inline void cppc_freq_invariance_exit(void)
}
#endif /* CONFIG_ACPI_CPPC_CPUFREQ_FIE */
+/* Precompute the affine mapping used by cppc_perf_to_khz(). */
+static void cppc_cpufreq_init_perf_map(struct cppc_perf_caps *caps,
+ struct cppc_perf_freq_map *map)
+{
+ if (caps->lowest_freq && caps->nominal_freq) {
+ if (caps->lowest_freq == caps->nominal_freq) {
+ map->multiplier = (u64)caps->nominal_freq * KHZ_PER_MHZ;
+ map->divisor = caps->nominal_perf;
+ map->offset = 0;
+ } else {
+ map->multiplier = (u64)(caps->nominal_freq -
+ caps->lowest_freq) * KHZ_PER_MHZ;
+ map->divisor = caps->nominal_perf - caps->lowest_perf;
+ map->offset = (s64)caps->nominal_freq * KHZ_PER_MHZ -
+ div64_u64(caps->nominal_perf * map->multiplier,
+ map->divisor);
+ }
+ } else {
+ map->multiplier = cppc_get_dmi_max_khz();
+ map->divisor = caps->highest_perf;
+ map->offset = 0;
+ }
+}
+
+static unsigned int
+cppc_cpufreq_perf_to_khz(const struct cppc_perf_freq_map *map, u32 perf)
+{
+ s64 freq = map->offset + div64_u64(perf * map->multiplier,
+ map->divisor);
+
+ return freq > 0 ? freq : 0;
+}
+
+/*
+ * Invert integer-kHz rounding: L finds the first level at or above the
+ * target frequency, H the last at or below it, within the supplied bounds.
+ */
+static u32 cppc_cpufreq_perf_for_freq(const struct cppc_perf_freq_map *map,
+ unsigned int target_freq,
+ u32 min_perf, u32 max_perf,
+ unsigned int relation)
+{
+ s64 scaled_freq = (s64)target_freq - map->offset;
+ u64 perf;
+
+ switch (relation) {
+ case CPUFREQ_RELATION_L:
+ if (!target_freq || scaled_freq <= 0)
+ perf = 0;
+ else
+ perf = mul_u64_u64_div_u64_roundup(scaled_freq,
+ map->divisor,
+ map->multiplier);
+ break;
+
+ case CPUFREQ_RELATION_H:
+ if (scaled_freq < 0) {
+ perf = 0;
+ } else {
+ perf = mul_u64_u64_div_u64_roundup(scaled_freq + 1,
+ map->divisor,
+ map->multiplier);
+ perf--;
+ }
+ break;
+
+ default:
+ WARN_ON_ONCE(1);
+ return min_perf;
+ }
+
+ return clamp_t(u64, perf, min_perf, max_perf);
+}
+
+static void cppc_cpufreq_perf_limits(struct cppc_perf_caps *caps,
+ struct cppc_cpufreq_data *data,
+ unsigned int min_freq,
+ unsigned int max_freq,
+ u32 *min_perf, u32 *max_perf)
+{
+ u32 policy_max_perf;
+
+ /* Do not include boosted levels that alias the nominal frequency. */
+ policy_max_perf = max_freq <= data->nominal_khz ?
+ caps->nominal_perf : caps->highest_perf;
+ *min_perf = cppc_cpufreq_perf_for_freq(&data->map, min_freq,
+ caps->lowest_perf,
+ policy_max_perf,
+ CPUFREQ_RELATION_L);
+ *max_perf = cppc_cpufreq_perf_for_freq(&data->map, max_freq,
+ caps->lowest_perf,
+ policy_max_perf,
+ CPUFREQ_RELATION_H);
+}
+
+static unsigned int
+cppc_cpufreq_resolve_freq(struct cpufreq_policy *policy,
+ unsigned int target_freq,
+ unsigned int min_freq,
+ unsigned int max_freq,
+ unsigned int relation)
+{
+ struct cppc_cpudata *cpu_data = policy->driver_data;
+ struct cppc_perf_caps *caps = &cpu_data->perf_caps;
+ struct cppc_cpufreq_data *data = cppc_cpufreq_data(cpu_data);
+ const struct cppc_perf_freq_map *map = &data->map;
+ u32 min_perf, max_perf, perf;
+
+ cppc_cpufreq_perf_limits(caps, data, min_freq, max_freq,
+ &min_perf, &max_perf);
+ if (WARN_ON_ONCE(min_perf > max_perf))
+ return cppc_cpufreq_perf_to_khz(map, max_perf);
+
+ switch (relation) {
+ case CPUFREQ_RELATION_L:
+ case CPUFREQ_RELATION_H:
+ perf = cppc_cpufreq_perf_for_freq(map, target_freq,
+ min_perf, max_perf, relation);
+ break;
+ case CPUFREQ_RELATION_C: {
+ u32 lower = cppc_cpufreq_perf_for_freq(map, target_freq,
+ min_perf, max_perf,
+ CPUFREQ_RELATION_H);
+ u32 upper = cppc_cpufreq_perf_for_freq(map, target_freq,
+ min_perf, max_perf,
+ CPUFREQ_RELATION_L);
+ unsigned int lower_freq = cppc_cpufreq_perf_to_khz(map, lower);
+ unsigned int upper_freq = cppc_cpufreq_perf_to_khz(map, upper);
+
+ if (lower_freq >= target_freq)
+ perf = lower;
+ else if (upper_freq <= target_freq)
+ perf = upper;
+ else if (target_freq - lower_freq < upper_freq - target_freq)
+ perf = lower;
+ else
+ perf = upper;
+ break;
+ }
+ default:
+ WARN_ON_ONCE(1);
+ return target_freq;
+ }
+
+ return cppc_cpufreq_perf_to_khz(map, perf);
+}
+
static void cppc_cpufreq_get_perf_limits(struct cppc_cpudata *cpu_data,
struct cpufreq_policy *policy,
u32 *min_perf, u32 *max_perf)
{
struct cppc_perf_caps *caps = &cpu_data->perf_caps;
+ struct cppc_cpufreq_data *data = cppc_cpufreq_data(cpu_data);
unsigned int min_freq, max_freq;
u32 min, max;
@@ -315,11 +483,11 @@ static void cppc_cpufreq_get_perf_limits(struct cppc_cpudata *cpu_data,
if (unlikely(min_freq > max_freq))
min_freq = max_freq;
- min = cppc_khz_to_perf(caps, min_freq);
- max = cppc_khz_to_perf(caps, max_freq);
+ cppc_cpufreq_perf_limits(caps, data, min_freq, max_freq,
+ &min, &max);
- *min_perf = clamp_t(u32, min, caps->lowest_perf, caps->highest_perf);
- *max_perf = clamp_t(u32, max, caps->lowest_perf, caps->highest_perf);
+ *min_perf = min(min, max);
+ *max_perf = max;
}
static void cppc_cpufreq_update_perf_limits(struct cppc_cpudata *cpu_data,
@@ -330,6 +498,26 @@ static void cppc_cpufreq_update_perf_limits(struct cppc_cpudata *cpu_data,
&cpu_data->perf_ctrls.max_perf);
}
+static unsigned int
+cppc_cpufreq_update_perf_ctrls(struct cppc_cpudata *cpu_data,
+ struct cpufreq_policy *policy,
+ unsigned int target_freq)
+{
+ struct cppc_cpufreq_data *data = cppc_cpufreq_data(cpu_data);
+ u32 min_perf, max_perf, desired_perf;
+
+ cppc_cpufreq_get_perf_limits(cpu_data, policy, &min_perf, &max_perf);
+ /* Use the first level when several share the same integer-kHz value. */
+ desired_perf = cppc_cpufreq_perf_for_freq(&data->map, target_freq,
+ min_perf, max_perf,
+ CPUFREQ_RELATION_L);
+ cpu_data->perf_ctrls.min_perf = min_perf;
+ cpu_data->perf_ctrls.max_perf = max_perf;
+ cpu_data->perf_ctrls.desired_perf = desired_perf;
+
+ return cppc_cpufreq_perf_to_khz(&data->map, desired_perf);
+}
+
static int cppc_cpufreq_set_target(struct cpufreq_policy *policy,
unsigned int target_freq,
unsigned int relation)
@@ -339,12 +527,9 @@ static int cppc_cpufreq_set_target(struct cpufreq_policy *policy,
struct cpufreq_freqs freqs;
int ret = 0;
- cpu_data->perf_ctrls.desired_perf =
- cppc_khz_to_perf(&cpu_data->perf_caps, target_freq);
- cppc_cpufreq_update_perf_limits(cpu_data, policy);
-
freqs.old = policy->cur;
- freqs.new = target_freq;
+ freqs.new = cppc_cpufreq_update_perf_ctrls(cpu_data, policy,
+ target_freq);
cpufreq_freq_transition_begin(policy, &freqs);
ret = cppc_set_perf(cpu, &cpu_data->perf_ctrls);
@@ -361,13 +546,12 @@ static unsigned int cppc_cpufreq_fast_switch(struct cpufreq_policy *policy,
unsigned int target_freq)
{
struct cppc_cpudata *cpu_data = policy->driver_data;
+ unsigned int resolved_freq;
unsigned int cpu = policy->cpu;
- u32 desired_perf;
int ret;
- desired_perf = cppc_khz_to_perf(&cpu_data->perf_caps, target_freq);
- cpu_data->perf_ctrls.desired_perf = desired_perf;
- cppc_cpufreq_update_perf_limits(cpu_data, policy);
+ resolved_freq = cppc_cpufreq_update_perf_ctrls(cpu_data, policy,
+ target_freq);
ret = cppc_set_perf(cpu, &cpu_data->perf_ctrls);
if (ret) {
@@ -376,12 +560,39 @@ static unsigned int cppc_cpufreq_fast_switch(struct cpufreq_policy *policy,
return 0;
}
- return target_freq;
+ return resolved_freq;
}
static int cppc_verify_policy(struct cpufreq_policy_data *policy)
{
+ struct cpufreq_policy *cur_policy;
+ struct cppc_cpudata *cpu_data;
+ struct cppc_cpufreq_data *data;
+ struct cppc_perf_caps *caps;
+ unsigned int min_freq, max_freq;
+ u32 min_perf, max_perf;
+
cpufreq_verify_within_cpu_limits(policy);
+
+ cur_policy = cpufreq_cpu_get_raw(policy->cpu);
+ if (WARN_ON_ONCE(!cur_policy || !cur_policy->driver_data))
+ return -ENODEV;
+
+ cpu_data = cur_policy->driver_data;
+ data = cppc_cpufreq_data(cpu_data);
+ caps = &cpu_data->perf_caps;
+ cppc_cpufreq_perf_limits(caps, data, policy->min, policy->max,
+ &min_perf, &max_perf);
+ min_freq = cppc_cpufreq_perf_to_khz(&data->map, min_perf);
+ max_freq = cppc_cpufreq_perf_to_khz(&data->map, max_perf);
+
+ /* Favor the maximum if no supported frequency lies in the interval. */
+ if (min_perf > max_perf || min_freq > policy->max ||
+ max_freq < policy->min)
+ min_freq = max_freq;
+
+ policy->min = min_freq;
+ policy->max = max_freq;
return 0;
}
@@ -618,12 +829,14 @@ static void populate_efficiency_class(void)
static struct cppc_cpudata *cppc_cpufreq_get_cpu_data(unsigned int cpu)
{
+ struct cppc_cpufreq_data *data;
struct cppc_cpudata *cpu_data;
int ret;
- cpu_data = kzalloc_obj(struct cppc_cpudata);
- if (!cpu_data)
+ data = kzalloc_obj(struct cppc_cpufreq_data);
+ if (!data)
goto out;
+ cpu_data = &data->cpu_data;
if (!zalloc_cpumask_var(&cpu_data->shared_cpu_map, GFP_KERNEL))
goto free_cpu;
@@ -640,6 +853,10 @@ static struct cppc_cpudata *cppc_cpufreq_get_cpu_data(unsigned int cpu)
goto free_mask;
}
+ cppc_cpufreq_init_perf_map(&cpu_data->perf_caps, &data->map);
+ data->nominal_khz = cppc_cpufreq_perf_to_khz(&data->map,
+ cpu_data->perf_caps.nominal_perf);
+
ret = cppc_get_perf(cpu, &cpu_data->perf_ctrls);
if (ret) {
pr_debug("Err reading CPU%d perf ctrls: ret:%d\n", cpu, ret);
@@ -651,7 +868,7 @@ static struct cppc_cpudata *cppc_cpufreq_get_cpu_data(unsigned int cpu)
free_mask:
free_cpumask_var(cpu_data->shared_cpu_map);
free_cpu:
- kfree(cpu_data);
+ kfree(data);
out:
return NULL;
}
@@ -659,9 +876,12 @@ static struct cppc_cpudata *cppc_cpufreq_get_cpu_data(unsigned int cpu)
static void cppc_cpufreq_put_cpu_data(struct cpufreq_policy *policy)
{
struct cppc_cpudata *cpu_data = policy->driver_data;
+ struct cppc_cpufreq_data *data;
+
+ data = container_of(cpu_data, struct cppc_cpufreq_data, cpu_data);
free_cpumask_var(cpu_data->shared_cpu_map);
- kfree(cpu_data);
+ kfree(data);
policy->driver_data = NULL;
}
@@ -1056,6 +1276,7 @@ static struct cpufreq_driver cppc_cpufreq_driver = {
.flags = CPUFREQ_CONST_LOOPS | CPUFREQ_NEED_UPDATE_LIMITS,
.verify = cppc_verify_policy,
.target = cppc_cpufreq_set_target,
+ .resolve_freq = cppc_cpufreq_resolve_freq,
.get = cppc_cpufreq_get_rate,
.fast_switch = cppc_cpufreq_fast_switch,
.init = cppc_cpufreq_cpu_init,
^ permalink raw reply [flat|nested] 6+ messages in thread
* [PATCH 3/3] cpufreq: Skip updates for unchanged resolved limits
2026-09-29 10:29 [PATCH 0/3] cpufreq: Resolve CPPC frequencies to performance levels Christian Loehle
2026-09-29 10:29 ` [PATCH 1/3] cpufreq: Add a driver frequency resolution callback Christian Loehle
2026-09-29 10:29 ` [PATCH 2/3] cpufreq: CPPC: Resolve frequencies to performance levels Christian Loehle
@ 2026-09-29 10:29 ` Christian Loehle
2026-09-30 20:06 ` [PATCH 0/3] cpufreq: Resolve CPPC frequencies to performance levels Mario Limonciello
2026-10-01 10:34 ` Peter Zijlstra
4 siblings, 0 replies; 6+ messages in thread
From: Christian Loehle @ 2026-09-29 10:29 UTC (permalink / raw)
To: rafael
Cc: zhenglifeng1, zhongqiu.han, viresh.kumar, linux-kernel,
linux-acpi, linux-arm-kernel, linux-doc, Mario Limonciello,
K Prateek Nayak, Huang Rui, Perry Yuan, Gautham R . Shenoy,
Vanshidhar Konda, Shubhang Kaushik, Pierre Gondois,
Beata Michalska, Dietmar Eggemann, Ionela Voinescu, Sudeep Holla,
Lukasz Luba, Jeremy Linton, Peter Zijlstra, jonathanh, zhanjie9,
Vincent Guittot, Jonathan Corbet, Shuah Khan, Randy Dunlap,
Christian Loehle
Commit 9801be8bef65 ("cpufreq: Avoid redundant target() calls for unchanged
limits") introduced policy->update_limits to skip ->target() when both the
frequency and effective limits are unchanged. schedutil still bypasses its
frequency cache on every limit notification for CPUFREQ_NEED_UPDATE_LIMITS
drivers, and fast switching never consumes the pending state.
Keep raw QoS/limit notifications triggering schedutil recomputation, but
only let CPUFREQ_NEED_UPDATE_LIMITS force a same-frequency callback when
resolved policy->{min,max} changes remain pending. Expose that test through
a cpufreq-core helper.
Consume pending state before both slow and fast callbacks, including
frequency changes, instead of clearing it only in the target == cur case.
Use release/acquire ordering for the limits and restore pending state on
failure, preserving concurrent updates and retries. Favor the maximum
if lockless limit reads observe an inverted pair.
This applies to cppc-cpufreq and amd-pstate's frequency-based paths; the
->adjust_perf() path is unchanged.
Signed-off-by: Christian Loehle <christian.loehle@arm.com>
---
drivers/cpufreq/cpufreq.c | 82 ++++++++++++++++++++++++++------
include/linux/cpufreq.h | 1 +
kernel/sched/cpufreq_schedutil.c | 19 ++------
3 files changed, 72 insertions(+), 30 deletions(-)
diff --git a/drivers/cpufreq/cpufreq.c b/drivers/cpufreq/cpufreq.c
index 44bda2f32fcf..9e932c2a986d 100644
--- a/drivers/cpufreq/cpufreq.c
+++ b/drivers/cpufreq/cpufreq.c
@@ -2058,6 +2058,21 @@ bool cpufreq_driver_test_flags(u16 flags)
return !!(cpufreq_driver->flags & flags);
}
+/**
+ * cpufreq_driver_needs_limits_update - Check for a pending driver limit update.
+ * @policy: CPU frequency policy to check.
+ *
+ * Return: Whether changed resolved limits require a driver callback.
+ */
+bool cpufreq_driver_needs_limits_update(struct cpufreq_policy *policy)
+{
+ if (!cpufreq_driver_test_flags(CPUFREQ_NEED_UPDATE_LIMITS))
+ return false;
+
+ /* Pairs with cpufreq_set_update_limits(). */
+ return smp_load_acquire(&policy->update_limits);
+}
+
/**
* cpufreq_get_current_driver - Return the current driver's name.
*
@@ -2183,6 +2198,30 @@ EXPORT_SYMBOL(cpufreq_unregister_notifier);
* GOVERNORS *
*********************************************************************/
+static bool cpufreq_test_and_clear_update_limits(struct cpufreq_policy *policy)
+{
+ if (!cpufreq_driver_needs_limits_update(policy))
+ return false;
+
+ return xchg(&policy->update_limits, false);
+}
+
+static void cpufreq_set_update_limits(struct cpufreq_policy *policy)
+{
+ /* Publish the new limits before making their update pending. */
+ smp_store_release(&policy->update_limits, true);
+}
+
+static void cpufreq_read_policy_limits(struct cpufreq_policy *policy,
+ unsigned int *min, unsigned int *max)
+{
+ /* Lockless reads can mix limit updates; favor max if inverted. */
+ *min = READ_ONCE(policy->min);
+ *max = READ_ONCE(policy->max);
+ if (unlikely(*min > *max))
+ *min = *max;
+}
+
/**
* cpufreq_driver_fast_switch - Carry out a fast CPU frequency switch.
* @policy: cpufreq policy to switch the frequency for.
@@ -2209,14 +2248,20 @@ EXPORT_SYMBOL(cpufreq_unregister_notifier);
unsigned int cpufreq_driver_fast_switch(struct cpufreq_policy *policy,
unsigned int target_freq)
{
- unsigned int freq;
+ bool update_limits;
+ unsigned int min, max, freq;
int cpu;
- target_freq = clamp_val(target_freq, policy->min, policy->max);
+ update_limits = cpufreq_test_and_clear_update_limits(policy);
+ cpufreq_read_policy_limits(policy, &min, &max);
+ target_freq = clamp_val(target_freq, min, max);
freq = cpufreq_driver->fast_switch(policy, target_freq);
- if (!freq)
+ if (!freq) {
+ if (update_limits)
+ cpufreq_set_update_limits(policy);
return 0;
+ }
policy->cur = freq;
arch_set_freq_scale(policy->related_cpus, freq,
@@ -2367,12 +2412,16 @@ int __cpufreq_driver_target(struct cpufreq_policy *policy,
unsigned int relation)
{
unsigned int old_target_freq = target_freq;
+ unsigned int min, max;
+ bool update_limits;
+ int ret;
if (cpufreq_disabled())
return -ENODEV;
- target_freq = __resolve_freq(policy, target_freq, policy->min,
- policy->max, relation);
+ update_limits = cpufreq_test_and_clear_update_limits(policy);
+ cpufreq_read_policy_limits(policy, &min, &max);
+ target_freq = __resolve_freq(policy, target_freq, min, max, relation);
pr_debug("CPU %u: cur %u kHz -> target %u kHz (req %u kHz, rel %u)\n",
policy->cpu, policy->cur, target_freq, old_target_freq, relation);
@@ -2384,11 +2433,8 @@ int __cpufreq_driver_target(struct cpufreq_policy *policy,
* calls.
*/
if (target_freq == policy->cur) {
- if (!(cpufreq_driver->flags & CPUFREQ_NEED_UPDATE_LIMITS) ||
- !policy->update_limits)
+ if (!update_limits)
return 0;
-
- policy->update_limits = false;
}
if (cpufreq_driver->target) {
@@ -2399,13 +2445,19 @@ int __cpufreq_driver_target(struct cpufreq_policy *policy,
if (!policy->efficiencies_available)
relation &= ~CPUFREQ_RELATION_E;
- return cpufreq_driver->target(policy, target_freq, relation);
+ ret = cpufreq_driver->target(policy, target_freq, relation);
+ } else if (cpufreq_driver->target_index) {
+ ret = __target_index(policy, policy->cached_resolved_idx);
+ } else {
+ if (update_limits)
+ cpufreq_set_update_limits(policy);
+ return -EINVAL;
}
- if (!cpufreq_driver->target_index)
- return -EINVAL;
+ if (ret && update_limits)
+ cpufreq_set_update_limits(policy);
- return __target_index(policy, policy->cached_resolved_idx);
+ return ret;
}
EXPORT_SYMBOL_GPL(__cpufreq_driver_target);
@@ -2681,7 +2733,7 @@ static int cpufreq_set_policy(struct cpufreq_policy *policy,
CPUFREQ_RELATION_H);
if (freq != policy->max) {
WRITE_ONCE(policy->max, freq);
- policy->update_limits = true;
+ cpufreq_set_update_limits(policy);
}
freq = __resolve_freq(policy, new_data.min, new_data.min, new_data.max,
@@ -2689,7 +2741,7 @@ static int cpufreq_set_policy(struct cpufreq_policy *policy,
freq = min(freq, policy->max);
if (freq != policy->min) {
WRITE_ONCE(policy->min, freq);
- policy->update_limits = true;
+ cpufreq_set_update_limits(policy);
}
trace_cpu_frequency_limits(policy);
diff --git a/include/linux/cpufreq.h b/include/linux/cpufreq.h
index a0a7619d11fd..acd60c2b30f1 100644
--- a/include/linux/cpufreq.h
+++ b/include/linux/cpufreq.h
@@ -499,6 +499,7 @@ int cpufreq_register_driver(struct cpufreq_driver *driver_data);
void cpufreq_unregister_driver(struct cpufreq_driver *driver_data);
bool cpufreq_driver_test_flags(u16 flags);
+bool cpufreq_driver_needs_limits_update(struct cpufreq_policy *policy);
const char *cpufreq_get_current_driver(void);
void *cpufreq_get_driver_data(void);
diff --git a/kernel/sched/cpufreq_schedutil.c b/kernel/sched/cpufreq_schedutil.c
index 49ccd6f1c185..43e448ee3c4c 100644
--- a/kernel/sched/cpufreq_schedutil.c
+++ b/kernel/sched/cpufreq_schedutil.c
@@ -123,22 +123,11 @@ static bool sugov_should_update_freq(struct sugov_policy *sg_policy, u64 time)
static bool sugov_update_next_freq(struct sugov_policy *sg_policy, u64 time,
unsigned int next_freq)
{
- if (sg_policy->need_freq_update) {
- sg_policy->need_freq_update = false;
- /*
- * The policy limits have changed, but if the return value of
- * cpufreq_driver_resolve_freq() after applying the new limits
- * is still equal to the previously selected frequency, the
- * driver callback need not be invoked unless the driver
- * specifically wants that to happen on every update of the
- * policy limits.
- */
- if (sg_policy->next_freq == next_freq &&
- !cpufreq_driver_test_flags(CPUFREQ_NEED_UPDATE_LIMITS))
- return false;
- } else if (sg_policy->next_freq == next_freq) {
+ sg_policy->need_freq_update = false;
+
+ if (sg_policy->next_freq == next_freq &&
+ !cpufreq_driver_needs_limits_update(sg_policy->policy))
return false;
- }
sg_policy->next_freq = next_freq;
sg_policy->last_freq_update_time = time;
^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [PATCH 0/3] cpufreq: Resolve CPPC frequencies to performance levels
2026-09-29 10:29 [PATCH 0/3] cpufreq: Resolve CPPC frequencies to performance levels Christian Loehle
` (2 preceding siblings ...)
2026-09-29 10:29 ` [PATCH 3/3] cpufreq: Skip updates for unchanged resolved limits Christian Loehle
@ 2026-09-30 20:06 ` Mario Limonciello
2026-10-01 10:34 ` Peter Zijlstra
4 siblings, 0 replies; 6+ messages in thread
From: Mario Limonciello @ 2026-09-30 20:06 UTC (permalink / raw)
To: Christian Loehle
Cc: zhenglifeng1, zhongqiu.han, viresh.kumar, linux-kernel,
linux-acpi, linux-arm-kernel, linux-doc, K Prateek Nayak,
Huang Rui, Perry Yuan, Gautham R . Shenoy, Vanshidhar Konda,
Shubhang Kaushik, Pierre Gondois, Beata Michalska,
Dietmar Eggemann, Ionela Voinescu, Sudeep Holla, Lukasz Luba,
Jeremy Linton, Peter Zijlstra, jonathanh, zhanjie9,
Vincent Guittot, Jonathan Corbet, Shuah Khan, Randy Dunlap,
rafael
On 9/29/26 05:29, Christian Loehle wrote:
> For table-based cpufreq drivers, the core resolves requests to supported
> frequency-table entries before governors such as schedutil compare them
> with their cached target. Different kHz requests selecting the same entry
> therefore need not reach the driver again.
>
> Table-less cppc-cpufreq lacks this resolution: the core returns the requested
> kHz value even if firmware exposes only a few CPPC performance levels.
> Different requests can miss the governor's cache yet program the same
> Desired Performance value. The tested ARM AGI CPU exposes only 61 levels,
> with bounds visible at:
>
> /sys/devices/system/cpu/cpuX/acpi_cppc/{lowest,highest}_perf
>
> Add ->resolve_freq() so table-less ->target() drivers can canonicalize
> requests without constructing a frequency table. CPPC resolves requests
> before sugov_update_next_freq(), allowing its existing cache to skip
> equivalent requests. Raw limit notifications still trigger
> reconsideration, but only changed resolved limits force a driver update.
> Failed limit writes remain pending for retry.
>
> Testing on an ARM AGI CPU (61 distinct CPPC levels) with schedutil gives
> the following end-to-end results across 16 iterations:
>
> schbench -m 2 -t 31 -F 256 -n 5 -R 18000 -r 60 -w 20 -i 60
>
> metric baseline resolve-freq change
> median p99 3364 us 3280 us -2.5%
> mean of run p99s 3664.2 us 3299.0 us -10.0%
> worst run p99 4360 us 3500 us -19.7%
> median throughput 16431.08 RPS 16460.69 RPS +0.2%
>
> In an instrumented run of the same workload, cppc_set_perf() calls fell
> from 1,136,868 to 866,703, a 23.8% reduction.
>
> Christian Loehle (3):
> cpufreq: Add a driver frequency resolution callback
> cpufreq: CPPC: Resolve frequencies to performance levels
> cpufreq: Skip updates for unchanged resolved limits
>
> Documentation/admin-guide/pm/cpufreq.rst | 4 +
> Documentation/cpu-freq/cpu-drivers.rst | 19 ++
> drivers/cpufreq/cppc_cpufreq.c | 257 +++++++++++++++++++++--
> drivers/cpufreq/cpufreq.c | 99 +++++++--
> include/linux/cpufreq.h | 11 +
> kernel/sched/cpufreq_schedutil.c | 19 +-
> 6 files changed, 359 insertions(+), 50 deletions(-)
>
>
> base-commit: 6375e61c01e93e35ee7acd336a689ac1fae4b509
Reviewed-by: Mario Limonciello (AMD) <superm1@kernel.org>
^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [PATCH 0/3] cpufreq: Resolve CPPC frequencies to performance levels
2026-09-29 10:29 [PATCH 0/3] cpufreq: Resolve CPPC frequencies to performance levels Christian Loehle
` (3 preceding siblings ...)
2026-09-30 20:06 ` [PATCH 0/3] cpufreq: Resolve CPPC frequencies to performance levels Mario Limonciello
@ 2026-10-01 10:34 ` Peter Zijlstra
4 siblings, 0 replies; 6+ messages in thread
From: Peter Zijlstra @ 2026-10-01 10:34 UTC (permalink / raw)
To: Christian Loehle
Cc: rafael, zhenglifeng1, zhongqiu.han, viresh.kumar, linux-kernel,
linux-acpi, linux-arm-kernel, linux-doc, Mario Limonciello,
K Prateek Nayak, Huang Rui, Perry Yuan, Gautham R . Shenoy,
Vanshidhar Konda, Shubhang Kaushik, Pierre Gondois,
Beata Michalska, Dietmar Eggemann, Ionela Voinescu, Sudeep Holla,
Lukasz Luba, Jeremy Linton, jonathanh, zhanjie9, Vincent Guittot,
Jonathan Corbet, Shuah Khan, Randy Dunlap
On Tue, Sep 29, 2026 at 11:29:54AM +0100, Christian Loehle wrote:
> For table-based cpufreq drivers, the core resolves requests to supported
> frequency-table entries before governors such as schedutil compare them
> with their cached target. Different kHz requests selecting the same entry
> therefore need not reach the driver again.
>
> Table-less cppc-cpufreq lacks this resolution: the core returns the requested
> kHz value even if firmware exposes only a few CPPC performance levels.
> Different requests can miss the governor's cache yet program the same
> Desired Performance value. The tested ARM AGI CPU exposes only 61 levels,
> with bounds visible at:
>
> /sys/devices/system/cpu/cpuX/acpi_cppc/{lowest,highest}_perf
>
> Add ->resolve_freq() so table-less ->target() drivers can canonicalize
> requests without constructing a frequency table. CPPC resolves requests
> before sugov_update_next_freq(), allowing its existing cache to skip
> equivalent requests. Raw limit notifications still trigger
> reconsideration, but only changed resolved limits force a driver update.
> Failed limit writes remain pending for retry.
>
> Testing on an ARM AGI CPU (61 distinct CPPC levels) with schedutil gives
> the following end-to-end results across 16 iterations:
>
> schbench -m 2 -t 31 -F 256 -n 5 -R 18000 -r 60 -w 20 -i 60
>
> metric baseline resolve-freq change
> median p99 3364 us 3280 us -2.5%
> mean of run p99s 3664.2 us 3299.0 us -10.0%
> worst run p99 4360 us 3500 us -19.7%
> median throughput 16431.08 RPS 16460.69 RPS +0.2%
>
> In an instrumented run of the same workload, cppc_set_perf() calls fell
> from 1,136,868 to 866,703, a 23.8% reduction.
>
> Christian Loehle (3):
> cpufreq: Add a driver frequency resolution callback
> cpufreq: CPPC: Resolve frequencies to performance levels
> cpufreq: Skip updates for unchanged resolved limits
>
> Documentation/admin-guide/pm/cpufreq.rst | 4 +
> Documentation/cpu-freq/cpu-drivers.rst | 19 ++
> drivers/cpufreq/cppc_cpufreq.c | 257 +++++++++++++++++++++--
> drivers/cpufreq/cpufreq.c | 99 +++++++--
> include/linux/cpufreq.h | 11 +
> kernel/sched/cpufreq_schedutil.c | 19 +-
> 6 files changed, 359 insertions(+), 50 deletions(-)
No objection to the kernel/sched/ change. I'm assuming rjw or other
cpufreq maintainer will take this?
^ permalink raw reply [flat|nested] 6+ messages in thread
end of thread, other threads:[~2026-10-01 10:35 UTC | newest]
Thread overview: 6+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-29 10:29 [PATCH 0/3] cpufreq: Resolve CPPC frequencies to performance levels Christian Loehle
2026-09-29 10:29 ` [PATCH 1/3] cpufreq: Add a driver frequency resolution callback Christian Loehle
2026-09-29 10:29 ` [PATCH 2/3] cpufreq: CPPC: Resolve frequencies to performance levels Christian Loehle
2026-09-29 10:29 ` [PATCH 3/3] cpufreq: Skip updates for unchanged resolved limits Christian Loehle
2026-09-30 20:06 ` [PATCH 0/3] cpufreq: Resolve CPPC frequencies to performance levels Mario Limonciello
2026-10-01 10:34 ` Peter Zijlstra
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®