* [PATCH v2] PM: QoS: Add a sysfs CPU latency request with a boot-time value
@ 2026-10-10 13:29 Roman Kagan
2026-10-10 13:40 ` sashiko-bot
0 siblings, 1 reply; 2+ messages in thread
From: Roman Kagan @ 2026-10-10 13:29 UTC (permalink / raw)
To: Rafael J. Wysocki, Daniel Lezcano, Christian Loehle,
Jonathan Corbet, Shuah Khan, Randy Dunlap, Len Brown,
Pavel Machek
Cc: linux-pm, linux-doc, linux-kernel, Pasha Tatashin, Mike Rapoport,
Pratyush Yadav, kexec, nh-open-source, Roman Kagan
The idle states offered by a platform trade wakeup latency for energy
savings, and there are situations where the trade is not worth making:
while a latency-sensitive workload is running, or during a live update
via kexec, where everything from the outgoing kernel stopping the
workload to the incoming kernel resuming it is downtime, and deep idle
states may lengthen it.
The CPU latency PM QoS limit is meant for that, but user space can only
take part in it through /dev/cpu_dma_latency, with a request that lives
only as long as the file stays open. It cannot be in effect while the
incoming kernel boots, and on the outgoing side the process holding the
file normally goes away with the rest of user space before the kexec.
The per-CPU pm_qos_resume_latency_us and per-state "disable" attributes
in sysfs are not tied to an open file, but they cannot be set before
user space is up either. The existing command line options restricting
idle states either cannot be undone (cpuidle.off=1, idle=poll,
idle=halt, processor.max_cstate= and intel_idle.max_cstate=) or are
specific to one driver (intel_idle.states_off=).
Add a CPU latency request shared by the entire user space, like the
resume latency request of each CPU, and expose it as
/sys/devices/system/cpu/pm_qos_cpu_latency_us with the same format as
the per-CPU attributes: "0" means no constraint and "n/a" means that no
latency is acceptable. Its initial value can be set with the
pm_qos_cpu_latency_us= command line parameter, and the request is added
before any cpuidle driver registers, much like the boot-time request
intel_idle holds until device_initcall_sync. For a live update, the
outgoing kernel can be given a limit through sysfs before the kexec, the
incoming kernel can be booted with the same limit, and user space can
write 0 to lift it once the workload is running again. As the limit is
the minimum over all requests, this one can only make it stricter.
PM QoS is used rather than the per-state "disable" flags, e.g. through
a command line option disabling the states above a given exit latency,
because the genpd governor applies the CPU latency limit to the idle
states of CPU PM domains too, while the flags only apply to the states
of individual CPUs, and because lifting the limit takes one write rather
than one per CPU and state. A boot-time value for the per-CPU resume
latency requests would take one write per CPU to lift and would depend
on CONFIG_PM.
Assisted-by: LLM
Signed-off-by: Roman Kagan <rkagan@amazon.de>
---
Changes in v2:
- Drop the shallow governors (Christian Loehle)
- Add a CPU latency PM QoS request settable from the kernel command line
and through sysfs instead; the commit message explains why PM QoS
rather than the per-state "disable" flags
- Link to v1: https://patch.msgid.link/20261008-b4-cpuidle-shallow-v1-1-c19e71127b14@amazon.de
---
Documentation/admin-guide/kernel-parameters.txt | 11 ++
kernel/power/qos.c | 111 +++++++++++++++++++++
Documentation/ABI/testing/sysfs-devices-system-cpu | 27 +++++
Documentation/admin-guide/pm/cpuidle.rst | 32 ++++--
Documentation/power/pm_qos_interface.rst | 17 +++-
5 files changed, 188 insertions(+), 10 deletions(-)
diff --git a/Documentation/admin-guide/kernel-parameters.txt b/Documentation/admin-guide/kernel-parameters.txt
index e75344f4e0cd..a01c8f9cb00f 100644
--- a/Documentation/admin-guide/kernel-parameters.txt
+++ b/Documentation/admin-guide/kernel-parameters.txt
@@ -5389,6 +5389,17 @@ Kernel parameters
Format: <bool>
Default value is set by CONFIG_DPM_WATCHDOG_ENABLED.
+ pm_qos_cpu_latency_us=
+ [CPU_IDLE] Initial value of the CPU latency limit
+ requested through
+ /sys/devices/system/cpu/pm_qos_cpu_latency_us, in
+ microseconds. It stays in effect until user space
+ writes a different value to that attribute.
+ Format: <integer> | n/a
+ "0" means no constraint (the default) and "n/a" means
+ that no latency is acceptable. Other values, from 1
+ to 1999999999, are limits in microseconds.
+
pnp.debug=1 [PNP]
Enable PNP debug messages (depends on the
CONFIG_PNP_DEBUG_MESSAGES option). Change at run-time
diff --git a/kernel/power/qos.c b/kernel/power/qos.c
index 1944dbeb0d4c..f38136193907 100644
--- a/kernel/power/qos.c
+++ b/kernel/power/qos.c
@@ -21,6 +21,7 @@
/*#define DEBUG*/
#include <linux/pm_qos.h>
+#include <linux/cpu.h>
#include <linux/sched.h>
#include <linux/spinlock.h>
#include <linux/slab.h>
@@ -335,6 +336,116 @@ void cpu_latency_qos_remove_request(struct pm_qos_request *req)
}
EXPORT_SYMBOL_GPL(cpu_latency_qos_remove_request);
+/* User space interface to the CPU latency QoS via sysfs. */
+
+/*
+ * Unlike the requests made through the misc device, this one is shared by the
+ * entire user space and is not tied to an open file, like the resume latency
+ * request of each CPU. Its initial value can be set in the kernel command
+ * line.
+ */
+static struct pm_qos_request cpu_latency_qos_sysfs_req;
+static s32 cpu_latency_qos_boot_value __initdata = PM_QOS_DEFAULT_VALUE;
+
+/*
+ * Use the format of the power/pm_qos_resume_latency_us attribute of CPUs: "0"
+ * means no constraint and "n/a" means that no latency is acceptable.
+ */
+static int cpu_latency_qos_parse(const char *buf, s32 *value)
+{
+ s32 val;
+
+ if (!kstrtos32(buf, 0, &val)) {
+ /*
+ * Prevent users from writing negative or "no constraint" values
+ * directly.
+ */
+ if (val < 0 || val >= PM_QOS_CPU_LATENCY_DEFAULT_VALUE)
+ return -EINVAL;
+
+ *value = val ? val : PM_QOS_DEFAULT_VALUE;
+ } else if (sysfs_streq(buf, "n/a")) {
+ *value = 0;
+ } else {
+ return -EINVAL;
+ }
+
+ return 0;
+}
+
+static int __init cpu_latency_qos_setup(char *str)
+{
+ if (cpu_latency_qos_parse(str, &cpu_latency_qos_boot_value))
+ pr_warn("Invalid pm_qos_cpu_latency_us= value: %s\n", str);
+
+ return 1;
+}
+__setup("pm_qos_cpu_latency_us=", cpu_latency_qos_setup);
+
+static ssize_t pm_qos_cpu_latency_us_show(struct device *dev,
+ struct device_attribute *attr,
+ char *buf)
+{
+ s32 value = READ_ONCE(cpu_latency_qos_sysfs_req.node.prio);
+
+ if (value == 0)
+ return sysfs_emit(buf, "n/a\n");
+ if (value == PM_QOS_CPU_LATENCY_DEFAULT_VALUE)
+ value = 0;
+
+ return sysfs_emit(buf, "%d\n", value);
+}
+
+static ssize_t pm_qos_cpu_latency_us_store(struct device *dev,
+ struct device_attribute *attr,
+ const char *buf, size_t n)
+{
+ s32 value;
+ int ret;
+
+ ret = cpu_latency_qos_parse(buf, &value);
+ if (ret)
+ return ret;
+
+ cpu_latency_qos_update_request(&cpu_latency_qos_sysfs_req, value);
+
+ return n;
+}
+
+static DEVICE_ATTR_RW(pm_qos_cpu_latency_us);
+
+/*
+ * The cpu subsystem root device is registered before any initcalls run, so
+ * both the request and its attribute can be set up before any cpuidle driver
+ * registers.
+ */
+static int __init cpu_latency_qos_sysfs_init(void)
+{
+ struct device *dev_root;
+ int ret = -ENODEV;
+
+ cpu_latency_qos_add_request(&cpu_latency_qos_sysfs_req,
+ cpu_latency_qos_boot_value);
+
+ dev_root = bus_get_dev_root(&cpu_subsys);
+ if (dev_root) {
+ ret = device_create_file(dev_root,
+ &dev_attr_pm_qos_cpu_latency_us);
+ put_device(dev_root);
+ }
+
+ if (ret) {
+ pr_err("%s: %s setup failed\n", __func__,
+ dev_attr_pm_qos_cpu_latency_us.attr.name);
+ /* Do not leave a limit in place that cannot be lifted. */
+ cpu_latency_qos_update_request(&cpu_latency_qos_sysfs_req,
+ PM_QOS_DEFAULT_VALUE);
+ }
+
+ return ret;
+}
+core_initcall(cpu_latency_qos_sysfs_init);
+
/* User space interface to the CPU latency QoS via misc device. */
static int cpu_latency_qos_open(struct inode *inode, struct file *filp)
diff --git a/Documentation/ABI/testing/sysfs-devices-system-cpu b/Documentation/ABI/testing/sysfs-devices-system-cpu
index 82d10d556cc8..08ffb3c47b32 100644
--- a/Documentation/ABI/testing/sysfs-devices-system-cpu
+++ b/Documentation/ABI/testing/sysfs-devices-system-cpu
@@ -141,6 +141,33 @@ Description: Discover cpuidle policy and mechanism
Documentation/driver-api/pm/cpuidle.rst for more information.
+What: /sys/devices/system/cpu/pm_qos_cpu_latency_us
+Date: October 2026
+Contact: Linux power management list <linux-pm@vger.kernel.org>
+Description:
+ (RW) The CPU latency limit requested by user space for all
+ CPUs, in microseconds. It is a single PM QoS request shared by
+ the entire user space, which stays in effect until it is
+ updated. The global CPU latency limit is the minimum of this
+ request and of the other ones, made by kernel code and through
+ /dev/cpu_dma_latency, so it may be lower than this value.
+ Reading the file returns the value of this request, not the
+ global limit.
+
+ As for the power/pm_qos_resume_latency_us attributes of CPUs,
+ "0" means no constraint and "n/a" means that no latency is
+ acceptable. Other values, from 1 to 1999999999, are limits in
+ microseconds. Writing anything else fails with -EINVAL.
+
+ The initial value can be set with the pm_qos_cpu_latency_us=
+ kernel command line parameter.
+
+ Present if CONFIG_CPU_IDLE is set.
+
+ See Documentation/admin-guide/pm/cpuidle.rst for more
+ information.
+
+
What: /sys/devices/system/cpu/cpuX/cpuidle/state<N>/name
/sys/devices/system/cpu/cpuX/cpuidle/stateN/latency
/sys/devices/system/cpu/cpuX/cpuidle/stateN/power
diff --git a/Documentation/admin-guide/pm/cpuidle.rst b/Documentation/admin-guide/pm/cpuidle.rst
index be4c1120e3f0..ea3bc4e4f29e 100644
--- a/Documentation/admin-guide/pm/cpuidle.rst
+++ b/Documentation/admin-guide/pm/cpuidle.rst
@@ -520,13 +520,16 @@ individual CPUs. Kernel code (e.g. device drivers) can set both of them with
the help of special internal interfaces provided by the PM QoS framework. User
space can modify the former by opening the :file:`cpu_dma_latency` special
device file under :file:`/dev/` and writing a binary value (interpreted as a
-signed 32-bit integer) to it. In turn, the resume latency constraint for a CPU
-can be modified from user space by writing a string (representing a signed
-32-bit integer) to the :file:`power/pm_qos_resume_latency_us` file under
+signed 32-bit integer) to it, or by writing a string to the
+:file:`pm_qos_cpu_latency_us` file under :file:`/sys/devices/system/cpu/` in
+``sysfs``. In turn, the resume latency constraint for a CPU can be modified
+from user space by writing a string (representing a signed 32-bit integer) to
+the :file:`power/pm_qos_resume_latency_us` file under
:file:`/sys/devices/system/cpu/cpu<N>/` in ``sysfs``, where the CPU number
-``<N>`` is allocated at the system initialization time. Negative values
-will be rejected in both cases and, also in both cases, the written integer
-number will be interpreted as a requested PM QoS constraint in microseconds.
+``<N>`` is allocated at the system initialization time. Negative values will
+be rejected in all of these cases and the written integer number will be
+interpreted as a requested PM QoS constraint in microseconds, except that ``0``
+written to the ``sysfs`` files means no constraint (see below).
The requested value is not automatically applied as a new constraint, however,
as it may be less restrictive (greater in this particular case) than another
@@ -574,6 +577,16 @@ determine the effective value to be set as the resume latency constraint for the
CPU in question every time the list of requests is updated this way or another
(there may be other requests coming from kernel code in that list).
+Similarly, there is one global CPU latency limit request associated with the
+:file:`pm_qos_cpu_latency_us` file under :file:`/sys/devices/system/cpu/` in
+``sysfs``. It is shared by the entire user space and is not tied to any file
+descriptor, so it stays in effect until it is updated. Its initial value can be
+set with the ``pm_qos_cpu_latency_us=`` kernel command line parameter, which
+puts the limit in place before any ``CPUIdle`` driver is registered, and user
+space can lift it later. As for the :file:`power/pm_qos_resume_latency_us`
+files of CPUs, ``0`` means no constraint and ``n/a`` means that no latency is
+acceptable. Other values, from 1 to 1999999999, are limits in microseconds.
+
CPU idle time governors are expected to regard the minimum of the global
(effective) CPU latency limit and the effective resume latency constraint for
the given CPU as the upper limit for the exit latency of the idle states that
@@ -615,6 +628,13 @@ governor will be used instead of the default one. It is possible to force
the ``menu`` governor to be used on the systems that use the ``ladder`` governor
by default this way, for example.
+The ``pm_qos_cpu_latency_us=`` kernel command line parameter sets the initial
+value of the global CPU latency limit request associated with the
+:file:`pm_qos_cpu_latency_us` file in ``sysfs`` (see `above <cpu-pm-qos_>`_), so
+it can be used to keep idle states with exit latency beyond the given value from
+being selected from the start, and the limit can be lifted at run time by
+writing to that file.
+
The other kernel command line parameters controlling CPU idle time management
described below are only relevant for the *x86* architecture and references
to ``intel_idle`` affect Intel processors only.
diff --git a/Documentation/power/pm_qos_interface.rst b/Documentation/power/pm_qos_interface.rst
index 4c008e2202f0..f5d53f01652f 100644
--- a/Documentation/power/pm_qos_interface.rst
+++ b/Documentation/power/pm_qos_interface.rst
@@ -57,11 +57,12 @@ From user space:
The infrastructure exposes two separate device nodes, /dev/cpu_dma_latency for
the CPU latency QoS and /dev/cpu_wakeup_latency for the CPU system wakeup
-latency QoS.
+latency QoS, and a sysfs file, /sys/devices/system/cpu/pm_qos_cpu_latency_us,
+for the CPU latency QoS.
-Only processes can register a PM QoS request. To provide for automatic
-cleanup of a process, the interface requires the process to register its
-parameter requests as follows.
+Requests made through the device nodes belong to the processes making them.
+To provide for automatic cleanup of a process, the interface requires the
+process to register its parameter requests as follows.
To register the default PM QoS target for the CPU latency QoS, the process must
open /dev/cpu_dma_latency. To register a CPU system wakeup QoS limit, the
@@ -77,6 +78,14 @@ using the 10 char long format e.g. "0x12345678".
To remove the user mode request for a target value simply close the device
node.
+The sysfs file is associated with a single CPU latency QoS request that is not
+tied to any process and is shared by the entire user space. Writing to the
+file updates that request, which stays in effect until the next update. The
+value is a number of microseconds from 1 to 1999999999, with the same special
+cases as for the per-CPU resume latency: "0" means no constraint and "n/a"
+means that no latency is acceptable. The initial value of that request can be
+set with the pm_qos_cpu_latency_us= kernel command line parameter.
+
2. PM QoS per-device latency and flags framework
================================================
---
base-commit: af32da41b0327b9c6a37856ba82b6760d6c8d10e
change-id: 20261007-b4-cpuidle-shallow-a4d6190bc8b1
Best regards,
--
Roman Kagan <rkagan@amazon.de>
^ permalink raw reply [flat|nested] 2+ messages in thread
* Re: [PATCH v2] PM: QoS: Add a sysfs CPU latency request with a boot-time value
2026-10-10 13:29 [PATCH v2] PM: QoS: Add a sysfs CPU latency request with a boot-time value Roman Kagan
@ 2026-10-10 13:40 ` sashiko-bot
0 siblings, 0 replies; 2+ messages in thread
From: sashiko-bot @ 2026-10-10 13:40 UTC (permalink / raw)
To: Roman Kagan, Pavel Machek, Jonathan Corbet, Len Brown,
Christian Loehle, Randy Dunlap, Daniel Lezcano,
Rafael J. Wysocki, Shuah Khan
Cc: linux-pm, linux-kernel, nh-open-source, Mike Rapoport, kexec,
Pasha Tatashin, Pratyush Yadav, linux-doc
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Low] locking: data race on req->node.prio in cpu_latency_qos_update_request
--
commit d19f16780c138f4d33dfa7e439454d2244724233
Author: Roman Kagan <rkagan@amazon.de>
PM: QoS: Add a sysfs CPU latency request with a boot-time value
This commit adds a CPU latency limit requested by user space for all CPUs,
exposed via a sysfs attribute. It operates similarly to the per-CPU resume
latency request but acts as a single PM QoS request shared by the entire
user space.
Link: https://patch.msgid.link/20261008-b4-cpuidle-shallow-v1-1-c19e71127b14@amazon.de
> diff --git a/kernel/power/qos.c b/kernel/power/qos.c
> index 1944dbeb0d4c9..f381361939079 100644
> --- a/kernel/power/qos.c
> +++ b/kernel/power/qos.c
[ ... ]
> +static ssize_t pm_qos_cpu_latency_us_show(struct device *dev,
> + struct device_attribute *attr,
> + char *buf)
> +{
> + s32 value = READ_ONCE(cpu_latency_qos_sysfs_req.node.prio);
> +
> + if (value == 0)
> + return sysfs_emit(buf, "n/a\n");
> + if (value == PM_QOS_CPU_LATENCY_DEFAULT_VALUE)
> + value = 0;
> +
> + return sysfs_emit(buf, "%d\n", value);
> +}
> +
> +static ssize_t pm_qos_cpu_latency_us_store(struct device *dev,
> + struct device_attribute *attr,
> + const char *buf, size_t n)
> +{
> + s32 value;
> + int ret;
> +
> + ret = cpu_latency_qos_parse(buf, &value);
> + if (ret)
> + return ret;
> +
> + cpu_latency_qos_update_request(&cpu_latency_qos_sysfs_req, value);
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[Severity: Low]
Can concurrent writes to this sysfs attribute trigger a data race?
This calls cpu_latency_qos_update_request() on the globally shared
cpu_latency_qos_sysfs_req object without external locking. Inside
cpu_latency_qos_update_request(), there is a fast-path check:
if (new_value == req->node.prio)
return;
This performs a plain lockless read of req->node.prio, which concurrently
races with modifications under pm_qos_lock via plist_node_init() in
pm_qos_update_target().
The pm_qos_cpu_latency_us_show() function correctly uses READ_ONCE() for
this field, but the plain read on the update fast-path could cause KCSAN
warnings or undefined behavior if multiple threads write to the sysfs
attribute simultaneously.
> +
> + return n;
> +}
> +
> +static DEVICE_ATTR_RW(pm_qos_cpu_latency_us);
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261010-b4-cpuidle-shallow-v2-1-adede7dd9077@amazon.de?part=1
^ permalink raw reply [flat|nested] 2+ messages in thread
end of thread, other threads:[~2026-10-10 13:40 UTC | newest]
Thread overview: 2+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-10 13:29 [PATCH v2] PM: QoS: Add a sysfs CPU latency request with a boot-time value Roman Kagan
2026-10-10 13:40 ` sashiko-bot
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®