* [PATCH] cpuidle: Add the shallow governors
@ 2026-10-08 18:06 Roman Kagan
2026-10-08 18:17 ` Christian Loehle
0 siblings, 1 reply; 5+ messages in thread
From: Roman Kagan @ 2026-10-08 18:06 UTC (permalink / raw)
To: Rafael J. Wysocki, Daniel Lezcano, Christian Loehle,
Jonathan Corbet, Shuah Khan, Randy Dunlap
Cc: Roman Kagan, linux-pm, linux-doc, linux-kernel, nh-open-source
The idle states offered by a platform trade wakeup latency for energy
savings, and there are situations where the trade is not worth making:
while a latency-sensitive workload is running, or during a live update
via kexec, where everything from the outgoing kernel stopping the
workload to the incoming kernel resuming it is downtime, and deep idle
states may lengthen it.
The mechanisms currently available for that are all one-way.
cpuidle.off=1, idle=poll and idle=halt can only be requested in the
kernel command line and cannot be undone, and the PM QoS interfaces
(/dev/cpu_dma_latency and the per-CPU pm_qos_resume_latency_us
attribute) can only be used once user space is up, so they cannot cover
the boot of the incoming kernel.
Add governors that keep the CPUs in the shallowest idle states and leave
the scheduler tick running. Being governors, they can be selected both
at run time, by writing a name to the current_governor attribute in
sysfs, and in the kernel command line, via cpuidle.governor=. For a
live update, that makes it possible to switch the outgoing kernel to one
of them before the kexec, pass the same governor to the incoming kernel
so that it is in effect from the very beginning of its boot, and switch
to an energy-efficient governor once the post-update work is done.
On x86 and on some powerpc platforms, the shallowest idle state is a
polling loop. It has the lowest wakeup latency, but also the side
effects documented for idle=poll: the CPU saves almost no energy,
competes with its SMT sibling for the core, and on Intel hardware keeps
the package from using the P-states that require some of its CPUs to be
idle. Whether that is acceptable is a policy decision, so there are two
governors and the user makes it by choosing between them. The shallow
governor always selects the shallowest enabled state. The
shallow_nopoll one skips polling states, unless the state it would
select is too slow to leave for the PM QoS latency limit or all of the
non-polling states are disabled.
A governor named in the kernel command line is not replaced by a
higher-rated one registering later, so cpuidle.governor= holds for the
entire boot. Both governors have rating 1, below all of the other
governors, so that neither is picked by default. Where the cpuidle
driver registers no polling states, such as on arm64, the two select the
same states, and neither is a substitute for idle=poll.
Assisted-by: LLM
Signed-off-by: Roman Kagan <rkagan@amazon.de>
---
drivers/cpuidle/governors/Makefile | 1 +
drivers/cpuidle/governors/shallow.c | 104 +++++++++++++++++++++++++++++++
Documentation/admin-guide/pm/cpuidle.rst | 53 +++++++++++++---
drivers/cpuidle/Kconfig | 20 ++++++
4 files changed, 169 insertions(+), 9 deletions(-)
diff --git a/drivers/cpuidle/governors/Makefile b/drivers/cpuidle/governors/Makefile
index 63abb5393a4d..10e64092c9ec 100644
--- a/drivers/cpuidle/governors/Makefile
+++ b/drivers/cpuidle/governors/Makefile
@@ -7,3 +7,4 @@ obj-$(CONFIG_CPU_IDLE_GOV_LADDER) += ladder.o
obj-$(CONFIG_CPU_IDLE_GOV_MENU) += menu.o
obj-$(CONFIG_CPU_IDLE_GOV_TEO) += teo.o
obj-$(CONFIG_CPU_IDLE_GOV_HALTPOLL) += haltpoll.o
+obj-$(CONFIG_CPU_IDLE_GOV_SHALLOW) += shallow.o
diff --git a/drivers/cpuidle/governors/shallow.c b/drivers/cpuidle/governors/shallow.c
new file mode 100644
index 000000000000..8d4ac46d302a
--- /dev/null
+++ b/drivers/cpuidle/governors/shallow.c
@@ -0,0 +1,104 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * shallow.c - the shallow and shallow_nopoll idle governors
+ *
+ * Always select the shallowest enabled idle state, trading idle energy
+ * savings for the lowest wakeup latency the platform can offer. The
+ * shallow_nopoll one skips polling states unless PM QoS requires them.
+ */
+
+#include <linux/cpuidle.h>
+#include <linux/init.h>
+
+/**
+ * shallow_select - select the shallowest enabled idle state
+ * @drv: cpuidle driver containing state data
+ * @dev: the CPU
+ * @stop_tick: indication on whether or not to stop the tick
+ *
+ * The state selected here is the least energy-efficient one offered by the
+ * driver, so there is no point in looking at either the timer events or the
+ * PM QoS latency limits: no constraint can justify going shallower than that.
+ * Leave the tick running, as the CPU is not expected to stay idle for long.
+ *
+ * Return: the index of the first idle state not disabled for @dev, or 0 if
+ * all of them are disabled.
+ */
+static int shallow_select(struct cpuidle_driver *drv,
+ struct cpuidle_device *dev, bool *stop_tick)
+{
+ int i;
+
+ *stop_tick = false;
+
+ /*
+ * cpuidle_enter_state() does not check whether the state it is asked
+ * to enter has been disabled, so the per-state "disable" attributes
+ * need to be honoured here. If all of the states are disabled, fall
+ * back to the first one, like the other governors do.
+ */
+ for (i = 0; i < drv->state_count; i++)
+ if (!dev->states_usage[i].disable)
+ return i;
+
+ return 0;
+}
+
+/**
+ * shallow_nopoll_select - select the shallowest enabled non-polling idle state
+ * @drv: cpuidle driver containing state data
+ * @dev: the CPU
+ * @stop_tick: indication on whether or not to stop the tick
+ *
+ * Same as shallow_select(), but skip the polling states. Unlike them, the
+ * shallowest non-polling state may take longer to leave than the PM QoS
+ * latency limit allows, so fall back to shallow_select() in that case, as
+ * well as when all of the non-polling states are disabled.
+ *
+ * Return: the index of the first idle state that is neither disabled for @dev
+ * nor a polling one, or what shallow_select() returns in the cases above.
+ */
+static int shallow_nopoll_select(struct cpuidle_driver *drv,
+ struct cpuidle_device *dev, bool *stop_tick)
+{
+ int i;
+
+ for (i = 0; i < drv->state_count; i++) {
+ if (dev->states_usage[i].disable ||
+ (drv->states[i].flags & CPUIDLE_FLAG_POLLING))
+ continue;
+
+ if (drv->states[i].exit_latency_ns >
+ cpuidle_governor_latency_req(dev->cpu))
+ break;
+
+ *stop_tick = false;
+ return i;
+ }
+
+ return shallow_select(drv, dev, stop_tick);
+}
+
+static struct cpuidle_governor shallow_governor = {
+ .name = "shallow",
+ .rating = 1,
+ .select = shallow_select,
+};
+
+static struct cpuidle_governor shallow_nopoll_governor = {
+ .name = "shallow_nopoll",
+ .rating = 1,
+ .select = shallow_nopoll_select,
+};
+
+static int __init init_shallow(void)
+{
+ int ret = cpuidle_register_governor(&shallow_governor);
+
+ if (ret)
+ return ret;
+
+ return cpuidle_register_governor(&shallow_nopoll_governor);
+}
+
+postcore_initcall(init_shallow);
diff --git a/Documentation/admin-guide/pm/cpuidle.rst b/Documentation/admin-guide/pm/cpuidle.rst
index be4c1120e3f0..8c6a438936d4 100644
--- a/Documentation/admin-guide/pm/cpuidle.rst
+++ b/Documentation/admin-guide/pm/cpuidle.rst
@@ -159,15 +159,16 @@ governor uses that information depends on what algorithm is implemented by it
and that is the primary reason for having more than one governor in the
``CPUIdle`` subsystem.
-There are four ``CPUIdle`` governors available, ``menu``, `TEO <teo-gov_>`_,
-``ladder`` and ``haltpoll``. Which of them is used by default depends on the
-configuration of the kernel and in particular on whether or not the scheduler
-tick can be `stopped by the idle loop <idle-cpus-and-tick_>`_. Available
-governors can be read from the :file:`available_governors`, and the governor
-can be changed at runtime. The name of the ``CPUIdle`` governor currently
-used by the kernel can be read from the :file:`current_governor_ro` or
-:file:`current_governor` file under :file:`/sys/devices/system/cpu/cpuidle/`
-in ``sysfs``.
+There are six ``CPUIdle`` governors available, ``menu``, `TEO <teo-gov_>`_,
+``ladder``, ``haltpoll``, `shallow <shallow-gov_>`_ and
+`shallow_nopoll <shallow-gov_>`_. Which of them is used by default depends on
+the configuration of the kernel and in particular on whether or not the
+scheduler tick can be `stopped by the idle loop <idle-cpus-and-tick_>`_.
+Available governors can be read from the
+:file:`available_governors`, and the governor can be changed at runtime. The
+name of the ``CPUIdle`` governor currently used by the kernel can be read from
+the :file:`current_governor_ro` or :file:`current_governor` file under
+:file:`/sys/devices/system/cpu/cpuidle/` in ``sysfs``.
Which ``CPUIdle`` driver is used, on the other hand, usually depends on the
platform the kernel is running on, but there are platforms with more than one
@@ -345,6 +346,40 @@ given conditions. However, it applies a different approach to that problem.
.. kernel-doc:: drivers/cpuidle/governors/teo.c
:doc: teo-description
+.. _shallow-gov:
+
+The Shallow Governors
+=====================
+
+The ``shallow`` governor always selects the shallowest enabled idle state and
+leaves the scheduler tick running. Unlike the governors described above, it
+makes no attempt to save energy at all; its purpose is to keep the CPU wakeup
+latency, and in particular the latency of inter-processor interrupts, at the
+minimum offered by the platform.
+
+The ``shallow_nopoll`` governor does the same, except that it skips polling
+idle states and selects the shallowest enabled idle state that is not a polling
+one. That avoids the side effects of idle CPUs spinning, which are described
+for ``idle=poll`` below, at the cost of a higher wakeup latency. If the exit
+latency of that state is above the `PM QoS <cpu-pm-qos_>`_ limit in effect for
+the CPU, or if all of the non-polling states are disabled, ``shallow_nopoll``
+selects the state that ``shallow`` would select.
+
+Neither of them is used by default, regardless of the configuration of the
+kernel, and they have to be requested explicitly. That can be done either by
+passing ``cpuidle.governor=shallow`` or ``cpuidle.governor=shallow_nopoll`` in
+the kernel command line, or by writing the governor name to the
+:file:`current_governor` file described `above <idle-loop_>`_. A governor
+requested in the kernel command line is not replaced by a higher-rated one
+registering later, so it holds for the entire boot and user space can hand the
+deep idle states back at any later point by switching to a different governor.
+
+Note that the shallowest idle state is not necessarily a polling loop. On the
+platforms whose ``CPUIdle`` driver does not register the generic polling state,
+it is the idle instruction of the CPU architecture (for instance, ``WFI`` on
+arm64), so ``shallow`` is not a substitute for ``idle=poll``. If the driver
+registers no polling states at all, the two governors select the same states.
+
.. _idle-states-representation:
Representation of Idle States
diff --git a/drivers/cpuidle/Kconfig b/drivers/cpuidle/Kconfig
index 00e2562041fd..c78aa07ca548 100644
--- a/drivers/cpuidle/Kconfig
+++ b/drivers/cpuidle/Kconfig
@@ -44,6 +44,26 @@ config CPU_IDLE_GOV_HALTPOLL
Some virtualized workloads benefit from using it.
+config CPU_IDLE_GOV_SHALLOW
+ bool "Shallow governors (for latency-sensitive systems)"
+ help
+ The shallow governor always selects the shallowest enabled idle
+ state, which keeps the CPU wakeup latency at the minimum offered by
+ the platform at the cost of giving up idle energy savings. The
+ shallow_nopoll governor selects the shallowest enabled idle state
+ that is not a polling one instead, unless PM QoS requires a lower
+ latency than that state offers.
+
+ They are never used by default and have to be requested explicitly,
+ either by passing cpuidle.governor=shallow or
+ cpuidle.governor=shallow_nopoll in the kernel command line, or by
+ writing the governor name to the current_governor attribute in
+ sysfs. That makes it possible to keep the CPUs out of deep idle
+ states while booting or while running a latency-sensitive workload
+ and to switch to an energy-efficient governor afterwards.
+
+ If unsure, say N.
+
config DT_IDLE_STATES
bool
---
base-commit: 08df884136f1c1197bab2a27814404fd329d9aac
change-id: 20261007-b4-cpuidle-shallow-a4d6190bc8b1
Best regards,
--
Roman Kagan <rkagan@amazon.de>
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH] cpuidle: Add the shallow governors
2026-10-08 18:06 [PATCH] cpuidle: Add the shallow governors Roman Kagan
@ 2026-10-08 18:17 ` Christian Loehle
2026-10-08 18:29 ` Christian Loehle
0 siblings, 1 reply; 5+ messages in thread
From: Christian Loehle @ 2026-10-08 18:17 UTC (permalink / raw)
To: Roman Kagan, Rafael J. Wysocki, Daniel Lezcano, Jonathan Corbet,
Shuah Khan, Randy Dunlap
Cc: linux-pm, linux-doc, linux-kernel, nh-open-source
On 10/8/26 19:06, Roman Kagan wrote:
> The idle states offered by a platform trade wakeup latency for energy
> savings, and there are situations where the trade is not worth making:
> while a latency-sensitive workload is running, or during a live update
> via kexec, where everything from the outgoing kernel stopping the
> workload to the incoming kernel resuming it is downtime, and deep idle
> states may lengthen it.
>
> The mechanisms currently available for that are all one-way.
> cpuidle.off=1, idle=poll and idle=halt can only be requested in the
> kernel command line and cannot be undone, and the PM QoS interfaces
> (/dev/cpu_dma_latency and the per-CPU pm_qos_resume_latency_us
> attribute) can only be used once user space is up, so they cannot cover
> the boot of the incoming kernel.
You can also disable all but the shallowest idle state in sysfs:
echo 1 > /sys/devices/system/cpu/cpuX/cpuidle/stateX/disable
>
> Add governors that keep the CPUs in the shallowest idle states and leave
> the scheduler tick running. Being governors, they can be selected both
> at run time, by writing a name to the current_governor attribute in
> sysfs, and in the kernel command line, via cpuidle.governor=. For a
> live update, that makes it possible to switch the outgoing kernel to one
> of them before the kexec, pass the same governor to the incoming kernel
> so that it is in effect from the very beginning of its boot, and switch
> to an energy-efficient governor once the post-update work is done.
>
> On x86 and on some powerpc platforms, the shallowest idle state is a
> polling loop. It has the lowest wakeup latency, but also the side
> effects documented for idle=poll: the CPU saves almost no energy,
> competes with its SMT sibling for the core, and on Intel hardware keeps
> the package from using the P-states that require some of its CPUs to be
> idle. Whether that is acceptable is a policy decision, so there are two
> governors and the user makes it by choosing between them. The shallow
> governor always selects the shallowest enabled state. The
> shallow_nopoll one skips polling states, unless the state it would
> select is too slow to leave for the PM QoS latency limit or all of the
> non-polling states are disabled.
>
> A governor named in the kernel command line is not replaced by a
> higher-rated one registering later, so cpuidle.governor= holds for the
> entire boot. Both governors have rating 1, below all of the other
> governors, so that neither is picked by default. Where the cpuidle
> driver registers no polling states, such as on arm64, the two select the
> same states, and neither is a substitute for idle=poll.
>
> Assisted-by: LLM
> Signed-off-by: Roman Kagan <rkagan@amazon.de>
> ---
> drivers/cpuidle/governors/Makefile | 1 +
> drivers/cpuidle/governors/shallow.c | 104 +++++++++++++++++++++++++++++++
> Documentation/admin-guide/pm/cpuidle.rst | 53 +++++++++++++---
> drivers/cpuidle/Kconfig | 20 ++++++
> 4 files changed, 169 insertions(+), 9 deletions(-)
>
> diff --git a/drivers/cpuidle/governors/Makefile b/drivers/cpuidle/governors/Makefile
> index 63abb5393a4d..10e64092c9ec 100644
> --- a/drivers/cpuidle/governors/Makefile
> +++ b/drivers/cpuidle/governors/Makefile
> @@ -7,3 +7,4 @@ obj-$(CONFIG_CPU_IDLE_GOV_LADDER) += ladder.o
> obj-$(CONFIG_CPU_IDLE_GOV_MENU) += menu.o
> obj-$(CONFIG_CPU_IDLE_GOV_TEO) += teo.o
> obj-$(CONFIG_CPU_IDLE_GOV_HALTPOLL) += haltpoll.o
> +obj-$(CONFIG_CPU_IDLE_GOV_SHALLOW) += shallow.o
> diff --git a/drivers/cpuidle/governors/shallow.c b/drivers/cpuidle/governors/shallow.c
> new file mode 100644
> index 000000000000..8d4ac46d302a
> --- /dev/null
> +++ b/drivers/cpuidle/governors/shallow.c
> @@ -0,0 +1,104 @@
> +// SPDX-License-Identifier: GPL-2.0
> +/*
> + * shallow.c - the shallow and shallow_nopoll idle governors
> + *
> + * Always select the shallowest enabled idle state, trading idle energy
> + * savings for the lowest wakeup latency the platform can offer. The
> + * shallow_nopoll one skips polling states unless PM QoS requires them.
> + */
> +
> +#include <linux/cpuidle.h>
> +#include <linux/init.h>
> +
> +/**
> + * shallow_select - select the shallowest enabled idle state
> + * @drv: cpuidle driver containing state data
> + * @dev: the CPU
> + * @stop_tick: indication on whether or not to stop the tick
> + *
> + * The state selected here is the least energy-efficient one offered by the
> + * driver, so there is no point in looking at either the timer events or the
> + * PM QoS latency limits: no constraint can justify going shallower than that.
> + * Leave the tick running, as the CPU is not expected to stay idle for long.
> + *
> + * Return: the index of the first idle state not disabled for @dev, or 0 if
> + * all of them are disabled.
> + */
> +static int shallow_select(struct cpuidle_driver *drv,
> + struct cpuidle_device *dev, bool *stop_tick)
> +{
> + int i;
> +
> + *stop_tick = false;
> +
> + /*
> + * cpuidle_enter_state() does not check whether the state it is asked
> + * to enter has been disabled, so the per-state "disable" attributes
> + * need to be honoured here. If all of the states are disabled, fall
> + * back to the first one, like the other governors do.
> + */
> + for (i = 0; i < drv->state_count; i++)
> + if (!dev->states_usage[i].disable)
> + return i;
> +
> + return 0;
> +}
> +
> +/**
> + * shallow_nopoll_select - select the shallowest enabled non-polling idle state
> + * @drv: cpuidle driver containing state data
> + * @dev: the CPU
> + * @stop_tick: indication on whether or not to stop the tick
> + *
> + * Same as shallow_select(), but skip the polling states. Unlike them, the
> + * shallowest non-polling state may take longer to leave than the PM QoS
> + * latency limit allows, so fall back to shallow_select() in that case, as
> + * well as when all of the non-polling states are disabled.
> + *
> + * Return: the index of the first idle state that is neither disabled for @dev
> + * nor a polling one, or what shallow_select() returns in the cases above.
> + */
> +static int shallow_nopoll_select(struct cpuidle_driver *drv,
> + struct cpuidle_device *dev, bool *stop_tick)
> +{
> + int i;
> +
> + for (i = 0; i < drv->state_count; i++) {
> + if (dev->states_usage[i].disable ||
> + (drv->states[i].flags & CPUIDLE_FLAG_POLLING))
> + continue;
> +
> + if (drv->states[i].exit_latency_ns >
> + cpuidle_governor_latency_req(dev->cpu))
> + break;
> +
> + *stop_tick = false;
> + return i;
> + }
> +
> + return shallow_select(drv, dev, stop_tick);
> +}
> +
> +static struct cpuidle_governor shallow_governor = {
> + .name = "shallow",
> + .rating = 1,
> + .select = shallow_select,
> +};
> +
> +static struct cpuidle_governor shallow_nopoll_governor = {
> + .name = "shallow_nopoll",
> + .rating = 1,
> + .select = shallow_nopoll_select,
> +};
> +
> +static int __init init_shallow(void)
> +{
> + int ret = cpuidle_register_governor(&shallow_governor);
> +
> + if (ret)
> + return ret;
> +
> + return cpuidle_register_governor(&shallow_nopoll_governor);
> +}
> +
> +postcore_initcall(init_shallow);
> diff --git a/Documentation/admin-guide/pm/cpuidle.rst b/Documentation/admin-guide/pm/cpuidle.rst
> index be4c1120e3f0..8c6a438936d4 100644
> --- a/Documentation/admin-guide/pm/cpuidle.rst
> +++ b/Documentation/admin-guide/pm/cpuidle.rst
> @@ -159,15 +159,16 @@ governor uses that information depends on what algorithm is implemented by it
> and that is the primary reason for having more than one governor in the
> ``CPUIdle`` subsystem.
>
> -There are four ``CPUIdle`` governors available, ``menu``, `TEO <teo-gov_>`_,
> -``ladder`` and ``haltpoll``. Which of them is used by default depends on the
> -configuration of the kernel and in particular on whether or not the scheduler
> -tick can be `stopped by the idle loop <idle-cpus-and-tick_>`_. Available
> -governors can be read from the :file:`available_governors`, and the governor
> -can be changed at runtime. The name of the ``CPUIdle`` governor currently
> -used by the kernel can be read from the :file:`current_governor_ro` or
> -:file:`current_governor` file under :file:`/sys/devices/system/cpu/cpuidle/`
> -in ``sysfs``.
> +There are six ``CPUIdle`` governors available, ``menu``, `TEO <teo-gov_>`_,
> +``ladder``, ``haltpoll``, `shallow <shallow-gov_>`_ and
> +`shallow_nopoll <shallow-gov_>`_. Which of them is used by default depends on
> +the configuration of the kernel and in particular on whether or not the
> +scheduler tick can be `stopped by the idle loop <idle-cpus-and-tick_>`_.
> +Available governors can be read from the
> +:file:`available_governors`, and the governor can be changed at runtime. The
> +name of the ``CPUIdle`` governor currently used by the kernel can be read from
> +the :file:`current_governor_ro` or :file:`current_governor` file under
> +:file:`/sys/devices/system/cpu/cpuidle/` in ``sysfs``.
>
> Which ``CPUIdle`` driver is used, on the other hand, usually depends on the
> platform the kernel is running on, but there are platforms with more than one
> @@ -345,6 +346,40 @@ given conditions. However, it applies a different approach to that problem.
> .. kernel-doc:: drivers/cpuidle/governors/teo.c
> :doc: teo-description
>
> +.. _shallow-gov:
> +
> +The Shallow Governors
> +=====================
> +
> +The ``shallow`` governor always selects the shallowest enabled idle state and
> +leaves the scheduler tick running. Unlike the governors described above, it
> +makes no attempt to save energy at all; its purpose is to keep the CPU wakeup
> +latency, and in particular the latency of inter-processor interrupts, at the
> +minimum offered by the platform.
> +
> +The ``shallow_nopoll`` governor does the same, except that it skips polling
> +idle states and selects the shallowest enabled idle state that is not a polling
> +one. That avoids the side effects of idle CPUs spinning, which are described
> +for ``idle=poll`` below, at the cost of a higher wakeup latency. If the exit
> +latency of that state is above the `PM QoS <cpu-pm-qos_>`_ limit in effect for
> +the CPU, or if all of the non-polling states are disabled, ``shallow_nopoll``
> +selects the state that ``shallow`` would select.
> +
> +Neither of them is used by default, regardless of the configuration of the
> +kernel, and they have to be requested explicitly. That can be done either by
> +passing ``cpuidle.governor=shallow`` or ``cpuidle.governor=shallow_nopoll`` in
> +the kernel command line, or by writing the governor name to the
> +:file:`current_governor` file described `above <idle-loop_>`_. A governor
> +requested in the kernel command line is not replaced by a higher-rated one
> +registering later, so it holds for the entire boot and user space can hand the
> +deep idle states back at any later point by switching to a different governor.
> +
> +Note that the shallowest idle state is not necessarily a polling loop. On the
> +platforms whose ``CPUIdle`` driver does not register the generic polling state,
> +it is the idle instruction of the CPU architecture (for instance, ``WFI`` on
> +arm64), so ``shallow`` is not a substitute for ``idle=poll``. If the driver
> +registers no polling states at all, the two governors select the same states.
> +
> .. _idle-states-representation:
>
> Representation of Idle States
> diff --git a/drivers/cpuidle/Kconfig b/drivers/cpuidle/Kconfig
> index 00e2562041fd..c78aa07ca548 100644
> --- a/drivers/cpuidle/Kconfig
> +++ b/drivers/cpuidle/Kconfig
> @@ -44,6 +44,26 @@ config CPU_IDLE_GOV_HALTPOLL
>
> Some virtualized workloads benefit from using it.
>
> +config CPU_IDLE_GOV_SHALLOW
> + bool "Shallow governors (for latency-sensitive systems)"
> + help
> + The shallow governor always selects the shallowest enabled idle
> + state, which keeps the CPU wakeup latency at the minimum offered by
> + the platform at the cost of giving up idle energy savings. The
> + shallow_nopoll governor selects the shallowest enabled idle state
> + that is not a polling one instead, unless PM QoS requires a lower
> + latency than that state offers.
> +
> + They are never used by default and have to be requested explicitly,
> + either by passing cpuidle.governor=shallow or
> + cpuidle.governor=shallow_nopoll in the kernel command line, or by
> + writing the governor name to the current_governor attribute in
> + sysfs. That makes it possible to keep the CPUs out of deep idle
> + states while booting or while running a latency-sensitive workload
> + and to switch to an energy-efficient governor afterwards.
> +
> + If unsure, say N.
> +
> config DT_IDLE_STATES
> bool
>
>
> ---
> base-commit: 08df884136f1c1197bab2a27814404fd329d9aac
> change-id: 20261007-b4-cpuidle-shallow-a4d6190bc8b1
>
> Best regards,
> --
> Roman Kagan <rkagan@amazon.de>
>
>
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH] cpuidle: Add the shallow governors
2026-10-08 18:17 ` Christian Loehle
@ 2026-10-08 18:29 ` Christian Loehle
2026-10-09 12:12 ` Roman Kagan
0 siblings, 1 reply; 5+ messages in thread
From: Christian Loehle @ 2026-10-08 18:29 UTC (permalink / raw)
To: Roman Kagan, Rafael J. Wysocki, Daniel Lezcano, Jonathan Corbet,
Shuah Khan, Randy Dunlap
Cc: linux-pm, linux-doc, linux-kernel, nh-open-source
On 10/8/26 19:17, Christian Loehle wrote:
> On 10/8/26 19:06, Roman Kagan wrote:
>> The idle states offered by a platform trade wakeup latency for energy
>> savings, and there are situations where the trade is not worth making:
>> while a latency-sensitive workload is running, or during a live update
>> via kexec, where everything from the outgoing kernel stopping the
>> workload to the incoming kernel resuming it is downtime, and deep idle
>> states may lengthen it.
>>
>> The mechanisms currently available for that are all one-way.
>> cpuidle.off=1, idle=poll and idle=halt can only be requested in the
>> kernel command line and cannot be undone, and the PM QoS interfaces
>> (/dev/cpu_dma_latency and the per-CPU pm_qos_resume_latency_us
>> attribute) can only be used once user space is up, so they cannot cover
>> the boot of the incoming kernel.
>
> You can also disable all but the shallowest idle state in sysfs:
> echo 1 > /sys/devices/system/cpu/cpuX/cpuidle/stateX/disable
>
And I'd probably prefer having that exposed via the cmdline rather than
two separate governors...
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH] cpuidle: Add the shallow governors
2026-10-08 18:29 ` Christian Loehle
@ 2026-10-09 12:12 ` Roman Kagan
2026-10-09 12:45 ` Christian Loehle
0 siblings, 1 reply; 5+ messages in thread
From: Roman Kagan @ 2026-10-09 12:12 UTC (permalink / raw)
To: Christian Loehle
Cc: Rafael J. Wysocki, Daniel Lezcano, Jonathan Corbet, Shuah Khan,
Randy Dunlap, linux-pm, linux-doc, linux-kernel, nh-open-source
On Thu, Oct 08, 2026 at 07:29:43PM +0100, Christian Loehle wrote:
> On 10/8/26 19:17, Christian Loehle wrote:
> > On 10/8/26 19:06, Roman Kagan wrote:
> >> The idle states offered by a platform trade wakeup latency for energy
> >> savings, and there are situations where the trade is not worth making:
> >> while a latency-sensitive workload is running, or during a live update
> >> via kexec, where everything from the outgoing kernel stopping the
> >> workload to the incoming kernel resuming it is downtime, and deep idle
> >> states may lengthen it.
> >>
> >> The mechanisms currently available for that are all one-way.
> >> cpuidle.off=1, idle=poll and idle=halt can only be requested in the
> >> kernel command line and cannot be undone, and the PM QoS interfaces
> >> (/dev/cpu_dma_latency and the per-CPU pm_qos_resume_latency_us
> >> attribute) can only be used once user space is up, so they cannot cover
> >> the boot of the incoming kernel.
> >
> > You can also disable all but the shallowest idle state in sysfs:
> > echo 1 > /sys/devices/system/cpu/cpuX/cpuidle/stateX/disable
> >
>
> And I'd probably prefer having that exposed via the cmdline rather than
> two separate governors...
Doing this cmdline configuration per-cpu per-state is non-realistic. I
guess you mean a single option that would express a policy, like "for
all cpus in the system, disable all but the shallowest state" or "...
all but the shallowest non-polling". But policy is exactly what
governors are for.
Why exactly does having two more separate simple, narrow-purpose
governors sound wrong to you?
Thanks,
Roman.
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH] cpuidle: Add the shallow governors
2026-10-09 12:12 ` Roman Kagan
@ 2026-10-09 12:45 ` Christian Loehle
0 siblings, 0 replies; 5+ messages in thread
From: Christian Loehle @ 2026-10-09 12:45 UTC (permalink / raw)
To: Roman Kagan, Rafael J. Wysocki, Daniel Lezcano, Jonathan Corbet,
Shuah Khan, Randy Dunlap, linux-pm, linux-doc, linux-kernel,
nh-open-source
On 10/9/26 13:12, Roman Kagan wrote:
> On Thu, Oct 08, 2026 at 07:29:43PM +0100, Christian Loehle wrote:
>> On 10/8/26 19:17, Christian Loehle wrote:
>>> On 10/8/26 19:06, Roman Kagan wrote:
>>>> The idle states offered by a platform trade wakeup latency for energy
>>>> savings, and there are situations where the trade is not worth making:
>>>> while a latency-sensitive workload is running, or during a live update
>>>> via kexec, where everything from the outgoing kernel stopping the
>>>> workload to the incoming kernel resuming it is downtime, and deep idle
>>>> states may lengthen it.
>>>>
>>>> The mechanisms currently available for that are all one-way.
>>>> cpuidle.off=1, idle=poll and idle=halt can only be requested in the
>>>> kernel command line and cannot be undone, and the PM QoS interfaces
>>>> (/dev/cpu_dma_latency and the per-CPU pm_qos_resume_latency_us
>>>> attribute) can only be used once user space is up, so they cannot cover
>>>> the boot of the incoming kernel.
>>>
>>> You can also disable all but the shallowest idle state in sysfs:
>>> echo 1 > /sys/devices/system/cpu/cpuX/cpuidle/stateX/disable
>>>
>>
>> And I'd probably prefer having that exposed via the cmdline rather than
>> two separate governors...
>
> Doing this cmdline configuration per-cpu per-state is non-realistic. I
> guess you mean a single option that would express a policy, like "for
> all cpus in the system, disable all but the shallowest state" or "...
> all but the shallowest non-polling". But policy is exactly what
> governors are for.
Yes, I had something like
cpuidle.max_exit_latency_us=<N>
in mind that then sets the disable attribute for the applicable states.
>
> Why exactly does having two more separate simple, narrow-purpose
> governors sound wrong to you?
Because it's a lot of duplicate code we have to maintain (and the
documentation).
^ permalink raw reply [flat|nested] 5+ messages in thread
end of thread, other threads:[~2026-10-09 12:45 UTC | newest]
Thread overview: 5+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-08 18:06 [PATCH] cpuidle: Add the shallow governors Roman Kagan
2026-10-08 18:17 ` Christian Loehle
2026-10-08 18:29 ` Christian Loehle
2026-10-09 12:12 ` Roman Kagan
2026-10-09 12:45 ` Christian Loehle
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®