From: Christian Loehle <christian.loehle@arm.com>
To: Roman Kagan <rkagan@amazon.de>,
"Rafael J. Wysocki" <rafael@kernel.org>,
Daniel Lezcano <daniel.lezcano@kernel.org>,
Jonathan Corbet <corbet@lwn.net>,
Shuah Khan <skhan@linuxfoundation.org>,
Randy Dunlap <rdunlap@infradead.org>
Cc: linux-pm@vger.kernel.org, linux-doc@vger.kernel.org,
linux-kernel@vger.kernel.org, nh-open-source@amazon.com
Subject: Re: [PATCH] cpuidle: Add the shallow governors
Date: Thu, 8 Oct 2026 19:17:25 +0100 [thread overview]
Message-ID: <e18bd763-0d6d-4bfd-b59e-78f952c94aa2@arm.com> (raw)
In-Reply-To: <20261008-b4-cpuidle-shallow-v1-1-c19e71127b14@amazon.de>
On 10/8/26 19:06, Roman Kagan wrote:
> The idle states offered by a platform trade wakeup latency for energy
> savings, and there are situations where the trade is not worth making:
> while a latency-sensitive workload is running, or during a live update
> via kexec, where everything from the outgoing kernel stopping the
> workload to the incoming kernel resuming it is downtime, and deep idle
> states may lengthen it.
>
> The mechanisms currently available for that are all one-way.
> cpuidle.off=1, idle=poll and idle=halt can only be requested in the
> kernel command line and cannot be undone, and the PM QoS interfaces
> (/dev/cpu_dma_latency and the per-CPU pm_qos_resume_latency_us
> attribute) can only be used once user space is up, so they cannot cover
> the boot of the incoming kernel.
You can also disable all but the shallowest idle state in sysfs:
echo 1 > /sys/devices/system/cpu/cpuX/cpuidle/stateX/disable
>
> Add governors that keep the CPUs in the shallowest idle states and leave
> the scheduler tick running. Being governors, they can be selected both
> at run time, by writing a name to the current_governor attribute in
> sysfs, and in the kernel command line, via cpuidle.governor=. For a
> live update, that makes it possible to switch the outgoing kernel to one
> of them before the kexec, pass the same governor to the incoming kernel
> so that it is in effect from the very beginning of its boot, and switch
> to an energy-efficient governor once the post-update work is done.
>
> On x86 and on some powerpc platforms, the shallowest idle state is a
> polling loop. It has the lowest wakeup latency, but also the side
> effects documented for idle=poll: the CPU saves almost no energy,
> competes with its SMT sibling for the core, and on Intel hardware keeps
> the package from using the P-states that require some of its CPUs to be
> idle. Whether that is acceptable is a policy decision, so there are two
> governors and the user makes it by choosing between them. The shallow
> governor always selects the shallowest enabled state. The
> shallow_nopoll one skips polling states, unless the state it would
> select is too slow to leave for the PM QoS latency limit or all of the
> non-polling states are disabled.
>
> A governor named in the kernel command line is not replaced by a
> higher-rated one registering later, so cpuidle.governor= holds for the
> entire boot. Both governors have rating 1, below all of the other
> governors, so that neither is picked by default. Where the cpuidle
> driver registers no polling states, such as on arm64, the two select the
> same states, and neither is a substitute for idle=poll.
>
> Assisted-by: LLM
> Signed-off-by: Roman Kagan <rkagan@amazon.de>
> ---
> drivers/cpuidle/governors/Makefile | 1 +
> drivers/cpuidle/governors/shallow.c | 104 +++++++++++++++++++++++++++++++
> Documentation/admin-guide/pm/cpuidle.rst | 53 +++++++++++++---
> drivers/cpuidle/Kconfig | 20 ++++++
> 4 files changed, 169 insertions(+), 9 deletions(-)
>
> diff --git a/drivers/cpuidle/governors/Makefile b/drivers/cpuidle/governors/Makefile
> index 63abb5393a4d..10e64092c9ec 100644
> --- a/drivers/cpuidle/governors/Makefile
> +++ b/drivers/cpuidle/governors/Makefile
> @@ -7,3 +7,4 @@ obj-$(CONFIG_CPU_IDLE_GOV_LADDER) += ladder.o
> obj-$(CONFIG_CPU_IDLE_GOV_MENU) += menu.o
> obj-$(CONFIG_CPU_IDLE_GOV_TEO) += teo.o
> obj-$(CONFIG_CPU_IDLE_GOV_HALTPOLL) += haltpoll.o
> +obj-$(CONFIG_CPU_IDLE_GOV_SHALLOW) += shallow.o
> diff --git a/drivers/cpuidle/governors/shallow.c b/drivers/cpuidle/governors/shallow.c
> new file mode 100644
> index 000000000000..8d4ac46d302a
> --- /dev/null
> +++ b/drivers/cpuidle/governors/shallow.c
> @@ -0,0 +1,104 @@
> +// SPDX-License-Identifier: GPL-2.0
> +/*
> + * shallow.c - the shallow and shallow_nopoll idle governors
> + *
> + * Always select the shallowest enabled idle state, trading idle energy
> + * savings for the lowest wakeup latency the platform can offer. The
> + * shallow_nopoll one skips polling states unless PM QoS requires them.
> + */
> +
> +#include <linux/cpuidle.h>
> +#include <linux/init.h>
> +
> +/**
> + * shallow_select - select the shallowest enabled idle state
> + * @drv: cpuidle driver containing state data
> + * @dev: the CPU
> + * @stop_tick: indication on whether or not to stop the tick
> + *
> + * The state selected here is the least energy-efficient one offered by the
> + * driver, so there is no point in looking at either the timer events or the
> + * PM QoS latency limits: no constraint can justify going shallower than that.
> + * Leave the tick running, as the CPU is not expected to stay idle for long.
> + *
> + * Return: the index of the first idle state not disabled for @dev, or 0 if
> + * all of them are disabled.
> + */
> +static int shallow_select(struct cpuidle_driver *drv,
> + struct cpuidle_device *dev, bool *stop_tick)
> +{
> + int i;
> +
> + *stop_tick = false;
> +
> + /*
> + * cpuidle_enter_state() does not check whether the state it is asked
> + * to enter has been disabled, so the per-state "disable" attributes
> + * need to be honoured here. If all of the states are disabled, fall
> + * back to the first one, like the other governors do.
> + */
> + for (i = 0; i < drv->state_count; i++)
> + if (!dev->states_usage[i].disable)
> + return i;
> +
> + return 0;
> +}
> +
> +/**
> + * shallow_nopoll_select - select the shallowest enabled non-polling idle state
> + * @drv: cpuidle driver containing state data
> + * @dev: the CPU
> + * @stop_tick: indication on whether or not to stop the tick
> + *
> + * Same as shallow_select(), but skip the polling states. Unlike them, the
> + * shallowest non-polling state may take longer to leave than the PM QoS
> + * latency limit allows, so fall back to shallow_select() in that case, as
> + * well as when all of the non-polling states are disabled.
> + *
> + * Return: the index of the first idle state that is neither disabled for @dev
> + * nor a polling one, or what shallow_select() returns in the cases above.
> + */
> +static int shallow_nopoll_select(struct cpuidle_driver *drv,
> + struct cpuidle_device *dev, bool *stop_tick)
> +{
> + int i;
> +
> + for (i = 0; i < drv->state_count; i++) {
> + if (dev->states_usage[i].disable ||
> + (drv->states[i].flags & CPUIDLE_FLAG_POLLING))
> + continue;
> +
> + if (drv->states[i].exit_latency_ns >
> + cpuidle_governor_latency_req(dev->cpu))
> + break;
> +
> + *stop_tick = false;
> + return i;
> + }
> +
> + return shallow_select(drv, dev, stop_tick);
> +}
> +
> +static struct cpuidle_governor shallow_governor = {
> + .name = "shallow",
> + .rating = 1,
> + .select = shallow_select,
> +};
> +
> +static struct cpuidle_governor shallow_nopoll_governor = {
> + .name = "shallow_nopoll",
> + .rating = 1,
> + .select = shallow_nopoll_select,
> +};
> +
> +static int __init init_shallow(void)
> +{
> + int ret = cpuidle_register_governor(&shallow_governor);
> +
> + if (ret)
> + return ret;
> +
> + return cpuidle_register_governor(&shallow_nopoll_governor);
> +}
> +
> +postcore_initcall(init_shallow);
> diff --git a/Documentation/admin-guide/pm/cpuidle.rst b/Documentation/admin-guide/pm/cpuidle.rst
> index be4c1120e3f0..8c6a438936d4 100644
> --- a/Documentation/admin-guide/pm/cpuidle.rst
> +++ b/Documentation/admin-guide/pm/cpuidle.rst
> @@ -159,15 +159,16 @@ governor uses that information depends on what algorithm is implemented by it
> and that is the primary reason for having more than one governor in the
> ``CPUIdle`` subsystem.
>
> -There are four ``CPUIdle`` governors available, ``menu``, `TEO <teo-gov_>`_,
> -``ladder`` and ``haltpoll``. Which of them is used by default depends on the
> -configuration of the kernel and in particular on whether or not the scheduler
> -tick can be `stopped by the idle loop <idle-cpus-and-tick_>`_. Available
> -governors can be read from the :file:`available_governors`, and the governor
> -can be changed at runtime. The name of the ``CPUIdle`` governor currently
> -used by the kernel can be read from the :file:`current_governor_ro` or
> -:file:`current_governor` file under :file:`/sys/devices/system/cpu/cpuidle/`
> -in ``sysfs``.
> +There are six ``CPUIdle`` governors available, ``menu``, `TEO <teo-gov_>`_,
> +``ladder``, ``haltpoll``, `shallow <shallow-gov_>`_ and
> +`shallow_nopoll <shallow-gov_>`_. Which of them is used by default depends on
> +the configuration of the kernel and in particular on whether or not the
> +scheduler tick can be `stopped by the idle loop <idle-cpus-and-tick_>`_.
> +Available governors can be read from the
> +:file:`available_governors`, and the governor can be changed at runtime. The
> +name of the ``CPUIdle`` governor currently used by the kernel can be read from
> +the :file:`current_governor_ro` or :file:`current_governor` file under
> +:file:`/sys/devices/system/cpu/cpuidle/` in ``sysfs``.
>
> Which ``CPUIdle`` driver is used, on the other hand, usually depends on the
> platform the kernel is running on, but there are platforms with more than one
> @@ -345,6 +346,40 @@ given conditions. However, it applies a different approach to that problem.
> .. kernel-doc:: drivers/cpuidle/governors/teo.c
> :doc: teo-description
>
> +.. _shallow-gov:
> +
> +The Shallow Governors
> +=====================
> +
> +The ``shallow`` governor always selects the shallowest enabled idle state and
> +leaves the scheduler tick running. Unlike the governors described above, it
> +makes no attempt to save energy at all; its purpose is to keep the CPU wakeup
> +latency, and in particular the latency of inter-processor interrupts, at the
> +minimum offered by the platform.
> +
> +The ``shallow_nopoll`` governor does the same, except that it skips polling
> +idle states and selects the shallowest enabled idle state that is not a polling
> +one. That avoids the side effects of idle CPUs spinning, which are described
> +for ``idle=poll`` below, at the cost of a higher wakeup latency. If the exit
> +latency of that state is above the `PM QoS <cpu-pm-qos_>`_ limit in effect for
> +the CPU, or if all of the non-polling states are disabled, ``shallow_nopoll``
> +selects the state that ``shallow`` would select.
> +
> +Neither of them is used by default, regardless of the configuration of the
> +kernel, and they have to be requested explicitly. That can be done either by
> +passing ``cpuidle.governor=shallow`` or ``cpuidle.governor=shallow_nopoll`` in
> +the kernel command line, or by writing the governor name to the
> +:file:`current_governor` file described `above <idle-loop_>`_. A governor
> +requested in the kernel command line is not replaced by a higher-rated one
> +registering later, so it holds for the entire boot and user space can hand the
> +deep idle states back at any later point by switching to a different governor.
> +
> +Note that the shallowest idle state is not necessarily a polling loop. On the
> +platforms whose ``CPUIdle`` driver does not register the generic polling state,
> +it is the idle instruction of the CPU architecture (for instance, ``WFI`` on
> +arm64), so ``shallow`` is not a substitute for ``idle=poll``. If the driver
> +registers no polling states at all, the two governors select the same states.
> +
> .. _idle-states-representation:
>
> Representation of Idle States
> diff --git a/drivers/cpuidle/Kconfig b/drivers/cpuidle/Kconfig
> index 00e2562041fd..c78aa07ca548 100644
> --- a/drivers/cpuidle/Kconfig
> +++ b/drivers/cpuidle/Kconfig
> @@ -44,6 +44,26 @@ config CPU_IDLE_GOV_HALTPOLL
>
> Some virtualized workloads benefit from using it.
>
> +config CPU_IDLE_GOV_SHALLOW
> + bool "Shallow governors (for latency-sensitive systems)"
> + help
> + The shallow governor always selects the shallowest enabled idle
> + state, which keeps the CPU wakeup latency at the minimum offered by
> + the platform at the cost of giving up idle energy savings. The
> + shallow_nopoll governor selects the shallowest enabled idle state
> + that is not a polling one instead, unless PM QoS requires a lower
> + latency than that state offers.
> +
> + They are never used by default and have to be requested explicitly,
> + either by passing cpuidle.governor=shallow or
> + cpuidle.governor=shallow_nopoll in the kernel command line, or by
> + writing the governor name to the current_governor attribute in
> + sysfs. That makes it possible to keep the CPUs out of deep idle
> + states while booting or while running a latency-sensitive workload
> + and to switch to an energy-efficient governor afterwards.
> +
> + If unsure, say N.
> +
> config DT_IDLE_STATES
> bool
>
>
> ---
> base-commit: 08df884136f1c1197bab2a27814404fd329d9aac
> change-id: 20261007-b4-cpuidle-shallow-a4d6190bc8b1
>
> Best regards,
> --
> Roman Kagan <rkagan@amazon.de>
>
>
next prev parent reply other threads:[~2026-10-08 18:17 UTC|newest]
Thread overview: 6+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-08 18:06 Roman Kagan
2026-10-08 18:17 ` Christian Loehle [this message]
2026-10-08 18:29 ` Christian Loehle
2026-10-09 12:12 ` Roman Kagan
2026-10-09 12:45 ` Christian Loehle
2026-10-09 16:50 ` Roman Kagan
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=e18bd763-0d6d-4bfd-b59e-78f952c94aa2@arm.com \
--to=christian.loehle@arm.com \
--cc=corbet@lwn.net \
--cc=daniel.lezcano@kernel.org \
--cc=linux-doc@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-pm@vger.kernel.org \
--cc=nh-open-source@amazon.com \
--cc=rafael@kernel.org \
--cc=rdunlap@infradead.org \
--cc=rkagan@amazon.de \
--cc=skhan@linuxfoundation.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®