* [REGRESSION] tick/sched: /proc/stat idle time exceeds wall time since v7.2
@ 2026-10-04 14:27 Stian Halseth
2026-10-04 18:47 ` [PATCH] sched/cputime: Don't account idle time twice after dyntick-idle Stian Halseth
0 siblings, 1 reply; 4+ messages in thread
From: Stian Halseth @ 2026-10-04 14:27 UTC (permalink / raw)
To: Frederic Weisbecker, Thomas Gleixner
Cc: Anna-Maria Behnsen, Ingo Molnar, Peter Zijlstra, Shrikanth Hegde,
regressions, linux-kernel
Hi,
Since v7.2-rc1, /proc/stat counts more idle time than wall time on
NO_HZ_IDLE kernels with TICK_CPU_ACCOUNTING. Monitoring that computes
CPU usage as 1 minus the idle rate shows negative usage on idle
machines.
Measured over 60 s with /proc/stat and the idle_sleeps counters in
/proc/timer_list:
SPARC T7-1, 7.3-rc5, HZ=100, 256 CPUs:
idle per wall second 1.0057 on average, 1.44 on the worst CPU
(92 tick stops/s, 4.8 ms of excess per tick stop)
amd64 (Opteron 1214), 7.2.7, HZ=1000, 2 CPUs:
idle machine: +0.03% and +0.06%
a task waking every 3.7 ms on CPU 1: +13.1% on CPU 1
(285 tick stops/s, 0.46 ms of excess per tick stop)
On both machines the excess is about half a tick per tick stop.
From reading the code (not bisected), I think the cause is that
cf6444c3e1bb7 ("tick/sched: Unify idle cputime accounting") made
dyntick-idle time and tick-sampled time share cpustat[CPUTIME_IDLE].
On idle exit, kcpustat_dyntick_stop() accounts idle time up to now.
The tick is then restarted on its old period, and the first tick
accounts a whole TICK_NSEC through account_process_tick(), although
the part of that period before the idle exit has just been accounted
as idle. Before v7.2, /proc/stat read ts->idle_sleeptime, and the tick
fed a counter it did not use, so nothing was counted twice.
I have a fix that records the overlap in kcpustat_dyntick_stop() and
leaves it out of the first tick. With it, idle per wall second is
1.0000 on both machines, and -0.1% under the wakeup load at 280 tick
stops/s. I am still testing it with IRQ_TIME_ACCOUNTING and with steal
time in a KVM guest, and will send it when that is done.
#regzbot introduced: cf6444c3e1bb7
Stian
^ permalink raw reply [flat|nested] 4+ messages in thread
* [PATCH] sched/cputime: Don't account idle time twice after dyntick-idle
2026-10-04 14:27 [REGRESSION] tick/sched: /proc/stat idle time exceeds wall time since v7.2 Stian Halseth
@ 2026-10-04 18:47 ` Stian Halseth
2026-10-05 4:25 ` Thorsten Leemhuis
2026-10-05 11:33 ` Frederic Weisbecker
0 siblings, 2 replies; 4+ messages in thread
From: Stian Halseth @ 2026-10-04 18:47 UTC (permalink / raw)
To: Frederic Weisbecker, Thomas Gleixner
Cc: Anna-Maria Behnsen, Ingo Molnar, Peter Zijlstra, Shrikanth Hegde,
regressions, linux-kernel
On idle exit the dyntick-idle accounting accounts the time up to now,
then the tick is restarted on its old period. The first tick accounts a
whole TICK_NSEC to whatever runs, although the part of that period
before the idle exit has just been accounted as idle time, or with
IRQ_TIME_ACCOUNTING as IRQ time. That is up to a full tick per idle
exit, and /proc/stat reports more idle time than wall time.
Record how much of the tick period had passed since the tick was
stopped, and leave it out of the first tick.
Fixes: cf6444c3e1bb7 ("tick/sched: Unify idle cputime accounting")
Link: https://lore.kernel.org/all/20261004142724.3896396-1-stian@itx.no/
Signed-off-by: Stian Halseth <stian@itx.no>
---
Tested by comparing /proc/stat with CLOCK_MONOTONIC per CPU over 30 to
60 s, with a task on one CPU that sleeps in a loop. Total CPU time per
wall second on that CPU:
before after
SPARC T7-1, HZ=100, busiest CPU* 1.44 1.0000
SPARC T7-1, HZ=100, 3.7 ms sleeps 1.0001
SPARC T7-1, HZ=100, 25 ms sleeps 0.9996
Opteron, HZ=1000, 3.7 ms sleeps 1.131 0.999
x86_64 KVM guest, HZ=250, 3.7 ms 1.53 0.998
same guest, 9 ms sleeps 1.195 0.982
* under its normal load, about 95 tick stops/s
The guest was tested with and without IRQ_TIME_ACCOUNTING, with
highres=off, with steal time from a busy loop on the host CPU, and with
a test-only change that forces tick_nohz_idle_restart_tick() on every
idle loop iteration with the tick stopped.
The -1.8% left with long sleeps in the guest is time between an idle
tick's expiry and its delivery to the halted vCPU, which nothing
accounts. The summed lateness of those ticks matches it within 2 ms/s,
or within 8 ms/s with steal time, also with highres=off. It is there
without this patch as well, where the double counting hides it. On the
T7-1 it is within noise.
With 3.7 ms sleeps and steal time, +0.3% to +1.3% is left, which I
have not explained.
include/linux/kernel_stat.h | 6 ++++--
kernel/sched/cputime.c | 31 +++++++++++++++++++++++++------
kernel/time/tick-sched.c | 23 ++++++++++++++++++++---
3 files changed, 49 insertions(+), 11 deletions(-)
diff --git a/include/linux/kernel_stat.h b/include/linux/kernel_stat.h
index 9ca6c2259dfea..950867deea29f 100644
--- a/include/linux/kernel_stat.h
+++ b/include/linux/kernel_stat.h
@@ -40,6 +40,8 @@ struct kernel_cpustat {
seqcount_t idle_sleeptime_seq;
u64 idle_entrytime;
u64 idle_stealtime[2];
+ u64 idle_dyntick_entry;
+ u64 idle_tick_overlap;
#endif
u64 cpustat[NR_STATS];
};
@@ -111,7 +113,7 @@ static inline unsigned long kstat_cpu_irqs_sum(unsigned int cpu)
#ifdef CONFIG_HAVE_VIRT_CPU_ACCOUNTING_IDLE
static inline void kcpustat_dyntick_start(u64 now) { }
-static inline void kcpustat_dyntick_stop(u64 now) { }
+static inline void kcpustat_dyntick_stop(u64 now, u64 tick_start) { }
static inline void kcpustat_irq_enter(u64 now) { }
static inline void kcpustat_irq_exit(u64 now) { }
static inline bool kcpustat_idle_dyntick(void) { return false; }
@@ -132,7 +134,7 @@ static inline u64 kcpustat_field_iowait(int cpu)
#else /* !CONFIG_HAVE_VIRT_CPU_ACCOUNTING_IDLE */
extern void kcpustat_dyntick_start(u64 now);
-extern void kcpustat_dyntick_stop(u64 now);
+extern void kcpustat_dyntick_stop(u64 now, u64 tick_start);
extern void kcpustat_irq_enter(u64 now);
extern void kcpustat_irq_exit(u64 now);
extern u64 kcpustat_field_idle(int cpu);
diff --git a/kernel/sched/cputime.c b/kernel/sched/cputime.c
index 06bddaa738e52..9205e92680943 100644
--- a/kernel/sched/cputime.c
+++ b/kernel/sched/cputime.c
@@ -358,6 +358,19 @@ void thread_group_cputime(struct task_struct *tsk, struct task_cputime *times)
}
}
+/*
+ * The first tick after dyntick-idle covers a period that the dyntick-idle
+ * accounting may already have accounted up to the idle exit.
+ */
+static u64 tick_cputime(void)
+{
+#if defined(CONFIG_NO_HZ_COMMON) && !defined(CONFIG_HAVE_VIRT_CPU_ACCOUNTING_IDLE)
+ return TICK_NSEC - __this_cpu_xchg(kernel_cpustat.idle_tick_overlap, 0);
+#else
+ return TICK_NSEC;
+#endif
+}
+
#ifdef CONFIG_IRQ_TIME_ACCOUNTING
/*
* Account a tick to a process and cpustat
@@ -381,9 +394,9 @@ void thread_group_cputime(struct task_struct *tsk, struct task_cputime *times)
* softirq as those do not count in task exec_runtime any more.
*/
static void irqtime_account_process_tick(struct task_struct *p, int user_tick,
- int ticks)
+ u64 cputime)
{
- u64 other, cputime = TICK_NSEC * ticks;
+ u64 other;
/*
* When returning from idle, many ticks can get accounted at
@@ -418,7 +431,7 @@ static void irqtime_account_process_tick(struct task_struct *p, int user_tick,
#else /* !CONFIG_IRQ_TIME_ACCOUNTING: */
static inline void irqtime_account_process_tick(struct task_struct *p, int user_tick,
- int nr_ticks) { }
+ u64 cputime) { }
#endif /* !CONFIG_IRQ_TIME_ACCOUNTING */
#if defined(CONFIG_NO_HZ_COMMON) && !defined(CONFIG_HAVE_VIRT_CPU_ACCOUNTING_IDLE)
@@ -468,12 +481,15 @@ static void kcpustat_idle_start(struct kernel_cpustat *kc, u64 now)
write_seqcount_end(&kc->idle_sleeptime_seq);
}
-void kcpustat_dyntick_stop(u64 now)
+void kcpustat_dyntick_stop(u64 now, u64 tick_start)
{
struct kernel_cpustat *kc = kcpustat_this_cpu;
if (!vtime_generic_enabled_this_cpu()) {
WARN_ON_ONCE(!kc->idle_dyntick);
+ tick_start = max(tick_start, kc->idle_dyntick_entry);
+ if (now > tick_start)
+ kc->idle_tick_overlap = now - tick_start;
kcpustat_idle_stop(kc, now);
kc->idle_dyntick = false;
vtime_dyntick_stop();
@@ -487,6 +503,8 @@ void kcpustat_dyntick_start(u64 now)
if (!vtime_generic_enabled_this_cpu()) {
vtime_dyntick_start();
kc->idle_dyntick = true;
+ kc->idle_dyntick_entry = now;
+ kc->idle_tick_overlap = 0;
kcpustat_idle_start(kc, now);
}
}
@@ -695,12 +713,13 @@ void account_process_tick(struct task_struct *p, int user_tick)
if (kcpustat_idle_dyntick())
return;
+ cputime = tick_cputime();
+
if (irqtime_enabled()) {
- irqtime_account_process_tick(p, user_tick, 1);
+ irqtime_account_process_tick(p, user_tick, cputime);
return;
}
- cputime = TICK_NSEC;
steal = steal_account_process_time(ULONG_MAX);
if (steal >= cputime)
diff --git a/kernel/time/tick-sched.c b/kernel/time/tick-sched.c
index 6c3fea3867139..1c5cefbb10812 100644
--- a/kernel/time/tick-sched.c
+++ b/kernel/time/tick-sched.c
@@ -763,13 +763,20 @@ static ktime_t tick_forward_now(ktime_t expires, ktime_t now)
return expires + TICK_NSEC;
}
-static void tick_nohz_restart(struct tick_sched *ts, ktime_t now)
+static ktime_t tick_nohz_restart_expires(struct tick_sched *ts, ktime_t now)
{
ktime_t expires = ts->last_tick;
if (now >= expires)
expires = tick_forward_now(expires, now);
+ return expires;
+}
+
+static void tick_nohz_restart(struct tick_sched *ts, ktime_t now)
+{
+ ktime_t expires = tick_nohz_restart_expires(ts, now);
+
if (tick_sched_flag_test(ts, TS_FLAG_HIGHRES)) {
hrtimer_start(&ts->sched_timer, expires, HRTIMER_MODE_ABS_PINNED_HARD);
} else {
@@ -1329,6 +1336,16 @@ unsigned long tick_nohz_get_idle_calls_cpu(int cpu)
return ts->idle_calls;
}
+static void tick_nohz_dyntick_stop(struct tick_sched *ts, ktime_t now)
+{
+ ktime_t tick_start = now;
+
+ if (!tick_nohz_full_cpu(smp_processor_id()))
+ tick_start = tick_nohz_restart_expires(ts, now) - TICK_NSEC;
+
+ kcpustat_dyntick_stop(now, tick_start);
+}
+
void tick_nohz_idle_restart_tick(void)
{
struct tick_sched *ts = this_cpu_ptr(&tick_cpu_sched);
@@ -1341,7 +1358,7 @@ void tick_nohz_idle_restart_tick(void)
* no tiny amount of idle time is accounted twice.
*/
ts->idle_entrytime = ktime_get();
- kcpustat_dyntick_stop(ts->idle_entrytime);
+ tick_nohz_dyntick_stop(ts, ts->idle_entrytime);
tick_nohz_restart_sched_tick(ts, ts->idle_entrytime);
}
}
@@ -1385,7 +1402,7 @@ void tick_nohz_idle_exit(void)
if (tick_sched_flag_test(ts, TS_FLAG_STOPPED)) {
now = ktime_get();
- kcpustat_dyntick_stop(now);
+ tick_nohz_dyntick_stop(ts, now);
tick_nohz_idle_update_tick(ts, now);
}
base-commit: 72d3fcf802c45d00b300f25b848a93c3a2bd7c7e
--
2.43.0
^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [PATCH] sched/cputime: Don't account idle time twice after dyntick-idle
2026-10-04 18:47 ` [PATCH] sched/cputime: Don't account idle time twice after dyntick-idle Stian Halseth
@ 2026-10-05 4:25 ` Thorsten Leemhuis
2026-10-05 11:33 ` Frederic Weisbecker
1 sibling, 0 replies; 4+ messages in thread
From: Thorsten Leemhuis @ 2026-10-05 4:25 UTC (permalink / raw)
To: Stian Halseth, Frederic Weisbecker, Thomas Gleixner
Cc: Anna-Maria Behnsen, Ingo Molnar, Peter Zijlstra, Shrikanth Hegde,
regressions, linux-kernel, Ahmed Shaltout
On 10/4/26 20:47, Stian Halseth wrote:
> On idle exit the dyntick-idle accounting accounts the time up to now,
> then the tick is restarted on its old period. The first tick accounts a
> whole TICK_NSEC to whatever runs, although the part of that period
> before the idle exit has just been accounted as idle time, or with
> IRQ_TIME_ACCOUNTING as IRQ time. That is up to a full tick per idle
> exit, and /proc/stat reports more idle time than wall time.
>
> Record how much of the tick period had passed since the tick was
> stopped, and leave it out of the first tick.
+CC Ahmed Shaltout, who reported a problem that sounded somewhat similar
to me and might be helped by this (but I might be wrong there, this is
not my area of expertise):
https://lore.kernel.org/all/E76BE609-83E4-4FB6-88B8-191F7816EFD6@gmail.com/
@Ahmed, see also:
https://lore.kernel.org/all/20261004142724.3896396-1-stian@itx.no/
Ciao, Thorsten
> Fixes: cf6444c3e1bb7 ("tick/sched: Unify idle cputime accounting")
> Link: https://lore.kernel.org/all/20261004142724.3896396-1-stian@itx.no/
> Signed-off-by: Stian Halseth <stian@itx.no>
> ---
> Tested by comparing /proc/stat with CLOCK_MONOTONIC per CPU over 30 to
> 60 s, with a task on one CPU that sleeps in a loop. Total CPU time per
> wall second on that CPU:
>
> before after
> SPARC T7-1, HZ=100, busiest CPU* 1.44 1.0000
> SPARC T7-1, HZ=100, 3.7 ms sleeps 1.0001
> SPARC T7-1, HZ=100, 25 ms sleeps 0.9996
> Opteron, HZ=1000, 3.7 ms sleeps 1.131 0.999
> x86_64 KVM guest, HZ=250, 3.7 ms 1.53 0.998
> same guest, 9 ms sleeps 1.195 0.982
>
> * under its normal load, about 95 tick stops/s
>
> The guest was tested with and without IRQ_TIME_ACCOUNTING, with
> highres=off, with steal time from a busy loop on the host CPU, and with
> a test-only change that forces tick_nohz_idle_restart_tick() on every
> idle loop iteration with the tick stopped.
>
> The -1.8% left with long sleeps in the guest is time between an idle
> tick's expiry and its delivery to the halted vCPU, which nothing
> accounts. The summed lateness of those ticks matches it within 2 ms/s,
> or within 8 ms/s with steal time, also with highres=off. It is there
> without this patch as well, where the double counting hides it. On the
> T7-1 it is within noise.
>
> With 3.7 ms sleeps and steal time, +0.3% to +1.3% is left, which I
> have not explained.
>
> include/linux/kernel_stat.h | 6 ++++--
> kernel/sched/cputime.c | 31 +++++++++++++++++++++++++------
> kernel/time/tick-sched.c | 23 ++++++++++++++++++++---
> 3 files changed, 49 insertions(+), 11 deletions(-)
>
> diff --git a/include/linux/kernel_stat.h b/include/linux/kernel_stat.h
> index 9ca6c2259dfea..950867deea29f 100644
> --- a/include/linux/kernel_stat.h
> +++ b/include/linux/kernel_stat.h
> @@ -40,6 +40,8 @@ struct kernel_cpustat {
> seqcount_t idle_sleeptime_seq;
> u64 idle_entrytime;
> u64 idle_stealtime[2];
> + u64 idle_dyntick_entry;
> + u64 idle_tick_overlap;
> #endif
> u64 cpustat[NR_STATS];
> };
> @@ -111,7 +113,7 @@ static inline unsigned long kstat_cpu_irqs_sum(unsigned int cpu)
> #ifdef CONFIG_HAVE_VIRT_CPU_ACCOUNTING_IDLE
>
> static inline void kcpustat_dyntick_start(u64 now) { }
> -static inline void kcpustat_dyntick_stop(u64 now) { }
> +static inline void kcpustat_dyntick_stop(u64 now, u64 tick_start) { }
> static inline void kcpustat_irq_enter(u64 now) { }
> static inline void kcpustat_irq_exit(u64 now) { }
> static inline bool kcpustat_idle_dyntick(void) { return false; }
> @@ -132,7 +134,7 @@ static inline u64 kcpustat_field_iowait(int cpu)
> #else /* !CONFIG_HAVE_VIRT_CPU_ACCOUNTING_IDLE */
>
> extern void kcpustat_dyntick_start(u64 now);
> -extern void kcpustat_dyntick_stop(u64 now);
> +extern void kcpustat_dyntick_stop(u64 now, u64 tick_start);
> extern void kcpustat_irq_enter(u64 now);
> extern void kcpustat_irq_exit(u64 now);
> extern u64 kcpustat_field_idle(int cpu);
> diff --git a/kernel/sched/cputime.c b/kernel/sched/cputime.c
> index 06bddaa738e52..9205e92680943 100644
> --- a/kernel/sched/cputime.c
> +++ b/kernel/sched/cputime.c
> @@ -358,6 +358,19 @@ void thread_group_cputime(struct task_struct *tsk, struct task_cputime *times)
> }
> }
>
> +/*
> + * The first tick after dyntick-idle covers a period that the dyntick-idle
> + * accounting may already have accounted up to the idle exit.
> + */
> +static u64 tick_cputime(void)
> +{
> +#if defined(CONFIG_NO_HZ_COMMON) && !defined(CONFIG_HAVE_VIRT_CPU_ACCOUNTING_IDLE)
> + return TICK_NSEC - __this_cpu_xchg(kernel_cpustat.idle_tick_overlap, 0);
> +#else
> + return TICK_NSEC;
> +#endif
> +}
> +
> #ifdef CONFIG_IRQ_TIME_ACCOUNTING
> /*
> * Account a tick to a process and cpustat
> @@ -381,9 +394,9 @@ void thread_group_cputime(struct task_struct *tsk, struct task_cputime *times)
> * softirq as those do not count in task exec_runtime any more.
> */
> static void irqtime_account_process_tick(struct task_struct *p, int user_tick,
> - int ticks)
> + u64 cputime)
> {
> - u64 other, cputime = TICK_NSEC * ticks;
> + u64 other;
>
> /*
> * When returning from idle, many ticks can get accounted at
> @@ -418,7 +431,7 @@ static void irqtime_account_process_tick(struct task_struct *p, int user_tick,
>
> #else /* !CONFIG_IRQ_TIME_ACCOUNTING: */
> static inline void irqtime_account_process_tick(struct task_struct *p, int user_tick,
> - int nr_ticks) { }
> + u64 cputime) { }
> #endif /* !CONFIG_IRQ_TIME_ACCOUNTING */
>
> #if defined(CONFIG_NO_HZ_COMMON) && !defined(CONFIG_HAVE_VIRT_CPU_ACCOUNTING_IDLE)
> @@ -468,12 +481,15 @@ static void kcpustat_idle_start(struct kernel_cpustat *kc, u64 now)
> write_seqcount_end(&kc->idle_sleeptime_seq);
> }
>
> -void kcpustat_dyntick_stop(u64 now)
> +void kcpustat_dyntick_stop(u64 now, u64 tick_start)
> {
> struct kernel_cpustat *kc = kcpustat_this_cpu;
>
> if (!vtime_generic_enabled_this_cpu()) {
> WARN_ON_ONCE(!kc->idle_dyntick);
> + tick_start = max(tick_start, kc->idle_dyntick_entry);
> + if (now > tick_start)
> + kc->idle_tick_overlap = now - tick_start;
> kcpustat_idle_stop(kc, now);
> kc->idle_dyntick = false;
> vtime_dyntick_stop();
> @@ -487,6 +503,8 @@ void kcpustat_dyntick_start(u64 now)
> if (!vtime_generic_enabled_this_cpu()) {
> vtime_dyntick_start();
> kc->idle_dyntick = true;
> + kc->idle_dyntick_entry = now;
> + kc->idle_tick_overlap = 0;
> kcpustat_idle_start(kc, now);
> }
> }
> @@ -695,12 +713,13 @@ void account_process_tick(struct task_struct *p, int user_tick)
> if (kcpustat_idle_dyntick())
> return;
>
> + cputime = tick_cputime();
> +
> if (irqtime_enabled()) {
> - irqtime_account_process_tick(p, user_tick, 1);
> + irqtime_account_process_tick(p, user_tick, cputime);
> return;
> }
>
> - cputime = TICK_NSEC;
> steal = steal_account_process_time(ULONG_MAX);
>
> if (steal >= cputime)
> diff --git a/kernel/time/tick-sched.c b/kernel/time/tick-sched.c
> index 6c3fea3867139..1c5cefbb10812 100644
> --- a/kernel/time/tick-sched.c
> +++ b/kernel/time/tick-sched.c
> @@ -763,13 +763,20 @@ static ktime_t tick_forward_now(ktime_t expires, ktime_t now)
> return expires + TICK_NSEC;
> }
>
> -static void tick_nohz_restart(struct tick_sched *ts, ktime_t now)
> +static ktime_t tick_nohz_restart_expires(struct tick_sched *ts, ktime_t now)
> {
> ktime_t expires = ts->last_tick;
>
> if (now >= expires)
> expires = tick_forward_now(expires, now);
>
> + return expires;
> +}
> +
> +static void tick_nohz_restart(struct tick_sched *ts, ktime_t now)
> +{
> + ktime_t expires = tick_nohz_restart_expires(ts, now);
> +
> if (tick_sched_flag_test(ts, TS_FLAG_HIGHRES)) {
> hrtimer_start(&ts->sched_timer, expires, HRTIMER_MODE_ABS_PINNED_HARD);
> } else {
> @@ -1329,6 +1336,16 @@ unsigned long tick_nohz_get_idle_calls_cpu(int cpu)
> return ts->idle_calls;
> }
>
> +static void tick_nohz_dyntick_stop(struct tick_sched *ts, ktime_t now)
> +{
> + ktime_t tick_start = now;
> +
> + if (!tick_nohz_full_cpu(smp_processor_id()))
> + tick_start = tick_nohz_restart_expires(ts, now) - TICK_NSEC;
> +
> + kcpustat_dyntick_stop(now, tick_start);
> +}
> +
> void tick_nohz_idle_restart_tick(void)
> {
> struct tick_sched *ts = this_cpu_ptr(&tick_cpu_sched);
> @@ -1341,7 +1358,7 @@ void tick_nohz_idle_restart_tick(void)
> * no tiny amount of idle time is accounted twice.
> */
> ts->idle_entrytime = ktime_get();
> - kcpustat_dyntick_stop(ts->idle_entrytime);
> + tick_nohz_dyntick_stop(ts, ts->idle_entrytime);
> tick_nohz_restart_sched_tick(ts, ts->idle_entrytime);
> }
> }
> @@ -1385,7 +1402,7 @@ void tick_nohz_idle_exit(void)
>
> if (tick_sched_flag_test(ts, TS_FLAG_STOPPED)) {
> now = ktime_get();
> - kcpustat_dyntick_stop(now);
> + tick_nohz_dyntick_stop(ts, now);
> tick_nohz_idle_update_tick(ts, now);
> }
>
>
> base-commit: 72d3fcf802c45d00b300f25b848a93c3a2bd7c7e
^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [PATCH] sched/cputime: Don't account idle time twice after dyntick-idle
2026-10-04 18:47 ` [PATCH] sched/cputime: Don't account idle time twice after dyntick-idle Stian Halseth
2026-10-05 4:25 ` Thorsten Leemhuis
@ 2026-10-05 11:33 ` Frederic Weisbecker
1 sibling, 0 replies; 4+ messages in thread
From: Frederic Weisbecker @ 2026-10-05 11:33 UTC (permalink / raw)
To: Stian Halseth
Cc: Thomas Gleixner, Anna-Maria Behnsen, Ingo Molnar, Peter Zijlstra,
Shrikanth Hegde, regressions, linux-kernel
Le Sun, Oct 04, 2026 at 08:47:01PM +0200, Stian Halseth a écrit :
> On idle exit the dyntick-idle accounting accounts the time up to now,
> then the tick is restarted on its old period. The first tick accounts a
> whole TICK_NSEC to whatever runs, although the part of that period
> before the idle exit has just been accounted as idle time, or with
> IRQ_TIME_ACCOUNTING as IRQ time. That is up to a full tick per idle
> exit, and /proc/stat reports more idle time than wall time.
>
> Record how much of the tick period had passed since the tick was
> stopped, and leave it out of the first tick.
>
> Fixes: cf6444c3e1bb7 ("tick/sched: Unify idle cputime accounting")
> Link: https://lore.kernel.org/all/20261004142724.3896396-1-stian@itx.no/
> Signed-off-by: Stian Halseth <stian@itx.no>
That makes sense. Some comments below:
> ---
> Tested by comparing /proc/stat with CLOCK_MONOTONIC per CPU over 30 to
> 60 s, with a task on one CPU that sleeps in a loop. Total CPU time per
> wall second on that CPU:
>
> before after
> SPARC T7-1, HZ=100, busiest CPU* 1.44 1.0000
> SPARC T7-1, HZ=100, 3.7 ms sleeps 1.0001
> SPARC T7-1, HZ=100, 25 ms sleeps 0.9996
> Opteron, HZ=1000, 3.7 ms sleeps 1.131 0.999
> x86_64 KVM guest, HZ=250, 3.7 ms 1.53 0.998
> same guest, 9 ms sleeps 1.195 0.982
>
> * under its normal load, about 95 tick stops/s
>
> The guest was tested with and without IRQ_TIME_ACCOUNTING, with
> highres=off, with steal time from a busy loop on the host CPU, and with
> a test-only change that forces tick_nohz_idle_restart_tick() on every
> idle loop iteration with the tick stopped.
>
> The -1.8% left with long sleeps in the guest is time between an idle
> tick's expiry and its delivery to the halted vCPU, which nothing
> accounts. The summed lateness of those ticks matches it within 2 ms/s,
> or within 8 ms/s with steal time, also with highres=off. It is there
> without this patch as well, where the double counting hides it. On the
> T7-1 it is within noise.
>
> With 3.7 ms sleeps and steal time, +0.3% to +1.3% is left, which I
> have not explained.
Perhaps because sometimes idle is reentered shortly after exiting and
kcpustart_dyntick_start() overwrites the previous overlap. Say we have:
TICK_NSEC=100
next_tick=X
tick stop()
entry = X-90
tick_restart()
exit = entry + 10 (which is X-80)
overlap = 10
schedule()
// tick still hasn't fired
tick_stop()
entry = X-10
tick_restart()
exit = entry + 5 (which is X-5)
overlap = 5
tick()
tick accounts TICK_NSEC - 5 but it should also consider the previous
overlap, so it should be TICK_NSEC - 15.
But beware as it only matters for tick X, not for the next-next one that will
fire at X + TICK_NSEC, in which case the previous overlap should be ignored.
But please double check what I'm saying while being sleep deprived :-)
>
> include/linux/kernel_stat.h | 6 ++++--
> kernel/sched/cputime.c | 31 +++++++++++++++++++++++++------
> kernel/time/tick-sched.c | 23 ++++++++++++++++++++---
> 3 files changed, 49 insertions(+), 11 deletions(-)
>
> diff --git a/include/linux/kernel_stat.h b/include/linux/kernel_stat.h
> index 9ca6c2259dfea..950867deea29f 100644
> --- a/include/linux/kernel_stat.h
> +++ b/include/linux/kernel_stat.h
> @@ -40,6 +40,8 @@ struct kernel_cpustat {
> seqcount_t idle_sleeptime_seq;
> u64 idle_entrytime;
> u64 idle_stealtime[2];
> + u64 idle_dyntick_entry;
> + u64 idle_tick_overlap;
> #endif
> u64 cpustat[NR_STATS];
> };
> @@ -111,7 +113,7 @@ static inline unsigned long kstat_cpu_irqs_sum(unsigned int cpu)
> #ifdef CONFIG_HAVE_VIRT_CPU_ACCOUNTING_IDLE
>
> static inline void kcpustat_dyntick_start(u64 now) { }
> -static inline void kcpustat_dyntick_stop(u64 now) { }
> +static inline void kcpustat_dyntick_stop(u64 now, u64 tick_start) { }
> static inline void kcpustat_irq_enter(u64 now) { }
> static inline void kcpustat_irq_exit(u64 now) { }
> static inline bool kcpustat_idle_dyntick(void) { return false; }
> @@ -132,7 +134,7 @@ static inline u64 kcpustat_field_iowait(int cpu)
> #else /* !CONFIG_HAVE_VIRT_CPU_ACCOUNTING_IDLE */
>
> extern void kcpustat_dyntick_start(u64 now);
> -extern void kcpustat_dyntick_stop(u64 now);
> +extern void kcpustat_dyntick_stop(u64 now, u64 tick_start);
> extern void kcpustat_irq_enter(u64 now);
> extern void kcpustat_irq_exit(u64 now);
> extern u64 kcpustat_field_idle(int cpu);
> diff --git a/kernel/sched/cputime.c b/kernel/sched/cputime.c
> index 06bddaa738e52..9205e92680943 100644
> --- a/kernel/sched/cputime.c
> +++ b/kernel/sched/cputime.c
> @@ -358,6 +358,19 @@ void thread_group_cputime(struct task_struct *tsk, struct task_cputime *times)
> }
> }
>
> +/*
> + * The first tick after dyntick-idle covers a period that the dyntick-idle
> + * accounting may already have accounted up to the idle exit.
> + */
> +static u64 tick_cputime(void)
> +{
> +#if defined(CONFIG_NO_HZ_COMMON) && !defined(CONFIG_HAVE_VIRT_CPU_ACCOUNTING_IDLE)
> + return TICK_NSEC - __this_cpu_xchg(kernel_cpustat.idle_tick_overlap, 0);
> +#else
Please use IS_ENABLED()
Thanks.
--
Frederic Weisbecker
SUSE Labs
^ permalink raw reply [flat|nested] 4+ messages in thread
end of thread, other threads:[~2026-10-05 11:33 UTC | newest]
Thread overview: 4+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-04 14:27 [REGRESSION] tick/sched: /proc/stat idle time exceeds wall time since v7.2 Stian Halseth
2026-10-04 18:47 ` [PATCH] sched/cputime: Don't account idle time twice after dyntick-idle Stian Halseth
2026-10-05 4:25 ` Thorsten Leemhuis
2026-10-05 11:33 ` Frederic Weisbecker
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®