From: Frederic Weisbecker <frederic@kernel.org>
To: Stian Halseth <stian@itx.no>
Cc: Thomas Gleixner <tglx@kernel.org>,
Anna-Maria Behnsen <anna-maria@linutronix.de>,
Ingo Molnar <mingo@redhat.com>,
Peter Zijlstra <peterz@infradead.org>,
Shrikanth Hegde <sshegde@linux.ibm.com>,
regressions@lists.linux.dev, linux-kernel@vger.kernel.org
Subject: Re: [PATCH] sched/cputime: Don't account idle time twice after dyntick-idle
Date: Mon, 5 Oct 2026 13:33:00 +0200 [thread overview]
Message-ID: <asOK7M4Re5aYqmsl@localhost.localdomain> (raw)
In-Reply-To: <20261004184701.4112237-1-stian@itx.no>
Le Sun, Oct 04, 2026 at 08:47:01PM +0200, Stian Halseth a écrit :
> On idle exit the dyntick-idle accounting accounts the time up to now,
> then the tick is restarted on its old period. The first tick accounts a
> whole TICK_NSEC to whatever runs, although the part of that period
> before the idle exit has just been accounted as idle time, or with
> IRQ_TIME_ACCOUNTING as IRQ time. That is up to a full tick per idle
> exit, and /proc/stat reports more idle time than wall time.
>
> Record how much of the tick period had passed since the tick was
> stopped, and leave it out of the first tick.
>
> Fixes: cf6444c3e1bb7 ("tick/sched: Unify idle cputime accounting")
> Link: https://lore.kernel.org/all/20261004142724.3896396-1-stian@itx.no/
> Signed-off-by: Stian Halseth <stian@itx.no>
That makes sense. Some comments below:
> ---
> Tested by comparing /proc/stat with CLOCK_MONOTONIC per CPU over 30 to
> 60 s, with a task on one CPU that sleeps in a loop. Total CPU time per
> wall second on that CPU:
>
> before after
> SPARC T7-1, HZ=100, busiest CPU* 1.44 1.0000
> SPARC T7-1, HZ=100, 3.7 ms sleeps 1.0001
> SPARC T7-1, HZ=100, 25 ms sleeps 0.9996
> Opteron, HZ=1000, 3.7 ms sleeps 1.131 0.999
> x86_64 KVM guest, HZ=250, 3.7 ms 1.53 0.998
> same guest, 9 ms sleeps 1.195 0.982
>
> * under its normal load, about 95 tick stops/s
>
> The guest was tested with and without IRQ_TIME_ACCOUNTING, with
> highres=off, with steal time from a busy loop on the host CPU, and with
> a test-only change that forces tick_nohz_idle_restart_tick() on every
> idle loop iteration with the tick stopped.
>
> The -1.8% left with long sleeps in the guest is time between an idle
> tick's expiry and its delivery to the halted vCPU, which nothing
> accounts. The summed lateness of those ticks matches it within 2 ms/s,
> or within 8 ms/s with steal time, also with highres=off. It is there
> without this patch as well, where the double counting hides it. On the
> T7-1 it is within noise.
>
> With 3.7 ms sleeps and steal time, +0.3% to +1.3% is left, which I
> have not explained.
Perhaps because sometimes idle is reentered shortly after exiting and
kcpustart_dyntick_start() overwrites the previous overlap. Say we have:
TICK_NSEC=100
next_tick=X
tick stop()
entry = X-90
tick_restart()
exit = entry + 10 (which is X-80)
overlap = 10
schedule()
// tick still hasn't fired
tick_stop()
entry = X-10
tick_restart()
exit = entry + 5 (which is X-5)
overlap = 5
tick()
tick accounts TICK_NSEC - 5 but it should also consider the previous
overlap, so it should be TICK_NSEC - 15.
But beware as it only matters for tick X, not for the next-next one that will
fire at X + TICK_NSEC, in which case the previous overlap should be ignored.
But please double check what I'm saying while being sleep deprived :-)
>
> include/linux/kernel_stat.h | 6 ++++--
> kernel/sched/cputime.c | 31 +++++++++++++++++++++++++------
> kernel/time/tick-sched.c | 23 ++++++++++++++++++++---
> 3 files changed, 49 insertions(+), 11 deletions(-)
>
> diff --git a/include/linux/kernel_stat.h b/include/linux/kernel_stat.h
> index 9ca6c2259dfea..950867deea29f 100644
> --- a/include/linux/kernel_stat.h
> +++ b/include/linux/kernel_stat.h
> @@ -40,6 +40,8 @@ struct kernel_cpustat {
> seqcount_t idle_sleeptime_seq;
> u64 idle_entrytime;
> u64 idle_stealtime[2];
> + u64 idle_dyntick_entry;
> + u64 idle_tick_overlap;
> #endif
> u64 cpustat[NR_STATS];
> };
> @@ -111,7 +113,7 @@ static inline unsigned long kstat_cpu_irqs_sum(unsigned int cpu)
> #ifdef CONFIG_HAVE_VIRT_CPU_ACCOUNTING_IDLE
>
> static inline void kcpustat_dyntick_start(u64 now) { }
> -static inline void kcpustat_dyntick_stop(u64 now) { }
> +static inline void kcpustat_dyntick_stop(u64 now, u64 tick_start) { }
> static inline void kcpustat_irq_enter(u64 now) { }
> static inline void kcpustat_irq_exit(u64 now) { }
> static inline bool kcpustat_idle_dyntick(void) { return false; }
> @@ -132,7 +134,7 @@ static inline u64 kcpustat_field_iowait(int cpu)
> #else /* !CONFIG_HAVE_VIRT_CPU_ACCOUNTING_IDLE */
>
> extern void kcpustat_dyntick_start(u64 now);
> -extern void kcpustat_dyntick_stop(u64 now);
> +extern void kcpustat_dyntick_stop(u64 now, u64 tick_start);
> extern void kcpustat_irq_enter(u64 now);
> extern void kcpustat_irq_exit(u64 now);
> extern u64 kcpustat_field_idle(int cpu);
> diff --git a/kernel/sched/cputime.c b/kernel/sched/cputime.c
> index 06bddaa738e52..9205e92680943 100644
> --- a/kernel/sched/cputime.c
> +++ b/kernel/sched/cputime.c
> @@ -358,6 +358,19 @@ void thread_group_cputime(struct task_struct *tsk, struct task_cputime *times)
> }
> }
>
> +/*
> + * The first tick after dyntick-idle covers a period that the dyntick-idle
> + * accounting may already have accounted up to the idle exit.
> + */
> +static u64 tick_cputime(void)
> +{
> +#if defined(CONFIG_NO_HZ_COMMON) && !defined(CONFIG_HAVE_VIRT_CPU_ACCOUNTING_IDLE)
> + return TICK_NSEC - __this_cpu_xchg(kernel_cpustat.idle_tick_overlap, 0);
> +#else
Please use IS_ENABLED()
Thanks.
--
Frederic Weisbecker
SUSE Labs
prev parent reply other threads:[~2026-10-05 11:33 UTC|newest]
Thread overview: 4+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-04 14:27 [REGRESSION] tick/sched: /proc/stat idle time exceeds wall time since v7.2 Stian Halseth
2026-10-04 18:47 ` [PATCH] sched/cputime: Don't account idle time twice after dyntick-idle Stian Halseth
2026-10-05 4:25 ` Thorsten Leemhuis
2026-10-05 11:33 ` Frederic Weisbecker [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=asOK7M4Re5aYqmsl@localhost.localdomain \
--to=frederic@kernel.org \
--cc=anna-maria@linutronix.de \
--cc=linux-kernel@vger.kernel.org \
--cc=mingo@redhat.com \
--cc=peterz@infradead.org \
--cc=regressions@lists.linux.dev \
--cc=sshegde@linux.ibm.com \
--cc=stian@itx.no \
--cc=tglx@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®