From: Frederic Weisbecker <frederic@kernel.org>
To: Ahmed Shaltout <ahmedshaltout.payment@gmail.com>
Cc: tglx@kernel.org, linux-kernel@vger.kernel.org
Subject: Re: [REGRESSION] /proc/stat idle time exceeds wall clock since v7.2
Date: Wed, 30 Sep 2026 14:35:27 +0200 [thread overview]
Message-ID: <ar0CD5n8yhJ2V1R_@localhost.localdomain> (raw)
In-Reply-To: <CAORRHKi+LauVxEaKopniK361qE18rV6j-3zdpKENvCWkRaE-HA@mail.gmail.com>
Le Wed, Sep 30, 2026 at 08:58:50AM +0400, Ahmed Shaltout a écrit :
> Hi Frederic,
>
>
> Le Tue, Sep 29, 2026, Frederic Weisbecker a écrit :
>
> > I ran on KVM too, overcommiting the vcpus like you did (16 vCPUs on 8
> CPUs)
>
> > and still nothing wrong.
>
> >
>
> > I'm wondering if this is specific to opensuse somehow. Can you try to
> build
>
> > the latest upstream kernel?
>
>
> I will, but first one difference that may explain why your guest stays
>
> clean, and a hypothesis that fits our numbers.
>
>
> Attached again is the 7.2.5 config, plus the 7.1.8 config from the control
>
> box. Both are the stock openSUSE kernel-default files from /boot,
> unmodified.
>
> Among the scheduler, tick, idle, accounting, preemption and RCU options they
>
> differ only in CONFIG_SCHED_CACHE=y (7.2.5). All seven affected nodes run a
>
> byte-identical config. None of the openSUSE-specific patches on top of
>
> upstream stable touch tick, cputime, nohz, cpuidle or /proc. Going by their
>
> names, they cover lockdown/secure boot, kABI, a few drivers and packaging.
>
>
> 1. There is no cpuidle driver on these guests
>
> ---------------------------------------------
>
>
> On every node, affected or not:
>
>
> /sys/devices/system/cpu/cpuidle/current_driver none
>
> /sys/devices/system/cpu/cpuidle/current_governor menu
>
> /sys/devices/system/cpu/cpu0/cpuidle/ no state* entries
>
>
> So cpuidle_idle_call() takes the cpuidle_not_available() branch and calls
>
> idle_call_stop_or_retain_tick(got_tick). On the first pass through the idle
>
> loop the tick is retained, and it is stopped only after a tick has fired in
>
> idle. What does your guest report there? If a driver's governor stops the
>
> tick at idle entry, that could be the difference.
I have the same cpuidle configuration: menu with no drivers.
> 2. Hypothesis: the tick after a dyntick-idle window re-counts part of it
>
> ------------------------------------------------------------------------
>
>
> This comes from reading v7.2, not from a test. Take the case with no cpuidle
>
> driver:
>
>
> - The retained idle tick is charged to CPUTIME_IDLE by
>
> account_process_tick(), since kcpustat_idle_dyntick() is still false.
>
> Its irq exit sets ts->idle_entrytime. The window then opens there via
>
> kcpustat_dyntick_start(ts->idle_entrytime), so the start is seamless.
>
>
> - On wakeup, tick_nohz_idle_exit() closes the window at "now"
>
> (kcpustat_dyntick_stop). tick_nohz_restart() re-arms the tick at the
>
> next jiffy boundary. That first tick charges a full TICK_NSEC, although
>
> the part of that jiffy before "now" is already in the window.
>
>
> - The over-count is the wakeup's phase within its jiffy, once per
>
> window. Up to v7.1, /proc/stat read get_cpu_idle_time_us() for online
>
> CPUs and ignored tick-charged idle, so the two never met in one counter.
>
>
> - When a governor stops the tick at idle entry instead, the window starts
>
> mid-jiffy. The fragment between the last tick and idle entry is then
>
> never charged, which would roughly offset the end seam on average. That
>
> might be why your guest looks right.
IIRC it seldom stops the tick at idle entry because it needs to wait for one
jiffy if this is a not short period.
> I have not verified this on a patched kernel, so please treat it as a lead
>
> only.
Ok it's possible but why then do you reliably observe an excess when I never do?
> 3. Measurements, 60 s windows taken today on each node
>
> -------------------------------------------------------
>
>
> These are the per-CPU cpuN lines of /proc/stat against /proc/uptime's first
>
> field, and .idle_sleeps from /proc/timer_list over the same window.
>
> "sum/wall" is all eight modes divided by uptime x nr_cpus.
>
>
> node cpus kernel sum/wall excess ms/s/cpu sleeps/s/cpu
> ms/sleep
>
> cp-fsn1 8 7.2.5 1.100 100 351 0.285
>
> cp-nbg1-pps 8 7.2.5 1.080 80 268 0.299
>
> cp-nbg1-waa 8 7.2.5 1.072 74 245 0.301
>
> cp-hel1 8 7.2.5 1.081 83 278 0.297
>
> cp-hel1-b 8 7.2.5 1.076 76 255 0.298
>
> worker-fsn1 8 7.2.5 1.063 58 200 0.292
>
> monitoring 4 7.2.5 1.093 95 282 0.338
>
> staging 8 7.1.8 0.985 -14 263 -
>
>
> - Across nodes the excess tracks the tick-stop rate at about 0.3 ms per
>
> .idle_sleeps event. Within one node it does not: a CPU with twice the
>
> sleeps shows the same excess. .idle_sleeps is incremented on every
>
> __tick_nohz_idle_stop_tick() call, including re-stops after an IRQ inside
>
> an already-stopped idle period. So it overstates the number of windows on
>
> IRQ-heavy CPUs. These guests stop the tick 200-350 times/s per CPU. A
>
> near-idle test guest would show much less.
>
>
> - /proc/stat idle and /proc/uptime idle still agree to within 0.02 s on
>
> every node, over the same window.
>
>
> - The seven affected nodes were on older kernels on the same VMs until
>
> 19 Sept: five on 7.1.8 and two on 7.0.12. Their daily all-mode sum read
>
> 0.979-0.993 of wall clock. Since rebooting into 7.2.5 it reads 1.04-1.10.
>
>
> 4. Corrections to my previous mail
>
> ----------------------------------
>
>
> - Not every node is AMD EPYC-Rome. cp-hel1 reports "Intel Xeon Processor
>
> (Skylake, IBRS, no TSX)" and shows the same excess (1.081), so it is not
>
> vendor-specific. All nodes are the same Hetzner cx43 VM type, except the
>
> 4-vCPU one (cx33).
>
>
> - cp-hel1 is tainted W, from a boot-time "CPA detected W^X violation"
>
> warning in __change_page_attr. The other six are untainted and show the
>
> same excess.
>
>
> - I wrote that 7.1.8 sums to 0.992-0.997 of the ceiling. Over the last
>
> three days it is 0.983-0.985, so 7.1.8 slightly under-counts rather than
>
> matching wall clock exactly.
>
>
> If the cpuidle difference does not let you reproduce it, I will boot
>
> openSUSE's kernel-vanilla (Kernel:HEAD, currently 7.3-rc5, no distro
> patches)
>
> on a VM of the same type, with 7.2.5 and 7.1.8 on the same VM as controls,
>
> and send the three results.
Yes please! Thanks a lot!
--
Frederic Weisbecker
SUSE Labs
prev parent reply other threads:[~2026-09-30 12:35 UTC|newest]
Thread overview: 6+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-19 23:14 Ahmed Shaltout
2026-09-22 12:53 ` Frederic Weisbecker
2026-09-22 15:53 ` Ahmed Shaltout
2026-09-29 13:50 ` Frederic Weisbecker
2026-09-30 4:58 ` Ahmed Shaltout
2026-09-30 12:35 ` Frederic Weisbecker [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=ar0CD5n8yhJ2V1R_@localhost.localdomain \
--to=frederic@kernel.org \
--cc=ahmedshaltout.payment@gmail.com \
--cc=linux-kernel@vger.kernel.org \
--cc=tglx@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®