mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Frederic Weisbecker <frederic@kernel.org>
To: Ahmed Shaltout <ahmedshaltout.payment@gmail.com>
Cc: tglx@kernel.org, linux-kernel@vger.kernel.org
Subject: Re: [REGRESSION] /proc/stat idle time exceeds wall clock since v7.2
Date: Wed, 30 Sep 2026 14:35:27 +0200	[thread overview]
Message-ID: <ar0CD5n8yhJ2V1R_@localhost.localdomain> (raw)
In-Reply-To: <CAORRHKi+LauVxEaKopniK361qE18rV6j-3zdpKENvCWkRaE-HA@mail.gmail.com>

Le Wed, Sep 30, 2026 at 08:58:50AM +0400, Ahmed Shaltout a écrit :
> Hi Frederic,
> 
> 
> Le Tue, Sep 29, 2026, Frederic Weisbecker a écrit :
> 
> > I ran on KVM too, overcommiting the vcpus like you did (16 vCPUs on 8
> CPUs)
> 
> > and still nothing wrong.
> 
> >
> 
> > I'm wondering if this is specific to opensuse somehow. Can you try to
> build
> 
> > the latest upstream kernel?
> 
> 
> I will, but first one difference that may explain why your guest stays
> 
> clean, and a hypothesis that fits our numbers.
> 
> 
> Attached again is the 7.2.5 config, plus the 7.1.8 config from the control
> 
> box. Both are the stock openSUSE kernel-default files from /boot,
> unmodified.
> 
> Among the scheduler, tick, idle, accounting, preemption and RCU options they
> 
> differ only in CONFIG_SCHED_CACHE=y (7.2.5). All seven affected nodes run a
> 
> byte-identical config. None of the openSUSE-specific patches on top of
> 
> upstream stable touch tick, cputime, nohz, cpuidle or /proc. Going by their
> 
> names, they cover lockdown/secure boot, kABI, a few drivers and packaging.
> 
> 
> 1. There is no cpuidle driver on these guests
> 
> ---------------------------------------------
> 
> 
> On every node, affected or not:
> 
> 
>   /sys/devices/system/cpu/cpuidle/current_driver    none
> 
>   /sys/devices/system/cpu/cpuidle/current_governor  menu
> 
>   /sys/devices/system/cpu/cpu0/cpuidle/             no state* entries
> 
> 
> So cpuidle_idle_call() takes the cpuidle_not_available() branch and calls
> 
> idle_call_stop_or_retain_tick(got_tick). On the first pass through the idle
> 
> loop the tick is retained, and it is stopped only after a tick has fired in
> 
> idle. What does your guest report there? If a driver's governor stops the
> 
> tick at idle entry, that could be the difference.

I have the same cpuidle configuration: menu with no drivers.

> 2. Hypothesis: the tick after a dyntick-idle window re-counts part of it
> 
> ------------------------------------------------------------------------
> 
> 
> This comes from reading v7.2, not from a test. Take the case with no cpuidle
> 
> driver:
> 
> 
>   - The retained idle tick is charged to CPUTIME_IDLE by
> 
>     account_process_tick(), since kcpustat_idle_dyntick() is still false.
> 
>     Its irq exit sets ts->idle_entrytime. The window then opens there via
> 
>     kcpustat_dyntick_start(ts->idle_entrytime), so the start is seamless.
> 
> 
>   - On wakeup, tick_nohz_idle_exit() closes the window at "now"
> 
>     (kcpustat_dyntick_stop). tick_nohz_restart() re-arms the tick at the
> 
>     next jiffy boundary. That first tick charges a full TICK_NSEC, although
> 
>     the part of that jiffy before "now" is already in the window.
> 
> 
>   - The over-count is the wakeup's phase within its jiffy, once per
> 
>     window. Up to v7.1, /proc/stat read get_cpu_idle_time_us() for online
> 
>     CPUs and ignored tick-charged idle, so the two never met in one counter.
> 
> 
>   - When a governor stops the tick at idle entry instead, the window starts
> 
>     mid-jiffy. The fragment between the last tick and idle entry is then
> 
>     never charged, which would roughly offset the end seam on average. That
> 
>     might be why your guest looks right.

IIRC it seldom stops the tick at idle entry because it needs to wait for one
jiffy if this is a not short period.

> I have not verified this on a patched kernel, so please treat it as a lead
> 
> only.

Ok it's possible but why then do you reliably observe an excess when I never do?
 
> 3. Measurements, 60 s windows taken today on each node
> 
> -------------------------------------------------------
> 
> 
> These are the per-CPU cpuN lines of /proc/stat against /proc/uptime's first
> 
> field, and .idle_sleeps from /proc/timer_list over the same window.
> 
> "sum/wall" is all eight modes divided by uptime x nr_cpus.
> 
> 
>   node         cpus kernel  sum/wall  excess ms/s/cpu  sleeps/s/cpu
> ms/sleep
> 
>   cp-fsn1        8  7.2.5    1.100        100              351        0.285
> 
>   cp-nbg1-pps    8  7.2.5    1.080         80              268        0.299
> 
>   cp-nbg1-waa    8  7.2.5    1.072         74              245        0.301
> 
>   cp-hel1        8  7.2.5    1.081         83              278        0.297
> 
>   cp-hel1-b      8  7.2.5    1.076         76              255        0.298
> 
>   worker-fsn1    8  7.2.5    1.063         58              200        0.292
> 
>   monitoring     4  7.2.5    1.093         95              282        0.338
> 
>   staging        8  7.1.8    0.985        -14              263          -
> 
> 
> - Across nodes the excess tracks the tick-stop rate at about 0.3 ms per
> 
>   .idle_sleeps event. Within one node it does not: a CPU with twice the
> 
>   sleeps shows the same excess. .idle_sleeps is incremented on every
> 
>   __tick_nohz_idle_stop_tick() call, including re-stops after an IRQ inside
> 
>   an already-stopped idle period. So it overstates the number of windows on
> 
>   IRQ-heavy CPUs. These guests stop the tick 200-350 times/s per CPU. A
> 
>   near-idle test guest would show much less.
> 
> 
> - /proc/stat idle and /proc/uptime idle still agree to within 0.02 s on
> 
>   every node, over the same window.
> 
> 
> - The seven affected nodes were on older kernels on the same VMs until
> 
>   19 Sept: five on 7.1.8 and two on 7.0.12. Their daily all-mode sum read
> 
>   0.979-0.993 of wall clock. Since rebooting into 7.2.5 it reads 1.04-1.10.
> 
> 
> 4. Corrections to my previous mail
> 
> ----------------------------------
> 
> 
> - Not every node is AMD EPYC-Rome. cp-hel1 reports "Intel Xeon Processor
> 
>   (Skylake, IBRS, no TSX)" and shows the same excess (1.081), so it is not
> 
>   vendor-specific. All nodes are the same Hetzner cx43 VM type, except the
> 
>   4-vCPU one (cx33).
> 
> 
> - cp-hel1 is tainted W, from a boot-time "CPA detected W^X violation"
> 
>   warning in __change_page_attr. The other six are untainted and show the
> 
>   same excess.
> 
> 
> - I wrote that 7.1.8 sums to 0.992-0.997 of the ceiling. Over the last
> 
>   three days it is 0.983-0.985, so 7.1.8 slightly under-counts rather than
> 
>   matching wall clock exactly.
> 
> 
> If the cpuidle difference does not let you reproduce it, I will boot
> 
> openSUSE's kernel-vanilla (Kernel:HEAD, currently 7.3-rc5, no distro
> patches)
> 
> on a VM of the same type, with 7.2.5 and 7.1.8 on the same VM as controls,
> 
> and send the three results.

Yes please! Thanks a lot!

-- 
Frederic Weisbecker
SUSE Labs

      reply	other threads:[~2026-09-30 12:35 UTC|newest]

Thread overview: 6+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-19 23:14 Ahmed Shaltout
2026-09-22 12:53 ` Frederic Weisbecker
2026-09-22 15:53   ` Ahmed Shaltout
2026-09-29 13:50     ` Frederic Weisbecker
2026-09-30  4:58       ` Ahmed Shaltout
2026-09-30 12:35         ` Frederic Weisbecker [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=ar0CD5n8yhJ2V1R_@localhost.localdomain \
    --to=frederic@kernel.org \
    --cc=ahmedshaltout.payment@gmail.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=tglx@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®