* [PATCH] watchdog/softlockup:Fix incorrect CPU utilization output during softlockup
@ 2025-08-12 8:25 yaozhenguo
2025-08-18 2:18 ` Andrew Morton
0 siblings, 1 reply; 3+ messages in thread
From: yaozhenguo @ 2025-08-12 8:25 UTC (permalink / raw)
To: tglx, yaoma, akpm
Cc: max.kellermann, lihuafei1, yaozhenguo, linux-kernel, ZhenguoYao
From: ZhenguoYao <yaozhenguo1@gmail.com>
Since we use 16-bit precision, the raw data will undergo
integer division, which may sometimes result in data loss.
This can lead to slightly inaccurate CPU utilization calculations.
Under normal circumstances, this isn’t an issue. However,
when CPU utilization reaches 100%, the calculated result might
exceed 100%. For example, with raw data like the following:
sample_period 400000134 new_stat 83648414036 old_stat 83247417494
sample_period=400000134/2^24=23
new_stat=83648414036/2^24=4985
old_stat=83247417494/2^24=4961
util=105%
Below log will output:
CPU#3 Utilization every 0s during lockup:
#1: 0% system, 0% softirq, 105% hardirq, 0% idle
#2: 0% system, 0% softirq, 105% hardirq, 0% idle
#3: 0% system, 0% softirq, 100% hardirq, 0% idle
#4: 0% system, 0% softirq, 105% hardirq, 0% idle
#5: 0% system, 0% softirq, 105% hardirq, 0% idle
To avoid confusion, we enforce a 100% display cap when
calculations exceed this threshold.
Signed-off-by: ZhenguoYao <yaozhenguo1@gmail.com>
---
kernel/watchdog.c | 7 +++++++
1 file changed, 7 insertions(+)
diff --git a/kernel/watchdog.c b/kernel/watchdog.c
index 9c7134f7d2c4..29787996c69c 100644
--- a/kernel/watchdog.c
+++ b/kernel/watchdog.c
@@ -444,6 +444,13 @@ static void update_cpustat(void)
old_stat = __this_cpu_read(cpustat_old[i]);
new_stat = get_16bit_precision(cpustat[tracked_stats[i]]);
util = DIV_ROUND_UP(100 * (new_stat - old_stat), sample_period_16);
+ /* Since we use 16-bit precision, the raw data will undergo
+ * integer division, which may sometimes result in data loss,
+ * and then result might exceed 100%. To avoid confusion,
+ * we enforce a 100% display cap when calculations exceed this threshold.
+ */
+ if (util > 100)
+ util = 100;
__this_cpu_write(cpustat_util[tail][i], util);
__this_cpu_write(cpustat_old[i], new_stat);
}
--
2.43.5
^ permalink raw reply [flat|nested] 3+ messages in thread* Re: [PATCH] watchdog/softlockup:Fix incorrect CPU utilization output during softlockup
2025-08-12 8:25 [PATCH] watchdog/softlockup:Fix incorrect CPU utilization output during softlockup yaozhenguo
@ 2025-08-18 2:18 ` Andrew Morton
2025-08-18 8:16 ` Zhenguo Yao
0 siblings, 1 reply; 3+ messages in thread
From: Andrew Morton @ 2025-08-18 2:18 UTC (permalink / raw)
To: yaozhenguo
Cc: tglx, yaoma, max.kellermann, lihuafei1, yaozhenguo, linux-kernel
On Tue, 12 Aug 2025 16:25:10 +0800 yaozhenguo <yaozhenguo1@gmail.com> wrote:
> From: ZhenguoYao <yaozhenguo1@gmail.com>
>
> Since we use 16-bit precision, the raw data will undergo
> integer division, which may sometimes result in data loss.
> This can lead to slightly inaccurate CPU utilization calculations.
> Under normal circumstances, this isn’t an issue. However,
> when CPU utilization reaches 100%, the calculated result might
> exceed 100%. For example, with raw data like the following:
>
> sample_period 400000134 new_stat 83648414036 old_stat 83247417494
>
> sample_period=400000134/2^24=23
> new_stat=83648414036/2^24=4985
> old_stat=83247417494/2^24=4961
> util=105%
>
> Below log will output:
>
> CPU#3 Utilization every 0s during lockup:
> #1: 0% system, 0% softirq, 105% hardirq, 0% idle
> #2: 0% system, 0% softirq, 105% hardirq, 0% idle
> #3: 0% system, 0% softirq, 100% hardirq, 0% idle
> #4: 0% system, 0% softirq, 105% hardirq, 0% idle
> #5: 0% system, 0% softirq, 105% hardirq, 0% idle
>
> To avoid confusion, we enforce a 100% display cap when
> calculations exceed this threshold.
>
> ...
>
> --- a/kernel/watchdog.c
> +++ b/kernel/watchdog.c
> @@ -444,6 +444,13 @@ static void update_cpustat(void)
> old_stat = __this_cpu_read(cpustat_old[i]);
> new_stat = get_16bit_precision(cpustat[tracked_stats[i]]);
> util = DIV_ROUND_UP(100 * (new_stat - old_stat), sample_period_16);
> + /* Since we use 16-bit precision, the raw data will undergo
/*
* Since ...
please.
> + * integer division, which may sometimes result in data loss,
> + * and then result might exceed 100%. To avoid confusion,
> + * we enforce a 100% display cap when calculations exceed this threshold.
> + */
> + if (util > 100)
> + util = 100;
> __this_cpu_write(cpustat_util[tail][i], util);
> __this_cpu_write(cpustat_old[i], new_stat);
> }
Can we do something to make this output more accurate? For example,
return (data_ns + (1 << 23)) >> 24LL;
would round to the nearest multiple of 16.8ms?
^ permalink raw reply [flat|nested] 3+ messages in thread
* Re: [PATCH] watchdog/softlockup:Fix incorrect CPU utilization output during softlockup
2025-08-18 2:18 ` Andrew Morton
@ 2025-08-18 8:16 ` Zhenguo Yao
0 siblings, 0 replies; 3+ messages in thread
From: Zhenguo Yao @ 2025-08-18 8:16 UTC (permalink / raw)
To: Andrew Morton
Cc: tglx, yaoma, max.kellermann, lihuafei1, yaozhenguo, linux-kernel
Andrew Morton <akpm@linux-foundation.org> 于2025年8月18日周一 10:18写道:
>
> On Tue, 12 Aug 2025 16:25:10 +0800 yaozhenguo <yaozhenguo1@gmail.com> wrote:
>
> > From: ZhenguoYao <yaozhenguo1@gmail.com>
> >
> > Since we use 16-bit precision, the raw data will undergo
> > integer division, which may sometimes result in data loss.
> > This can lead to slightly inaccurate CPU utilization calculations.
> > Under normal circumstances, this isn’t an issue. However,
> > when CPU utilization reaches 100%, the calculated result might
> > exceed 100%. For example, with raw data like the following:
> >
> > sample_period 400000134 new_stat 83648414036 old_stat 83247417494
> >
> > sample_period=400000134/2^24=23
> > new_stat=83648414036/2^24=4985
> > old_stat=83247417494/2^24=4961
> > util=105%
> >
> > Below log will output:
> >
> > CPU#3 Utilization every 0s during lockup:
> > #1: 0% system, 0% softirq, 105% hardirq, 0% idle
> > #2: 0% system, 0% softirq, 105% hardirq, 0% idle
> > #3: 0% system, 0% softirq, 100% hardirq, 0% idle
> > #4: 0% system, 0% softirq, 105% hardirq, 0% idle
> > #5: 0% system, 0% softirq, 105% hardirq, 0% idle
> >
> > To avoid confusion, we enforce a 100% display cap when
> > calculations exceed this threshold.
> >
> > ...
> >
> > --- a/kernel/watchdog.c
> > +++ b/kernel/watchdog.c
> > @@ -444,6 +444,13 @@ static void update_cpustat(void)
> > old_stat = __this_cpu_read(cpustat_old[i]);
> > new_stat = get_16bit_precision(cpustat[tracked_stats[i]]);
> > util = DIV_ROUND_UP(100 * (new_stat - old_stat), sample_period_16);
> > + /* Since we use 16-bit precision, the raw data will undergo
>
> /*
> * Since ...
>
> please.
>
> > + * integer division, which may sometimes result in data loss,
> > + * and then result might exceed 100%. To avoid confusion,
> > + * we enforce a 100% display cap when calculations exceed this threshold.
> > + */
> > + if (util > 100)
> > + util = 100;
> > __this_cpu_write(cpustat_util[tail][i], util);
> > __this_cpu_write(cpustat_old[i], new_stat);
> > }
>
> Can we do something to make this output more accurate? For example,
>
> return (data_ns + (1 << 23)) >> 24LL;
>
> would round to the nearest multiple of 16.8ms?
>
>
Yes.
^ permalink raw reply [flat|nested] 3+ messages in thread
end of thread, other threads:[~2025-08-18 8:16 UTC | newest]
Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2025-08-12 8:25 [PATCH] watchdog/softlockup:Fix incorrect CPU utilization output during softlockup yaozhenguo
2025-08-18 2:18 ` Andrew Morton
2025-08-18 8:16 ` Zhenguo Yao
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®