From: "Doug Smythies" <dsmythies@telus.net>
To: "'Peter Zijlstra'" <peterz@infradead.org>
Cc: <linux-kernel@vger.kernel.org>, <vincent.guittot@linaro.org>,
"'Ingo Molnar'" <mingo@kernel.org>, <wuyun.abel@bytedance.com>,
"Doug Smythies" <dsmythies@telus.net>
Subject: RE: [REGRESSION] Re: [PATCH 00/24] Complete EEVDF
Date: Sat, 18 Jan 2025 16:09:02 -0800 [thread overview]
Message-ID: <004901db6a06$59b12050$0d1360f0$@telus.net> (raw)
In-Reply-To: <20250114105847.GC8385@noisy.programming.kicks-ass.net>
Hi Peter,
An update.
On 2025.01.14 02:59 Peter Zijlstra wrote:
> On Mon, Jan 13, 2025 at 12:03:12PM +0100, Peter Zijlstra wrote:
>> On Sun, Jan 12, 2025 at 03:14:17PM -0800, Doug Smythies wrote:
>>> means that there were 19 occurrences of turbostat interval times
>>> between 1.016 and 1.016999 seconds.
>>
>> OK, let me lower my threshold to 10ms and change the turbostat
>> invocation -- see if I can catch me some wabbits :-)
>
> I've had it run overnight and have not caught a single >10ms event :-(
Okay, so both you and I have many many hours of testing and
never see >= 10ms in that area of the turbostat code anymore.
The lingering >= 10ms (but I have never seen more than 25 ms)
is outside of that timing. As previously reported, I thought it might
be in the sampling interval sleep step, but I did a bunch of testing
and it doesn't appear to be there. That leaves:
delta_platform(&platform_counters_even, &platform_counters_odd);
compute_average(ODD_COUNTERS);
format_all_counters(ODD_COUNTERS);
flush_output_stdout();
I modified your tracing trigger thing in turbostat to this:
doug@s19:~/kernel/linux/tools/power/x86/turbostat$ git diff turbostat.c
diff --git a/tools/power/x86/turbostat/turbostat.c b/tools/power/x86/turbostat/turbostat.c
index 58a487c225a7..777efb64a754 100644
--- a/tools/power/x86/turbostat/turbostat.c
+++ b/tools/power/x86/turbostat/turbostat.c
@@ -67,6 +67,7 @@
#include <stdbool.h>
#include <assert.h>
#include <linux/kernel.h>
+#include <sys/syscall.h>
#define UNUSED(x) (void)(x)
@@ -2704,7 +2705,7 @@ int format_counters(struct thread_data *t, struct core_data *c, struct pkg_data
struct timeval tv;
timersub(&t->tv_end, &t->tv_begin, &tv);
- outp += sprintf(outp, "%5ld\t", tv.tv_sec * 1000000 + tv.tv_usec);
+ outp += sprintf(outp, "%7ld\t", tv.tv_sec * 1000000 + tv.tv_usec);
}
/* Time_Of_Day_Seconds: on each row, print sec.usec last timestamp taken */
@@ -2713,6 +2714,11 @@ int format_counters(struct thread_data *t, struct core_data *c, struct pkg_data
interval_float = t->tv_delta.tv_sec + t->tv_delta.tv_usec / 1000000.0;
+ double requested_interval = (double) interval_tv.tv_sec + (double) interval_tv.tv_usec / 1000000.0;
+
+ if(interval_float >= (requested_interval + 0.01)) /* was the last interval over by more than 10 mSec? */
+ syscall(__NR_gettimeofday, &tv_delta, (void*)1);
+
tsc = t->tsc * tsc_tweak;
/* topo columns, print blanks on 1st (average) line */
@@ -4570,12 +4576,14 @@ int get_counters(struct thread_data *t, struct core_data *c, struct pkg_data *p)
int i;
int status;
+ gettimeofday(&t->tv_begin, (struct timezone *)NULL); /* doug test */
+
if (cpu_migrate(cpu)) {
fprintf(outf, "%s: Could not migrate to CPU %d\n", __func__, cpu);
return -1;
}
- gettimeofday(&t->tv_begin, (struct timezone *)NULL);
+// gettimeofday(&t->tv_begin, (struct timezone *)NULL);
if (first_counter_read)
get_apic_id(t);
And so that I could prove a correlation with the trace times
and to my graph times I also did not turn off tracing upon a hit:
doug@s19:~/kernel/linux$ git diff kernel/time/time.c
diff --git a/kernel/time/time.c b/kernel/time/time.c
index 1b69caa87480..fb84915159cc 100644
--- a/kernel/time/time.c
+++ b/kernel/time/time.c
@@ -149,6 +149,12 @@ SYSCALL_DEFINE2(gettimeofday, struct __kernel_old_timeval __user *, tv,
return -EFAULT;
}
if (unlikely(tz != NULL)) {
+ if (tz == (void*)1) {
+ trace_printk("WHOOPSIE!\n");
+// tracing_off();
+ return 0;
+ }
+
if (copy_to_user(tz, &sys_tz, sizeof(sys_tz)))
return -EFAULT;
}
I ran a test for about 1 hour and 28 minutes.
The data in the trace correlates with turbostat line by line
TOD differentials. Trace got:
turbostat-1370 [011] ..... 751.738151: __x64_sys_gettimeofday: WHOOPSIE!
turbostat-1370 [011] ..... 760.763184: __x64_sys_gettimeofday: WHOOPSIE!
turbostat-1370 [011] ..... 1362.788298: __x64_sys_gettimeofday: WHOOPSIE!
turbostat-1370 [011] ..... 1365.815332: __x64_sys_gettimeofday: WHOOPSIE!
turbostat-1370 [011] ..... 1366.836340: __x64_sys_gettimeofday: WHOOPSIE!
turbostat-1370 [011] ..... 1367.856355: __x64_sys_gettimeofday: WHOOPSIE!
turbostat-1370 [011] ..... 1368.867365: __x64_sys_gettimeofday: WHOOPSIE!
turbostat-1370 [011] ..... 1373.893423: __x64_sys_gettimeofday: WHOOPSIE!
turbostat-1370 [011] ..... 1374.910439: __x64_sys_gettimeofday: WHOOPSIE!
turbostat-1370 [011] ..... 1377.928469: __x64_sys_gettimeofday: WHOOPSIE!
turbostat-1370 [011] ..... 1378.941483: __x64_sys_gettimeofday: WHOOPSIE!
turbostat-1370 [011] ..... 1379.959490: __x64_sys_gettimeofday: WHOOPSIE!
turbostat-1370 [011] ..... 1382.982525: __x64_sys_gettimeofday: WHOOPSIE!
turbostat-1370 [011] ..... 1385.005548: __x64_sys_gettimeofday: WHOOPSIE!
turbostat-1370 [011] ..... 1386.019561: __x64_sys_gettimeofday: WHOOPSIE!
turbostat-1370 [011] ..... 1387.030572: __x64_sys_gettimeofday: WHOOPSIE!
turbostat-1370 [011] ..... 1398.097683: __x64_sys_gettimeofday: WHOOPSIE!
turbostat-1370 [011] ..... 1620.752963: __x64_sys_gettimeofday: WHOOPSIE!
turbostat-1370 [011] ..... 1621.772969: __x64_sys_gettimeofday: WHOOPSIE!
turbostat-1370 [011] ..... 1622.788972: __x64_sys_gettimeofday: WHOOPSIE!
turbostat-1370 [011] ..... 1697.022098: __x64_sys_gettimeofday: WHOOPSIE!
turbostat-1370 [011] ..... 1703.071104: __x64_sys_gettimeofday: WHOOPSIE!
turbostat-1370 [011] ..... 1704.088103: __x64_sys_gettimeofday: WHOOPSIE!
turbostat-1370 [011] ..... 1705.105107: __x64_sys_gettimeofday: WHOOPSIE!
turbostat-1370 [011] ..... 1706.116106: __x64_sys_gettimeofday: WHOOPSIE!
turbostat-1370 [011] ..... 1707.126107: __x64_sys_gettimeofday: WHOOPSIE!
Going back to some old test data from when the CPU migration in turbostat
often took up to 6 seconds. If I subtract that migration time from the measured
interval time, I get a lot of samples between 10 and 23 ms.
I am saying there were 2 different issues. The 2nd was hidden by the 1st
because its magnitude was about 260 times less.
I do not know if my trace is any use. I'll compress it and send it to you only, off list.
My trace is as per this older email:
https://lore.kernel.org/all/20240727105030.226163742@infradead.org/T/#m453062b267551ff4786d33a2eb5f326f92241e96
next prev parent reply other threads:[~2025-01-19 0:09 UTC|newest]
Thread overview: 51+ messages / expand[flat|nested] mbox.gz Atom feed top
2024-12-29 22:51 Doug Smythies
2025-01-06 11:57 ` Peter Zijlstra
2025-01-06 15:01 ` Doug Smythies
2025-01-06 16:59 ` Peter Zijlstra
2025-01-06 17:04 ` Peter Zijlstra
2025-01-06 17:14 ` Peter Zijlstra
2025-01-07 1:24 ` Doug Smythies
2025-01-07 10:49 ` Peter Zijlstra
2025-01-06 22:28 ` Doug Smythies
2025-01-07 11:26 ` Peter Zijlstra
2025-01-07 15:04 ` Doug Smythies
2025-01-07 16:25 ` Doug Smythies
2025-01-07 19:23 ` Peter Zijlstra
2025-01-08 5:15 ` Doug Smythies
2025-01-08 13:12 ` Peter Zijlstra
2025-01-08 15:48 ` Doug Smythies
2025-01-09 10:59 ` Peter Zijlstra
2025-01-09 12:18 ` [tip: sched/urgent] sched/fair: Fix EEVDF entity placement bug causing scheduling lag tip-bot2 for Peter Zijlstra
2025-04-17 9:56 ` Alexander Egorenkov
2025-04-22 5:40 ` ll"RE: " Doug Smythies
2025-04-24 7:56 ` Alexander Egorenkov
2025-04-26 15:09 ` Doug Smythies
2025-01-10 5:09 ` [REGRESSION] Re: [PATCH 00/24] Complete EEVDF Doug Smythies
2025-01-10 11:57 ` Peter Zijlstra
2025-01-12 23:14 ` Doug Smythies
2025-01-13 11:03 ` Peter Zijlstra
2025-01-14 10:58 ` Peter Zijlstra
2025-01-14 15:15 ` Doug Smythies
2025-01-15 2:08 ` Len Brown
2025-01-15 16:47 ` Doug Smythies
2025-01-19 0:09 ` Doug Smythies [this message]
2025-01-20 3:55 ` Doug Smythies
2025-01-21 11:06 ` Peter Zijlstra
2025-01-21 8:49 ` Peter Zijlstra
2025-01-21 11:21 ` Peter Zijlstra
2025-01-21 15:58 ` Doug Smythies
2025-01-24 4:34 ` Doug Smythies
2025-01-24 11:04 ` Peter Zijlstra
2025-01-13 11:05 ` Peter Zijlstra
2025-01-13 16:01 ` Doug Smythies
2025-01-13 12:58 ` [tip: sched/urgent] sched/fair: Fix update_cfs_group() vs DELAY_DEQUEUE tip-bot2 for Peter Zijlstra
2025-01-12 19:59 ` [REGRESSION] Re: [PATCH 00/24] Complete EEVDF Doug Smythies
-- strict thread matches above, loose matches on Subject: below --
2024-07-27 10:27 Peter Zijlstra
2024-11-28 10:32 ` [REGRESSION] " Marcel Ziswiler
2024-11-28 10:58 ` Peter Zijlstra
2024-11-28 11:37 ` Marcel Ziswiler
2024-11-29 9:08 ` Peter Zijlstra
2024-12-02 18:46 ` Marcel Ziswiler
2024-12-09 9:49 ` Peter Zijlstra
2024-12-10 16:05 ` Marcel Ziswiler
2024-12-10 16:13 ` Steven Rostedt
2024-12-10 8:45 ` Luis Machado
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to='004901db6a06$59b12050$0d1360f0$@telus.net' \
--to=dsmythies@telus.net \
--cc=linux-kernel@vger.kernel.org \
--cc=mingo@kernel.org \
--cc=peterz@infradead.org \
--cc=vincent.guittot@linaro.org \
--cc=wuyun.abel@bytedance.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
Powered by JetHome