From: Madadi Vineeth Reddy <vineethr@linux.ibm.com>
To: Chen Yu <yu.c.chen@intel.com>,
Peter Zijlstra <peterz@infradead.org>,
Mathieu Desnoyers <mathieu.desnoyers@efficios.com>,
Ingo Molnar <mingo@redhat.com>,
Vincent Guittot <vincent.guittot@linaro.org>,
Juri Lelli <juri.lelli@redhat.com>
Cc: Tim Chen <tim.c.chen@intel.com>, Aaron Lu <aaron.lu@intel.com>,
Dietmar Eggemann <dietmar.eggemann@arm.com>,
Steven Rostedt <rostedt@goodmis.org>,
Ben Segall <bsegall@google.com>, Mel Gorman <mgorman@suse.de>,
Daniel Bristot de Oliveira <bristot@redhat.com>,
Valentin Schneider <vschneid@redhat.com>,
K Prateek Nayak <kprateek.nayak@amd.com>,
"Gautham R . Shenoy" <gautham.shenoy@amd.com>,
linux-kernel@vger.kernel.org, Chen Yu <yu.chen.surf@gmail.com>
Subject: Re: [PATCH 0/2] Introduce SIS_CACHE to choose previous CPU during task wakeup
Date: Tue, 17 Oct 2023 15:19:24 +0530 [thread overview]
Message-ID: <3f98806b-fd74-cfba-b48c-2526109d10a3@linux.ibm.com> (raw)
In-Reply-To: <cover.1695704179.git.yu.c.chen@intel.com>
Hi Chen Yu,
On 26/09/23 10:40, Chen Yu wrote:
> RFC -> v1:
> - drop RFC
> - Only record the short sleeping time for each task, to better honor the
> burst sleeping tasks. (Mathieu Desnoyers)
> - Keep the forward movement monotonic for runqueue's cache-hot timeout value.
> (Mathieu Desnoyers, Aaron Lu)
> - Introduce a new helper function cache_hot_cpu() that considers
> rq->cache_hot_timeout. (Aaron Lu)
> - Add analysis of why inhibiting task migration could bring better throughput
> for some benchmarks. (Gautham R. Shenoy)
> - Choose the first cache-hot CPU, if all idle CPUs are cache-hot in
> select_idle_cpu(). To avoid possible task stacking on the waker's CPU.
> (K Prateek Nayak)
>
> Thanks for your comments and review!
>
> ----------------------------------------------------------------------
Regarding making the scan for finding an idle cpu longer vs cache benefits,
I ran some benchmarks.
Tested the patch on power system with 12 cores. Total of 96 CPU's.
System has two NUMA nodes.
Below are some of the benchmark results
schbench 99.0th latency (lower is better)
========
case load baseline[pct imp](std%) SIS_CACHE[pct imp]( std%)
normal 1-mthreads 1.00 [ 0.00]( 3.66) 1.00 [ 0.00]( 1.71)
normal 2-mthreads 1.00 [ 0.00]( 4.55) 1.02 [ -2.00]( 3.00)
normal 4-mthreads 1.00 [ 0.00]( 4.77) 0.96 [ +4.00]( 4.27)
normal 6-mthreads 1.00 [ 0.00]( 60.37) 2.66 [ -166.00]( 23.67)
schbench results are showing that there is not much impact in wakeup latencies due to more iterations
in search for an idle cpu in the select_idle_cpu code path and interestingly numbers are slightly better
for SIS_CACHE in case of 4-mthreads. I think we can ignore the last case due to huge run to run variations.
producer_consumer avg time/access (lower is better)
========
loads per consumer iteration baseline[pct imp](std%) SIS_CACHE[pct imp]( std%)
5 1.00 [ 0.00]( 0.00) 0.87 [ +13.0]( 1.92)
20 1.00 [ 0.00]( 0.00) 0.92 [ +8.00]( 0.00)
50 1.00 [ 0.00]( 0.00) 1.00 [ 0.00]( 0.00)
100 1.00 [ 0.00]( 0.00) 1.00 [ 0.00]( 0.00)
The main goal of the patch of improving cache locality is reflected as SIS_CACHE only improves in this workload,
mainly when loads per consumer iteration is lower.
hackbench normalized time in seconds (lower is better)
========
case load baseline[pct imp](std%) SIS_CACHE[pct imp]( std%)
process-pipe 1-groups 1.00 [ 0.00]( 1.50) 1.02 [ -2.00]( 3.36)
process-pipe 2-groups 1.00 [ 0.00]( 4.76) 0.99 [ +1.00]( 5.68)
process-sockets 1-groups 1.00 [ 0.00]( 2.56) 1.00 [ 0.00]( 0.86)
process-sockets 2-groups 1.00 [ 0.00]( 0.50) 0.99 [ +1.00]( 0.96)
threads-pipe 1-groups 1.00 [ 0.00]( 3.87) 0.71 [ +29.0]( 3.56)
threads-pipe 2-groups 1.00 [ 0.00]( 1.60) 0.97 [ +3.00]( 3.44)
threads-sockets 1-groups 1.00 [ 0.00]( 7.65) 0.99 [ +1.00]( 1.05)
threads-sockets 2-groups 1.00 [ 0.00]( 3.12) 1.03 [ -3.00]( 1.70)
hackbench results are similar in both kernels except the case where there is an improvement of
29% in case of threads-pipe case with 1 groups.
Daytrader throughput (higher is better)
========
As per Ingo suggestion, ran a real life workload daytrader
baseline:
===================================================================================
Instance 1
Throughputs Ave. Resp. Time Min. Resp. Time Max. Resp. Time
================ =============== =============== ===============
10124.5 2 0 3970
SIS_CACHE:
===================================================================================
Instance 1
Throughputs Ave. Resp. Time Min. Resp. Time Max. Resp. Time
================ =============== =============== ===============
10319.5 2 0 5771
In the above run, daytrader perfomance was 2% better in case of SIS_CACHE.
Thanks and Regards
Madadi Vineeth Reddy
next prev parent reply other threads:[~2023-10-17 9:52 UTC|newest]
Thread overview: 18+ messages / expand[flat|nested] mbox.gz Atom feed top
2023-09-26 5:10 Chen Yu
2023-09-26 5:11 ` [PATCH 1/2] sched/fair: Record the short sleeping time of a task Chen Yu
2023-09-26 10:29 ` Mathieu Desnoyers
2023-09-27 6:17 ` Chen Yu
2023-09-27 7:53 ` Aaron Lu
2023-09-28 7:57 ` Chen Yu
2023-09-26 5:11 ` [PATCH 2/2] sched/fair: skip the cache hot CPU in select_idle_cpu() Chen Yu
2023-09-27 16:11 ` Mathieu Desnoyers
2023-09-28 7:52 ` Chen Yu
2023-09-27 8:00 ` [PATCH 0/2] Introduce SIS_CACHE to choose previous CPU during task wakeup Ingo Molnar
2023-09-27 21:34 ` Tim Chen
2023-09-28 8:23 ` Chen Yu
2023-10-05 6:22 ` K Prateek Nayak
2023-10-07 3:23 ` Chen Yu
2023-10-17 9:49 ` Madadi Vineeth Reddy [this message]
2023-10-17 11:09 ` Chen Yu
2023-10-18 19:32 ` Madadi Vineeth Reddy
2023-10-19 10:57 ` Chen Yu
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=3f98806b-fd74-cfba-b48c-2526109d10a3@linux.ibm.com \
--to=vineethr@linux.ibm.com \
--cc=aaron.lu@intel.com \
--cc=bristot@redhat.com \
--cc=bsegall@google.com \
--cc=cover.1695704179.git.yu.c.chen@intel.com \
--cc=dietmar.eggemann@arm.com \
--cc=gautham.shenoy@amd.com \
--cc=juri.lelli@redhat.com \
--cc=kprateek.nayak@amd.com \
--cc=linux-kernel@vger.kernel.org \
--cc=mathieu.desnoyers@efficios.com \
--cc=mgorman@suse.de \
--cc=mingo@redhat.com \
--cc=peterz@infradead.org \
--cc=rostedt@goodmis.org \
--cc=tim.c.chen@intel.com \
--cc=vincent.guittot@linaro.org \
--cc=vschneid@redhat.com \
--cc=yu.c.chen@intel.com \
--cc=yu.chen.surf@gmail.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®