From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org X-Spam-Level: X-Spam-Status: No, score=-10.3 required=3.0 tests=BAYES_00, HEADER_FROM_DIFFERENT_DOMAINS,MAILING_LIST_MULTI,MENTIONS_GIT_HOSTING, NICE_REPLY_A,SPF_HELO_NONE,SPF_PASS,USER_AGENT_SANE_1 autolearn=unavailable autolearn_force=no version=3.4.0 Received: from mail.kernel.org (mail.kernel.org [198.145.29.99]) by smtp.lore.kernel.org (Postfix) with ESMTP id 62D1CC43460 for ; Wed, 5 May 2021 12:30:11 +0000 (UTC) Received: from vger.kernel.org (vger.kernel.org [23.128.96.18]) by mail.kernel.org (Postfix) with ESMTP id 22432611AC for ; Wed, 5 May 2021 12:30:11 +0000 (UTC) Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S233269AbhEEMbF (ORCPT ); Wed, 5 May 2021 08:31:05 -0400 Received: from foss.arm.com ([217.140.110.172]:43808 "EHLO foss.arm.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S231633AbhEEMbD (ORCPT ); Wed, 5 May 2021 08:31:03 -0400 Received: from usa-sjc-imap-foss1.foss.arm.com (unknown [10.121.207.14]) by usa-sjc-mx-foss1.foss.arm.com (Postfix) with ESMTP id AE4D231B; Wed, 5 May 2021 05:30:06 -0700 (PDT) Received: from [192.168.178.6] (unknown [172.31.20.19]) by usa-sjc-imap-foss1.foss.arm.com (Postfix) with ESMTPSA id 394AA3F70D; Wed, 5 May 2021 05:30:00 -0700 (PDT) Subject: Re: [RFC PATCH v6 3/4] scheduler: scan idle cpu in cluster for tasks within one LLC To: "Song Bao Hua (Barry Song)" , Vincent Guittot Cc: "tim.c.chen@linux.intel.com" , "catalin.marinas@arm.com" , "will@kernel.org" , "rjw@rjwysocki.net" , "bp@alien8.de" , "tglx@linutronix.de" , "mingo@redhat.com" , "lenb@kernel.org" , "peterz@infradead.org" , "rostedt@goodmis.org" , "bsegall@google.com" , "mgorman@suse.de" , "msys.mizuma@gmail.com" , "valentin.schneider@arm.com" , "gregkh@linuxfoundation.org" , Jonathan Cameron , "juri.lelli@redhat.com" , "mark.rutland@arm.com" , "sudeep.holla@arm.com" , "aubrey.li@linux.intel.com" , "linux-arm-kernel@lists.infradead.org" , "linux-kernel@vger.kernel.org" , "linux-acpi@vger.kernel.org" , "x86@kernel.org" , "xuwei (O)" , "Zengtao (B)" , "guodong.xu@linaro.org" , yangyicong , "Liguozhu (Kenneth)" , "linuxarm@openeuler.org" , "hpa@zytor.com" References: <20210420001844.9116-1-song.bao.hua@hisilicon.com> <20210420001844.9116-4-song.bao.hua@hisilicon.com> <80f489f9-8c88-95d8-8241-f0cfd2c2ac66@arm.com> <8b5277d9-e367-566d-6bd1-44ac78d21d3f@arm.com> <185746c4d02a485ca8f3509439328b26@hisilicon.com> <4d1f063504b1420c9f836d1f1a7f8e77@hisilicon.com> From: Dietmar Eggemann Message-ID: <142c7192-cde8-6dbe-bb9d-f0fce21ec959@arm.com> Date: Wed, 5 May 2021 14:29:54 +0200 User-Agent: Mozilla/5.0 (X11; Linux x86_64; rv:78.0) Gecko/20100101 Thunderbird/78.7.1 MIME-Version: 1.0 In-Reply-To: <4d1f063504b1420c9f836d1f1a7f8e77@hisilicon.com> Content-Type: text/plain; charset=utf-8 Content-Language: en-US Content-Transfer-Encoding: 7bit Precedence: bulk List-ID: X-Mailing-List: linux-kernel@vger.kernel.org On 03/05/2021 13:35, Song Bao Hua (Barry Song) wrote: [...] >> From: Song Bao Hua (Barry Song) [...] >>> From: Dietmar Eggemann [mailto:dietmar.eggemann@arm.com] [...] >>> On 29/04/2021 00:41, Song Bao Hua (Barry Song) wrote: >>>> >>>> >>>>> -----Original Message----- >>>>> From: Dietmar Eggemann [mailto:dietmar.eggemann@arm.com] >>> >>> [...] >>> >>>>>>>> From: Dietmar Eggemann [mailto:dietmar.eggemann@arm.com] >>>>> >>>>> [...] >>>>> >>>>>>>> On 20/04/2021 02:18, Barry Song wrote: [...] > > On the other hand, according to "sched: Implement smarter wake-affine logic" > https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=62470419 > > Proper factor in wake_wide is mainly beneficial of 1:n tasks like postgresql/pgbench. > So using the smaller cluster size as factor might help make wake_affine false so > improve pgbench. > > From the commit log, while clients = 2*cpus, the commit made the biggest > improvement. In my case, It should be clients=48 for a machine whose LLC > size is 24. > > In Linux, I created a 240MB database and ran "pgbench -c 48 -S -T 20 pgbench" > under two different scenarios: > 1. page cache always hit, so no real I/O for database read > 2. echo 3 > /proc/sys/vm/drop_caches > > For case 1, using cluster_size and using llc_size will result in similar > tps= ~108000, all of 24 cpus have 100% cpu utilization. > > For case 2, using llc_size still shows better performance. > > tps for each test round(cluster size as factor in wake_wide): > 1398.450887 1275.020401 1632.542437 1412.241627 1611.095692 1381.354294 1539.877146 > avg tps = 1464 > > tps for each test round(llc size as factor in wake_wide): > 1718.402983 1443.169823 1502.353823 1607.415861 1597.396924 1745.651814 1876.802168 > avg tps = 1641 (+12%) > > so it seems using cluster_size as factor in "slave >= factor && master >= slave * > factor" isn't a good choice for my machine at least. So SD size = 4 (instead of 24) seems to be too small for `-c 48`. Just curious, have you seen the benefit of using wake wide on SD size = 24 (LLC) compared to not using it at all?