From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from vger.kernel.org (vger.kernel.org [23.128.96.18]) by smtp.lore.kernel.org (Postfix) with ESMTP id DE3B1CDB482 for ; Mon, 16 Oct 2023 13:14:04 +0000 (UTC) Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S232155AbjJPNOE (ORCPT ); Mon, 16 Oct 2023 09:14:04 -0400 Received: from lindbergh.monkeyblade.net ([23.128.96.19]:44682 "EHLO lindbergh.monkeyblade.net" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S233600AbjJPMy2 (ORCPT ); Mon, 16 Oct 2023 08:54:28 -0400 Received: from szxga01-in.huawei.com (szxga01-in.huawei.com [45.249.212.187]) by lindbergh.monkeyblade.net (Postfix) with ESMTPS id 630FDB4 for ; Mon, 16 Oct 2023 05:54:25 -0700 (PDT) Received: from canpemm500009.china.huawei.com (unknown [172.30.72.55]) by szxga01-in.huawei.com (SkyGuard) with ESMTP id 4S8H4h32jxzvQ6D; Mon, 16 Oct 2023 20:49:40 +0800 (CST) Received: from [10.67.121.177] (10.67.121.177) by canpemm500009.china.huawei.com (7.192.105.203) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_128_GCM_SHA256) id 15.1.2507.31; Mon, 16 Oct 2023 20:54:22 +0800 CC: , , , , , , , , , , , , , , , , , , , , <21cnbao@gmail.com>, , Subject: Re: [PATCH v10 3/3] sched/fair: Use candidate prev/recent_used CPU if scanning failed for cluster wakeup To: Vincent Guittot References: <20231012121707.51368-1-yangyicong@huawei.com> <20231012121707.51368-4-yangyicong@huawei.com> From: Yicong Yang Message-ID: <33d8d0c1-da40-278b-5b84-ecb983ee9d34@huawei.com> Date: Mon, 16 Oct 2023 20:54:22 +0800 User-Agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:78.0) Gecko/20100101 Thunderbird/78.5.1 MIME-Version: 1.0 In-Reply-To: Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: 7bit X-Originating-IP: [10.67.121.177] X-ClientProxiedBy: dggems702-chm.china.huawei.com (10.3.19.179) To canpemm500009.china.huawei.com (7.192.105.203) X-CFilter-Loop: Reflected Precedence: bulk List-ID: X-Mailing-List: linux-kernel@vger.kernel.org Hi Vincent, On 2023/10/13 23:04, Vincent Guittot wrote: > On Thu, 12 Oct 2023 at 14:19, Yicong Yang wrote: >> >> From: Yicong Yang >> >> Chen Yu reports a hackbench regression of cluster wakeup when >> hackbench threads equal to the CPU number [1]. Analysis shows >> it's because we wake up more on the target CPU even if the >> prev_cpu is a good wakeup candidate and leads to the decrease >> of the CPU utilization. >> >> Generally if the task's prev_cpu is idle we'll wake up the task >> on it without scanning. On cluster machines we'll try to wake up >> the task in the same cluster of the target for better cache >> affinity, so if the prev_cpu is idle but not sharing the same >> cluster with the target we'll still try to find an idle CPU within >> the cluster. This will improve the performance at low loads on >> cluster machines. But in the issue above, if the prev_cpu is idle >> but not in the cluster with the target CPU, we'll try to scan an >> idle one in the cluster. But since the system is busy, we're >> likely to fail the scanning and use target instead, even if >> the prev_cpu is idle. Then leads to the regression. >> >> This patch solves this in 2 steps: >> o record the prev_cpu/recent_used_cpu if they're good wakeup >> candidates but not sharing the cluster with the target. >> o on scanning failure use the prev_cpu/recent_used_cpu if >> they're still idle >> >> [1] https://lore.kernel.org/all/ZGzDLuVaHR1PAYDt@chenyu5-mobl1/ >> Reported-by: Chen Yu >> Signed-off-by: Yicong Yang >> --- >> kernel/sched/fair.c | 19 ++++++++++++++++++- >> 1 file changed, 18 insertions(+), 1 deletion(-) >> >> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c >> index 4039f9b348ec..f1d94668bd71 100644 >> --- a/kernel/sched/fair.c >> +++ b/kernel/sched/fair.c >> @@ -7392,7 +7392,7 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target) >> bool has_idle_core = false; >> struct sched_domain *sd; >> unsigned long task_util, util_min, util_max; >> - int i, recent_used_cpu; >> + int i, recent_used_cpu, prev_aff = -1; >> >> /* >> * On asymmetric system, update task utilization because we will check >> @@ -7425,6 +7425,8 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target) >> >> if (cpus_share_resources(prev, target)) >> return prev; >> + >> + prev_aff = prev; >> } >> >> /* >> @@ -7457,6 +7459,8 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target) >> >> if (cpus_share_resources(recent_used_cpu, target)) >> return recent_used_cpu; >> + } else { >> + recent_used_cpu = -1; >> } >> >> /* >> @@ -7497,6 +7501,19 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target) >> if ((unsigned)i < nr_cpumask_bits) >> return i; >> >> + /* >> + * For cluster machines which have lower sharing cache like L2 or >> + * LLC Tag, we tend to find an idle CPU in the target's cluster >> + * first. But prev_cpu or recent_used_cpu may also be a good candidate, >> + * use them if possible when no idle CPU found in select_idle_cpu(). >> + */ >> + if ((unsigned int)prev_aff < nr_cpumask_bits && >> + (available_idle_cpu(prev_aff) || sched_idle_cpu(prev_aff))) > > Hasn't prev_aff (i.e. prev) been already tested as idle ? > >> + return prev_aff; >> + if ((unsigned int)recent_used_cpu < nr_cpumask_bits && >> + (available_idle_cpu(recent_used_cpu) || sched_idle_cpu(recent_used_cpu))) >> + return recent_used_cpu; > > same here > It was thought that there maybe a small potential race window here that the prev/recent_used CPU becoming non-idle after scanning, discussed in [1]. I think the check here won't be expensive so added it here. It should be redundant and can be removed. [1] https://lore.kernel.org/all/ZIams6s+qShFWhfQ@BLR-5CG11610CF.amd.com/ Thanks. > >> + >> return target; >> } >> >> -- >> 2.24.0 >> >