From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org X-Spam-Level: X-Spam-Status: No, score=-3.7 required=3.0 tests=BAYES_00, HEADER_FROM_DIFFERENT_DOMAINS,MAILING_LIST_MULTI,SPF_HELO_NONE,SPF_PASS, URIBL_BLOCKED autolearn=no autolearn_force=no version=3.4.0 Received: from mail.kernel.org (mail.kernel.org [198.145.29.99]) by smtp.lore.kernel.org (Postfix) with ESMTP id 7F8C5C433E0 for ; Mon, 1 Feb 2021 21:50:09 +0000 (UTC) Received: from vger.kernel.org (vger.kernel.org [23.128.96.18]) by mail.kernel.org (Postfix) with ESMTP id 403A06024A for ; Mon, 1 Feb 2021 21:50:09 +0000 (UTC) Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S229634AbhBAVuD convert rfc822-to-8bit (ORCPT ); Mon, 1 Feb 2021 16:50:03 -0500 Received: from szxga02-in.huawei.com ([45.249.212.188]:3412 "EHLO szxga02-in.huawei.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S229498AbhBAVuC (ORCPT ); Mon, 1 Feb 2021 16:50:02 -0500 Received: from dggeme759-chm.china.huawei.com (unknown [172.30.72.55]) by szxga02-in.huawei.com (SkyGuard) with ESMTP id 4DV1lN1sQ8z5Mdc; Tue, 2 Feb 2021 05:48:00 +0800 (CST) Received: from dggemi761-chm.china.huawei.com (10.1.198.147) by dggeme759-chm.china.huawei.com (10.3.19.105) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_128_CBC_SHA256_P256) id 15.1.2106.2; Tue, 2 Feb 2021 05:49:17 +0800 Received: from dggemi761-chm.china.huawei.com ([10.9.49.202]) by dggemi761-chm.china.huawei.com ([10.9.49.202]) with mapi id 15.01.2106.006; Tue, 2 Feb 2021 05:49:17 +0800 From: "Song Bao Hua (Barry Song)" To: Valentin Schneider , "vincent.guittot@linaro.org" , "mgorman@suse.de" , "mingo@kernel.org" , "peterz@infradead.org" , "dietmar.eggemann@arm.com" , "morten.rasmussen@arm.com" , "linux-kernel@vger.kernel.org" CC: "linuxarm@openeuler.org" , "xuwei (O)" , "Liguozhu (Kenneth)" , "tiantao (H)" , wanghuiqiang , "Zengtao (B)" , Jonathan Cameron , "guodong.xu@linaro.org" , Meelis Roos Subject: RE: [PATCH] sched/topology: fix the issue groups don't span domain->span for NUMA diameter > 2 Thread-Topic: [PATCH] sched/topology: fix the issue groups don't span domain->span for NUMA diameter > 2 Thread-Index: AQHW+EyH3+RsPKpAu0uyCnuQYKixoapDFJoAgACzNkA= Date: Mon, 1 Feb 2021 21:49:17 +0000 Message-ID: <8cfd37e2617248f4b008f5564eb854a9@hisilicon.com> References: <20210201033830.15040-1-song.bao.hua@hisilicon.com> In-Reply-To: Accept-Language: en-GB, en-US Content-Language: en-US X-MS-Has-Attach: X-MS-TNEF-Correlator: x-originating-ip: [10.126.202.106] Content-Type: text/plain; charset="us-ascii" Content-Transfer-Encoding: 8BIT MIME-Version: 1.0 X-CFilter-Loop: Reflected Precedence: bulk List-ID: X-Mailing-List: linux-kernel@vger.kernel.org > -----Original Message----- > From: Valentin Schneider [mailto:valentin.schneider@arm.com] > Sent: Tuesday, February 2, 2021 7:11 AM > To: Song Bao Hua (Barry Song) ; > vincent.guittot@linaro.org; mgorman@suse.de; mingo@kernel.org; > peterz@infradead.org; dietmar.eggemann@arm.com; morten.rasmussen@arm.com; > linux-kernel@vger.kernel.org > Cc: linuxarm@openeuler.org; xuwei (O) ; Liguozhu (Kenneth) > ; tiantao (H) ; wanghuiqiang > ; Zengtao (B) ; Jonathan > Cameron ; guodong.xu@linaro.org; Song Bao Hua > (Barry Song) ; Meelis Roos > Subject: Re: [PATCH] sched/topology: fix the issue groups don't span > domain->span for NUMA diameter > 2 > > > Hi, > > On 01/02/21 16:38, Barry Song wrote: > > A tricky thing is that we shouldn't use the sgc of the 1st CPU of node2 > > for the sched_group generated by grandchild, otherwise, when this cpu > > becomes the balance_cpu of another sched_group of cpus other than node0, > > our sched_group generated by grandchild will access the same sgc with > > the sched_group generated by child of another CPU. > > > > So in init_overlap_sched_group(), sgc's capacity be overwritten: > > build_balance_mask(sd, sg, mask); > > cpu = cpumask_first_and(sched_group_span(sg), mask); > > > > sg->sgc = *per_cpu_ptr(sdd->sgc, cpu); > > > > And WARN_ON_ONCE(!cpumask_equal(group_balance_mask(sg), mask)) will > > also be triggered: > > static void init_overlap_sched_group(struct sched_domain *sd, > > struct sched_group *sg) > > { > > if (atomic_inc_return(&sg->sgc->ref) == 1) > > cpumask_copy(group_balance_mask(sg), mask); > > else > > WARN_ON_ONCE(!cpumask_equal(group_balance_mask(sg), mask)); > > } > > > > So here move to use the sgc of the 2nd cpu. For the corner case, if NUMA > > has only one CPU, we will still trigger this WARN_ON_ONCE. But It is > > really unlikely to be a real case for one NUMA to have one CPU only. > > > > Well, it's trivial to boot this with QEMU, and it's actually the example > the comment atop that WARN_ONCE() is based on. Also, you could end up with > a single CPU on a node during hotplug operations... Hi Valentin, The qemu topology is just a reflection of real kunpeng920 case, and pls also note Meelis has also tested on another real hardware "8-node Sun Fire X4600-M2" and gave the tested-by. It might not a perfect fix, but it is the simplest way to fix for this moment and for real cases. A "perfect" fix will require major refactoring of topology.c. I don't think hotplug is much relevant as even some cpus are unplugged and only one cpu is left in the sched_group of the sched_domain, the related domain and group are still getting right settings. On the other hand, the corner could literally be fixed, but will get some very ugly code involved. I mean, two sched_group can result in using the same sgc: 1. the sched_group generated by grandchild with only one numa 2. the sched_group generated by child with more than one numa Right now, I'm moving to the 2nd cpu for sched_group1, if we move to use 2nd cpu for sched_group2, then having only one cpu in one NUMA wouldn't be a problem anymore. But the code will be very ugly. So I would prefer to keep this assumption and just ignore the unreal corner case. > > I am not entirely sure whether having more than one CPU per node is a > sufficient condition. I'm starting to *think* it is, but I'm not entirely > convinced yet - and now I need a new notebook. Me too. Some extremely complicated topology might break the assumption. Really need a new notebook to draw this kind of complicated topology to break the assumption :-) But it is sufficient for the existing real cases which need fixing. When someday a real case in which each numa has more than one CPU wakes up the below warning: WARN_ON_ONCE(!cpumask_equal(group_balance_mask(sg), mask)). It might be the right time to consider major refactoring of topology.c. Thanks Barry