From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from vger.kernel.org (vger.kernel.org [23.128.96.18]) by smtp.lore.kernel.org (Postfix) with ESMTP id 13418CDB46E for ; Thu, 12 Oct 2023 13:14:16 +0000 (UTC) Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1378767AbjJLNOP (ORCPT ); Thu, 12 Oct 2023 09:14:15 -0400 Received: from lindbergh.monkeyblade.net ([23.128.96.19]:53782 "EHLO lindbergh.monkeyblade.net" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1347216AbjJLNOM (ORCPT ); Thu, 12 Oct 2023 09:14:12 -0400 Received: from mgamail.intel.com (mgamail.intel.com [192.55.52.151]) by lindbergh.monkeyblade.net (Postfix) with ESMTPS id 7F202CC for ; Thu, 12 Oct 2023 06:14:09 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1697116449; x=1728652449; h=from:to:cc:subject:references:date:in-reply-to: message-id:mime-version; bh=eiWMQcDXi/+eJkMwzq0GtXmInqEDzIHYJMGyPtFb05g=; b=VcoY+xz6Pkww76vGO/KkzgrniybATqpL6KEVH3Wjy5BnhwyCigjCCxro S3fU5rnLF1AmmJ/JgVafWaRqNYuSEXeKZ3IrLHUjKUICSvfBcMJ4nfMv9 0L2/Ca1Thh+wnkaup7sYDhuXP6ZlE9pUrL8fdHuz3aQ1MlyTcN1pn92wG F0vBeO4eB814vJfrULuGo2Ks3qmTus4i+CaiNOcZithtWSSS4OX4KP4V7 q12IM9+lDQ+FzjfM68Yorqlxp+tPB0EUlK41r1O9eqWTa31FB9Mfj+LGz rRu+gHAQvsr3Y7K75pGFlybK4hm0geDD0k9HfiJWHJhRbcMMGNw90Yqe3 g==; X-IronPort-AV: E=McAfee;i="6600,9927,10861"; a="365187183" X-IronPort-AV: E=Sophos;i="6.03,219,1694761200"; d="scan'208";a="365187183" Received: from orsmga004.jf.intel.com ([10.7.209.38]) by fmsmga107.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 12 Oct 2023 06:14:09 -0700 X-ExtLoop1: 1 X-IronPort-AV: E=McAfee;i="6600,9927,10861"; a="878098259" X-IronPort-AV: E=Sophos;i="6.03,219,1694761200"; d="scan'208";a="878098259" Received: from yhuang6-desk2.sh.intel.com (HELO yhuang6-desk2.ccr.corp.intel.com) ([10.238.208.55]) by orsmga004-auth.jf.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 12 Oct 2023 06:14:05 -0700 From: "Huang, Ying" To: Mel Gorman Cc: , , Arjan Van De Ven , Sudeep Holla , Andrew Morton , Vlastimil Babka , "David Hildenbrand" , Johannes Weiner , "Dave Hansen" , Michal Hocko , "Pavel Tatashin" , Matthew Wilcox , Christoph Lameter Subject: Re: [PATCH 02/10] cacheinfo: calculate per-CPU data cache size References: <20230920061856.257597-1-ying.huang@intel.com> <20230920061856.257597-3-ying.huang@intel.com> <20231011122027.pw3uw32sdxxqjsrq@techsingularity.net> <87h6mwf3gf.fsf@yhuang6-desk2.ccr.corp.intel.com> <20231012125253.fpeehd6362c5v2sj@techsingularity.net> Date: Thu, 12 Oct 2023 21:12:00 +0800 In-Reply-To: <20231012125253.fpeehd6362c5v2sj@techsingularity.net> (Mel Gorman's message of "Thu, 12 Oct 2023 13:52:53 +0100") Message-ID: <87v8bcdly7.fsf@yhuang6-desk2.ccr.corp.intel.com> User-Agent: Gnus/5.13 (Gnus v5.13) Emacs/28.2 (gnu/linux) MIME-Version: 1.0 Content-Type: text/plain; charset=ascii Precedence: bulk List-ID: X-Mailing-List: linux-kernel@vger.kernel.org Mel Gorman writes: > On Thu, Oct 12, 2023 at 08:08:32PM +0800, Huang, Ying wrote: >> Mel Gorman writes: >> >> > On Wed, Sep 20, 2023 at 02:18:48PM +0800, Huang Ying wrote: >> >> Per-CPU data cache size is useful information. For example, it can be >> >> used to determine per-CPU cache size. So, in this patch, the data >> >> cache size for each CPU is calculated via data_cache_size / >> >> shared_cpu_weight. >> >> >> >> A brute-force algorithm to iterate all online CPUs is used to avoid >> >> to allocate an extra cpumask, especially in offline callback. >> >> >> >> Signed-off-by: "Huang, Ying" >> > >> > It's not necessarily relevant to the patch, but at least the scheduler >> > also stores some per-cpu topology information such as sd_llc_size -- the >> > number of CPUs sharing the same last-level-cache as this CPU. It may be >> > worth unifying this at some point if it's common that per-cpu >> > information is too fine and per-zone or per-node information is too >> > coarse. This would be particularly true when considering locking >> > granularity, >> > >> >> Cc: Sudeep Holla >> >> Cc: Andrew Morton >> >> Cc: Mel Gorman >> >> Cc: Vlastimil Babka >> >> Cc: David Hildenbrand >> >> Cc: Johannes Weiner >> >> Cc: Dave Hansen >> >> Cc: Michal Hocko >> >> Cc: Pavel Tatashin >> >> Cc: Matthew Wilcox >> >> Cc: Christoph Lameter >> >> --- >> >> drivers/base/cacheinfo.c | 42 ++++++++++++++++++++++++++++++++++++++- >> >> include/linux/cacheinfo.h | 1 + >> >> 2 files changed, 42 insertions(+), 1 deletion(-) >> >> >> >> diff --git a/drivers/base/cacheinfo.c b/drivers/base/cacheinfo.c >> >> index cbae8be1fe52..3e8951a3fbab 100644 >> >> --- a/drivers/base/cacheinfo.c >> >> +++ b/drivers/base/cacheinfo.c >> >> @@ -898,6 +898,41 @@ static int cache_add_dev(unsigned int cpu) >> >> return rc; >> >> } >> >> >> >> +static void update_data_cache_size_cpu(unsigned int cpu) >> >> +{ >> >> + struct cpu_cacheinfo *ci; >> >> + struct cacheinfo *leaf; >> >> + unsigned int i, nr_shared; >> >> + unsigned int size_data = 0; >> >> + >> >> + if (!per_cpu_cacheinfo(cpu)) >> >> + return; >> >> + >> >> + ci = ci_cacheinfo(cpu); >> >> + for (i = 0; i < cache_leaves(cpu); i++) { >> >> + leaf = per_cpu_cacheinfo_idx(cpu, i); >> >> + if (leaf->type != CACHE_TYPE_DATA && >> >> + leaf->type != CACHE_TYPE_UNIFIED) >> >> + continue; >> >> + nr_shared = cpumask_weight(&leaf->shared_cpu_map); >> >> + if (!nr_shared) >> >> + continue; >> >> + size_data += leaf->size / nr_shared; >> >> + } >> >> + ci->size_data = size_data; >> >> +} >> > >> > This needs comments. >> > >> > It would be nice to add a comment on top describing the limitation of >> > CACHE_TYPE_UNIFIED here in the context of >> > update_data_cache_size_cpu(). >> >> Sure. Will do that. >> > > Thanks. > >> > The L2 cache could be unified but much smaller than a L3 or other >> > last-level-cache. It's not clear from the code what level of cache is being >> > used due to a lack of familiarity of the cpu_cacheinfo code but size_data >> > is not the size of a cache, it appears to be the share of a cache a CPU >> > would have under ideal circumstances. >> >> Yes. And it isn't for one specific level of cache. It's sum of per-CPU >> shares of all levels of cache. But the calculation is inaccurate. More >> details are in the below reply. >> >> > However, as it appears to also be >> > iterating hierarchy then this may not be accurate. Caches may or may not >> > allow data to be duplicated between levels so the value may be inaccurate. >> >> Thank you very much for pointing this out! The cache can be inclusive >> or not. So, we cannot calculate the per-CPU slice of all-level caches >> via adding them together blindly. I will change this in a follow-on >> patch. >> > > Please do, I would strongly suggest basing this on LLC only because it's > the only value you can be sure of. This change is the only change that may > warrant a respin of the series as the history will be somewhat confusing > otherwise. I am still checking whether it's possible to get cache inclusive information via cpuid. If there's no reliable way to do that. We can use the max value of per-CPU share of each level of cache. For inclusive cache, that will be the value of LLC. For non-inclusive cache, the value will be more accurate. For example, on Intel Sapphire Rapids, the L2 cache is 2 MB per core, while LLC is 1.875 MB per core according to [1]. [1] https://www.intel.com/content/www/us/en/developer/articles/technical/fourth-generation-xeon-scalable-family-overview.html I will respin the series. Thanks a lot for review! -- Best Regards, Huang, Ying