From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1755927AbcH3Fyc (ORCPT ); Tue, 30 Aug 2016 01:54:32 -0400 Received: from mga01.intel.com ([192.55.52.88]:61767 "EHLO mga01.intel.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1752040AbcH3Fyb (ORCPT ); Tue, 30 Aug 2016 01:54:31 -0400 X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="5.30,254,1470726000"; d="scan'208";a="872456040" Subject: Re: [PATCH] thp: reduce usage of huge zero page's atomic counter To: Anshuman Khandual , Andrew Morton References: <20160829155021.2a85910c3d6b16a7f75ffccd@linux-foundation.org> <36b76a95-5025-ac64-0862-b98b2ebdeaf7@intel.com> <20160829203916.6a2b45845e8fb0c356cac17d@linux-foundation.org> <57C50F29.4070309@linux.vnet.ibm.com> Cc: Linux Memory Management List , "'Kirill A. Shutemov'" , Dave Hansen , Tim Chen , Huang Ying , Vlastimil Babka , Jerome Marchand , Andrea Arcangeli , Mel Gorman , Ebru Akagunduz , linux-kernel@vger.kernel.org From: Aaron Lu Message-ID: <0342377a-26b8-16b9-5817-1964fac0e12d@intel.com> Date: Tue, 30 Aug 2016 13:54:27 +0800 MIME-Version: 1.0 In-Reply-To: <57C50F29.4070309@linux.vnet.ibm.com> Content-Type: text/plain; charset=utf-8 Content-Transfer-Encoding: 7bit Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org On 08/30/2016 12:44 PM, Anshuman Khandual wrote: > On 08/30/2016 09:09 AM, Andrew Morton wrote: >> On Tue, 30 Aug 2016 11:09:15 +0800 Aaron Lu wrote: >> >>>>> Case used for test on Haswell EP: >>>>> usemem -n 72 --readonly -j 0x200000 100G >>>>> Which spawns 72 processes and each will mmap 100G anonymous space and >>>>> then do read only access to that space sequentially with a step of 2MB. >>>>> >>>>> perf report for base commit: >>>>> 54.03% usemem [kernel.kallsyms] [k] get_huge_zero_page >>>>> perf report for this commit: >>>>> 0.11% usemem [kernel.kallsyms] [k] mm_get_huge_zero_page >>>> >>>> Does this mean that overall usemem runtime halved? >>> >>> Sorry for the confusion, the above line is extracted from perf report. >>> It shows the percent of CPU cycles executed in a specific function. >>> >>> The above two perf lines are used to show get_huge_zero_page doesn't >>> consume that much CPU cycles after applying the patch. >>> >>>> >>>> Do we have any numbers for something which is more real-wordly? >>> >>> Unfortunately, no real world numbers. >>> >>> We think the global atomic counter could be an issue for performance >>> so I'm trying to solve the problem. >> >> So, umm, we don't actually know if the patch is useful to anyone? > > On a POWER system it improves the CPU consumption of the above mentioned > function a little bit. Dont think its going to improve actual throughput > of the workload substantially. > > 0.07% usemem [kernel.vmlinux] [k] mm_get_huge_zero_page I guess this is the base commit? But there shouldn't be the new mm_get_huge_zero_page symbol before this patch. A typo perhaps? Regards, Aaron > to > > 0.01% usemem [kernel.vmlinux] [k] mm_get_huge_zero_page >