From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1756922AbdEKUrc (ORCPT ); Thu, 11 May 2017 16:47:32 -0400 Received: from userp1040.oracle.com ([156.151.31.81]:46992 "EHLO userp1040.oracle.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1752459AbdEKUra (ORCPT ); Thu, 11 May 2017 16:47:30 -0400 Subject: Re: [v3 0/9] parallelized "struct page" zeroing From: Pasha Tatashin To: Michal Hocko Cc: linux-kernel@vger.kernel.org, sparclinux@vger.kernel.org, linux-mm@kvack.org, linuxppc-dev@lists.ozlabs.org, linux-s390@vger.kernel.org, borntraeger@de.ibm.com, heiko.carstens@de.ibm.com, davem@davemloft.net References: <1494003796-748672-1-git-send-email-pasha.tatashin@oracle.com> <20170509181234.GA4397@dhcp22.suse.cz> <20170510072419.GC31466@dhcp22.suse.cz> <3f5f1416-aa91-a2ff-cc89-b97fcaa3e4db@oracle.com> <20170510145726.GM31466@dhcp22.suse.cz> Message-ID: <9088ad7e-8b3b-8eba-2fdf-7b0e36e4582e@oracle.com> Date: Thu, 11 May 2017 16:47:05 -0400 User-Agent: Mozilla/5.0 (X11; Linux x86_64; rv:52.0) Gecko/20100101 Thunderbird/52.1.0 MIME-Version: 1.0 In-Reply-To: Content-Type: text/plain; charset=utf-8; format=flowed Content-Language: en-US Content-Transfer-Encoding: 7bit X-Source-IP: aserv0022.oracle.com [141.146.126.234] Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org >> >> Have you measured that? I do not think it would be super hard to >> measure. I would be quite surprised if this added much if anything at >> all as the whole struct page should be in the cache line already. We do >> set reference count and other struct members. Almost nobody should be >> looking at our page at this time and stealing the cache line. On the >> other hand a large memcpy will basically wipe everything away from the >> cpu cache. Or am I missing something? >> Here is data for single thread (deferred struct page init is disabled): Intel CPU E7-8895 v3 @ 2.60GHz 1T memory ----------------------------------------- time to memset "struct pages in memblock: 11.28s time to init "struct pag"es: 4.90s Moving memset into __init_single_page() time to init and memset "struct page"es: 8.39s SPARC M6 @ 3600 MHz 1T memory ----------------------------------------- time to memset "struct pages in memblock: 1.60s time to init "struct pag"es: 3.37s Moving memset into __init_single_page() time to init and memset "struct page"es: 12.99s So, moving memset() into __init_single_page() benefits Intel. I am actually surprised why memset() is so slow on intel when it is called from memblock. But, hurts SPARC, I guess these membars at the end of memset() kills the performance. Also, when looking at these values, remeber that Intel has twice as many "struct page" for the same amount of memory. Pasha