From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1752616AbdHJN1Q (ORCPT ); Thu, 10 Aug 2017 09:27:16 -0400 Received: from mx1.redhat.com ([209.132.183.28]:40522 "EHLO mx1.redhat.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1751914AbdHJN1P (ORCPT ); Thu, 10 Aug 2017 09:27:15 -0400 DMARC-Filter: OpenDMARC Filter v1.3.2 mx1.redhat.com DAF9F24B641 Authentication-Results: ext-mx09.extmail.prod.ext.phx2.redhat.com; dmarc=none (p=none dis=none) header.from=redhat.com Authentication-Results: ext-mx09.extmail.prod.ext.phx2.redhat.com; spf=fail smtp.mailfrom=longman@redhat.com Subject: Re: [RESEND PATCH v5] locking/pvqspinlock: Relax cmpxchg's to improve performance on some archs To: Peter Zijlstra Cc: Ingo Molnar , linux-kernel@vger.kernel.org, Pan Xinhui , Boqun Feng , Andrea Parri References: <1495633108-12818-1-git-send-email-longman@redhat.com> <20170810115034.ie65wfxepiq6noew@hirez.programming.kicks-ass.net> From: Waiman Long Organization: Red Hat Message-ID: <945c28c3-5779-c8c8-13bb-40477abd1f0e@redhat.com> Date: Thu, 10 Aug 2017 09:27:10 -0400 User-Agent: Mozilla/5.0 (X11; Linux x86_64; rv:52.0) Gecko/20100101 Thunderbird/52.2.0 MIME-Version: 1.0 In-Reply-To: <20170810115034.ie65wfxepiq6noew@hirez.programming.kicks-ass.net> Content-Type: text/plain; charset=utf-8 Content-Transfer-Encoding: 7bit Content-Language: en-US X-Greylist: Sender IP whitelisted, not delayed by milter-greylist-4.5.16 (mx1.redhat.com [10.5.110.38]); Thu, 10 Aug 2017 13:27:15 +0000 (UTC) Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org On 08/10/2017 07:50 AM, Peter Zijlstra wrote: > On Wed, May 24, 2017 at 09:38:28AM -0400, Waiman Long wrote: >> # of thread w/o patch with patch % Change >> ----------- --------- ---------- -------- >> 4 4053.3 Mop/s 4223.7 Mop/s +4.2% >> 8 3310.4 Mop/s 3406.0 Mop/s +2.9% >> 12 2576.4 Mop/s 2674.6 Mop/s +3.8% > Waiman, could you run those numbers again but with the below 'fixed' ? > >> @@ -361,6 +361,13 @@ static void pv_kick_node(struct qspinlock *lock, struct mcs_spinlock *node) >> * observe its next->locked value and advance itself. >> * >> * Matches with smp_store_mb() and cmpxchg() in pv_wait_node() >> + * >> + * The write to next->locked in arch_mcs_spin_unlock_contended() >> + * must be ordered before the read of pn->state in the cmpxchg() >> + * below for the code to work correctly. However, this is not >> + * guaranteed on all architectures when the cmpxchg() call fails. >> + * Both x86 and PPC can provide that guarantee, but other >> + * architectures not necessarily. >> */ > smp_mb(); > >> if (cmpxchg(&pn->state, vcpu_halted, vcpu_hashed) != vcpu_halted) >> return; > Ideally this Power CPU can optimize back-to-back SYNC instructions, but > who knows... Yes, I can run the numbers again. However, the changes here is in the slowpath. My current patch optimizes the fast path only and my original test doesn't stress the slowpath at all, I think. I will have to make some changes to stress the slowpath. Cheers, Longman