From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S933098AbdHVPfa convert rfc822-to-8bit (ORCPT ); Tue, 22 Aug 2017 11:35:30 -0400 Received: from mx1.redhat.com ([209.132.183.28]:47784 "EHLO mx1.redhat.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S932991AbdHVPf3 (ORCPT ); Tue, 22 Aug 2017 11:35:29 -0400 DMARC-Filter: OpenDMARC Filter v1.3.2 mx1.redhat.com 6878F7E42C Authentication-Results: ext-mx03.extmail.prod.ext.phx2.redhat.com; dmarc=none (p=none dis=none) header.from=redhat.com Authentication-Results: ext-mx03.extmail.prod.ext.phx2.redhat.com; spf=fail smtp.mailfrom=longman@redhat.com Subject: Re: [RESEND PATCH v5] locking/pvqspinlock: Relax cmpxchg's to improve performance on some archs To: Peter Zijlstra , Will Deacon Cc: Ingo Molnar , linux-kernel@vger.kernel.org, Pan Xinhui , Boqun Feng , Andrea Parri , Paul McKenney References: <20170810161524.2wzocpcxrliy7nt6@hirez.programming.kicks-ass.net> <7cb318a8-d5b9-0019-a537-1720fc5222cc@redhat.com> <73daa6e6-537e-b0ce-e1e0-7afa75334509@redhat.com> <20170811090601.2owslxi4lgv3kond@hirez.programming.kicks-ass.net> <20170814120121.GA24249@arm.com> <20170814184711.GL6524@worktop.programming.kicks-ass.net> <20170815184034.GD10801@arm.com> <20170821105508.j3p4zv7mdojpyb7e@hirez.programming.kicks-ass.net> <20170821180001.GA22335@arm.com> <20170821192550.3dbj3jbgl33v2eeg@hirez.programming.kicks-ass.net> <20170821194246.GA32112@worktop.programming.kicks-ass.net> From: Waiman Long Organization: Red Hat Message-ID: <64b4957e-34f3-cc54-94bd-7ab64334f590@redhat.com> Date: Tue, 22 Aug 2017 11:35:28 -0400 User-Agent: Mozilla/5.0 (X11; Linux x86_64; rv:52.0) Gecko/20100101 Thunderbird/52.2.0 MIME-Version: 1.0 In-Reply-To: <20170821194246.GA32112@worktop.programming.kicks-ass.net> Content-Type: text/plain; charset=utf-8 Content-Transfer-Encoding: 8BIT Content-Language: en-US X-Greylist: Sender IP whitelisted, not delayed by milter-greylist-4.5.16 (mx1.redhat.com [10.5.110.27]); Tue, 22 Aug 2017 15:35:29 +0000 (UTC) Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org On 08/21/2017 03:42 PM, Peter Zijlstra wrote: > On Mon, Aug 21, 2017 at 09:25:50PM +0200, Peter Zijlstra wrote: >> On Mon, Aug 21, 2017 at 07:00:02PM +0100, Will Deacon wrote: >>>> No, I meant _from_ the LL load, not _to_ a later load. >>> Sorry, I'm still not following enough to give you a definitive answer on >>> that. Could you give an example, please? These sequences usually run in >>> a loop, so the conditional branch back (based on the status flag) is where >>> the read-after-read comes in. >>> >>> Any control dependencies from the loaded data exist regardless of the status >>> flag. >> Basically what Waiman ended up doing, something like: >> >> if (cmpxchg_relaxed(&pn->state, vcpu_halted, vcpu_hashed) != vcpu_halted) >> return; >> >> WRITE_ONCE(l->locked, _Q_SLOW_VAL); >> >> Where the STORE depends on the LL value being 'complete'. >> pn->state == vcpu_halted is the prerequisite of putting _Q_SLOW_VAL into the lock. The order of writing vcpu_hashed into pn->state doesn't really matter. The cmpxchg_relaxed() here should synchronize with the cmpxchg() in pv_wait_node(). Cheers, Longman