From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1751510AbdFFNI7 (ORCPT ); Tue, 6 Jun 2017 09:08:59 -0400 Received: from mx1.redhat.com ([209.132.183.28]:57022 "EHLO mx1.redhat.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1751305AbdFFNI5 (ORCPT ); Tue, 6 Jun 2017 09:08:57 -0400 DMARC-Filter: OpenDMARC Filter v1.3.2 mx1.redhat.com 56A79811A7 Authentication-Results: ext-mx03.extmail.prod.ext.phx2.redhat.com; dmarc=none (p=none dis=none) header.from=redhat.com Authentication-Results: ext-mx03.extmail.prod.ext.phx2.redhat.com; spf=pass smtp.mailfrom=pbonzini@redhat.com DKIM-Filter: OpenDKIM Filter v2.11.0 mx1.redhat.com 56A79811A7 Subject: Re: [PATCH RFC tip/core/rcu 1/2] srcu: Allow use of Tiny/Tree SRCU from both process and interrupt context To: Peter Zijlstra , "Paul E. McKenney" Cc: linux-kernel@vger.kernel.org, mingo@kernel.org, jiangshanlai@gmail.com, dipankar@in.ibm.com, akpm@linux-foundation.org, mathieu.desnoyers@efficios.com, josh@joshtriplett.org, tglx@linutronix.de, rostedt@goodmis.org, dhowells@redhat.com, edumazet@google.com, fweisbec@gmail.com, oleg@redhat.com, kvm@vger.kernel.org, Linus Torvalds References: <20170605220919.GA27820@linux.vnet.ibm.com> <1496700591-30177-1-git-send-email-paulmck@linux.vnet.ibm.com> <20170606105343.ibhzrk6jwhmoja5t@hirez.programming.kicks-ass.net> From: Paolo Bonzini Message-ID: Date: Tue, 6 Jun 2017 15:08:50 +0200 User-Agent: Mozilla/5.0 (X11; Linux x86_64; rv:52.0) Gecko/20100101 Thunderbird/52.1.0 MIME-Version: 1.0 In-Reply-To: <20170606105343.ibhzrk6jwhmoja5t@hirez.programming.kicks-ass.net> Content-Type: text/plain; charset=utf-8 Content-Language: en-US Content-Transfer-Encoding: 8bit X-Greylist: Sender IP whitelisted, not delayed by milter-greylist-4.5.16 (mx1.redhat.com [10.5.110.27]); Tue, 06 Jun 2017 13:08:57 +0000 (UTC) Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org On 06/06/2017 12:53, Peter Zijlstra wrote: > On Mon, Jun 05, 2017 at 03:09:50PM -0700, Paul E. McKenney wrote: >> There would be a slowdown if 1) fast this_cpu_inc is not available and >> cannot be implemented (this usually means that atomic_inc has implicit >> memory barriers), > > I don't get this. > > How is per-cpu crud related to being strongly ordered? > > this_cpu_ has 3 forms: > > x86: single instruction > arm64,s390: preempt_disable()+atomic_op > generic: local_irq_save()+normal_op > > Only s390 is TSO, arm64 is very much a weak arch. Right, and thus arm64 can implement a fast this_cpu_inc using LL/SC. s390 cannot because its atomic_inc has implicit memory barriers. s390's this_cpu_inc is *faster* than the generic one, but still pretty slow. >> and 2) local_irq_save/restore is slower than disabling >> preemption. The main architecture with these constraints is s390, which >> however is already paying the price in __srcu_read_unlock and has not >> complained. > > IIRC only PPC (and hopefully soon x86) has a local_irq_save() that is as > fast as preempt_disable(). 1 = arch-specific this_cpu_inc is available 2 = local_irq_save/restore as fast as preempt_disable/enable If either 1 or 2 are true, this patch makes SRCU faster or equal x86 (single instruction): 1 = true, 2 = false -> ok arm64 (weakly ordered): 1 = true, 2 = false -> ok powerpc: 1 = false, 2 = true -> ok s390: 1 = false, 2 = false -> slower For other LL/SC architectures, notably arm, fast this_cpu_* ops not yet available, but could be written pretty easily. >> A valid optimization on s390 would be to skip the smp_mb; >> AIUI, this_cpu_inc implies a memory barrier (!) due to its implementation. > > You mean the s390 this_cpu_inc() in specific, right? Because > this_cpu_inc() in general does not imply any such thing. Yes, of course, this is only for s390. Alternatively, we could change the counters to atomic_t and use smp_mb__{before,after}_atomic, as in the (unnecessary) srcutiny patch. That should shave a few cycles on x86 too, since "lock inc" is faster than "inc; mfence". For srcuclassic (and stable) however I'd rather keep the simple __this_cpu_inc -> this_cpu_inc change. Paolo