From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S935024AbbCPSxq (ORCPT ); Mon, 16 Mar 2015 14:53:46 -0400 Received: from mail.efficios.com ([78.47.125.74]:50539 "EHLO mail.efficios.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1754010AbbCPSxm (ORCPT ); Mon, 16 Mar 2015 14:53:42 -0400 Date: Mon, 16 Mar 2015 18:53:35 +0000 (UTC) From: Mathieu Desnoyers To: Peter Zijlstra Cc: linux-kernel@vger.kernel.org, KOSAKI Motohiro , Steven Rostedt , "Paul E. McKenney" , Nicholas Miell , Linus Torvalds , Ingo Molnar , Alan Cox , Lai Jiangshan , Stephen Hemminger , Andrew Morton , Josh Triplett , Thomas Gleixner , David Howells , Nick Piggin Message-ID: <1003922584.10662.1426532015839.JavaMail.zimbra@efficios.com> In-Reply-To: <20150316172104.GH21418@twins.programming.kicks-ass.net> References: <1426447459-28620-1-git-send-email-mathieu.desnoyers@efficios.com> <20150316141939.GE21418@twins.programming.kicks-ass.net> <1203077851.9491.1426520636551.JavaMail.zimbra@efficios.com> <20150316172104.GH21418@twins.programming.kicks-ass.net> Subject: Re: [RFC PATCH] sys_membarrier(): system/process-wide memory barrier (x86) (v12) MIME-Version: 1.0 Content-Type: text/plain; charset=utf-8 Content-Transfer-Encoding: 7bit X-Originating-IP: [173.246.22.116] X-Mailer: Zimbra 8.0.7_GA_6021 (ZimbraWebClient - FF36 (Linux)/8.0.7_GA_6021) Thread-Topic: sys_membarrier(): system/process-wide memory barrier (x86) (v12) Thread-Index: M0redrOaf4FfJheMbj743NBEqT6RAw== Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org ----- Original Message ----- > From: "Peter Zijlstra" > To: "Mathieu Desnoyers" > Cc: linux-kernel@vger.kernel.org, "KOSAKI Motohiro" , "Steven Rostedt" > , "Paul E. McKenney" , "Nicholas Miell" , > "Linus Torvalds" , "Ingo Molnar" , "Alan Cox" > , "Lai Jiangshan" , "Stephen Hemminger" > , "Andrew Morton" , "Josh Triplett" , > "Thomas Gleixner" , "David Howells" , "Nick Piggin" > Sent: Monday, March 16, 2015 1:21:04 PM > Subject: Re: [RFC PATCH] sys_membarrier(): system/process-wide memory barrier (x86) (v12) > > On Mon, Mar 16, 2015 at 03:43:56PM +0000, Mathieu Desnoyers wrote: > > > On which; I absolutely hate that rq->lock thing in there. What is > > > 'wrong' with doing a lockless compare there? Other than not actually > > > being able to deref rq->curr of course, but we need to fix that anyhow. > > > > If we can make sure rq->curr deref could be done without holding the rq > > lock, then I think all we would need is to ensure that updates to rq->curr > > are surrounded by memory barriers. Therefore, we would have the following: > > > > * When a thread is scheduled out, a memory barrier would be issued before > > rq->curr is updated to the next thread task_struct. > > > > * Before a thread is scheduled in, a memory barrier needs to be issued > > after rq->curr is updated to the incoming thread. > > I'm not entirely awake atm but I'm not seeing why it would need to be > that strict; I think the current single MB on task switch is sufficient > because if we're in the middle of schedule, userspace isn't actually > running. > > So from the point of userspace the task switch is atomic. Therefore even > if we do not get a barrier before setting ->curr, the expedited thing > missing us doesn't matter as userspace cannot observe the difference. AFAIU, atomicity is not what matters here. It's more about memory ordering. What is guaranteeing that upon entry in kernel-space, all prior memory accesses (loads and stores) are ordered prior to following loads/stores ? The same applies when returning to user-space: what is guaranteeing that all prior loads/stores are ordered before the user-space loads/stores performed after returning to user-space ? > > > In order to be able to dereference rq->curr->mm without holding the > > rq->lock, do you envision we should protect task reclaim with RCU-sched ? > > A recent discussion had Linus suggest SLAB_DESTROY_BY_RCU, although I > think Oleg did mention it would still be 'interesting'. I've not yet had > time to really think about that. This might be an "interesting" modification. :) This could perhaps come as an optimization later on ? By the way, I now remember why we start from the mm_cpumask, and then double-check the mm: using the mm_cpumask serves as an approximation of the CPUs we need to double-check. Therefore, rather than grabbing the rq lock for all CPUs, we only need to grab it for CPUs that are in the mm_cpumask. Thanks, Mathieu -- Mathieu Desnoyers EfficiOS Inc. http://www.efficios.com