From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S933894AbcI2Oqm (ORCPT ); Thu, 29 Sep 2016 10:46:42 -0400 Received: from Galois.linutronix.de ([146.0.238.70]:55698 "EHLO Galois.linutronix.de" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S933575AbcI2Oqe (ORCPT ); Thu, 29 Sep 2016 10:46:34 -0400 Date: Thu, 29 Sep 2016 10:43:54 -0400 (EDT) From: Thomas Gleixner To: Peter Zijlstra cc: Steven Rostedt , mingo@kernel.org, juri.lelli@arm.com, xlpang@redhat.com, bigeasy@linutronix.de, linux-kernel@vger.kernel.org, mathieu.desnoyers@efficios.com, jdesfossez@efficios.com, bristot@redhat.com, Ingo Molnar Subject: Re: [PATCH -v2 1/9] rtmutex: Deboost before waking up the top waiter In-Reply-To: <20160926154112.GH5016@twins.programming.kicks-ass.net> Message-ID: References: <20160926123213.851818224@infradead.org> <20160926124127.863639194@infradead.org> <20160926111511.1d963075@grimm.local.home> <20160926152228.GE5016@twins.programming.kicks-ass.net> <20160926113503.7d0528de@grimm.local.home> <20160926113727.4c08c58c@grimm.local.home> <20160926154112.GH5016@twins.programming.kicks-ass.net> User-Agent: Alpine 2.20 (DEB 67 2015-01-07) MIME-Version: 1.0 Content-Type: text/plain; charset=US-ASCII Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org On Mon, 26 Sep 2016, Peter Zijlstra wrote: > On Mon, Sep 26, 2016 at 11:37:27AM -0400, Steven Rostedt wrote: > > On Mon, 26 Sep 2016 11:35:03 -0400 > > Steven Rostedt wrote: > > > > > Especially now that the code after the spin_unlock(&hb->lock) is now a > > > critical section (preemption is disable). There's nothing obvious in > > > futex.c that says it is. > > > > Not to mention, this looks like it will break PREEMPT_RT as wake_up_q() > > calls sleepable spin locks. > > What locks would that be? None :) It still breaks RT in the futex case due to: deboost = rt_mutex_futex_unlock(); spin_unlock(&hb->lock); .... migrate_enable(); if (in_atomic()) return; So the migrate_disable() which was emitted by spin_lock(&hb->lock) will not be cleaned up and we leak the migrate disable count. We can work around that, but it's not pretty. As a related note, Sebastian decoded another possible priority inversion issue in the futex mess. T1 holds futex T2 blocks on futex and boosts T1 T1 unlocks futex and holds hb->lock T1 unlocks rt mutex, so T1 has no more pi waiters T3 blocks on hb->lock and adds itself to the pi waiters list of T1 T1 unlocks hb->lock and deboosts itself T4 preempts T1 so the wakeup of T2 gets delayed ..... We tried to fix it with a preempt_disable() and that's where we ran into that migrate_enable() hickup. We have a non deboosting variant for spin_unlock() for now, but we'll have to revisit that anyway ... Thanks, tglx